Early AI · AI developer tools · Sole designer · since 2024
Made an AI’s reasoning readable enough for McAfee to adopt it.
AI writes more code than teams can verify. Early generates the tests. My job was to make the model's reasoning legible enough that an engineer would accept it, and visible enough that a manager would sign off on it.
- Role
- Product designer, end to end. Sole designer.
- Timeframe
- Since 2024
- Team
- COO · Founder · CTO · Engineers (3)
- Surfaces
- IDE extension · Web platform · Brand · Marketing site
The short version
Plenty of tools generate code. Almost nothing validates it, so enterprises absorb the difference as regression, and fragmented workflows kill adoption before a pilot ends. I joined a raw VS Code extension with no UX foundation, no research artefacts and no visibility into what the model had decided, while the company pivoted from selling to developers to selling to enterprises.
Finding out where the trust actually breaks.
There were no research artefacts when I joined, and a highly technical audience with unclear expectations. Four tracks in parallel: interviews with CTOs, engineers and pilot clients; workflow analysis across GitHub, pull requests and CI/CD; teardowns of Cursor, Claude, Diffblue and Copilot; and a review of error logs to find the exact moment a developer stops believing the output.
Insights from pilot calls, grouped rather than counted
Recurring themes beat the loudest single complaint. The affinity map is how I stopped designing for whoever spoke last.
Two journey maps, because there were two jobs
Developers and managers fail at different points. One map would have hidden the manager's problem entirely, and the manager is the person who signs.
Nine tools torn down, one gap in all of them
Claude, Copilot, Cursor, Tabnine, Qodo, CodeRabbit and three more, read for what they do to a test rather than to a line of code. Every one of them generated. None of them reasoned about coverage. I ran the same pass across their interfaces and their marketing sites, because the pivot needed patterns an enterprise buyer would already recognise.
I still read every generated test line by line before committing.Developer
Code can look fully tested while still hiding risk.Developer
I can't see what the AI assumed or how developers rely on it.Manager
If production breaks, it's still on me, AI or not.Manager
One call decided the rest of the work.
Stay a developer tool inside the IDE, or build a web platform for team visibility and governance? I was in the room with the CEO and engineering leadership. We chose both, and I designed both. The extension keeps developers in flow; the platform carries what a buyer needs and a developer never opens. Two surfaces from one designer meant the platform had to stay deliberately flat, and I built the design system and its documentation before the first product screen.
The MVP platform, flat on purpose
Four areas only, matching the four questions a buyer actually asks: who is on my team and what may they do, what does it cost, what is it covering, can I prove it upward. Designed with the CEO and CTO, for a single developer and a whole department on the same screens.
Entry kept to a single screen
Sign-in only. Everything that could wait until after first value waited until after first value.
Team creation, end to end
A three-step setup wizard, plan selection, checkout and confirmation. The full flow at the length it really is.
Billing, seats and credit usage
Current month, seat management, credits, invoices, subscription. The screen a finance approver opens once and judges the product by.
Coverage analytics through GitHub
Monitor repository coverage, trigger test runs. This is the screen that made the governance conversation possible, and with it the McAfee deal.
Make the AI a junior team member showing their work.
Not a black box producing output. Every decision renders as a readable step with a confidence signal, low confidence is made visible rather than smoothed over, and the developer can edit or reject at any point. Getting this wrong produces one of two failures, both bad: users ignore the AI entirely, or they follow it without thinking.
The extension, rebuilt around transparency
Tests arrive with the reasoning attached and controls to refine, inside the editor where the work already happens.
Before, during and after the model writes
Pre-generation prompts and smart defaults to steer it. Error states that say what failed and why. Post-generation refinement and immediate regeneration.
A brand that stops reading as a side project
The original kit was built for developers: minimal to the point of thin, with a typeface that did not read as infrastructure. The typeface went first, since it did the most damage per pixel, then patterns and graphics, then one system across product, marketing and site.
A marketing site that could be re-cut as positioning moved
I designed and launched it in Webflow to reposition the company for an enterprise buyer. It has been through several iterations since, each one tracking a shift in what we were selling.
What it moved, and what it did not.
The governance and coverage analytics gave enterprise buyers enough transparency to proceed, and the platform helped the company secure additional funding and its first enterprise contracts, including McAfee. Post-launch feedback confirmed demand, and the same conversations exposed real weaknesses around AI trust, visibility and day-to-day usability, which fed the next iteration.
Then the differentiator shipped inside someone else's product.
Claude Code released native unit-test generation directly into developer workflows, which attacked the thing Early was differentiated on. We opened a new discovery cycle to find unmet needs beyond test generation and redefine the value in an AI-native landscape. That work is live and unresolved. I keep it in the case because a case that ends with everything working has stopped being useful to read.
The reassessment, as written at the time
Kept verbatim rather than rewritten with hindsight.