Product Design | AI UX | Design System | Live Prototype

CONTEXT
A rep owns 100+ accounts, and each account's agent can surface two or three tasks a day — more work than anyone can do. The agent already does 80% of it; the rep still has to get back up to speed, review what was produced, correct it where it's wrong, and cover the last mile. At 8:55 in the morning the question isn't "what did the agent do?" It's "what am I doing today?" — and a longer list makes that harder, not easier. A day holds maybe ten slots of real attention; the plan that follows recommends eight, sized to fit.
100+ accounts
One rep, one always-on agent per account — reading email, calls, Salesforce, the calendar.
200+ candidates a day
The agents' overnight output. A first pass collapses it to 23 reviewable packages — still not a day.
10–15 minutes
The window a rep has, over coffee, to commit to a day with confidence.
The existing agent inbox is good at producing work, but no help in choosing it — every morning it hands the rep a pile, and the pile keeps growing. The brief asked for something that feels less like an AI-driven task list and more like working with a Chief of Staff. I took that literally: the last mile isn't generating more work — it's composing the day.
Scope I chose: an account executive reviewing at the desktop, 10–15 minutes. I designed the contract the ranking must fulfill — evidence, uncertainty, time cost — not the ranking model itself. And I assumed the agent is good but imperfect: sources are patchy, the CRM is messy, and the rep doesn't trust it yet.
SOLUTION
01 · THE FRAME
The sidebar looks like plain navigation — but it exposes the whole architecture. The product wears its own structure, one to one: read it top-down and it's the rep's day; read it bottom-up and it's the machine. And trust needs a floor you can inspect: every number opens downward — the "8" opens the day's cards, the "23" opens the inbox, every account name opens the record.
Today’s Mix
The front — where attention gets allocated.
Agent
Where the agent works — open the hood anytime.
Records
A familiar CRM underneath.
On top of the machine, two doors: Push — the agent hands you its overnight results each morning — and Pull — you ask for anything, anytime. In the product the tabs keep plain names; a rep at 8:55 shouldn't have to decode a metaphor.
02 · THE HOME
A feed pushes you to keep scrolling; a cockpit puts you in charge. The day opens as a three-sentence brief, and each sentence has a job. The agent proposes; the rep disposes.
The controls
Every word is clickable — flip 90 minutes to 45, switch to "between meetings" — and the day re-plans itself.
The track record
The agent shows what it already did overnight — before it asks you for anything.
The recommendation
Where to start, sized to your time and priced in pipeline — "I'd start with these 8, covering $594K."
The brief — context-aware; this is just what a morning looks like.
Below the brief sit the tasks. The agent prepared 23 and recommends starting with 8 — picked by matching the rep's situation against each task's time, impact, and confidence. The rep and the AI read the same dashboard. And like Spotify's Daily Mix, tasks come as ready-made sets, one per work scenario, grouped by the kind of judgment they need — not by account, not by recency. Each card carries proof, not promises: the draft's opening lines, the exact CRM fields to change, time cost, and the agent's own confidence.
Focus Moves
For deep attention — Warby Parker's stalled procurement, Vox Media's business case.
Quick Clears
Between meetings — Webflow's pricing reply, Chobani's opener. High confidence, two minutes each.
Judgment Calls
What only a human can weigh — Oscar Health's new VP: promising but unproven, so nothing is drafted yet.
Grainger's inconclusive noise is deliberately not in the mix — it lives in a one-line "the agent is watching…" note. Silence reads as coverage, not neglect. And the full library stays one click away: curation is a default, not a wall.
Each card is built to be read in a second — and I want to show how it got there. After defining the product, I gave Claude Code the framework as a markdown file, and it built the whole finished-looking screen in minutes. The structure was right — but the task card came back like this: plenty of information, not good design.
V1 · The AI draft
Crowded — pills fighting for attention; the hierarchy backwards; every card repeating a label the section header already gives.
V2 · The IA pass
I directed the next round to improve the information architecture — key numbers up front as a snapshot, icons to make it scan fast.
V3 · The shipped card
Then one more push: the icons became a thumbnail of the real content — people recognize it faster than a label; it's why Google Docs and Figma show large thumbnails — and the main button moved to the bottom right, where people expect it.
AI built the finished-looking screen in minutes; the details it has no feel for are where the design work lives.
03 · THE REVIEW
Everything on a task page was made by the agent — the rep's AI teammate — so the question becomes how to review a teammate's work without redoing it. Engineers solved this twenty years ago with the pull request: the change, the why, the evidence, and the unknowns, on one page. The drafted email is the diff, the CRM updates are the changeset, and what the agent is not sure about sits right beside the Approve button — not in a footnote.
Uncertainty at the decision
"Unsure about: whether James owns procurement" sits beside the button. The agent admits what it doesn't know before you find out.
Evidence one click deep
The why and the sources are always there — the email, the call, the CRM field — never a wall in the way.
Exception-only labels
Nothing gets a badge for simply running. A label appears only when something needs you — so silence means healthy.
Honest side effects
Everything an approval touches is stated at the button, and one click excludes it. Trust is mechanical, not rhetorical.
Request changes revises the draft live; Mark wrong files the miss. And correction comes with a visible receipt — "Preference learned," and what changes next time. Today's one-line correction buys tomorrow's trust. When something still feels off, the reasoning chain is one click away: which agents ran, what each did, what each read. The result alone can't tell you whether to trust it — the process carries the reliability.
One deliberate omission: there is no "agent accuracy score" anywhere. A rep doesn't trust a percentage — they trust a system that shows its work, states its doubts, and takes correction gracefully.
The reasoning chain — an audit trail, not required reading.
Even Defer is a feedback channel — postponed work returns at the right moment.
04 · THE OTHER DOOR
Everything so far runs one direction — the agent brings work to the rep. This door flips it. With a colleague you don't describe the thing, you point at it: hit Select — or drag the agent onto any card or draft — and that exact element binds to the question as context. Point first, then direct: "make this softer" · "what changed since last week?" No copy-pasting context, no chat buttons sprinkled on every card — one entry, one gesture, works on everything.
Select mode — the page dims, any element becomes a target, and the question binds to exactly that context.
Or drag the agent itself onto the target — same binding, one gesture.
05 · THE CHIEF OF STAFF
The agent produces more than anyone can check, so it works in layers: what needs no judgment, it just does — and logs. One signal it caught, it didn't touch: it wasn't sure, so it flagged the change and waited. History keeps two ledgers side by side — the agent's autonomous actions and the rep's own calls — so nobody loses the big picture.
Does the low-judgment work itself
Routine logging, enrichment, monitoring — handled overnight, receipts in the ledger.
Brings a short, efficient plan
Opinionated, sized to your time, priced in pipeline — a proposal, not an assignment.
Takes pushback in one sentence
Correcting it never costs more than a line — and the correction visibly sticks.
History — two ledgers, one page: the agent's autonomous actions beside the rep's own calls.
TRADE-OFFS
The ranking can be wrong
Every action is one-click reversible, every claim shows its source — I designed the rule the ranking must obey, not the model.
Autonomy could overreach
V1 does nothing outbound on its own; anything under the rep's name needs a click, and the ledger ships the same day.
Categories could mislabel
Judgment Calls are the "I'm not sure" lane — forced confidence is the real failure mode.
Ship the trust core first. Autonomy grows only as fast as trust.
METHOD
Research first — a seller's morning, and what Salesforce, HubSpot, Pipedrive, and Gong each do.
HubSpot
Pipedrive
Gong
Then ideation — naming where the morning actually breaks, sketching the Chief-of-Staff answer, and testing it against real scenarios until the shape held.
Everything was written down before any pixels: the product guide — the three layers, the record model, the decision categories, the tone — plus the task briefs and the build plan, a folder of .md specs that became the project's shared context.
Then a design system before any screens: a two-layer token model, primitives under semantic aliases, one action hue, with the documentation rendered from the product's own CSS so it can never drift from what ships.
Light and dark fall out of the same two-layer tokens — slide to compare.
Only then the build — with those specs and the design system loaded as context, the design landed through Claude Code: agents draft, panels score, and every change walks a real browser before it lands. All in, about ten hours: fifteen live routes, all seven scenarios from the brief. My judgment stayed the scarce resource.
See it run — the whole loop is live: the brief, the mix, the pull-request review, and the two-ledger history.