Blog
When agents build the UI, taste becomes pass/fail rules

When agents build the UI, taste becomes pass/fail rules

We wrote the design system before designing any of the fourteen prototypes in our gallery: a fourteen-step ink ramp, one indigo accent, four fixed status meanings, and a composer on every surface, each rule phrased so a screen passes it or fails it. Coding agents write most of the prototypes’ UI, and a builder with no memory of last week’s taste decisions drifts on any rule that leaves room for judgment. The system travels as a shadcn registry, so an agent scaffolding a new screen pulls the real components and tokens instead of improvising its own. Fourteen separate explorations open looking like one studio made them because of it.

Updated
July 21, 2026
Reading Time
8 min
AI Hero Design System — an interactive prototype. Click to open it live.

Two good screens that don’t belong to the same product

A coding agent that builds you a settings screen on Monday and another on Thursday will hand you two good screens that don’t belong to the same product. Different spacing, different button treatments, a new colour that seemed helpful in the moment. The agent isn’t wrong either time. It has no memory of the taste decisions you made last week, so it makes fresh ones, and fresh decisions drift. The same failure hides inside any rule with a judgment call in it. A design system written for human designers could afford “use elevation sparingly,” because a designer reads that roughly the same way twice; a builder that writes code and starts fresh on every generation cannot. A rule that passes or fails doesn’t drift.

So when agents build the UI, taste has to be compiled down to assertions a screen can pass or fail, and served where the builder actually pulls from. That is what the AI Hero design system is, and the receipt is the gallery: fourteen interactive prototypes, designed as separate explorations, that open looking like one studio made them. To be plain about its status, the system is a design-sprint artifact. It is the shared language behind those prototypes, not a product we ship to customers. Coding agents wrote most of the prototypes’ UI, and at agent speed a human can’t sit in review catching every spacing choice on every generated screen, so we settled the system before designing the first app. First the four principles, then the status vocabulary, then how the system reaches the agents.

Four principles, each one line long

Each principle is one line long and stated as a rule you can check a screen against.

Restraint over decoration. If a divider, shadow, or colour can be removed without loss, remove it. In practice: no drop shadows piled on cards, elevation is a single one-pixel hairline reserved for modals and the composer, and a cool neutral ramp of fourteen steps carries almost the entire interface. One type family, Geist, runs from the largest heading to the smallest label, with its monospace cut for the little uppercase tags that organize a screen without adding a box. Motion is rationed to genuine moments of confirmation.

One accent per surface. A single saturated indigo marks the primary action, the focused element, and links. The rule allows it on at most one or two affordances per surface and nothing else. In the drafting prototype the one indigo thing is “New draft”; in the engineering view it’s “New spec PR.” Someone who has used one prototype can open another they’ve never seen and find the primary action in the first second. Add a second accent and that signal is gone.

Write like a consulting partner. The voice rules are as checkable as the visual ones. Short sentences, full stops, no exclamation marks, specifics over adjectives. When a prototype reports what its agent did, it reads like a colleague: “Paused the schedule so it stops retrying and drafted an incident note for finance.” A small lexicon is simply banned: seamless, revolutionary, game-changing, and even AI-powered. Software that takes actions on your behalf should spend its words reporting what it did.

The composer is the page. Every surface gives the user a way to send a paragraph. The instrument is the voice bar, one object that takes text, files, and voice, with the default prompt “Send a paragraph…” It’s the one element granted real elevation off the page, and the layout is arranged around it. Software built around agents needs a place to state intent in plain language more than it needs another toolbar, and making that place the centerpiece is what separates these designs from a chat widget parked in a corner.

Rose belongs to the machine

On top of the ramp and the accent sits a tiny status vocabulary, and every colour in it has one fixed meaning across every prototype. Green means live and healthy. Amber means pay attention: a threshold approaching, a decision that can wait but shouldn’t wait long. Beta borrows the indigo. And rose marks the autonomous agent, the agent speaking or acting, as distinct from you. Rose belongs to the machine and nothing else. Status never appears as a bare coloured dot; it’s always a bordered, tinted badge that reads without a legend.

Three rows of swatches: a fourteen-step monochrome ink ramp from paper white to near-black, a single indigo accent swatch, and four status badges: green for live and healthy, amber for pay attention, rose for actions the agent took, and an indigo tint for beta.

The visual vocabulary: one ramp, one accent, four fixed status meanings. Same swatches, same meanings, every prototype, so a coding agent can apply them without asking.

The rose convention does the most work, because it answers the question agentic software raises constantly: did I do this, or did the agent? A rose badge answers it before you’ve read a word of the entry, and a user who learned the convention in one prototype reads it identically in the next without being taught. For the coding agents the fixed meanings matter just as much. A palette this small, with meanings this rigid, is something an agent can apply correctly on a screen nobody had imagined a week earlier. A rainbow with per-surface conventions is something it would have to guess at, and it would guess differently each time.

Served where the builder pulls from

Writing the rules down isn’t enough if they live in a document the agent never opens. The system is packaged the way our tooling already works: as a shadcn registry, the CLI’s distribution format in which a registry.json manifest lists components, tokens, and files that the shadcn add command installs into a project, with support for private, authenticated registries. When a coding agent scaffolds a new screen, it pulls the real components and tokens from the system instead of improvising its own. The principles travel as the written rules above; the pixels travel as code the agent installs. A new prototype starts from the system by default, and drifting away from it takes deliberate effort rather than a moment of inattention.

We’re not the only ones moving the rules to where the code gets written. Figma’s design-systems team describes the same move: mapping components to code with Code Connect and generating structured rules files so that an agent generating code reuses existing components and applies design tokens automatically, and a category of MCP servers now exists whose whole job is handing design-system context to coding agents. The common thread is the delivery: the system has to reach the builder in the builder’s own channel, or it doesn’t govern anything.

System first, designs second

The ordering is the part we’d repeat even if everything else changed: settle the system first, then design. The custom part of each prototype is the workflow (what the software does, how the agent behaves, where the person stays in charge), and none of that time should go to re-deciding what a button looks like. With the system in place, none of it does. This is also the supervision inversion from the fourteen-prototypes essay applied to our own build: the coding agent does the building, and the person’s control lives at the boundary, in rules written in advance, rather than in reviewing every generated screen after the fact.

What the refusals buy

Every rule above says no to something: the second accent, the decorative shadow, the exclamation mark. The refusals are what keep generated surfaces calm and consistent. When the agent has just paused a running job, the interface should carry composure, and a quiet monochrome surface with one point of colour does that better than a dashboard fighting for your attention.

The payoff sits in the gallery. The triage console, the eval dashboard, the drafting tool, and the engineering view were built as separate explorations, yet they open looking like one body of work, and the point of settling the system early is that software we build for a customer six months from now will still feel like the same hands made it. The business argument for fitted software, and for the team that keeps it current, lives in the bespoke SaaS post; this one is the part that decides what the screens look like. Open any two prototypes in the gallery side by side and check: the grammar stays the same even as the nouns change.

Article byRahul Parundekar

Rahul Parundekar

San Francisco-based consultant specializing in cutting-edge Generative AI (GenAI). I partner with organizations to pinpoint high-impact opportunities, streamline AI operations, and accelerate the launch of innovative products—efficiently, cost-effectively, and with controlled risk. Founder of Elevate.do and A.I. Hero, Inc.