# TheDesignAgent field notes

> Working notes on the patterns, anti-patterns and judgment calls behind agent-built UI, from the team building TheDesignAgent, the judgment layer for agent-built UI. Overview, install and tools: https://www.thedesignagent.ai/llms.txt

## The best AI design agents in 2026, and which job each one does

> AI design tools now do five different jobs: generating UI from a prompt, bridging design files and code, shaping how a site looks, checking one layer such as accessibility, and working as a design agent inside the coding agent's loop. Here are ten worth knowing, grouped by job, with where each fits and where it stops.

Source: https://www.thedesignagent.ai/field-notes/best-ai-design-agents-2026 · Concepts · 2026-10-03

"AI design agent" now means at least five different things. Some tools generate UI from a prompt. Some move designs between Figma and code. Some shape how the coding agent styles a site. Some check a single layer, like accessibility. And a few work as a design agent inside the coding agent's loop, on every screen it builds. Asking which one is best is like asking whether a linter beats a framework. They do different jobs, and most teams shipping UI with Claude Code, Codex or Cursor end up using one from more than one group.

So this list is grouped by job, not ranked. We build one of the tools on it (TheDesignAgent, the last group). We've tried to describe every tool the way its own docs do, and to say plainly where each stops, ours included. Facts were checked against each product's documentation in October 2026; this space moves monthly, so we'll update the note when it does.

> **Short answer:** to *generate* a first screen, use v0, Lovable or Google Stitch. To *carry a design system* from Figma into code, use the Figma MCP server or Builder.io. To *shape how the site looks*, add a design skill like Anthropic's frontend-design or Impeccable. To *check a single layer*, add axe or Lighthouse. To give the coding agent *a design agent in its loop*, briefing every screen and reviewing it for UX and visual quality, add TheDesignAgent.

## Generate UI from a prompt

These tools turn a description into a working screen or app. They're the fastest way from nothing to something, and the most likely to produce the generic look people now call "AI-generated".

**1. v0 (Vercel).** Builds React and Next.js apps and UI from a prompt, and deploys them straight to Vercel. It now has an API and an official MCP server, so a coding agent can hand v0 a task and read back the result. *Where it fits:* fast, production-shaped React scaffolds. *Where it stops:* it generates; it doesn't critique, and it's most at home on the Vercel stack.

**2. Lovable.** A full-stack app builder with a backend and deploys built in. Its MCP server (launched May 2026) lets Claude Code, Cursor and ChatGPT drive it. *Where it fits:* getting a whole working app, not just screens. *Where it stops:* your project lives in Lovable's environment, and the UI follows its defaults unless you push back.

**3. Google Stitch.** Google Labs' Gemini-powered design canvas, the successor to Galileo AI. Since March 2026 it has an SDK, an official MCP server and DESIGN.md support, and its design agent can critique on the canvas. *Where it fits:* exploring visual directions quickly. *Where it stops:* the output is a design; a coding agent still has to build it.

**4. Claude Design (Anthropic Labs).** Launched in April 2026 as a research preview in claude.ai. It makes prototypes, slides and one-pagers with Claude, picks up a brand system from your code and design files, and hands off to Claude Code. *Where it fits:* designing *before* the build, in the same family as the agent that will build it. *Where it stops:* it's a separate surface from the build loop, not a tool the coding agent calls.

## Bridge design files and code

These tools keep a real design system in play: components, variables and tokens rather than invented ones.

**5. Figma MCP server.** Gives Claude Code, Codex and other agents the components, variables and layout of a real Figma file, and can now write back to the canvas (beta). *Where it fits:* teams whose source of truth is Figma. *Where it stops:* the agent's output is only as good as the file. It tells the agent what the design *is*, not whether the screen serves the user's job.

**6. Builder.io (Fusion).** A visual canvas on top of your repo and design system, with a Figma plugin, a VS Code extension and an MCP server. *Where it fits:* designers and developers editing the same real components. *Where it stops:* it's a larger platform to adopt, and usage is credit-metered.

**7. Magic Patterns.** Prototypes in React and Tailwind, with an official MCP server so a coding agent can read a prototype and diff it against your code. *Where it fits:* turning a prototype into an implementation brief. *Where it stops:* prototype first; production code is a second step.

## Shape how the site looks

These are skills that live in your repo and change how the coding agent designs while it writes code. They're about the craft of the site itself: aesthetic direction, typography, color, layout and polish.

**8. Anthropic's frontend-design skill.** A short, open-source Claude Code skill that makes Claude commit to a clear aesthetic direction before it writes UI code. It's one of the cheapest ways to cut the default "AI look". *Where it fits:* every Claude Code project that builds UI. *Where it stops:* it guides the styling but doesn't score or check the result, and the same guidance applies to every project.

**9. Impeccable.** An open-source design skill pack for Claude Code, Codex, Cursor and Copilot, focused on designing the site itself, with commands that include critique and audit plus deterministic detector rules. *Where it fits:* teams who want stronger visual craft from their agent with no service in the loop. *Where it stops:* its focus is how the site looks; it doesn't keep a model of your users and their jobs, so it can't tell you whether a well-made screen is the right screen.

## Check one layer

Some of the most useful tools in the loop check exactly one thing, and do it well.

**Accessibility and quality audits.** Deque's axe MCP server finds and fixes accessibility violations from inside Claude Code, Copilot and Cursor. Chrome DevTools MCP runs Lighthouse audits for accessibility, SEO and best practices. Both are essential and both are narrow: they tell you whether a screen is *valid*, not whether it's *right* for the person using it. There are also a growing number of community MCP servers that score screenshots for "AI slop".

## A design agent in the loop

This is the newest group and the one most teams are missing: a design agent that the coding agent works with on every screen, the way it works with a test runner or a type checker. Not one layer, and not just the look: the whole design pass, before and after the build.

**10. TheDesignAgent.** That's us. It's an MCP server, with a Claude Code plugin, that works in layers:

- **Brief, before the build.** `Discover` gives the agent the user's job, the heuristics and patterns that apply to this screen, and your design tokens.
- **UX, after the build.** `Ux` scores the screen from 0 to 10 for job fit, UX heuristics, cognitive load and pattern conformance, with findings and fixes.
- **Visual, after the build.** `Visual` scores brand compliance, hierarchy and aesthetics from a real screenshot of the running page.
- **Project memory, across builds.** Jobs, personas, tokens and past findings are kept server-side and shared by every agent on the repo, so each review shapes the next brief.

Findings cite the user's job rather than a generic rubric. *Where it fits:* teams who want the coding agent to catch the dashboard-for-everything, modal-for-everything defaults before a human has to. *Where it stops:* it doesn't generate designs or replace a designer; it needs an API key; and reviewing pages on localhost needs the local server (the hosted one screenshots public and preview URLs). Pair it with an accessibility checker.

## How to combine them

Most stacks that ship good UI with a coding agent combine three things:

1. **A source of truth** for the design system: a Figma file through the Figma MCP, or a DESIGN.md in the repo.
2. **A design agent in the loop** that briefs each screen and reviews it for UX and visual quality: TheDesignAgent.
3. **Single-layer checks** where they matter: axe or Lighthouse for accessibility and validity, and a style skill if you want a stronger aesthetic voice.

Generators sit beside this rather than inside it. They're great for the first screen, and the first screen is exactly when a review pays off most.

> **Rule of thumb:** if your agent's UI looks fine in a screenshot but users stall on it, you're missing the design agent in the loop, not a better generator. Read [the dashboard default](/field-notes/the-dashboard-default) for the most common case, and [adding design review to your agent](/field-notes/adding-design-review-to-your-agent) for the setup.

---

## Adding design review to your agent: what the workflow looks like

> How to add a design judgment step to Claude Code, Cursor, or any agent that builds UI. What to call, when to call it, how to read the response, and what the loop looks like across multiple builds.

Source: https://www.thedesignagent.ai/field-notes/adding-design-review-to-your-agent · Workflow · 2026-05-06

Most agent workflows that build UI don't have a design review step. The agent builds, lints, type-checks, and opens a PR. Design feedback comes later — in a human review cycle that runs at human speed, after the code is already written and the layout is already embedded in the codebase.

Adding a design judgment step to the build loop changes when feedback arrives. Instead of catching a dashboard-for-a-deciding-job after the PR is open, you catch it before the agent finishes the first build. The cost of the fix drops significantly: the agent patches against structured feedback rather than a human reviewer's comments, and it can loop — score, patch, re-evaluate — until the fit is good.

Here's what that workflow looks like with TheDesignAgent.

## The setup: one file in the repo

Adding TheDesignAgent to an existing project starts with a single call. The agent invokes the `/ux` or `/visual` MCP tool with the task description, the artifact (the UI code or a description of the UI being built), and optionally some codebase context. On the first call with no prior project context, discovery runs: the system infers the domain, personas, jobs, and day-shape from the input and returns a `.thedesignagent` file.

The agent commits that file to the repo root. Every subsequent call from any agent working on the repo — whether that's Claude Code, Cursor, or anything else — will pick up the project context automatically. Discovery doesn't re-run; the profile is there.

This is the only configuration step. There's no dashboard to set up, no schema to define, no settings to configure. The file is the config.

## The call: when and what

The design review call fits naturally into the agent's build loop at the point where the agent has a working UI artifact to evaluate. This is typically after the first build is functional but before the PR is opened.

```
Task: build a review interface for [description]
→ Agent builds first version
→ Agent calls /ux: { task, artifact, project_id }
→ Receives: fit_score, patches[], flags[]
→ Agent applies patches
→ Agent calls /ux again: { task, updated_artifact, project_id }
→ Receives: higher score, remaining patches
→ Loop until fit or diminishing returns
→ Agent opens PR
```

The `/ux` call evaluates whether the layout serves the job. The `/visual` call evaluates whether the design is well-crafted. Most builds benefit from both. The order matters: structural problems (JTBD, cognitive load) should be addressed before visual ones, because fixing the structure changes what visual patches are relevant.

## Reading the response

The response contains three things that drive the next action.

**The fit score** is a 0-100 number that tells you where the build is. It's directional — use it to understand the delta between calls, not to optimize toward 100. An 87 that came from 42 is a better signal than an 87 that's been at 86 for two calls.

**The patches** are the actionable output. They're ordered by priority — critical patches should be applied first because they often change what subsequent patches are even relevant. Each patch has a change (what to do) and a reason (why). The reason is important: it tells you whether the patch is addressing a structural problem or a visual one, which helps you understand whether the score will move significantly on the next call.

**The flags** are named anti-patterns. `layout_fights_action` means the layout structure is working against the job. `high_extraneous_load` means the interface is adding overhead the job doesn't require. Flags are diagnostic — they tell you the category of problem more quickly than reading the full patch list.

## What the loop looks like

A typical first-build evaluation for a non-trivial UI returns a score in the 45-65 range and a set of three to six patches with two or three at critical or high priority. The agent applies all of them and calls again.

The second evaluation typically scores 20-35 points higher. The remaining patches shift toward medium and low priority — visual craft issues, accessibility refinements, edge cases in the interaction model. If the score is still below 75 after the second pass, the remaining gap is usually in one or two specific sub-scores, and the patches for those are the ones to focus on.

Three calls is a reasonable target for most builds: init (discovery + first score), patch pass (major structural and cognitive fixes), refinement (visual and accessibility). Beyond three calls, the marginal gain per loop typically drops into single digits, and the remaining work is better addressed in a human review.

## Adding it to CLAUDE.md

If you want Claude Code to run the design review step automatically on UI builds — without being asked each time — add a hook to your project's CLAUDE.md:

```markdown
When building or significantly modifying UI components, call the TheDesignAgent 
/ux MCP tool after the first working build and apply critical and high-priority 
patches before opening a PR. Use the project_id from .thedesignagent if it exists.
```

This makes the design review step part of the agent's standard workflow for UI work rather than something you have to remember to request. The agent knows when it's building UI (it's working on components, pages, layouts); the instruction tells it to add the quality gate at that point.

The result is a build loop that includes design judgment alongside the other quality checks already in the pipeline — not as an afterthought, but as a step in the workflow.

---

## Affordance density: how much belongs on a row before it stops being readable

> Agents solve row density in one of two ways: everything visible (unreadable) or almost nothing visible and everything behind a click (high-ceremony). The right answer is a calibrated three-tier model.

Source: https://www.thedesignagent.ai/field-notes/affordance-density · Patterns · 2026-05-06

A data row is a unit of information. How much it holds before it stops being readable is a calibration question, and agents tend to land at one of two extremes: everything visible at once, or almost nothing visible with the rest behind clicks.

The dense version collapses into noise. The sparse version adds ceremony — the user has to expand or click into every row to find the detail that should have been glanceable. Neither is right, and both fail in the same underlying way: they're calibrated to the schema rather than to the user's triage pattern.

## The three-tier model

A well-calibrated row operates on three tiers, each corresponding to how quickly the user needs the information and how often they need it.

**Tier one — always visible, no interaction required.** The primary identifier (what is this record?), the key status or priority signal (does this record need my attention?), and the primary action (what would I do to this record right now?). This is the scanning layer. The user should be able to triage the entire list at this tier — moving quickly past the records that don't need attention and stopping at the ones that do — without a single click.

**Tier two — one interaction away.** Secondary context: related metadata, the most recent activity, a confidence signal with its supporting reasoning, secondary actions that apply to a subset of records. This tier reveals on row hover, expand, or a secondary button — something that requires intent but not navigation. The user who needs it gets it quickly. The user who doesn't need it isn't paying the visual cost of seeing it.

**Tier three — full detail on demand.** The full record, edit state, history, linked entities, audit trail. This is a navigation event — a click into a detail view or an in-place expansion that replaces the row with a richer panel. It represents a context switch and should be designed as one.

> **Pattern:** a well-calibrated row shows everything the user needs to decide at a glance and nothing they don't need until they ask. Tier one enables triage. Tier two enables context. Tier three enables deep work.

## Calibrating to the triage pattern

The right content for tier one depends on the job. For an expense approvals queue, tier one is: vendor name, amount, submitter, approve/reject. The approver is scanning for items that require judgment — anomalous amounts, unfamiliar vendors — and passing quickly through the routine ones. Showing the full expense details at tier one would require reading every row rather than scanning them.

For a support ticket queue, tier one is: ticket title, customer name, priority, assigned agent, age. The support lead is looking for anything urgent and unassigned. Everything else can wait for tier two.

The mistake agents make is using the same tier-one content regardless of job. A generic "name, status, date, actions" row is not calibrated to anything — it's the data model projected onto the row. What goes in tier one has to come from the job, not from the available columns.

## The hover pattern and its limits

Many table designs use hover states to reveal tier-two content. This is a reasonable pattern for desktop workflows with a predictable interaction model. Its limits are worth understanding.

Hover reveals create invisible affordances — the user has to discover that hovering reveals more. For users new to a tool, this creates a learning cost. For users switching between tools, the inconsistency creates re-learning cost. And for touch or stylus-primary contexts, hover doesn't exist.

A better pattern for tier-two actions is a contextual action bar that appears at the row level when the row has focus (keyboard or cursor selection), not on hover. It's more discoverable, works across input methods, and keeps the primary scan uncluttered. The visual difference is subtle. The interaction difference is significant.

## How TheDesignAgent scores it

The `/visual` lens scores `density_fit` — whether the information density of the layout matches the cognitive profile of the job. A low `density_fit` score on a dense row means elements are competing for attention. A low score on a sparse row means the user is paying a click-tax to access information they need frequently.

Patches on density tend to be paired: reduce tier-one content to primary identifier, key signal, and primary action; move secondary context to a hover or expand state. The specific fields moved are named in the patch. The agent gets "move [columns] to tier two" rather than "simplify the row."

---

## Cognitive load is three things, not one

> Most design principles treat cognitive load as a single dial to turn down. It isn't. CLT has three sources that behave differently — and reducing the wrong one makes a screen worse.

Source: https://www.thedesignagent.ai/field-notes/cognitive-load-is-three-things · Concepts · 2026-05-06

Most design principles treat cognitive load as a single dial. More complexity equals bad; less equals good. Reduce elements, simplify the layout, cut the features. This framing is both true and useless — true because load matters, useless because it treats all load as the same thing.

Cognitive Load Theory distinguishes three sources of mental effort. They don't just differ in name. They differ in cause, in whether they're worth reducing, and in what actually helps. Confuse them and you'll strip a screen of the friction it needed while leaving the friction it didn't.

## Intrinsic load: the task itself

Intrinsic load is the cognitive work the task inherently requires. Reviewing twenty reorder recommendations is intrinsically complex — there are twenty decisions to make, each with a confidence signal, a quantity, and a reasoning thread. No layout change eliminates that complexity. The job has twenty decisions; the user has to make twenty decisions.

This matters because intrinsic load is often mistaken for a design problem. Agents and designers reach for simplification when the task is actually just hard. The right response to high intrinsic load isn't to hide information — it's to structure it so the user can work through it without losing their place. Sequence, grouping, and chunking reduce the effort of navigating a complex task. They don't reduce the task.

## Extraneous load: friction the interface adds

Extraneous load is the cognitive cost of the interface itself — the effort that has no relationship to the job. A modal that interrupts a batch workflow adds extraneous load. A table that requires four clicks to get to the action adds extraneous load. A confirmation dialog for a low-stakes action that takes three seconds to undo adds extraneous load.

This is the load worth eliminating. It contributes nothing and costs focus. In high-throughput jobs — warehouse operators, support queues, anything with a batch cadence — extraneous load is fatal. It compounds across repetitions. Twenty interruptions per hundred rows isn't a minor friction; it's an hour of lost throughput per shift.

> **Pattern:** extraneous load is the only load worth eliminating unconditionally. Intrinsic load is the job; germane load builds the mental model. Extraneous load is just overhead.

## Germane load: the load that builds understanding

Germane load is the cognitive effort that produces something. Learning a workflow, building a mental model of how prices relate to inventory, understanding what a confidence score means — these require effort, but the effort is the point. Spent once, it pays off in speed every time after.

This is the load designers most commonly destroy by accident. Onboarding flows that over-explain every step, empty states that narrate exactly what to do, tooltips that activate on hover — these feel helpful but prevent the user from developing the model that would actually make them fast. Good germane load investment upfront produces expert users. Good intentions with bad execution produce users who are always dependent on the scaffolding.

## How agents misread the signal

Agents optimize for visible complexity. A screen with twenty rows, inline actions, a confidence column, and a reasoning toggle looks heavy. An agent flagged on "simplicity" will collapse the reasoning into a click, move the actions behind a modal, and add tabs to reduce visible row count. The screen is now simpler-looking and worse. The intrinsic load (twenty decisions) hasn't changed. The extraneous load has increased. The germane load for learning the pattern has been replaced with re-scaffolding via modals.

The misread happens because agents — like most tooling — can see layout complexity but can't see load type. They can count elements. They can't distinguish intrinsic from extraneous.

## How TheDesignAgent scores it

The `/ux` lens scores four dimensions. Three of them map directly to this framework: `intrinsic_load`, `extraneous_load`, `germane_load`. A low score on `extraneous_load` means the interface is adding overhead the job doesn't require. A low score on `germane_load` means the interface is either preventing model-building or front-loading expertise it should be developing.

These scores don't average into a single "load" number. A screen can score well on intrinsic and germane while failing hard on extraneous — that's the modal reflex pattern. A screen can score well on all three while failing on `jtbd_alignment` — the job is wrong even if the load is managed well.

The distinction is what makes the score actionable. "Reduce cognitive load" is too vague to patch against. "Collapse the modal — it's adding extraneous load on a batch workflow" is a patch.

---

## Component grammar: when shadcn defaults stop being neutral

> Shadcn's defaults are a position, not a blank slate. When agents apply them unchanged, the result reads as AI-built before the user consciously registers why.

Source: https://www.thedesignagent.ai/field-notes/component-grammar · Concepts · 2026-05-06

Every component library ships with defaults. Padding values, border radii, font sizes, spacing scales, color tokens. These defaults are not neutral — they're a position. They represent the library author's judgment about what looks competent at a general-purpose level. Shadcn's defaults are well-calibrated for that goal: they produce interfaces that look professional without looking distinctive.

The problem isn't that they're bad. The problem is that when an agent applies them unchanged across a product, the result reads as AI-built before any individual user could explain why. Something about it is familiar in the wrong way. It looks like the last three tools you used.

That's component grammar breaking down.

## What component grammar means

Component grammar is the visual logic that makes a product's components feel like they belong to each other and to the product. Not just "consistent" in the sense that buttons look alike — but coherent in the sense that the spacing, the weight distribution, the radius choices, and the density all communicate the same underlying sensibility.

A product with strong component grammar feels like it was designed by one hand, or at least by a team with a shared point of view. The components feel specific to the context: a dense internal operations tool feels different from a consumer-facing onboarding flow even when they share the same underlying component kit. Component grammar is what produces that difference.

When grammar is weak, the product feels assembled rather than designed. Each component looks fine individually. Together, they don't read as a system.

## How shadcn defaults flatten it

Shadcn's defaults cluster around a particular aesthetic: `0.5rem` border radius, `slate` palette, 14px body text, generous but not generous whitespace, medium border weight. These are defensible choices. They also appear in roughly half the agent-built products shipped in 2025.

The issue isn't the choices themselves — it's that they're applied uniformly, without reference to what the product is, who uses it, or what it needs to communicate. A warehouse operations tool and a consumer budgeting app have different grammar requirements. They need different density levels, different emphasis weight, different radius sensibilities. Shadcn defaults produce the same output for both.

> **Pattern:** defaults are a position. If the agent applies them unchanged, the product looks like the agent built it — because every other agent is making the same choices. Component grammar is how you encode the product's actual sensibility into the build.

## The signal that grammar is failing

Users don't usually say "your component grammar is weak." They say things like "it looks kind of generic," "it doesn't feel like our product," or "it's fine but I'm not excited about it." Designers say "the shadcn defaults are showing" and mean it as a diagnosis.

The more precise signal is in the details. Button padding that doesn't feel native to the content density. Input borders that read as default web rather than considered choice. A card radius that works at 12px in every other context but reads as wrong here. Typography hierarchy that's functional but not expressive. These are all component grammar failures. Each one is small. Together they produce a product that reads as assembled.

The hardest part is that the failures are invisible until you look for them. A product with weak grammar passes a quick scan. It's only when you sit with it — or when you put it next to a product with strong grammar — that the difference becomes obvious.

## How TheDesignAgent scores it

The `/visual` lens includes `component_grammar` as one of seven sub-scores. It asks: do the components feel like they belong to this product, or do they read as defaults applied without modification?

Low scores on `component_grammar` produce patches that are unusually specific: radius adjustments calibrated to the product's density, color token replacements that shift from slate to a more specific palette, spacing scale changes that match the product's actual content weight. These are changes the agent could make on its own — but won't, without being told what the product's sensibility actually is.

That specificity is what makes `component_grammar` one of the hardest scores to hit on a first build. The agent knows how to apply components. It doesn't know what the product is supposed to feel like. Project memory helps — once a product's visual decisions are captured, future builds get compared against them. But the first build is always working from the corpus, and the corpus can only tell the agent what's generally wrong. It can't tell it what's specifically right for you.

We'll write more about how to encode visual taste into project memory so that agents build to a higher baseline from the start.

---

## Confirm everything: why agents wrap routine actions in "Are you sure?"

> The confirmation dialog is the agent's safety net. Applied to every action, not just destructive ones, it creates an interrupt-heavy experience that slows down every workflow it was meant to protect.

Source: https://www.thedesignagent.ai/field-notes/confirm-everything · Anti-patterns · 2026-05-06

The most defensive button an agent will add to any interface is "Are you sure?" Confirm dialogs appear on archive actions, delete actions, publish actions, submit actions, and occasionally on close actions that navigate away from a form. Sometimes they appear with a second confirmation for good measure. The agent is being careful.

The problem is that care applied uniformly produces the same interface as carelessness — one where the user has tuned out the confirmation text entirely and clicks confirm on reflex. When that happens, the dialog isn't providing safety. It's providing the appearance of safety while adding a click to every action in the product.

## When confirmation is the right answer

Confirmation dialogs have a legitimate use. The criteria are specific: the action is destructive AND irreversible, and the consequence of the mistake is significant.

Permanently deleting a record with no recovery path. Cancelling a subscription with immediate effect. Revoking access to a system with no reinstatement flow. These warrant a confirm step — ideally one that requires the user to demonstrate intent rather than just click a button. Typing the item's name, or the word "delete," forces a brief moment of active decision. It slows down the accidental and doesn't slow down the intentional.

Note what's absent from that list: archiving a draft, submitting a form, approving a recommendation, sending a message that can be recalled, and anything that routes to a review queue before taking final effect. These don't qualify. They're routine actions in a workflow, not one-way doors.

> **Pattern:** confirmation dialogs are for actions that are both destructive and irreversible. For everything else, build undo instead — it provides real safety without interrupting the workflow.

## Undo is the right answer for reversible actions

For actions that can be undone, the correct safety mechanism is an undo affordance — typically a dismissible toast that appears immediately after the action completes and offers a brief window to reverse it. Gmail's send undo. Notion's delete undo. The pattern is well-established because it's genuinely better: the action completes immediately (fast, no friction), the user has a real recovery path if they made a mistake, and the workflow is only interrupted if there was actually a mistake.

This is a better safety net than a confirm dialog for reversible actions. Confirm dialogs interrupt before the action, requiring a decision about whether to proceed. Undo affordances interrupt only if the action was wrong, requiring a decision to reverse. For routine actions in high-volume workflows, the confirm approach is strictly worse.

## Why agents over-confirm

Agents default to confirmation dialogs for the same reason they default to modals: safety feels like a feature, and adding a confirm step is easy to scaffold. The agent sees a delete button and knows that deleting things can be bad. The confirm dialog is the obvious protective response.

What the agent doesn't model is the frequency at which the action occurs. Confirming a one-time account deletion is different from confirming every record archive in a batch workflow. The former is reasonable; the latter compiles into a significant time cost across a workday. An operator archiving fifty completed records confirms fifty times. None of those confirmations were necessary. All of them were interruptions.

## How TheDesignAgent flags it

The `/ux` lens flags confirmation dialogs on actions that don't meet the destructive-and-irreversible bar. The `extraneous_load` sub-score reflects the interrupt cost, and the patch is direct: replace the confirm dialog with an undo toast on this action, reserve confirm step for the delete-permanently flow. 

The distinction between "archive" and "delete permanently" is where agents most commonly conflate the two. Archive is reversible and should get undo. Delete permanently is not and should get a confirm. When both get a confirm dialog, the serious one reads the same as the routine one — and users learn to dismiss both without reading them.

---

## Day-shape: why one persona has many jobs, and your UI needs to know which one is loaded

> Most users don't have one job — they have several, and they shift between them across the workday. Designing for 'the user' without day-shape means designing for the average of all their jobs, which usually fits none of them well.

Source: https://www.thedesignagent.ai/field-notes/day-shape · Concepts · 2026-05-06

Most design work treats the persona as a person with a job. The warehouse manager. The inventory planner. The support lead. The job defines the cognitive requirements — time pressure, working memory load, expertise level, stakes of error — and the design follows from those requirements.

The problem is that most users don't have one job. They have several, and they cycle through them across a workday. The warehouse manager who reviews exceptions in the early morning is doing a different job than the same manager who runs a capacity check before lunch or handles a supplier escalation in the afternoon. Same person. Different job. Different cognitive profile. Different interface requirements.

This is what day-shape captures: the sequence of jobs a persona moves through in a typical day, the transitions between them, and how the cognitive state shifts at each transition.

## Why day-shape changes what "good design" means

The design implications of day-shape are concrete. A manager doing exception review at 8am is in a triage mode: high volume, fast decisions, low deliberation per item. The layout that supports this is a decision surface — inline actions, clear priority signals, minimal ceremony. The goal is throughput.

The same manager doing capacity planning at noon is in a deliberation mode: lower volume, more time per item, higher consequence per decision. The layout that supports this has more context, more space for analysis, less pressure toward quick commitment. The goal is accuracy.

If the product conflates these jobs — building one layout that tries to serve both — it typically serves neither. The decision surface is too fast for deliberation. The analysis view is too slow for triage. Users adapt by building workarounds: they export to spreadsheets, they screenshot data for reference, they context-switch out of the tool rather than through it.

> **Pattern:** day-shape is the sequence of jobs a persona cycles through in a typical workday. Each job has different cognitive requirements. A product that knows the day-shape can surface the right interface at the right moment instead of averaging across all of them.

## How day-shape is captured in project memory

During discovery — the first call where TheDesignAgent infers the product's context — day-shape is one of the fields built into the persona profile. For each persona, the profile includes: who they are, their cognitive characteristics (working memory, expertise, time pressure, multitasking demands), and their day-shape — an ordered or loosely-ordered map of the jobs they do and when.

A warehouse manager's day-shape might look like: exception review (morning, high urgency, triage mode) → fulfillment oversight (late morning, monitoring mode) → inventory reconciliation (noon, deliberation mode) → supplier coordination (afternoon, communication mode) → shift close (end of day, summary mode). Each of these has a different load profile. The discovery output captures all of them.

The `active_job_id` in each subsequent call then identifies which job in that map the current artifact is serving. The score reflects how well the artifact serves *this specific job*, not the average of all the jobs in the day-shape. A layout that's excellent for exception review will score differently when evaluated against inventory reconciliation — because the job requirements are different.

## Implications for navigation and layout switching

Day-shape awareness opens a design direction agents rarely reach: interfaces that shift state rather than sending the user to a different tool. If the product knows the user's day-shape, and the user's context shifts from triage mode to deliberation mode, the interface can adapt: surface more context, change the action prominence, show the analysis view instead of the decision surface.

This isn't a magic "mode" toggle. It's a design decision about which jobs the product can serve and how the transition between them is handled. Some products legitimately serve one job and should be designed for it. Others serve the full day-shape and should be designed for transitions.

The first step is knowing the day-shape exists. Most agent-built products don't reflect it because the prompt didn't include it. The build agent received "build a tool for the warehouse manager" and built for the warehouse manager as a single undifferentiated job. Day-shape is what's missing from that prompt — and what TheDesignAgent's discovery process tries to recover from context, even when it wasn't explicitly stated.

---

## Decision surfaces: putting the action where the eye already is

> A decision surface is a layout where each row is a unit of decision and the action is on the record. It's what agents should build when the job is deciding, not watching.

Source: https://www.thedesignagent.ai/field-notes/decision-surfaces · Patterns · 2026-05-06

The most common recommendation TheDesignAgent's patches produce is some variation of the same structural change: collapse the modal into the row, move the action inline, surface the confidence signal next to the record it applies to, let the running total update as decisions are made. These patches keep appearing across different domains and different agents because they're all solving the same underlying problem.

There's a name for what they're building toward: a decision surface. It's not a component or a widget — it's a layout pattern for jobs where the user's work is to decide on a list of things.

## What a decision surface looks like

The anatomy is specific. Each row is a decision unit — the record plus everything needed to act on it in one place. The primary action (approve, reject, override, escalate, accept, flag) is exposed directly on the row, not behind a click. A confidence signal or priority indicator sits inline with the record so the user can triage at a glance. Secondary context — reasoning, history, linked data — is one step away but out of the critical path. A running summary at the top reflects the cumulative state of decisions made so far, not the aggregate state of the dataset.

Keyboard shortcuts are table stakes, not enhancements. Batch workflows have a cadence. Users process tens or hundreds of rows in a session. A layout that requires a mouse click per action, per row, makes that cadence impossible to maintain.

> **Pattern:** a decision surface is a layout where each row is a decision unit — the action should be where the eye already is. The running summary reflects what the user has decided, not what the data says.

## How it differs from a dashboard

Dashboards and decision surfaces can look similar — both have tables, both may have summary numbers at the top. The difference is behavioral, not visual.

A dashboard is for monitoring. The numbers at the top reflect system state: how many items are in each status, what the current capacity is, whether something is trending in the wrong direction. The table is a drill-down into that state. The user reads, maybe filters, maybe exports. Decisions are secondary.

A decision surface is for acting. The numbers at the top reflect the user's progress through their own work — how many they've approved, how many remaining, what the accepted total adds up to. The table is the primary work surface. Every row is a prompt for a decision. Monitoring the aggregate state is useful as a side-effect of working through the list, not as the primary orientation.

The agent builds a dashboard when it can scaffold from the schema. It builds a decision surface when it knows the job. The schema doesn't tell the agent which one is right. The job does.

## When to use it

Decision surfaces fit jobs that share a structure: a list of records that each require a single discrete judgment, where the user is processing the full list in one session, and where the action per row is quick enough that the session feels like a batch.

Reorder recommendations. Expense approvals. Content moderation. Vendor quote comparisons. Support ticket triage. Loan application reviews. The specific domain varies; the shape of the job is the same. The user is the deciding agent. The interface's job is to make each individual decision as frictionless as possible while keeping the cumulative picture visible.

This pattern is distinct from a form (one record, many fields, deliberate attention per field) and distinct from a pure data table (browse and filter, decisions are incidental). It's the pattern for jobs where speed and volume matter and each row is a complete unit of work.

## How it shows up in patches

When TheDesignAgent flags a dashboard built for a deciding job, the patches follow the anatomy above: move actions inline, expose confidence or priority signals directly on the row, replace the aggregate KPI strip with a progress summary, add keyboard bindings for the primary actions.

The before is usually functionally correct — all the data is there, all the actions are accessible. The after is the same data and actions with the ceremony removed. Both look professional. Only one lets the user process fifty rows before lunch.

See the inventory planning example in the [homepage receipts](/#receipts) for the side-by-side. The before is a clean dashboard; the after is a decision surface. The structural difference is about four patches.

---

## Why order fulfillment punishes friction harder than financial review

> Extraneous load is near-fatal in warehouse operations and tolerable in loan underwriting. Good UX isn't universal — the same pattern can be right in one domain and wrong in another.

Source: https://www.thedesignagent.ai/field-notes/domain-weightings · Concepts · 2026-05-06

"Reduce cognitive load" is one of the most repeated principles in UX design. It's also, on its own, almost useless — because it treats all cognitive load as equally bad and all reduction as equally good. Whether a given amount of friction is acceptable depends on what happens when the friction compounds.

In a warehouse picking station, an extra modal dialog on each scanned item might cost two seconds per scan. Across a four-hundred-item shift, that's thirteen minutes of lost throughput, multiplied by ten operators, multiplied by two shifts. The modal wasn't adding safety — it was adding thirteen minutes of delay to a system where throughput is the primary success metric.

In a loan underwriting tool, the same extra modal on each application review might be entirely acceptable. The underwriter is spending eight minutes per application anyway. Two extra seconds for a confirmation step is noise. And if that step prevents a single incorrectly marked application from moving to the wrong queue, the math is different.

Same friction. Different consequence. Different design decision.

## Domain load weightings

TheDesignAgent's corpus includes load weightings by domain type. These weightings reflect how much damage different kinds of cognitive load do in different operational contexts. They're not arbitrary — they're derived from the observable structure of the job.

**Throughput-critical domains** (warehouse operations, support queues, content moderation, data entry): extraneous load is heavily penalised. The job is repetitive and high-volume; the user performs the same action hundreds of times per session. Friction that costs one second per action costs minutes per hour, and those minutes are the product's primary cost. Germane load — the kind that builds expertise — is fine and expected early in the user's tenure, but should reduce as they become expert. The interface should accelerate toward low load for experienced users.

**Deliberation-critical domains** (loan underwriting, legal review, clinical decision support, compliance auditing): intrinsic load is expected and appropriate. The job requires deep engagement with each record. Extraneous load is tolerable as long as it contributes to accuracy — a confirmation step that forces a brief review is acceptable. Germane load for building domain expertise in the tool itself is a feature, not a bug.

**Mixed-cadence domains** (SaaS operations, account management, customer success): the load profile shifts across the day-shape. Some tasks are triage; others are deliberate. The weighting model has to account for which task is active at evaluation time.

> **Pattern:** cognitive load costs aren't universal. The domain determines what friction is tolerable and what's fatal. "Reduce friction" is the right directive in throughput jobs; "support deliberation" is the right directive in review jobs. The same modal can be wrong in one and right in the other.

## How this changes scoring

When TheDesignAgent evaluates a build, the domain weighting is applied before the score is calculated. A `extraneous_load` sub-score of 65 means something different in a warehouse picking tool than in a financial review tool.

In the warehouse context, 65 on extraneous load is a significant concern — the patching priority is high, and the patches are aggressive. Strip ceremony from routine actions, replace confirms with undo, move actions inline. The goal is to make the frequent path frictionless.

In the financial review context, 65 on extraneous load might not trigger a high-priority patch at all if the friction is deliberate — confirmation steps on high-consequence decisions, required field validation that forces completeness, a step that makes the user read a key piece of context before proceeding. The domain weighting doesn't excuse bad friction; it changes what counts as bad.

## Why agents miss domain context

Agents apply UX principles uniformly because the principles they've learned are usually presented without domain qualification. "Reduce clicks" is presented as a universal good. "Fewer fields in forms" is standard advice. "Minimize modals" appears in every pattern library.

These principles are correct in their domain of origin — usually consumer-facing products with intermittent use, where friction causes abandonment. They're less obviously correct in enterprise operational software with daily power users and high task volume. And they can be actively wrong in compliance-heavy domains where slowing down is the safety mechanism.

Domain weightings are how TheDesignAgent translates "this is right" into "this is right here, for this job, for this kind of user." Without them, every evaluation collapses to the same general-purpose advice — which is how generic-looking, generic-feeling products get built.

---

## JTBD lives in three places: corpus, project profile, and score

> Jobs-to-be-Done isn't a single lookup. In TheDesignAgent it shows up in three distinct places — and the evaluation only makes sense when all three are in play.

Source: https://www.thedesignagent.ai/field-notes/jtbd-lives-in-three-places · Concepts · 2026-05-06

Jobs-to-be-Done is often described like a research method — interview users, find the job, design for it. That's fine as far as it goes. But in a system that evaluates agent-built UI on every build, JTBD has to be operational, not just a framing. It has to be something a pipeline can ask against.

In TheDesignAgent, JTBD appears in three distinct places. Understanding all three is what makes the `jtbd_alignment` score mean something instead of just saying "does this serve the user."

## The corpus layer: generic job archetypes

The first place JTBD lives is the corpus — a body of domain-specific job patterns built from working examples across industries. Warehouse operations, inventory planning, financial review, SaaS operations, customer support queues. For each domain, the corpus holds: what the primary job archetypes are, what cognitive requirements those jobs carry, which UI patterns fit them, and which patterns consistently fail.

This layer is universal. It doesn't know your product. It knows that a planner reviewing reorder recommendations is doing a deciding job, not a monitoring job, and that deciding jobs need inline actions, not dashboards. It knows that warehouse fulfillment is sequential, not categorical, and that sequential jobs need progress rails, not tabs.

When an agent calls without any project context — on the very first call — the corpus is the only thing grounding the evaluation. It's still specific enough to be useful, because domain knowledge is real and transferable.

## The project memory layer: your actual jobs

The second place JTBD lives is project memory — the profile built for your specific product from the first time an agent called with real context.

Discovery — the `init` call type — is when this gets built. The system infers the domain, creates a persona profile (working memory capacity, time pressure, expertise level, multitasking demands), and builds a job registry: every job the persona does, with priority, frequency, complexity, cognitive requirements, and success criteria. Critically, it captures `day_shape` — the sequence of jobs a user moves through across a workday, because most users don't have one job, they have several, and the layout needs to know which one is loaded.

From that point on, every evaluation is anchored to this profile. The active job on each call is matched against the job registry. The score reflects whether the artifact serves *this persona doing this specific job* — not whether it serves a generic user.

> **Why it matters:** a score against generic JTBD tells you the artifact fits a category. A score against your project's job registry tells you whether it fits your users' actual day.

## The rubric layer: the jtbd_alignment sub-score

The third place JTBD lives is the evaluation rubric itself — specifically the `jtbd_alignment` sub-score in the `/ux` lens.

`jtbd_alignment` asks a pointed question: does the layout's shape match the job's shape? A sequential job needs a sequential interface — the layout should mirror the workflow, not the data model. A deciding job needs inline decision units — the action should be on the record, not behind it. A monitoring job, one of the rare ones that actually calls for a dashboard, needs aggregates and trend lines up front.

The sub-score is what makes JTBD actionable in a patch. If `jtbd_alignment` is low, the patches name the mismatch: "layout is shaped like the data (tabs over state categories); job is sequential (pick → pack → ship → confirm); switch to progress rail." The agent gets a concrete structural change, not a note to think about user needs.

## Why all three have to be in play

A score against only the corpus tells you if the general pattern fits the domain. Useful, but generic. A score against only the project memory without corpus knowledge misses the accumulated domain wisdom. A score that has both but can't express the mismatch in a structural patch is information without action.

The three layers braid: corpus knowledge sets the domain baseline, project memory makes it specific to your users, and the rubric translates both into a score the agent can act on. The first call is where the braid starts — which is why the init call matters more than any individual score that follows it.

We'll write more about what discovery actually captures and how to shape the init call to get the most specific profile in a future note.

---

## Modals all the way down: why agents bury decisions you'd put on the row

> The modal reflex is one of the most consistent patterns in agent-built UI. A button opens a dialog. The action lives in the dialog. Save closes it. It compounds in batch workflows.

Source: https://www.thedesignagent.ai/field-notes/modals-all-the-way-down · Anti-patterns · 2026-05-06

Ask an agent to make an action available in a data table and it will give you a button that opens a modal. The form lives in the modal. Save closes the modal. Sometimes there's a confirmation dialog inside the modal. The user clicks through three surfaces to complete one action on one row, then starts over on the next.

This happens so consistently it deserves a name: the modal reflex. It's not a bug or a misunderstanding — it's the path of least resistance when you're scaffolding from a data model without knowing the job.

## When modals are right

Modals have a legitimate role. They're right when the action is a genuine context switch — when completing it requires information or focus that would be disruptive to have inline. Creating a new entity from scratch is a modal candidate. Configuring a complex rule that references data from multiple sources is a modal candidate. A destructive action with real consequences — delete, cancel, revoke — warrants a modal plus a confirmation.

What these have in common: the user is *leaving the current context* to do something different. The modal signals that boundary. It's appropriate when the boundary is real.

## When modals are wrong

Modals are wrong when the action belongs to the row. If the user is going through twenty recommendations and approving, rejecting, or editing each one, the action lives on the row. The row is the decision unit. Opening a modal to approve a single recommendation and then closing it before moving to the next one adds three clicks of ceremony to every decision — sixty extra clicks for twenty rows.

In batch workflows — reorder decisions, support ticket triage, expense approvals, inventory updates — this compounds fast. It turns a tool for deciding into an obstacle course. The user's job is to process the batch. The interface's job should be to make that frictionless. A modal on every row is the opposite of that.

> **Pattern:** modals are for leaving the context. If the action lives in the same context as the row, keep it there. Approve, reject, edit-in-place, and confirm all belong inline when the user is working through a list.

## Why agents reach for modals

The modal reflex comes from scaffold-first thinking. When an agent sees a table and needs to make a record editable, the simplest structure is: button triggers a form, form lives in a container that overlays the table, save dismisses the container. That's a modal. It requires no knowledge of how many records the user processes per session, what the cadence is, or whether the action is the whole job or a sub-step.

Agents also reach for modals defensively. Modals feel safe — the user confirms before committing, the form is contained, there's a clear cancel path. The problem is that this safety is real for destructive or irreversible actions and unnecessary for everything else. Approving a low-stakes recommendation is not a dangerous action. Wrapping it in a modal adds ceremony without adding safety.

## How TheDesignAgent flags it

The `/ux` lens raises a `layout_fights_action` flag when a modal pattern is detected on what appears to be a batch or repetitive job. The `extraneous_load` sub-score reflects the per-action overhead — if every decision in a list requires a modal open and close, the score drops accordingly.

Patches on this pattern are consistent: collapse the modal into the row, expose the primary action directly, add an expand gesture for any secondary context the user might want. The agent gets explicit before/after: "move approve/reject/override inline; reserve modal for edit scenarios where quantity and reasoning need simultaneous view."

The difference is visible in the [homepage receipts](/#receipts) — the inventory planning case shows the before and after. The before is functionally correct. The after is the same logic with the ceremony stripped out.

---

## Outputs that update as you decide: the summary is a side-effect, not a dashboard

> Decision surfaces need a running total. Not a KPI strip showing system state — a summary that reflects what the user has decided so far. These are different numbers, and agents build the wrong one.

Source: https://www.thedesignagent.ai/field-notes/outputs-that-update-as-you-decide · Patterns · 2026-05-06

In a decision surface, there are two kinds of numbers that could appear at the top of the screen. The first kind reflects system state: how many records are in each status, what the total dataset looks like, what the trend has been over the last period. The second kind reflects the user's progress through their own work session: how many decisions they've made, what the accepted total adds up to, how many remain.

These numbers look similar. They're often about the same dataset. But they answer different questions — and agents almost always build the first kind when the job needs the second.

## Monitoring numbers vs. decision numbers

A KPI strip that shows "47 pending, 12 approved, 3 rejected" is monitoring output. It tells the user where the overall system stands. It would be equally at home on a manager's overview screen, updating as other people make decisions. It's a dashboard element.

A running total that shows "Approved: $14,200 · Remaining: 8" is decision output. It reflects only what this user has done in this session. It doesn't update when someone else approves something. It doesn't represent the whole queue — only the portion the user has worked through. It changes because the user is deciding, not because the underlying data is changing.

> **Pattern:** in a decision surface, the summary should reflect what the user has decided, not what the data says. It's a progress tracker for the current work session, not a dashboard widget sitting above the table.

## Why this matters for high-volume decision work

In a batch decision workflow — reorder approvals, expense reviews, content moderation, loan triage — the user processes a queue. Their session has a beginning and an end: they start with N items, they work through them, they reach zero. The thing they need to know as they work is: how am I doing against my own work session? What have I committed to so far? How much is left?

Monitoring numbers don't answer that question. If the queue shows "47 pending" and the user has approved 15 and rejected 3, the "47 pending" number is still reflecting the full queue minus whoever-else's approvals. The user has no signal of their own progress. They don't know how far into their session they are. They can count, but the interface isn't helping them count.

The decision-output summary makes the session legible. The user can see at a glance how much they've committed to, whether they're approaching their own limit (budget, capacity, time), and how many records remain in their personal queue. This is information they need. The system state KPIs are information that could be available if wanted, but they're not what the job requires at this moment.

## When system state IS the right number

The monitoring summary is right when the job is monitoring. A manager reviewing team performance across a period needs to know the aggregate queue state. An operations lead checking in on daily throughput needs trend data and totals by status. These are genuinely monitoring jobs, and a dashboard with system-state KPIs is the correct response.

The mistake is applying that response to an individual decision-maker's working tool. The planner approving reorders is not a manager monitoring the queue. They're a worker processing their portion of it. The numbers at the top of the screen should reflect their work, not the queue's state.

This is the distinction that decides whether the layout is a dashboard or a decision surface. Not the visual design — the type of number at the top.

## How agents miss it

Agents build monitoring KPIs because monitoring KPIs are derived directly from the schema: count by status, sum of a column, distinct values. These are natural aggregates that fall out of any table query. Decision-progress numbers are harder to compute: they depend on user-session state, track what this specific user has acted on, and update only when this user takes an action.

The session-state model requires thinking about how the interface is used over time. Agents operating from a schema don't have that context. They build what the data can produce, not what the job needs.

The `/ux` lens surfaces this when a decision-job layout includes system-state monitoring at the top and no user-progress tracking. The patch is specific: replace the KPI strip with a session progress summary (items decided, cumulative value committed, items remaining), and add the monitoring numbers as an accessible secondary view for context when wanted.

---

## Progress rails: layouts that mirror the sequence of work, not its categories

> A progress rail is the shape that emerges when a sequential job gets the interface it needs. Not a wizard, not a stepper, not tabs — something closer to a spatial map of the work already done and the work still ahead.

Source: https://www.thedesignagent.ai/field-notes/progress-rails · Patterns · 2026-05-06

The tab problem has a positive answer. When a sequential job gets a tabbed layout — categories imposed on a workflow that needs to move in one direction — the fix isn't just "remove the tabs." It's to replace them with something that mirrors what the job actually is: a sequence with a beginning, a current position, and an end.

That shape is a progress rail. Not a wizard that hides previous steps. Not a stepper that asks for navigation decisions at each transition. Something closer to a spatial map: here's where you've been, here's where you are, here's what's ahead.

## Anatomy of a rail

A well-built progress rail has a specific structure. The steps are visible simultaneously — the user can see the full sequence at a glance. But they're not equal: the current step gets the primary working area of the screen, and the others are rendered smaller, in the rail itself, as spatial context rather than navigation targets.

Completed steps show completion state and are accessible if the user needs to reference or revise them. They don't disappear. Upcoming steps are visible but not interactive until they're unlocked by completing the current one. The rail communicates progress through position — the user knows where they are in the sequence by looking at the rail, not by reading a step counter or a breadcrumb.

Status is implicit. The warehouse operator who's at the Pack step doesn't need a status badge that says "In Progress." They're at Pack. That's the status.

> **Pattern:** a progress rail mirrors the job's sequence. The current step gets the screen. Completed steps provide spatial memory. Upcoming steps set expectation. Navigation is through completion, not clicking.

## How it differs from tabs

Tabs and rails look similar at a glance — both show multiple labeled sections. The difference is behavioral and cognitive.

Tabs are for non-sequential browsing. The user can jump between them in any order; none implies a path through the others. The right tab for inventory status doesn't follow from the left tab for purchase orders. Tabs work when the sections are genuinely parallel — multiple views of related data with no prescribed sequence.

Rails are for sequential progress. The user moves through them in order because the work requires it. You can't pack before you pick; you can't ship before you pack. The rail enforces this and makes it visible. Trying to use tabs for a sequential job breaks the spatial relationship between layout and work — the operator has to mentally reconstruct the sequence the tabs took apart.

## How it differs from a wizard

Wizards are multi-step forms for one-time or infrequent tasks. They're linear and typically hide completed steps to reduce visual complexity. The model is "we're going to collect a lot of information across several screens, and you only need to focus on the current screen."

Rails are for operational workflows that repeat. An operator fulfills many orders per shift. A planner processes many recommendations per week. The rail is part of the working surface, not a temporary guide through a setup process. Completed steps remain visible because the operator may need to reference what they packed before confirming shipment. The spatial context of the full sequence is load-reducing, not load-adding, because the job is familiar.

## When to use it

Progress rails fit jobs that are sequential (steps must happen in order), repeated (the same workflow runs multiple times per session), and referential (the user benefits from seeing earlier steps while at a later one).

Order fulfillment. Expense approvals with a multi-stage review. Customer onboarding flows for account managers. Manufacturing quality checks. Claims processing. The domain varies; the job shape is the same — a fixed sequence of steps, each of which passes context to the next.

## How TheDesignAgent patches to it

When the `/ux` lens identifies a tab layout applied to a sequential job, the patch names the swap directly: "replace tabbed navigation with progress rail; steps: [extracted from the workflow]; current step gets full working area; completed steps accessible in rail." The agent gets the step names, the navigation model, and the layout prescription in one patch.

The score delta on `jtbd_alignment` is typically large when this patch is applied — sequential job on a tab layout is one of the highest-scoring mismatches the system identifies. The complementary score, `extraneous_load`, also improves: the cognitive cost of mentally stitching the split-screen-across-tabs job back together is eliminated.

---

## Reading the score trajectory: why convergence beats absolute fit

> A single fit score tells you where a build is. A score trajectory tells you whether it's getting better, plateaued, or regressing — and which dimensions are moving. The second question is usually the more important one.

Source: https://www.thedesignagent.ai/field-notes/reading-the-score-trajectory · Workflow · 2026-05-06

A fit score of 78 is information. A trajectory of 42 → 65 → 78 is a different kind of information — and usually the more useful kind. The score tells you where the build is. The trajectory tells you whether the agent is converging on something good, stuck on a plateau, or introducing new problems while fixing old ones.

These are different questions. The absolute score is an answer to "is this good?" The trajectory is an answer to "is this getting better, and how fast?" For a multi-call build loop, the second question matters more.

## What convergence looks like

A converging trajectory climbs with meaningful steps between calls — 42 → 65 → 78, or 58 → 74 → 82. Each call produces a materially better score. The patches are landing; the agent is applying them correctly and the layout is improving on the dimensions that were flagged.

The step sizes matter. A trajectory of 78 → 79 → 80 looks like convergence but isn't — it's a plateau. Three-point increments across three calls means the patches from the previous call are producing minimal improvement. Either the agent isn't applying them, the high-value patches have been applied and what remains is marginal, or the remaining issues are structural and the current patching isn't reaching them.

A trajectory of 84 → 81 → 79 is regression. The agent is making changes — but the changes are introducing new problems faster than they're fixing old ones. This is often a sign that the agent is applying visual patches to a layout that has a structural problem. Adjusting typography and spacing on a layout that's shaped wrong for the job produces visual polish on a fundamentally misaligned artifact.

> **Pattern:** the trajectory matters more than the number. A score of 84 that came from 42 is a healthy build. A score of 84 that's been at 84 for three calls is stuck. A score of 84 that came from 87 is regressing. Read all three.

## Sub-score trajectories are where the diagnosis lives

The aggregate fit score averages across sub-scores, which means a rising aggregate can mask a sub-score that's falling. The diagnostic value is in the per-dimension trajectory, not just the overall number.

If `jtbd_alignment` is climbing but `extraneous_load` is flat, the layout is getting structurally better but the friction is persisting. The agent is applying the shape patches but not the ceremony-reduction patches. If `hierarchy` is improving but `density_fit` is declining, visual patches are being applied selectively — the important-versus-secondary distinction is improving while the overall density is getting worse.

The most common plateau pattern is `jtbd_alignment` near-ceiling, `extraneous_load` stuck. The structure is right. The friction remains. This usually means the layout is correctly shaped for the job but individual interactions — confirms, modals, multi-click paths — haven't been addressed. The patches for these are specific and mechanical; they shouldn't be hard to apply. When they're not landing, it's usually because the agent doesn't have clear enough instructions about what "inline" means for this specific context.

## How to break a plateau

A build that's been at 78 → 79 → 80 for three calls is telling you something: the patches being applied are small improvements on areas that were already adequate, while the real problem hasn't been touched.

Read the sub-score breakdown from the most recent call. Find the lowest sub-score. Read the patches that target it. If those patches haven't been applied between the last two calls, that's the blocker — apply them. If they have been applied and the score hasn't moved, the patch may have been applied incorrectly, or the issue may be more structural than the patch described.

Providing more codebase context on the next call helps. If the agent can see the actual component structure, not just a description of it, the patches can be more precise. "Move the action inline" is a valid patch. "Move the action from the modal trigger in `row-actions.tsx` to the row component in `data-table.tsx` as a conditional render on row focus" is a patch that the agent can apply without ambiguity.

## Reading the first build

The init call score sets expectations. A first-build score of 42 on a complex B2B tool isn't alarming — it means discovery ran, the JTBD was identified, and the layout has significant room to improve. The patches from this call should produce a 20-30 point gain on the second call if applied.

A first-build score of 78 on a simple admin form is expected. There's not much structural work to do; the remaining gains are in visual craft and friction reduction. The trajectory from here will be flatter, and that's fine — the ceiling is closer.

The score trajectory is what makes the build loop useful. Each call tells you more than the previous one, and the pattern across calls tells you more than any individual call.

---

## The component library was the right answer to a problem agents no longer have

> Component libraries exist because building UI was expensive and inconsistency was hard to prevent. AI agents changed the cost equation. The case for building unencumbered — and what you need instead.

Source: https://www.thedesignagent.ai/field-notes/the-component-library-constraint · Concepts · 2026-05-06

An agent is building a contract review interface. The job is clear: a legal ops team needs to scan a hundred vendor agreements and flag anything that deviates from standard terms. The ideal interaction is a comparison row — the clause on the left, the standard on the right, a flag inline, an approve/reject action on the row itself.

The component library has a modal. The modal is what you use when something needs reviewing.

The agent uses the modal. It's wrong for the job — the user has to open and close it a hundred times instead of scanning down the page — but it's what the library supports. Alternatively, the agent files a request: can we add a comparison row component? The request goes into the backlog. The backlog is reviewed quarterly. The interface ships with the modal, the legal ops team spends twice as long on reviews, and three months later someone closes the backlog ticket as won't-fix because the product roadmap has moved on.

Here's what's notable about that story: building the comparison row from scratch would have taken the agent about ninety seconds.

## What component libraries were for

Component libraries exist because, not long ago, building UI was expensive and inconsistency was a real bug.

A custom date picker cost a day or two. A multi-select with async search cost three. A row with an inline action and a conditional disclosure took a half-sprint to build well and a full sprint to build well and accessibly. If you had a team of developers building a product with fifty features, the math was clear: shared components with shared defaults, built once and maintained in one place, paid for themselves within the first few months.

Libraries also encoded judgment. The decisions about padding values and border radii and focus states and keyboard behavior — someone made those decisions carefully, and putting them in a library meant you didn't make them again badly on every new feature. That's genuinely valuable. A well-built component library is accumulated design intelligence.

The case for the component library was real. It solved a real problem at real cost.

## The cost equation inverted

AI agents can build a purpose-built primitive — a component designed for exactly the job at hand — in seconds.

Not a close approximation. Not a library component with overrides stacked on top. A component built from scratch, to spec, for the specific interaction the job requires. The comparison row. The inline flag with contextual expand. The multi-step confirmation that flows forward instead of interrupting. Whatever the job asks for.

The marginal cost of that purpose-built primitive is close to zero.

The marginal cost of working around a library constraint is real and it compounds. The agent compromises the design — uses the modal because the library has a modal. Or the agent files the update request and waits. Or the agent ships an override, `!important` and `z-index` stacked until it looks right, until the next version bump breaks it. Each of these outcomes is worse than building fresh, and each one happens because the library was treated as a constraint rather than raw material.

"Can we update the component library?" is the symptom. The disease is the assumption that the library controls what gets built.

> **Pattern:** component libraries solved a consistency problem at a time when custom components were expensive. When an agent can build and customize any component in seconds, the library's value proposition inverts — instead of absorbing cost, it adds constraint. The right unit of judgment shifts from "does this conform to the library?" to "does this serve the job?"

## The ephemeral argument

There's a second thing happening alongside the cost shift.

UIs are becoming more distilled. The thick, multi-surface applications that dominated the last decade of SaaS are giving way to something more situational: interfaces built for a specific job, a specific session, sometimes a specific user. The contract review interface for the legal ops team doesn't need to share a component vocabulary with the expense reporting tool or the vendor onboarding flow. It needs to serve the job it was built for.

Component libraries are optimized for permanence. Build once, use everywhere, update in one place. That model works well when the interface is expected to last years and serve many jobs. It works poorly when the interface is expected to serve one job and may be rebuilt or retired before the next major library version.

The right model for agent-built UI is different: the library as a starting position, not a destination. Defaults as a floor, not a ceiling. An agent working on a new interface should start from whatever baseline the library provides — it's free, it's already solved, it encodes that accumulated judgment — and then build from there, unencumbered, to the shape the job actually requires.

The shift isn't "ignore the library." It's "don't let the library be the ceiling."

## What you need instead of library conformance

When agents build unencumbered, the question "is this component right?" doesn't disappear. It stops being answerable by checking the library spec.

Library conformance is a proxy for quality. A component that exists in the library, used as specified, is probably fine — it's been reviewed, it's accessible, it behaves predictably. That proxy is useful when the alternative is a developer building something new without review. It's less useful when the alternative is an agent building something new in ninety seconds.

The question behind the proxy is: does this component serve the user's job? Does it add extraneous load — steps, clicks, cognitive overhead — that the job doesn't require? Does it fit the way the user thinks about the task, or does it impose a different mental model?

Those are design judgment questions. They can't be answered by checking whether the component exists in the library. They require knowing the job, the user, and what good looks like for that specific combination. That's a different kind of quality gate — one that runs on understanding rather than conformance.

---

The component library isn't going away. It still provides defaults worth starting from and judgment worth inheriting. What's changing is its role: from the thing that controls what gets built to the thing that provides the raw material for building.

The control, in an agent-built world, moves somewhere else. To the question that actually matters — not "is this in the library?" but "does this serve the job?" That question needs an answer. The library isn't the place to look for it.

---

## The 8-field form: why agents collect what the schema has, not what the job needs

> Agents build forms from data models. A schema with twelve fields produces a form with twelve fields. The job usually requires three. This is how to close that gap.

Source: https://www.thedesignagent.ai/field-notes/the-eight-field-form · Anti-patterns · 2026-05-06

Give an agent a schema with twelve columns and ask it to build an intake form. You'll get twelve fields. Sometimes thirteen — agents like to add a notes field on principle. The form maps directly to the model: every column that accepts user input becomes a form field, laid out in the order the columns appear in the schema.

This is the eight-field form problem. Not literally eight — the number is whatever the schema happens to have. The problem is the mapping logic: form fields are derived from database columns instead of from what the user needs to supply right now.

## A database field is not the same as a form field

A database column is a slot for data. A form field is a prompt for a human to supply something. These overlap, but they're not the same set.

Some columns should be defaults. The agent creating a record probably knows the user's organization, the current date, and a handful of status values that don't need human input at creation time. Some columns are derived — calculated from other values and should never appear in a form. Some columns are only relevant at a later stage of the record's lifecycle: fields that matter when closing a deal don't belong in the form for opening one.

When agents skip this analysis, every column becomes a form field. The user is asked to fill in values the system already knows, values that won't be needed until much later, and values that don't meaningfully change between records. Most of them get left blank or receive a default typed by habit. The form is noise with a few signal fields buried in it.

> **Pattern:** a form should collect the minimum set of inputs the job requires at this moment. Everything the system can derive, default, or collect later is overhead — and overhead in a form creates abandonment.

## The minimum viable form

For any given job, there's a minimum set of fields the user actually needs to supply. Finding it requires knowing the job: what decision does this form enable? What does the user need to input for that decision to be sound? What can wait?

A reorder request might need three fields at creation — item, quantity, urgency — and fifteen fields at approval. Building both as the same form is the wrong answer. The creation form should be three fields. The approval form should surface the additional context in read-only display, with only the fields that require input at the approval stage made editable.

This is progressive disclosure: show the fields the current stage needs, reveal the rest when they're relevant. Not as a UX trick but as an accurate model of how work actually unfolds.

## Why agents build the opposite

Agents don't have a concept of stage. They see a schema, a task, and a user. They build the form that would let the user fill in everything the schema accepts. This isn't wrong as far as it goes — the form is complete. But complete and useful aren't the same thing.

The deeper issue is that agents optimise for correctness of data capture rather than reduction of cognitive cost to the user. A form that collects everything needed to fully define a record is correct. A form that asks the user to supply only what they know right now and leaves the rest for later is better — but it requires knowing what "right now" means for this stage of this job.

## How TheDesignAgent flags it

On the `/ux` lens, a form with more fields than the active job requires at its current stage triggers a `extraneous_load` flag. The patch is specific: identify the minimum set for the current stage, move the rest to a secondary panel, detail view, or later-stage form. The agent receives the field names to remove from the primary form, not a general instruction to simplify.

The score on `jtbd_alignment` drops when the form is shaped like the data model rather than the job. That pairing — high extraneous load plus low JTBD alignment — is the signature of the eight-field form. Both sub-scores have to move for the patch to be considered applied.

---

## The filter buffet: every column gets a filter, none of them get used

> Agents build filters from columns. A table with ten columns gets ten filter controls. The user's actual filtering question involves two of them. The rest is noise that teaches the interface to be ignored.

Source: https://www.thedesignagent.ai/field-notes/the-filter-buffet · Anti-patterns · 2026-05-06

Give an agent a table with ten columns and ask it to make the data filterable. You'll get ten filter controls, one per column, probably displayed simultaneously in a collapsible sidebar. Every dimension of the data — status, date range, assigned user, category, priority, region, account type — gets its own input. The filter buffet: an exhaustive representation of how the data could theoretically be sliced, with no opinion about how the user actually searches.

The problem isn't that these filters are wrong. For power users doing exploratory analysis, some combination of them might be useful. The problem is that they're all treated equally, all visible at once, and none of them are shaped around the question the user is actually asking when they open this table.

## Filtering is answering a question

Users filter data because they're looking for something specific. The filtering action is an implicit question: "show me X." In a support queue, the question is usually "show me the urgent open tickets assigned to me." In an inventory table, it's "show me the items that will go out of stock this week." In an expense report, it's "show me anything over the approval threshold."

These questions have answers. And the answer to each question involves a specific set of fields in a specific configuration. The agent knows the data model but not the question — so it exposes the entire data model as filter controls and leaves the question to the user.

The result is a filter panel that requires the user to translate their job-shaped question into data-model-shaped inputs every time they open the screen. That translation is extraneous load. It's not part of the job; it's a tax the interface charges for access to the job.

> **Pattern:** filters should answer the questions the user is actually asking, not expose every column that exists. The right filter set is derived from the job, not from the schema.

## Job-shaped filters vs. schema-shaped filters

Job-shaped filters look different from schema-shaped ones. Instead of "Status: [dropdown of all statuses]," they might be "Needs action" — a single toggle that filters to records where user input is required, regardless of what underlying status values produce that condition. Instead of "Date range: [from] [to]," they might be "Due this week" — a pre-configured range that matches the planner's actual review cadence.

This doesn't mean advanced filtering doesn't belong. It means the common-case filtering should be designed around what the user is actually trying to isolate, and the advanced filtering (including full column access) should be available but not the primary interface.

Saved views earn their place here. When users filter to the same configuration repeatedly — "my open urgent tickets" — a saved view with a name that describes the question is more useful than the same filter combination re-applied manually every session. Agents rarely add saved views unprompted, because they're not a feature that falls out of the data model.

## Why agents reach for the buffet

The filter buffet is cheap to build and defensible on its face. "Every column is filterable" sounds like a feature. It's harder to argue against than "we filtered out most of the columns because we thought you'd only use three." The exhaustive option feels safer.

Agents also genuinely don't know which filters matter without knowing the job. The buffet is the answer you get when the agent has access to a schema but not to an understanding of how the table is actually used. It's not wrong — it just exposes the absence of job knowledge rather than making any decision.

## How TheDesignAgent flags it

On the `/ux` lens, a filter panel with more than five or six controls simultaneously visible, derived directly from the schema without apparent job-shaping, triggers a `jtbd_alignment` flag. The patch is specific: identify the two or three filtering questions that account for the majority of use, redesign the primary filter bar around those, and move the rest behind an "advanced filters" disclosure.

The `extraneous_load` sub-score drops when the filter-question translation is visible — when the user has to know that "needs action" corresponds to statuses `pending_review` and `awaiting_approval` before they can use the filter correctly. That translation belongs inside the filter, not in the user's head.

---

## The first call is discovery: why the init prompt sets the frame for everything after

> The first time a build agent calls TheDesignAgent with no prior context, something different happens. Not just a score — an inference of domain, persona, jobs, and day-shape. Every evaluation that follows is calibrated to what was learned here.

Source: https://www.thedesignagent.ai/field-notes/the-first-call-is-discovery · Workflow · 2026-05-06

The first time a build agent calls TheDesignAgent with no prior project context — no stored profile, no previous calls — the response type is `call_type: "init"`. Something different happens on this call. Not just a score and a set of patches. A discovery: the system infers the domain, builds a persona profile, constructs a job registry, and produces a `.thedesignagent` file for the agent to commit to the repo. Every score that follows is calibrated to what was learned here.

This makes the init call the most consequential one. A generic init produces generic scores. A rich init produces scores that are specific to the actual people using the actual product.

## What discovery infers

From the task description and artifact the agent provides, discovery builds out three things.

**The domain profile**: what kind of product is this, what operational context does it operate in, what cognitive costs matter most in this domain. Domain is used to set the load weightings — the relative severity of extraneous load penalties, the tolerance for deliberation, the expected expertise level of users. A warehouse ops tool gets a different baseline than a financial review tool.

**The personas**: who uses this product, what are their cognitive characteristics, and what does their workday look like. Each persona gets a cognitive profile (working memory capacity, time pressure, expertise level, multitasking demand, risk tolerance) and a day-shape — the map of jobs they move between across a typical day. This is the frame that makes "is this good for the user" a specific question rather than a general one.

**The job registry**: every job the personas do in the product, with priority, frequency, and cognitive requirements per job. Each call after init specifies which job the current artifact is serving. The `jtbd_alignment` score is evaluated against the requirements of that specific job — not against a generic "user job" abstraction.

> **Pattern:** the init call is not a config step. It's the moment the system learns who your users are and what they're actually trying to do. The richer the context you provide here, the more specific every score will be.

## What the agent gets back

On an init call, the response includes everything that a graded call includes — fit score, patches, flags — plus two additional outputs.

`captured_context` is the full discovery output: the domain profile, personas, job registry, and active job identification. It's surfaced for transparency so you can review what the system inferred and correct anything that's wrong. If the persona profile missed a key cognitive characteristic, or the job registry missed a job that matters, this is where you see it.

`tda_file` is the content to write to `.thedesignagent` in the repo root. This file is the project identity — it's how subsequent calls from any agent on this repo get associated with the same project context without re-running discovery. The build agent is expected to commit this file. Once it's in the repo, every agent that clones or pulls the repo inherits the project context automatically.

## How to get more from the first call

Discovery infers from what it receives. A task description that says "build a tool for the warehouse team" produces a narrower inference than one that says "build an order fulfillment interface for warehouse operators doing pick-pack-ship workflows, approximately 200 orders per shift, time pressure is high." Both produce a usable discovery. The second one produces a more accurate persona with more specific load weightings.

Providing `codebase_context` on the init call — a summary of the existing product, the tech stack, the current user base, the domain — produces the richest inferences. If you're adding a new feature to an existing product, include a brief description of what the product does and who uses it. The discovery will fold that into the domain profile and produce a persona that reflects the actual user rather than an archetype of the domain.

## What happens when it's wrong

Discovery is inference. It will occasionally get something wrong — a domain assumption that doesn't fit, a persona cognitive profile that's off, a job registry that missed a key job. The `captured_context` output is designed for exactly this: review it, find what's off, and update the `.thedesignagent` file directly. The file is plain JSON; the fields map directly to the captured context output.

Future calls will pick up the updated file. The next score after a correction will reflect the more accurate profile. The system doesn't require a reset — the correction propagates forward.

---

## The full design process, end to end: what an evaluation actually looks at

> TheDesignAgent doesn't evaluate one thing. It runs through the same process a skilled human design reviewer would — JTBD, cognitive patterns, visual craft, visual cognitive load, component grammar, accessibility — encoded as a rubric an agent can call mid-build.

Source: https://www.thedesignagent.ai/field-notes/the-full-design-process · Workflow · 2026-05-06

A skilled human design reviewer doing a thorough critique of an interface doesn't ask one question. They ask a sequence of questions, each building on the previous. Does this serve the job? Is the cognitive load appropriate? Is the visual hierarchy clear? Is the visual parsing cost low? Does the component grammar feel intentional? Is it accessible? Are the patterns consistent with the decisions already made in this product?

TheDesignAgent runs the same process — encoded as a rubric, executed on every call. Understanding each stage makes scores and patches more legible: they're not arbitrary numbers and suggestions, they're the outputs of a structured review that's asking specific questions in a specific order.

## Stage one: JTBD alignment

The first question is structural: does this artifact serve the job it's supposed to serve, in the shape the job requires?

A deciding job needs a decision surface — inline actions, confidence signals on the record, session-progress tracking at the top. A sequential job needs a layout that mirrors the sequence — a progress rail, not tabs. A monitoring job needs aggregates and trend visibility at the primary level.

JTBD failures are foundational. If the layout is wrong for the job, applying visual fixes to it produces a polished version of the wrong thing. This is why JTBD comes first: structural problems should be diagnosed and fixed before craft-level work begins. The patches for JTBD failures are the largest structural changes in the set — switch layout type, move actions inline, restructure primary navigation.

## Stage two: cognitive patterns

The second question is behavioral: how much cognitive effort does this layout require, and is that effort justified?

This stage looks at intrinsic load (is the task genuinely complex, and is the layout helping the user manage that complexity?), extraneous load (is the interface adding overhead the job doesn't require?), and germane load (is the interface supporting the development of expertise, or preventing it?).

Domain weightings are applied here. A confirmation dialog on a routine action in a high-throughput job has a different cost than the same dialog in a deliberation job. The same feature, different penalty.

Cognitive pattern failures produce patches that address ceremony: replace confirms with undo, collapse modals into inline actions, reduce the number of steps between the user and the action they perform most frequently.

## Stage three: visual craft

The third stage asks: is the design well-made? This is where the `/visual` lens focuses — on spacing, hierarchy, contrast, typography, and density. Not "does it look good" as a subjective judgment, but "does it apply visual design principles correctly."

This stage is independent of JTBD and cognitive load. A layout can be perfectly shaped for the job and poorly crafted — right structure, weak execution. Visual craft failures produce patches that are precise: adjust the font weight of secondary labels, reduce the border-radius to match the density of the surrounding context, establish a primary focal point per section, reduce the competing high-contrast elements from four to one.

## Stage four: visual cognitive load

Stage four asks a related but distinct question: what is the parsing cost of this layout? A screen that's well-crafted on individual elements can still be expensive to read if the elements don't form a coherent visual path.

Visual cognitive load is the tax paid before any task thinking starts. It's high when density exists without hierarchy, when there are competing focal points at similar visual weight, or when the typography creates reading cost rather than reducing it. The `/visual` lens scores this through `hierarchy`, `density_fit`, and `critical_path` — three dimensions that together describe whether the eye is guided or left to do its own work.

## Stage five: component grammar

Stage five asks: do the components feel like they belong to this product? This is distinct from whether they work correctly or are visually well-crafted in isolation. It's about the coherence of the visual system.

Shadcn defaults produce a defensible starting point but not a specific one. Component grammar is what distinguishes a product that feels intentionally designed from one that feels assembled. Low `component_grammar` scores produce patches that are specific to the product's context: radius adjustments, palette refinements, density calibrations that encode the product's own visual sensibility rather than the library's default.

## Stage six: accessibility

Accessibility is evaluated in parallel, not after the fact. Contrast ratios are checked against WCAG 2.1 AA. Interactive elements are evaluated for keyboard reachability. Focus states are checked for visibility. Touch targets are evaluated for minimum size.

Accessibility failures surface as flags, not just patches — because some accessibility failures are blockers, not improvements. A form with no accessible labels isn't "slightly worse"; it's unusable for a portion of the user base. The flag distinction signals severity.

> **Pattern:** the stages aren't evaluated in strict sequence — they run in parallel — but the patches are ordered by stage. Structural patches (JTBD, cognitive) come first. Visual patches (craft, VCL, grammar) come after. Fixing visual issues on a structurally wrong layout wastes the fix.

## Reading a result with failures across multiple stages

When a build has failures across multiple stages, the patch order in the response reflects the priority: JTBD failures are critical priority, cognitive failures are high, visual failures are medium, grammar and accessibility are medium to low unless they're severe.

A result with a critical JTBD failure and several medium visual failures is telling you: fix the structure first, then revisit the visual. If you apply the visual patches before the structural one, the next call will re-run the same structural diagnosis on the new layout, and the visual patches may no longer apply after the structural change.

The design process is a sequence because that's how design review works. The stages build on each other. TheDesignAgent's job is to run that sequence faster and more consistently than it would happen in a human review cycle.

---

## The missing design layer: why AI coding agents ship UI without a judgment step

> The AI coding stack has linters, test runners, type checkers, and deployment gates. It has no design review step. This isn't an oversight — it's a hard problem. Here's why it's still unsolved and what solving it actually requires.

Source: https://www.thedesignagent.ai/field-notes/the-missing-design-layer · Concepts · 2026-05-06

Search for MCP servers available to Claude Code. You'll find tools for GitHub, Jira, Linear, Slack, Supabase, Stripe, Postgres, filesystem access, web search, code execution. You'll find tools that let agents read documentation, manage deployments, query databases, and send messages.

You won't find a design review tool. There isn't one.

This isn't an oversight in the ecosystem. It's a reflection of how much harder design judgment is to encode than the other kinds of review the stack already handles.

## What the stack can already do

The non-design parts of the quality stack are mature because they're automatable in a straightforward way. A linter checks code against rules that are explicit, deterministic, and domain-general. A type checker proves correctness against a formal specification. A test runner executes code and checks outputs against expected values. A CI pipeline gates deployment against passing checks.

These tools work because the thing they're evaluating has a clear, agreed-upon definition of "correct." Code either type-checks or it doesn't. Tests either pass or they fail. The evaluation is binary, deterministic, and doesn't require knowledge of who will use the software or what they'll be trying to do with it.

Design doesn't work that way.

## Why design is harder to automate

A well-designed interface is one that serves the person doing the job, in their domain, with their cognitive constraints, on their device, in their workflow context. There's no formal specification for "serves the person doing the job." There's no grammar that type-checks against it. There's no test suite that passes or fails against it.

This is why design review has historically required human judgment. A senior designer reviews a build not by running a checker but by understanding who the users are, what they're trying to accomplish, and whether the layout helps or hinders that. The review is a synthesis across multiple frameworks — JTBD, cognitive load theory, visual hierarchy principles, domain-specific pattern knowledge — applied to a specific artifact in a specific context.

Encoding that synthesis into something an agent can call is the hard part. The easy version is a linter — check that buttons have sufficient contrast, that interactive elements have accessible labels, that touch targets meet minimum size. Accessibility checkers do this already. But contrast ratios and label presence are necessary conditions for a good interface, not sufficient ones. A layout can pass every accessibility check and still be wrong for the job.

> **The gap:** the AI coding stack can verify that UI is technically correct. It has no tool for verifying that UI is experientially correct — that it serves the person doing the job, in the shape the job requires.

## What a real design judgment layer requires

For automated design review to be more than a linter, it needs several things that most tools don't have.

**Domain knowledge.** Design patterns that are correct in a warehouse picking station are wrong in a loan underwriting tool. Generic advice — "reduce cognitive load," "keep forms short" — is true enough to be quoted and not specific enough to be useful. A judgment layer needs to know what domain it's operating in and what that domain's cognitive requirements look like.

**User knowledge.** The evaluation has to be against a specific persona with specific constraints — not "the user" as an abstraction. A 14px font is fine in a desktop analytics tool and a problem in a warehouse handheld context. Touch target sizes that work for occasional users are too small for operators who use the interface hundreds of times per shift. Without knowing who the user is, these distinctions don't exist.

**Job knowledge.** The most fundamental design question — does this layout fit the job? — requires knowing what the job is. Not what the data is, not what the features are: what is the user actually trying to accomplish, in what sequence, with what constraints. This is JTBD at the operational level, and it's what most automated tools don't have access to.

**Pattern knowledge.** Knowing that tabs are wrong for sequential workflows and modals are wrong for batch decisions requires a body of domain-specific pattern knowledge — not just general UX principles, but knowledge of which patterns consistently fail for which jobs.

## Why this gap exists now

The tools that exist for agents today are built around data: code, documents, APIs, databases. These are structures agents can read, query, and modify with well-defined semantics.

Design judgment is about meaning — specifically, the meaning of a layout to a specific human doing a specific job. That's harder to represent as a structured query. You can't `SELECT` the extraneous load from a React component. You can't `grep` for JTBD misalignment.

The gap will close, but it requires a different kind of tool than the ones the stack has built so far. One that holds domain knowledge, user knowledge, and job knowledge as context — and applies that context to evaluate an artifact against the actual requirements of the job. That's not a linter. It's closer to a specialized reviewer that the agent calls the same way it calls a test suite.

That's what TheDesignAgent is. And right now, it's one of the only things in the space that's doing it.

---

## The tools agents reach for (and the design layer most setups still miss)

> Claude Code, Codex and Cursor can reach dozens of MCP tools for code, data, deployment and communication, and a growing set of design tools. Most of those design tools cover one layer: accessibility, a Figma file, or how the site looks. This is what agents have available, and what's still missing from most setups.

Source: https://www.thedesignagent.ai/field-notes/the-tools-agents-reach-for · Concepts · 2026-10-03

When a developer sets up Claude Code, Codex or Cursor for a project, they configure the tools available to the agent. A typical setup for a modern web application might include: a GitHub MCP for reading PRs and issues, a Supabase or Postgres MCP for querying the database, a filesystem tool for reading and writing code, a web search tool for documentation lookup, and possibly a Slack or Linear integration for surfacing relevant context.

With these tools, the agent can read the codebase, understand the ticket it's working on, check the database schema, look up library documentation, and query relevant issues. It has access to everything needed to produce technically correct code.

What it usually doesn't have: anything that checks whether the UI it's building will serve the people who will use it.

## What the MCP ecosystem covers today

The MCP ecosystem has grown quickly and covers most of the surfaces a development agent needs to touch. Categories that are well-served:

**Data and APIs.** Database tools (Postgres, Supabase, SQLite), REST API testing, GraphQL introspection. Agents can query real data, test endpoints, and understand schemas.

**Code and version control.** GitHub, GitLab, filesystem access. Agents can read and write code, create PRs, review diffs.

**Project management.** Linear, Jira, Notion. Agents can read tickets, update status, understand project context.

**Communication.** Slack, email. Agents can surface relevant thread context, post updates.

**Deployment and infrastructure.** Vercel, AWS, deployment logs. Agents can deploy, check status, read error logs.

**Documentation.** Web search, docs crawlers. Agents can look up how a library works.

Each of these categories reflects something a developer would reach for when doing their work — and the tools make that reach shorter for the agent.

## The design tools that exist now

When we first wrote this note in May, we called the design category empty. That's no longer true. Scan the MCP registries and skill directories today and you'll find tools that each cover one layer of design:

**Accessibility.** Deque's axe MCP server finds and fixes accessibility violations from inside Claude Code, Copilot and Cursor. Essential, and narrow by design: it tells you whether a screen is *valid* for assistive technology.

**Quality audits.** Chrome DevTools MCP runs Lighthouse, which scores accessibility, SEO and best practices. Again essential, again one layer.

**Design files.** The Figma MCP server gives the agent the components, variables and layout of a real Figma file. It tells the agent what the design *is*, if one exists.

**How the site looks.** Skills like Anthropic's frontend-design and Impeccable shape the agent's styling choices: aesthetic direction, typography, color, polish. They're about designing the site itself.

**Screenshot critics.** A growing number of community MCP servers score a screenshot for "AI slop" or run a generic heuristic pass.

Every one of these is worth having. None of them answers the question that decides whether the UI works: is this the right screen for the person using it, given the job they're trying to do?

> **The gap now:** the ecosystem has tools for accessibility, audits, design files and styling, each covering one layer. What most setups still miss is a design agent in the loop: something that briefs the agent before it builds a screen, reviews the result for job fit, UX and visual quality after, and remembers the project between calls.

## Why single-layer tools came first

The tools that exist for agents map onto surfaces that already had a programmatic answer. Accessibility has WCAG and a mature rules engine. Lighthouse already produced scores. Figma already had an API. Styling guidance fits in a skill file. Wrapping any of these for an agent is relatively straightforward — the data structure exists, the rules exist, the semantics are defined.

Design judgment doesn't have an existing API. There's no endpoint that takes a React component and returns whether it fits the user's job. There's no database of "this layout is appropriate for this job in this domain" that an MCP could query. The knowledge that would power a design judgment tool exists in designers' heads, in pattern libraries, in accumulated domain expertise — not in a structured API waiting to be wrapped.

Building that layer required building the knowledge layer first: a corpus of job patterns, cognitive patterns, and design principles organized by domain; a project memory that holds the personas and jobs for a specific product; a rubric that produces specific, actionable fixes rather than generic advice. That's not an MCP wrapper around an existing tool — it's a design agent.

## What it looks like in the workflow

A design agent works in the loop the way a test runner or a type checker does, but on both sides of the build:

1. **Before the build:** the agent asks for a brief for the screen it's about to build: the user's job, the heuristics and patterns that apply, the project's design tokens.
2. **Build:** the agent writes the UI from the brief.
3. **After the build:** the agent gets a UX review (job fit, heuristics, cognitive load, patterns) and a visual review (brand, hierarchy, aesthetics from a real screenshot), each scored with fixes, applies the fixes and re-checks once.
4. **Across builds:** what the reviews find is remembered, so the next brief starts smarter.

This is the same pattern as test-driven development, except the test is "does this serve the person doing the job" rather than "does this return the expected output." The loop is the same. The feedback mechanism is the same. The thing being checked is different.

Most agent workflows don't have this step today. The agent builds, lints, type-checks, maybe runs an accessibility check, and opens a PR. The design review happens later — in a human review cycle that moves much slower than the build loop. Or it doesn't happen at all, and the UI ships as the agent built it.

That's the gap. TheDesignAgent was built to close it: a design agent in your coding agent's loop, alongside the single-layer tools rather than instead of them. For a wider map of what each tool does, see [the best AI design agents in 2026](/field-notes/best-ai-design-agents-2026).

---

## Visual cognitive load on a financial review screen: when the data was right but the layout fought the eye

> The JTBD was correct. The workflow was sound. The screen still felt like work to look at. A visual lens evaluation on a loan review tool where the problem wasn't structure — it was parse cost.

Source: https://www.thedesignagent.ai/field-notes/vcl-financial-review · Case studies · 2026-05-06

Not every evaluation finds a structural problem. This one found a well-structured interface with a visual problem — and that distinction matters, because the fix is entirely different.

The interface was a loan application review tool for underwriters. Each application was a record; the underwriter's job was to review the key risk signals, read the applicant summary, and make a recommendation — approve, flag for committee review, or decline. The workflow was correctly shaped: one application at a time, the key signals visible without navigation, the recommendation action at the end of the review flow.

The JTBD sub-score was 81. Solid. The `extraneous_load` score was 76. The job was right and the ceremony was low.

The `/visual` scores told a different story. `hierarchy`: 51. `density_fit`: 48. `critical_path`: 44.

## What the visual evaluation found

The interface had seven distinct visual regions: an application header with key identifiers, a risk signal panel with six badge-style indicators, a financial summary table, an applicant background section, a property details panel, a notes and history section, and an action bar at the bottom. All seven regions were rendered at approximately equal visual weight — similar card styling, similar typography scale, similar border treatment. The page had no dominant focal point.

An underwriter opening this screen had no visual guidance about where to look first. The risk signals — the highest-value information for a quick triage decision — were in a panel that looked identical to the property details panel. A first-time user would scan the entire page to find them. An experienced user would have memorized the layout and compensated for it with habit. Neither outcome is what the design intended.

The `critical_path` score of 44 named this directly: the most important information for the job — risk signal status — was not the most visually dominant element on the screen. Something else was claiming the primary visual focus: the application header, which had the largest type on the page, used for identifiers the underwriter already knew from clicking into the application.

## The parse cost in practice

Visual cognitive load compounds with volume. An underwriter reviewing thirty applications in a day pays the parse cost thirty times. If each screen takes three seconds longer to orient to than it should — because the hierarchy doesn't guide the eye to the risk signals immediately — that's ninety seconds of accumulated overhead per day, across a job where deliberation time is already at a premium.

The subtler cost is attention quality. A screen that requires active scanning to find the important information is consuming working memory before the actual review starts. The underwriter arrives at the risk signals having already spent cognitive resources locating them. A screen where the hierarchy puts the risk signals first — visually dominant, immediately readable — lets the underwriter start the review at full attention.

> **Pattern:** visual cognitive load is paid before the user does anything. On a screen visited thirty times a day, even small parse costs accumulate into real time and real cognitive overhead. Hierarchy is the fix, not reduction.

## The patches

The visual patches were specific. They didn't reduce information — the underwriter needed all seven regions. They restructured visual weight so the hierarchy expressed priority.

**Risk signals to primary visual prominence.** The six risk badge indicators were pulled out of their equal-weight panel and placed at the top of the content area in a larger, more visually distinct treatment. High-risk and flagged signals received strong color differentiation — not decorative color, but semantic color that conveyed meaning at a glance. The panel became the first thing the eye landed on.

**Application header de-emphasized.** The applicant name and ID were real information, but not review-critical. The header type size was reduced from the largest on the page to a secondary level. The identifiers were still present and findable; they were no longer competing for primary visual attention.

**Financial summary table: hierarchy within the table.** The table had twelve rows of financial data at equal weight. Four of them were the key signals underwriters use for threshold decisions: LTV, DTI, income verification status, and property appraisal gap. These four received a distinct typographic treatment — slightly heavier weight and a subtle left-border accent on the row — so an experienced underwriter could locate the threshold values without reading all twelve rows every time.

**Reduced competing focal points.** Three other elements on the page had been styled with calls-to-action that competed visually with the primary action bar. Two were secondary links; one was a "print" button with icon treatment identical to the primary recommendation buttons. Visual weight was redistributed so the action bar at the bottom read as the singular destination for decision.

## The result

The `/visual` scores after the patches: `hierarchy`: 84. `density_fit`: 79. `critical_path`: 81. The JTBD scores were unchanged — the structural changes hadn't affected the workflow fit, only the parsing ease.

The practical difference was observable in user sessions: orientation time to the risk signals dropped. Experienced underwriters confirmed they were spending less time scanning and more time on the actual risk assessment. The information hadn't changed. The hierarchy now expressed what mattered.

This case is useful precisely because the structural problem wasn't there. When `jtbd_alignment` is already high and the score is still low, the remaining work is visual — and visual problems require a different kind of analysis than structural ones. The `/visual` lens is where that work happens.

---

## Visual cognitive load: when a screen is well-designed and still hard to look at

> A screen can be JTBD-aligned, cognitively appropriate for the job, and still be hard to parse. Visual cognitive load is the mental cost of reading the layout itself — paid before any task thinking starts.

Source: https://www.thedesignagent.ai/field-notes/visual-cognitive-load · Concepts · 2026-05-06

There's a class of design problem that passes a functional review and fails a visual one. The JTBD is right — the layout fits the job. The workflow is well-shaped. The actions are in the right places. And the screen is still, somehow, hard to look at.

Users describe this as "a lot," "cluttered," or "I need a second to find what I'm looking for." They're not describing task difficulty — they could complete the task fine. They're describing the cost of reading the screen before they get to the task. That cost is visual cognitive load.

## What visual cognitive load is

Visual cognitive load is the mental effort required to parse the visual layout itself — separate from the effort of doing the task the layout was built for. It's the tax paid before any work starts.

The distinction matters because task cognitive load and visual cognitive load have different causes and different fixes. Task load is high when the work is genuinely complex — twenty decisions to make, each with dependencies. Visual load is high when the screen is difficult to decode: too many elements competing for attention, hierarchy that doesn't guide the eye, typography that creates reading cost rather than reducing it.

A screen can have low task load and high visual load — a simple action buried in a cluttered layout. It can have high task load and low visual load — a dense but well-organized table of complex financial data where the hierarchy is clear and the eye knows exactly where to go. The two loads are related but independent.

## Sources of visual cognitive load

**Density without hierarchy.** When every element on the screen has approximately equal visual weight, the eye has no guidance about where to start. The user has to scan everything to find anything. Adding more whitespace doesn't help if the weights are still equal — the items are just further apart. The fix is hierarchy: establishing clear primary, secondary, and tertiary visual weight so the eye follows a path rather than performs an exhaustive search.

**Competing focal points.** A screen has one primary focal point — the thing the eye lands on first. When agents add multiple high-contrast elements (colored status badges, bold text, icon clusters, call-to-action buttons) without establishing one clear primary, the eye bounces between them trying to determine what matters most. This is visually expensive before a single word has been read. Reducing to one dominant focal point per section and treating secondary elements as genuinely secondary is the fix.

**Typography that adds reading cost.** Font size, weight, line height, and letter spacing all affect how hard text is to read. A body text that's 13px on a high-density display, or 14px with insufficient line height on a long paragraph, requires more effort to read than one that's been calibrated to the display context. When multiple text sizes are used with insufficient contrast between them (body at 14px, labels at 13px — one point of difference that creates no meaningful hierarchy), the typography is adding work rather than reducing it.

**Decoration that competes with content.** Background colors, border lines, and icon use intended to create structure can create noise when they carry more visual weight than the content they're organizing. A card with a strong border, a header background, an icon, and a title all at similar weights requires sorting through four visual signals to read one piece of content.

> **Pattern:** visual cognitive load is paid before the user does anything. High visual load makes fast users slow and slow users stuck — not because the task is harder, but because reading the screen costs more than it should.

## The difference from "busy"

Visual cognitive load is often described as "busyness," but they're not the same. A screen can be dense and have low visual cognitive load if the hierarchy is strong — a well-organized financial dashboard with many data points can be fast to read if each section has a clear dominant element and the relationships between sections are visually expressed. A screen can be sparse and have high visual cognitive load if the hierarchy is absent — a minimalist layout where nothing signals importance creates an undifferentiated field the eye doesn't know how to navigate.

Busyness is a visual quantity. Visual cognitive load is a parsing cost. The fix for busyness is reduction. The fix for visual cognitive load is hierarchy — and sometimes the right hierarchy requires more elements, not fewer.

## How TheDesignAgent scores it

The `/visual` lens evaluates five dimensions that map directly onto visual cognitive load: `hierarchy`, `density_fit`, `contrast`, `typography`, and `critical_path`. Together these produce a picture of whether the layout guides the eye or makes the eye work.

`critical_path` is the most direct measure: does the visual design lead the user to the most important element first? A low `critical_path` score means the most important action or information doesn't read as the most important — something else is claiming the primary visual focus.

Patches on visual cognitive load are among the most specific in the system: adjust the heading weight on X, reduce the contrast of secondary labels from Y to Z, collapse the status badge palette from six colors to three, establish one dominant font size for body content and reduce the number of intermediate sizes. These are changes the agent can make precisely. The specificity is what makes the difference between "clean it up" and an actual fix.

---

## What Figma access gives agents — and what it doesn't

> Agents can now read Figma files, extract components, and pull design tokens. This is useful for implementation fidelity. It doesn't solve design judgment. Here's why the gap between 'access to designs' and 'understanding of design intent' is still wide.

Source: https://www.thedesignagent.ai/field-notes/what-figma-access-gives-agents · Concepts · 2026-05-06

Figma has an API. Several MCP servers now wrap it, letting agents read file structures, extract component properties, pull design tokens, and get the spacing values off a specific frame. When you're using Claude Code to implement a design that lives in Figma, this is genuinely useful. The agent can check its implementation against the spec rather than asking you to eyeball the result.

But implementation fidelity and design judgment are different problems. Figma access solves the first. It doesn't touch the second.

## What Figma access actually gives you

A Figma MCP server lets an agent answer questions like: what is the border radius on this component? What are the approved color tokens? What does the approved button variant look like? What spacing does the design spec call for between these elements?

These are questions about matching — does the implementation match the design file? That's a real and useful problem to solve. Designers spend real time reviewing implementations for pixel-level fidelity to specs. If an agent can check its own implementation against a Figma source before sending a PR, that's a meaningful efficiency gain.

The Figma MCP is, in this framing, a spec-checker. The spec exists in Figma; the agent can read it; the implementation can be compared against it. Mismatches surface early.

## What it doesn't give you

Figma access tells an agent what the approved design looks like. It doesn't tell the agent whether the approved design is right for the job.

Design files contain decisions, not the reasoning behind the decisions. A Figma component library tells you that the approved button has these dimensions and this color. It doesn't tell you when to use a button versus a link, or when a button shouldn't exist at all because the action belongs inline on the row rather than in a separate control. It doesn't tell you that the modal pattern in the component library is appropriate for configuration flows and wrong for batch-decision workflows.

The reasoning lives in designers' heads, in documentation that rarely gets written, in institutional knowledge that doesn't transfer when the design team turns over. It doesn't live in the Figma file.

> **The distinction:** Figma access tells agents what approved design looks like. Design judgment tells agents whether the right design was chosen for the job. Both matter; only one is available today.

## The spec-checker problem

There's a deeper problem with spec-checking as the primary design quality gate: it validates compliance with decisions that may themselves be wrong.

If a design file specifies a tabbed layout for a sequential workflow — because the designer didn't know the job was sequential, or because the design was done before the operational context was clear — a Figma-checking agent will implement the tab layout faithfully. The implementation will match the spec. The spec will be wrong. The spec-checker will pass.

This isn't a failure of the Figma MCP. It's a limitation of spec-checking as a model. Spec-checking is a downstream quality gate: it catches implementation errors against an existing decision. Design judgment is an upstream quality gate: it catches decision errors before they become specs or implementations.

Both are needed. Most teams have the downstream gate (designers review implementations). Very few have the upstream gate in their agent workflow (judgment applied to the design decision itself, before it becomes a spec the agent implements).

## Where Figma fits in an agentic design workflow

Figma access and design judgment aren't competing — they're complementary layers in a complete design quality stack.

Design judgment (does this layout fit the job?) runs early, on the agent's proposed approach or first build. It catches structural errors before they get designed and specified.

Figma compliance (does this implementation match the approved spec?) runs late, after design decisions are made and codified. It catches implementation errors before they ship.

The gap in the current ecosystem is the first layer. Figma MCPs have built the second layer well. There's nothing for the first layer beyond what teams build manually through design review cycles — which don't happen on every build, and don't happen at the speed agents need to loop.

What's missing isn't access to design files. It's a judgment layer that runs before the files exist — that catches the decision to build a dashboard for a deciding job before the dashboard gets designed, specified, and implemented.

---

## Why there's no UX linter (and what it would take to build one)

> ESLint catches code that violates rules. There's no equivalent for UI that violates design principles. The reason isn't that nobody tried — it's that design correctness isn't rule-shaped.

Source: https://www.thedesignagent.ai/field-notes/why-there-is-no-ux-linter · Concepts · 2026-05-06

ESLint runs on every save. It catches unused variables, missing dependencies in useEffect, components that violate accessibility rules. It's fast, deterministic, and deeply integrated into every serious development workflow. The output is binary: a violation or not.

There's no equivalent for UX. No tool that runs on every build and catches: tabs used for a sequential workflow, a dashboard built for a deciding job, a modal on a batch-decision row. The accessibility parts of a linter come closest — contrast ratios, missing ARIA labels, insufficient touch targets — but these are the hygiene baseline, not the judgment layer.

The absence isn't surprising once you understand why linters work and why UX judgment doesn't share the same properties.

## Why linters work

A linter works because the rules it enforces are:

**Deterministic.** A missing `alt` attribute either exists or it doesn't. An `onClick` without a corresponding keyboard event either has one or it doesn't. The check is computable from the code alone.

**Context-independent.** An unused variable is an unused variable in any program. A missing dependency in `useEffect` is always a bug. The rule applies without needing to know what the software is for or who will use it.

**Formally specifiable.** The rules can be written down precisely enough that a computer can check them. "Every image element must have an alt attribute" is machine-checkable. "The layout should serve the user's job" is not.

Most design principles fail on all three counts. "Reduce cognitive load" is not deterministic — cognitive load depends on who the user is, what their expertise level is, and what the domain requires. "Don't bury the primary action" is not context-independent — the primary action depends on the job, and what's primary in one workflow is secondary in another. "Match the layout to the workflow" is not formally specifiable without knowing what the workflow is.

> **The core problem:** a linter checks code against rules. A design judgment tool has to check UI against meaning — the meaning of the layout to a specific user doing a specific job. Meaning isn't rule-shaped.

## What the accessible parts of a linter actually check

It's worth being precise about what accessibility linters do and don't catch, because they're often cited as "the closest thing to a UX linter."

Accessibility linters check for WCAG compliance on dimensions that are formally specifiable: contrast ratios above a threshold, interactive elements with accessible names, ARIA roles used correctly, focus order that follows DOM order. These are necessary conditions for an accessible interface.

They don't check whether the interface is actually usable by the people it's supposed to serve. A form with perfect ARIA labels and sufficient contrast can be unusable for a user with low working memory capacity if it presents twelve fields that need to be filled before submission. The accessibility linter would pass it. A design review would not.

The formally checkable accessibility rules represent roughly the bottom 20% of what makes a UI good or bad. They're the floor, not the ceiling.

## What a real design check would need to know

To catch the kinds of design failures agents commonly produce — dashboard for a deciding job, tabs for a sequential workflow, modal on a batch-decision row — a tool would need to know:

**What the job is.** A tool that doesn't know whether the user is deciding or monitoring can't distinguish a dashboard from a decision surface. Both are valid layouts; the question is which one fits the job.

**Who the user is.** A tool that doesn't know the user's cognitive profile can't evaluate whether the extraneous load is tolerable or fatal. In a warehouse picking context, extraneous load compounds into real throughput loss. In a deliberation context, the same amount of friction is acceptable.

**What domain patterns apply.** A tool that doesn't have domain-specific pattern knowledge can't flag that tabs are wrong for fulfillment workflows while being appropriate for settings pages. Generic "avoid tabs" advice is wrong; domain-specific "tabs are wrong for sequential jobs" is right.

This is why design judgment requires context that a traditional linter doesn't have. The check isn't against a rule; it's against a job-specific standard that varies by domain, by user, and by the nature of the task.

## What building toward it looks like

The path toward automated design judgment is not more rules. It's richer context.

If a tool knows the domain, the persona, and the job — and has accumulated knowledge of which patterns fail for which combinations of those three things — it can make job-specific evaluations that go beyond formal rules. This is closer to a specialized reviewer than a linter: it has expertise the agent doesn't, and it applies that expertise to a specific context.

The difference from a linter is that the tool has to learn what the product is, who uses it, and what they're trying to do. That learning happens at the first call, when discovery builds the project profile. Every evaluation after that is grounded in that context.

That's not a linter. It's something new. And it's why the design judgment layer has taken this long to exist.

---

## Two lenses, one pipeline: why UX and visual review are different jobs

> A single rubric kept producing the same shape of answer regardless of what the agent was actually getting wrong. Splitting it in two fixed that.

Source: https://www.thedesignagent.ai/field-notes/two-lenses-one-pipeline · Concepts · 2026-04-22

When we started TheDesignAgent, we tried to build a single rubric. One score, one set of patches. It kept producing the same shape of answer regardless of what the agent was actually getting wrong. A layout-fights-workflow problem looked the same as a contrast-too-low problem. Both got "mid-fit score" and a vague set of suggestions.

The fix was to split the rubric in two. Not as a feature or a tier — as an admission that **UX and visual review are answering different questions** and shouldn't be averaged.

## The two questions

`/ux` answers: *does this serve the person doing their job?* Its concerns are information architecture, the shape of the workflow, what the user already knows, what they need to do next. Its evidence is JTBD, cognitive load, and the affordances on screen.

`/visual` answers: *is this design well-crafted?* Its concerns are spacing, hierarchy, type, color, restraint. Its evidence is layout grids, contrast ratios, density, and the integrity of the component grammar.

> **Why the split:** a screen can be visually beautiful and still wrong for the job. A screen can be on-spec for the job and still ugly. Averaging hides which problem you have.

## Different rubrics, same envelope

Both lenses return the same response shape: a fit score, prioritized patches, and flags. That consistency matters because the agent is the one calling, applying, and looping. It expects a contract. We just have two rubrics behind the same contract.

The UX rubric scores four dimensions: `jtbd_alignment`, `intrinsic_load`, `extraneous_load`, `germane_load`. The visual rubric scores seven: `critical_path`, `density_fit`, `hierarchy`, `contrast`, `typography`, `spacing`, `component_grammar`. The dimensions aren't leaked to the agent — only the score and patches. The dimensions are how we ground the score, not how we communicate it.

## What this means in practice

For a build that's shaped wrong for the workflow but visually competent, `/ux` flags it; `/visual` doesn't. For a build that uses the right pattern but flattens it into shadcn defaults, `/visual` flags it; `/ux` doesn't. Most builds run both, and the patches stack: the agent rebuilds the shape first, then the craft.

This is why the homepage's receipts section [includes three different cases](/#receipts) — warehouse, onboarding, inventory. Each one is primarily a UX failure or a visual failure, not both. Splitting the lenses makes that legible to the agent and to you.

## What we're still figuring out

The hardest part isn't the rubrics. It's the project memory that grounds them. A persona's cognitive profile, the active job, the day-shape — these have to be inferred and updated, and they're what makes the score specific to *this* product instead of generic. We'll write more about that in a future note.

---

## The dashboard default: why agents build viewers when the job is deciding

> Four KPI cards, a table, a date filter. The dashboard is the agent's lowest-effort answer to 'build me a tool for this data.' It's also wrong most of the time.

Source: https://www.thedesignagent.ai/field-notes/the-dashboard-default · Anti-patterns · 2026-04-08

Ask an AI coding agent for a tool and you'll get a dashboard. Four KPI cards in a strip across the top, a table below, a date filter, an export button. It looks competent. It usually isn't.

The problem is that **most jobs aren't monitoring jobs**. They're deciding jobs. Watching is what dashboards are for. Doing is what most users were hired to do.

## Watching vs. doing

Consider a planner reviewing weekly reorder recommendations. The agent ships a dashboard: out-of-stock count, reorder-soon count, capacity used. A row table. A row click opens a modal, the planner edits, saves, closes, clicks the next row. Twenty rows takes ten minutes of ceremony.

The job, though, is to decide on twenty recommendations. Each row IS the decision. The KPIs are useful as a running side-effect — but the planner doesn't need them to *start*. They need: a confidence signal per row, the AI's reasoning if they want it, and inline accept / override / reject. The capacity bar at the top is the cumulative cost of their decisions, not a measurement of system state.

> **Pattern:** when the user is going to take an action on every row of data they see, put the action on the row, not behind a click. The dashboard pattern hides the decision behind ceremony.

## Why agents reach for it anyway

The dashboard is the lowest-effort answer to "build me a tool for this data." It maps cleanly to the data shape: each KPI is an aggregate, each row is a record. The agent doesn't need to know the user, the job, or the workflow — just the schema. That's why it's the default. It's also why it's wrong most of the time.

For a JTBD that's about deciding, the right pattern is a *decision surface*. Each row is a unit of decision. Confidence is in line. AI reasoning is one click away, not gone. Keyboard shortcuts because batches are the actual usage pattern. A summary that updates as you decide, not as the data updates.

## How TheDesignAgent flags it

On the `/ux` lens, a dashboard built for a deciding job triggers two things: a `layout_fights_action` cognitive flag, and patches that reframe each row as the decision unit. The agent gets back "collapse modal into row" and "move actions inline" before the build is in your PR.

We use this case in the [homepage receipts](/#receipts) as the inventory planning example. The before is a clean dashboard. The after is a decision surface. Both look professional. Only one fits the actual job.
