# Design conformance  -  the component walk, and why a glance is not a pass

<!-- toc -->
- [0. Why this runs as a gate](#0-why-this-runs-as-a-gate)
- [1. Enumerate first, then fill every cell](#1-enumerate-first-then-fill-every-cell)
- [2. Measure, never read the token](#2-measure-never-read-the-token)
- [3. Adaptive per-component convergence](#3-adaptive-per-component-convergence)
- [4. Scope tags  -  how an item is verified](#4-scope-tags---how-an-item-is-verified)
- [5. The catalog](#5-the-catalog)
- [6. Output](#6-output)
- [7. What a token catalog is, and why none ships here](#7-what-a-token-catalog-is-and-why-none-ships-here)
<!-- /toc -->

A design audit that reads a screen top to bottom and reports what looks wrong
catches roughly a single defect per element: it spots a label's wrong font and
misses the wrong colour and wrong inset on that very label. This file is the catalog, and the three
mechanisms that turn "I looked at it" into a countable result.

The doctrine is generic. It carries no design-token catalog and no product
examples - section 7 says what stands in for them.

Consumers: `/multi-agent:design-check` (the runner), Phase 3 review when a UI
diff is under review, and `features/visual-evidence.md` when a capture has to
prove a fix. Gate: `smoke-design-conformance.sh`.

## 0. Why this runs as a gate

A conformance check that no phase invokes leaves the user opening the app as the
only barrier between a build and visual drift: 16pt padding where the frame says
`Spacing/12` ships, and is found after the fact.

The reason it cannot be advice is structural, not historical: a reviewer reading a diff
cannot see spacing. Every other Phase 4 check reads text and reasons about text; this
one is the only thing in the pipeline that compares a rendered result against the
design it was drawn from. Left optional, it is the check that gets skipped on exactly
the runs that are in a hurry, which are the runs that produce drift.

## 1. Enumerate first, then fill every cell

Do **not** walk by finding, and do not walk by headline. Build the inventory
once, before any comparison:

1. Enumerate every visible element of the design frame - every text, icon,
   image, field, **button**, container. An element nested inside a button, field
   or card is **its own row**, marked as nested (`> in <Button>`), never folded
   into its parent.
2. For each row, pre-fill the design side of every rule that applies and leave
   the implementation side blank.
3. Complete every empty cell. **Any blank cell is a fail-to-verify** (never a pass),
   and the report lists it that way.

The inventory is the coverage guarantee; the rule groups in section 5 are the
per-cell recipe. A truncated inventory makes the denominator wrong, so it is
reported as incomplete rather than as a clean percentage of a partial set.

## 2. Measure, never read the token

Every `[MEASURE]` item compares the design value against **measured pixels** in
the render, or against the implementation's resolved value in code. Never
against:

- the design tool's node box, which includes invisible padding, or
- the implementation's own layout token (a spacing or size constant, a
  `.frame(height:)`, a Compose `Modifier.height`).

The token is the thing under test. Measuring it against itself is how a 48pt
field ships against a 56pt design with every check green. Button and field
**height** are the two a token lies about most often, because internal padding
and safe-area insets are not in it.

## 3. Adaptive per-component convergence

One pass over a component catches roughly one defect on it. So each component is
checked again and again, stopping once **two consecutive attempts find nothing new** there.

- Soft cap: if attempt 8 arrives and the last two attempts were not both clean, the
  component ends with a `did-not-converge` warning. A ninth attempt never runs.
- A component with a font, a margin and a colour defect needs about three
  finding attempts plus two clean ones. A correct component needs the two clean
  attempts and nothing more.
- The screen is finished once **each** component in the inventory has either two
  clean attempts in a row or has hit its cap and carries the warning. Not when the walk
  reached the bottom.

On a fixing run, each attempt that found something is fixed, re-captured, and
its changes appended to one accumulating record, so the final artefact shows
what each attempt changed rather than one undifferentiated diff.

## 4. Scope tags  -  how an item is verified

| Tag | Meaning |
|---|---|
| `[MEASURE]` | Verified in the primary capture: design value against measured pixels or resolved code value (section 2) |
| `[CAPTURE+]` | Requires one more capture - a second appearance, a mirrored layout, a different width or state. A pass may be skipped, provided the log records the skip |
| `[A11Y]` | Accessibility lane: the platform accessibility tree (labels, element frames) plus contrast math |
| `[DYNAMIC]` | Behavioural or temporal; a static capture cannot show it. Drive the device to verify it, or report `not verified (dynamic)`. **Never silently pass a `[DYNAMIC]` item**; it is either exercised or reported |

## 5. The catalog

Run every applicable item for **every** component, then every gap through
group 2, then the whole-screen passes (17, 21, 22, 24, 25). Any check left unrun
counts as a fail-to-verify.

**1. Layout** `[MEASURE]` - x measured from the container (y is never taken as an
absolute value, since the status bar shifts it), y relative to an in-content span, L/R/T/B margin to parent,
centring, baseline alignment across a row, shared leading edge for siblings,
label-left/value-right rows aligned at both ends, safe-area and system-bar
clearance.

> **Horizontal-inset lane** - the most-skipped one. For every element record
> **both** sides of the leading inset: the design value and the measured
> implementation value. If only one of the two is filled, that is a fail-to-verify.
> Two inexpensive checks follow that require no per-element arithmetic:
> (a) **leading rail** - loose text outside any card, stacked vertically, lines up
> on a single leading x; a subtitle the implementation places on the card's outer
> edge, where the design indents it, breaks the rail; (b) **sibling consistency** -
> if the screen has siblings, compare the inset of one element across all of them;
> the one that differs is usually the defect. Text that wraps at another point
> often signals a changed margin, and not only a changed font.

**2. Spacing** `[MEASURE]` - outer margins, inner padding, screen edge margin
(L=R symmetric), text-to-icon and icon-to-button and card-to-card and
section-to-section gaps, list/grid/row spacing, the vertical gap **between**
neighbouring components (compute `next.y - (prev.y + prev.h)` and compare it with the stack spacing), and
every gap resolving to a spacing token - an off-scale value is a deviation.

**3. Size** `[MEASURE]` - width and height against the design (measured), min and
max honoured, `[CAPTURE+]` responsive widths (group 17), `[A11Y]` interactive tap
target at least 44x44pt on iOS / 48x48dp on Android, read from the accessibility
tree rather than assumed.

**4. Typography** `[MEASURE]` - family, size, weight, style, letter spacing, line
height, paragraph spacing, alignment, text transform, text decoration, foreground
colour, and truncation: ellipsis position, wrap and line limit matching the
design.

**5. Colours** `[MEASURE]` - background colour and opacity, gradient direction and
stops, blur effects, foreground and secondary text, disabled and placeholder,
border and divider and icon tint, brand accent and the status set
(success/error/warning/info), field prefix or affix colour, and no hardcoded
value: every colour resolves to a design token.

**6. Border** `[MEASURE]` - width, corner radius (scalar and per-corner), style
(solid vs dashed) and stroke alignment, colour and opacity.

**7. Shadow / elevation** `[MEASURE]` - present or not, colour and opacity, blur
radius, spread, x/y offset, and parity between the iOS shadow and the Android
elevation for the same surface.

**8. Shape** `[MEASURE]` - the drawn shape matches: radius about half the height
means a pill; radius about half the smaller side with a square box means a
circle.

**9. Icons** `[MEASURE]` - correct icon and variant, **size measured from the
rendered bounding box in both images**, never the node box or a size token,
weight (line vs filled), colour and tint, padding and alignment and icon-to-text
spacing, rotation and mirroring, rendering mode (a multicolour icon rendered as a
single-tint template is a defect), and vector crispness.

**10. Images** `[MEASURE]` - width and height, aspect ratio preserved, crop mode
(fill vs fit), no distortion, no unintended blur and sufficient resolution,
`[CAPTURE+]` placeholder and loading state.

**11. Buttons** `[MEASURE]` - width and **measured height** (from the rendered
fill rect, never a frame token), L/R/T/B spacing with L=R symmetry, the inner
label as a **separate row** (font, weight, colour, alignment), the inner icon
(presence, size, colour, side, spacing), centring, fill and border and text
colour for the current state, `[CAPTURE+]` the state set (group 21).

> A button instance matches no text, icon or field kind, so an inventory built by
> kind drops it silently - and with it the button height and its nested label.
> So each button gets a row, and each child inside it gets a row of its own too.

**12. Input fields** `[MEASURE]` - **measured height** (the rendered field box,
never a frame or size token, because internal padding is not in the token),
width, border, radius, placeholder text and its colour and alignment, background
shade (editable vs read-only), the prefix or affix with its colour, cursor and
selection colour, character counter behaviour, `[CAPTURE+]` focus/error/disabled/
filled states with the error message's colour, position and icon, `[DYNAMIC]`
keyboard type.

**13. Card** `[MEASURE]` - radius, padding, border, background, shadow, and
content grouping: **sample the pixel behind every text block** - is it inside a
card or bare on the page background? A design block in a card that the
implementation renders bare (or the reverse) is a wrong-container finding
(group 27) and is invisible to the text checks alone. Diff the design's container
list one-to-one against the implementation's card wrappers; the implementation
side is code-verifiable (inside a card helper vs placed bare in a stack with
padding). Never infer this from a glance.

**14. Lists** `[MEASURE]` - item height and internal padding, divider colour and
thickness and insets, spacing between items, how section headers group and order,
`[DYNAMIC]` scroll behaviour and lazy loading.

**15. Navigation / header** `[MEASURE]` - bar height, title text and alignment,
back affordance present and correct and mirrored under RTL, right-side actions
present and ordered, bottom navigation items with icons, labels and selected
state, `[DYNAMIC]` transition style.

**16. Scroll** - `[DYNAMIC]` scroll smoothness, bounce and overscroll, a sticky
header, the scroll indicator; `[MEASURE]` the pinned header or footer position, which is the
static part a capture can check.

**17. Responsive** `[CAPTURE+]` - small, medium and large phone widths, tablet,
landscape. No overflow, clipping or distortion at any of them. Re-capture per
width rather than reasoning about it.

**18. Alignment** `[MEASURE]` - horizontal, vertical, and baseline across a row.

**19. Animation and motion** `[DYNAMIC]` - duration, curve, delay,
fade/scale/slide/rotation, loading and skeleton animation, haptics. Verify by
recording the device or report `not verified (dynamic)`.

**20. Visibility** `[CAPTURE+]` - both the visible and the hidden variant render as
designed; the expanded and collapsed states each get their own capture.

**21. States** `[CAPTURE+]` per interactive component - default, pressed, focus,
selected, active, disabled (colour, opacity, and genuinely not tappable),
loading, empty, error, success.

> **State fidelity gate**: before comparing, put the implementation into precisely
> the state shown by the design frame - the same toggle position, the same filled or empty
> content, the same validation result. A capture in the wrong state is redone,
> not compared.

**22. Accessibility** `[A11Y]` - contrast at least 4.5:1 computed from sampled
colours, tap target at least 44x44pt / 48x48dp from the accessibility tree,
screen-reader label present and meaningful plus a meaningful hint, correct
traits, an accessibility identifier on each element a user can interact with, and
a logical focus order. Whether to also audit scaled text sizes depends on the project: a
design system with fixed, non-scaling typography makes a forced-scale render
meaningless noise, so state which of the two the project is instead of assuming.

**23. Platform** - iOS: safe area, notch and Dynamic Island clearance, home
indicator, native navigation behaviour. Android: Material compliance, status and
navigation bar treatment, ripple.

**24. Dark appearance** `[CAPTURE+]` - capture a second image in the dark
appearance and re-run groups 3 to 9 against the dark design frame. Background,
text, icon, border, shadow, divider and gradient all adapt; no colour stuck at
its light value; contrast preserved.

**25. RTL and localization** - `[CAPTURE+]` a mirrored-locale capture: layout
mirrored, leading and trailing swapped, directional icons flipped; `[MEASURE]` no
hardcoded strings, every user-visible string from a localization key; a raw or
undefined key rendered on screen (the design shows copy, the implementation shows
`Screen.SomeKey`) reported in its own localization section; `[CAPTURE+]` a
long-language pass with no clipping. Compare the string's *presence and shape*,
not a translation of the same phrase against itself.

**26. Content and data formatting** `[MEASURE]` - number grouping, currency symbol
and position, date format, pluralization, and a long dynamic value that does not
overflow. The value itself (a name, a number, a masked field) is not under test.

**27. Component presence and order  -  hard gate** - the group the whole walk
exists for.

> Done mechanically, never by eye. Give a number to each element in the design
> list - the text **and** each icon, avatar and badge, **and each button** - and
> beside it name the matching implementation element, or write `MISSING`. State
> the totals outright: "N design components -> N implementation counterparts". "Complete" may not be concluded
> without that list. A button and the label nested inside it are two presence
> checks, not one.
>
> - **Presence checks miss icons more than anything else.** Find each design-list
>   icon in the render by measuring the bounding box of its coloured pixels. A design
>   icon with roughly no matching pixels is a missing-element finding.
> - **Descend into nested blocks.** Rows within cards (masked, list, key/value)
>   fall under the gate just as top-level sections do. A card that "seems fine"
>   has not passed.
> - **Bidirectional, and code-side too.** Diff design-to-implementation
>   (`MISSING`) *and* implementation-to-design (`EXTRA`). The most-missed extra is
>   a neutral decoration the implementation adds - a trailing chevron, an
>   underline, a strikethrough - which a colour-based, one-directional comparison
>   cannot see. Diff the implementation's explicit decoration modifiers against
>   the design's own decoration nodes, independently of the string.
> - **Missing localization never ends this gate early.** Raw keys form a single
>   finding; presence, icon, decoration and structure checks still run, because
>   they read code and spec rather than the rendered string.

Findings: `MISSING` · `EXTRA` · `MISSING-BETWEEN` (absent from inside the sequence,
not at its ends) · `REORDERED` · `WRONG-STATE` · `WRONG-CONTAINER` (present but grouped
differently, loose vs carded).

**28. Keyboard and input behaviour** `[DYNAMIC]` - keyboard avoidance with the
active field not covered, dismiss behaviour, and return/next moving to the
correct field.

**29. Analytics** - out of visual scope; checked only when the spec requires it.

**30. Cross-check matrix** - for each component confirm x, y, w and h, all four
margins and all four paddings, the font's family, size and weight, letter spacing,
line height, text colour, background colour, the border's width, radius and
colour, shadow, opacity, icon size, colour and position, image crop and size,
divider, alignment, responsive behaviour, state,
animation, safe area, and overflow. **Mark only deviations** on the image;
matches belong in the report text.

## 6. Output

Draw thin lines on the image around problems and nothing else. All other
content belongs in the report, whose section keys are stable and whose prose follows `outputLanguage`:

| Key | Content |
|---|---|
| `deviations` | Numbered, each one a design value against the measured or coded implementation value |
| `localization` | Raw keys, missing keys, hardcoded strings |
| `matches` | What was checked and agreed |
| `ignored` | What was deliberately out of scope, with the reason |
| `not_verified` | **Every** `[DYNAMIC]` and `[CAPTURE+]` item left unrun, which makes the coverage gap visible instead of silent |

Each finding is `[what] + [why] + [how]` with `file:line` wherever the
implementation side is code-resolvable. A finding without a file reference is
still a finding; a finding without a "how" is a complaint.

## 7. What a token catalog is, and why none ships here

Comparing typography and colour exactly needs the project's own catalogs - the
resolved name, size, weight and line height of each typography style, and the
resolved value of each colour token. Those are **project assets**, not pipeline
content: they carry brand colours, licensed font names and product vocabulary.

So this contract defines the *interface*, never the data. The runner resolves
catalogs in this order and states which one it used:

1. A path in preferences (`designCheck.tokenCatalogs`), when the project exports
   them.
2. Generated from the repo's own token definitions at run time.
3. Neither available -> typography and colour items degrade from `[MEASURE]` to
   sampled-pixel comparison, and the report says so. They are not silently
   passed.

A vendored catalog inside the pipeline would be one project's design system
shipped to every other project: wrong for all of them, and stale for its owner.
