# Design conformance  -  the component walk, and why a glance is not a pass

A design audit that reads a screen top to bottom and reports what looks wrong
finds about one defect per element: the wrong font on a label, and not that same
label's wrong colour and wrong inset. This file is the catalog, and the three
mechanisms that turn "I looked at it" into a countable result.

Derived generically from an upstream corporate screen-QA tool (recorded in
`derivedSkillSources`). The doctrine is portable. Its vendored design-token
catalogs and product examples are not, and were not taken - section 7 says what
replaces them.

Consumers: `/multi-agent:design-check` (the runner), Phase 4 review when a UI
diff is under review, and `features/visual-evidence.md` when a capture has to
prove a fix. Gate: `smoke-design-conformance.sh`.

## 1. Enumerate first, then fill every cell

Do **not** walk by finding, and do not walk by headline. Build the inventory
once, before any comparison:

1. Enumerate every visible element of the design frame - every text, icon,
   image, field, **button**, container. An element nested inside a button, field
   or card is **its own row**, marked as nested (`> in <Button>`), never folded
   into its parent.
2. For each row, pre-fill the design side of every rule that applies and leave
   the implementation side blank.
3. Fill every blank. **A blank cell is a fail-to-verify, not a pass**, and it is
   reported as one.

The inventory is the coverage guarantee; the rule groups in section 5 are the
per-cell recipe. A truncated inventory makes the denominator wrong, so it is
reported as incomplete rather than as a clean percentage of a partial set.

## 2. Measure, never read the token

Every `[MEASURE]` item compares the design value against **measured pixels** in
the render, or against the implementation's resolved value in code. Never
against:

- the design tool's node box, which includes invisible padding, or
- the implementation's own layout token (a spacing or size constant, a
  `.frame(height:)`, a Compose `Modifier.height`).

The token is the thing under test. Measuring it against itself is how a 48pt
field ships against a 56pt design with every check green. Button and field
**height** are the two a token lies about most often, because internal padding
and safe-area insets are not in it.

## 3. Adaptive per-component convergence

One pass over a component catches roughly one defect on it. So each component is
re-checked until **two consecutive attempts find nothing new on it**.

- Soft cap: reaching attempt 8 without two consecutive clean attempts stops that
  component with a `did-not-converge` warning. There is never a ninth attempt.
- A component with a font, a margin and a colour defect needs about three
  finding attempts plus two clean ones. A correct component needs the two clean
  attempts and nothing more.
- The screen is done when **every** inventoried component has reached two
  consecutive clean attempts or its cap with a warning. Not when the walk
  reached the bottom.

On a fixing run, each attempt that found something is fixed, re-captured, and
its changes appended to one accumulating record, so the final artefact shows
what each attempt changed rather than one undifferentiated diff.

## 4. Scope tags  -  how an item is verified

| Tag | Meaning |
|---|---|
| `[MEASURE]` | Verified in the primary capture: design value against measured pixels or resolved code value (section 2) |
| `[CAPTURE+]` | Needs an extra capture pass - a second appearance, a mirrored layout, another width, another state. Skipping one is allowed only if the skip is logged |
| `[A11Y]` | Accessibility lane: the platform accessibility tree (labels, element frames) plus contrast math |
| `[DYNAMIC]` | Behavioural or temporal; a static capture cannot show it. Verify by driving the device, or report `not verified (dynamic)`. **Never silently pass a `[DYNAMIC]` item** |

## 5. The catalog

Run every applicable item for **every** component, then every gap through
group 2, then the whole-screen passes (17, 21, 22, 24, 25). A check that was not
run is a fail-to-verify.

**1. Layout** `[MEASURE]` - x relative to container (never absolute y: the status
bar offsets it), y relative to an in-content span, L/R/T/B margin to parent,
centring, baseline alignment across a row, shared leading edge for siblings,
label-left/value-right rows aligned at both ends, safe-area and system-bar
clearance.

> **Horizontal-inset lane** - the most-skipped one. For every element record
> **both** sides of the leading inset: the design value and the measured
> implementation value. One side filled is a fail-to-verify. Then two cheap
> catches that need no per-element math: (a) **leading rail** - every loose,
> non-carded text in a vertical group must share one leading x; if the
> implementation puts a subtitle at the card's outer edge while the design insets
> it, the rail is broken; (b) **sibling consistency** - when sibling screens
> exist, diff the same element's inset across them; the outlier is usually the
> bug. A different text-wrap point is a symptom of a different margin, not only
> of a different font.

**2. Spacing** `[MEASURE]` - outer margins, inner padding, screen edge margin
(L=R symmetric), text-to-icon and icon-to-button and card-to-card and
section-to-section gaps, list/grid/row spacing, the vertical gap **between**
adjacent components (`next.y - (prev.y + prev.h)` against the stack spacing), and
every gap resolving to a spacing token - an off-scale value is a deviation.

**3. Size** `[MEASURE]` - width and height against the design (measured), min and
max honoured, `[CAPTURE+]` responsive widths (group 17), `[A11Y]` interactive tap
target at least 44x44pt on iOS / 48x48dp on Android, read from the accessibility
tree rather than assumed.

**4. Typography** `[MEASURE]` - family, size, weight, style, letter spacing, line
height, paragraph spacing, alignment, text transform, text decoration, foreground
colour, and truncation: ellipsis position, wrap and line limit matching the
design.

**5. Colours** `[MEASURE]` - background colour and opacity, gradient direction and
stops, blur effects, foreground and secondary text, disabled and placeholder,
border and divider and icon tint, brand accent and the status set
(success/error/warning/info), field prefix or affix colour, and no hardcoded
value: every colour resolves to a design token.

**6. Border** `[MEASURE]` - width, corner radius (scalar and per-corner), style
(solid vs dashed) and stroke alignment, colour and opacity.

**7. Shadow / elevation** `[MEASURE]` - present or not, colour and opacity, blur
radius, spread, x/y offset, and parity between the iOS shadow and the Android
elevation for the same surface.

**8. Shape** `[MEASURE]` - the drawn shape matches: radius about half the height
means a pill; radius about half the smaller side with a square box means a
circle.

**9. Icons** `[MEASURE]` - correct icon and variant, **size measured from the
rendered bounding box in both images**, never the node box or a size token,
weight (line vs filled), colour and tint, padding and alignment and icon-to-text
spacing, rotation and mirroring, rendering mode (a multicolour icon rendered as a
single-tint template is a defect), and vector crispness.

**10. Images** `[MEASURE]` - width and height, aspect ratio preserved, crop mode
(fill vs fit), no distortion, no unintended blur and sufficient resolution,
`[CAPTURE+]` placeholder and loading state.

**11. Buttons** `[MEASURE]` - width and **measured height** (from the rendered
fill rect, never a frame token), L/R/T/B spacing with L=R symmetry, the inner
label's font/weight/colour/alignment **as its own row**, the inner icon
(presence, size, colour, side, spacing), centring, fill and border and text
colour for the current state, `[CAPTURE+]` the state set (group 21).

> A button instance matches no text, icon or field kind, so an inventory built by
> kind drops it silently - and with it the button height and its nested label.
> The button is its own row and its children are exploded into their own rows.

**12. Input fields** `[MEASURE]` - **measured height** (the rendered field box,
never a frame or size token, because internal padding is not in the token),
width, border, radius, placeholder text and its colour and alignment, background
shade (read-only vs editable), prefix or affix and its colour, cursor and
selection colour, character counter behaviour, `[CAPTURE+]` focus/error/disabled/
filled states with the error message's colour, position and icon, `[DYNAMIC]`
keyboard type.

**13. Card** `[MEASURE]` - padding, radius, background, border, shadow, and
content grouping: **sample the pixel behind every text block** - is it inside a
card or bare on the page background? A design block in a card that the
implementation renders bare (or the reverse) is a wrong-container finding
(group 27) and is invisible to the text checks alone. Diff the design's container
list one-to-one against the implementation's card wrappers; the implementation
side is code-verifiable (inside a card helper vs placed bare in a stack with
padding). Never infer this from a glance.

**14. Lists** `[MEASURE]` - item height and internal padding, divider colour and
thickness and insets, inter-item spacing, section header grouping and order,
`[DYNAMIC]` scroll behaviour and lazy loading.

**15. Navigation / header** `[MEASURE]` - bar height, title text and alignment,
back affordance present and correct and mirrored under RTL, right-side actions
present and ordered, bottom navigation items with icons, labels and selected
state, `[DYNAMIC]` transition style.

**16. Scroll** - `[DYNAMIC]` smooth scroll, bounce, overscroll, sticky header,
scroll indicator; `[MEASURE]` the pinned header or footer position, which is the
static part a capture can check.

**17. Responsive** `[CAPTURE+]` - small, medium and large phone widths, tablet,
landscape. No overflow, clipping or distortion at any of them. Re-capture per
width rather than reasoning about it.

**18. Alignment** `[MEASURE]` - horizontal, vertical, and baseline across a row.

**19. Animation and motion** `[DYNAMIC]` - duration, curve, delay,
fade/scale/slide/rotation, loading and skeleton animation, haptics. Verify by
recording the device or report `not verified (dynamic)`.

**20. Visibility** `[CAPTURE+]` - hidden vs visible renders correctly; collapse
and expand states captured separately.

**21. States** `[CAPTURE+]` per interactive component - default, pressed, focus,
selected, active, disabled (colour, opacity, and genuinely not tappable),
loading, empty, error, success.

> **State fidelity gate**: the captured implementation must be in the exact state
> the design frame depicts - the same toggle position, the same filled or empty
> content, the same validation result. A capture in the wrong state is redone,
> not compared.

**22. Accessibility** `[A11Y]` - contrast at least 4.5:1 computed from sampled
colours, tap target at least 44x44pt / 48x48dp from the accessibility tree,
screen-reader label present and meaningful plus a meaningful hint, correct
traits, an accessibility identifier on every interactive element, and a logical
focus order. Whether to also audit scaled text sizes depends on the project: a
design system with fixed, non-scaling typography makes a forced-scale render
meaningless noise, so state which of the two the project is instead of assuming.

**23. Platform** - iOS: safe area, notch and Dynamic Island clearance, home
indicator, native navigation behaviour. Android: Material compliance, status and
navigation bar treatment, ripple.

**24. Dark appearance** `[CAPTURE+]` - capture a second image in the dark
appearance and re-run groups 3 to 9 against the dark design frame. Background,
text, icon, border, shadow, divider and gradient all adapt; no colour stuck at
its light value; contrast preserved.

**25. RTL and localization** - `[CAPTURE+]` a mirrored-locale capture: layout
mirrored, leading and trailing swapped, directional icons flipped; `[MEASURE]` no
hardcoded strings, every user-visible string from a localization key; a raw or
undefined key rendered on screen (the design shows copy, the implementation shows
`Screen.SomeKey`) reported in its own localization section; `[CAPTURE+]` a
long-language pass with no clipping. Compare the string's *presence and shape*,
not a translation of the same phrase against itself.

**26. Content and data formatting** `[MEASURE]` - number grouping, currency symbol
and position, date format, pluralization, and a long dynamic value that does not
overflow. The value itself (a name, a number, a masked field) is not under test.

**27. Component presence and order  -  hard gate** - the group the whole walk
exists for.

> Mechanical, not a glance. Number every element of the design list - text **and**
> every icon, avatar, badge, **and every button** - and for each write its
> implementation counterpart or `MISSING`. Report the count explicitly: "N design
> components -> N implementation counterparts". "Complete" may not be concluded
> without that list. A button and the label nested inside it are two presence
> checks, not one.
>
> - **Icons are the most-missed presence class.** For every icon in the design
>   list, locate it in the render by measuring its coloured bounding box. A design
>   icon with roughly no matching pixels is a missing-element finding.
> - **Recurse into nested blocks.** The gate applies to rows inside cards - masked
>   rows, list rows, key/value rows - not only to top-level sections. "That card
>   looks right" is not a pass.
> - **Bidirectional, and code-side too.** Diff design-to-implementation
>   (`MISSING`) *and* implementation-to-design (`EXTRA`). The most-missed extra is
>   a neutral decoration the implementation adds - a trailing chevron, an
>   underline, a strikethrough - which a colour-based, one-directional comparison
>   cannot see. Diff the implementation's explicit decoration modifiers against
>   the design's own decoration nodes, independently of the string.
> - **A localization gap does not short-circuit this gate.** Raw keys are one
>   finding; presence, icon, decoration and structure checks still run, because
>   they read code and spec rather than the rendered string.

Findings: `MISSING` · `EXTRA` · `MISSING-BETWEEN` (dropped from the middle of the
sequence) · `REORDERED` · `WRONG-STATE` · `WRONG-CONTAINER` (present but grouped
differently, loose vs carded).

**28. Keyboard and input behaviour** `[DYNAMIC]` - keyboard avoidance with the
active field not covered, dismiss behaviour, and return/next moving to the
correct field.

**29. Analytics** - out of visual scope; checked only when the spec requires it.

**30. Cross-check matrix** - per component confirm x, y, w, h, the four margins,
the four paddings, font family/size/weight, line height, letter spacing, text
colour, background colour, border width/radius/colour, shadow, opacity, icon
size/colour/position, image size/crop, divider, alignment, responsive, state,
animation, safe area, and overflow. **Mark only deviations** on the image;
matches belong in the report text.

## 6. Output

Mark only problems on the image, with thin lines. Everything else goes in the
report, whose section keys are stable and whose prose follows `outputLanguage`:

| Key | Content |
|---|---|
| `deviations` | Numbered, each one a design value against the measured or coded implementation value |
| `localization` | Raw keys, missing keys, hardcoded strings |
| `matches` | What was checked and agreed |
| `ignored` | What was deliberately out of scope, with the reason |
| `not_verified` | **Every** `[DYNAMIC]` and `[CAPTURE+]` item that was not run, so the coverage gap is explicit rather than absent |

Each finding is `[what] + [why] + [how]` with `file:line` wherever the
implementation side is code-resolvable. A finding without a file reference is
still a finding; a finding without a "how" is a complaint.

## 7. What a token catalog is, and why none ships here

Comparing typography and colour exactly needs the project's own catalogs - the
resolved name, size, weight and line height of each typography style, and the
resolved value of each colour token. Those are **project assets**, not pipeline
content: they carry brand colours, licensed font names and product vocabulary.

So this contract defines the *interface*, never the data. The runner resolves
catalogs in this order and states which one it used:

1. A path in preferences (`designCheck.tokenCatalogs`), when the project exports
   them.
2. Generated from the repo's own token definitions at run time.
3. Neither available -> typography and colour items degrade from `[MEASURE]` to
   sampled-pixel comparison, and the report says so. They are not silently
   passed.

A vendored catalog inside the pipeline would be one project's design system
shipped to every other project: wrong for all of them, and stale for its owner.
