# Rubric — composed AdiaUI surface (COMPOSE / REALIZE)

The output standard a **composed AdiaUI surface** is graded against — a screen, flow, or shell
wired from AdiaUI primitives. A builder (`screen-composition` / `shell-selection`) self-checks against it;
the `surface-qa-agent` agent scores it adversarially — so a surface gets the same verdict
whoever ran the build. This is the composition-QUALITY lens, distinct from `surface-qa`'s
mechanical browser gate (which proves it *renders*): this proves it was *composed right*.

Two axes are scored **separately** — **COMPOSE** (is the surface assembled correctly, to the
*resolved* intent? — intent · primitives · pattern · wiring) and **REALIZE** (is the assembly
real, disciplined, verified, trusted — not asserted? — tokens · verification · trust). This is
the **defect quadrant**: a surface that composes perfectly on paper cannot ship if it was never
verified against the rendered surface, leaks side-effect CSS, or baked in unsanitized model
content — so a high COMPOSE score cannot offset a sub-3 REALIZE dimension, and vice-versa. Scale
1–5; 1 = failure, 3 = adequate, 5 = excellent. Each dimension is **[gate]** (mechanically
checkable — a named audit/grep/artifact is the evidence) or **[review]** (judgment grounded in
the resolved plan + the rendered surface). The a2ui MCP / component yaml is the source of truth
for primitives and slots — this rubric never restates it.

(Claims re-verified against the kit's live source 2026-08-26; re-verify on a MINOR cut.)

## COMPOSE axis (D1–D4) — assembled correctly, to the resolved intent

| # | Dimension | Type | What it checks | 1 → 3 → 5 |
|---|---|---|---|---|
| D1 | Intent fidelity | [review] | Realises the **resolved** plan (the reasoning-ladder output + wireframe checkpoint from `domain-planning/references/spec-to-ui-reasoning.md`, gh#1207 — app-planning-agent's Domain Plan block), not a keyword-driven premature render. Every plan region present; nothing material invented or dropped. | 1: pattern-matched from prompt keywords — a "dashboard" rendered because the word appeared, no resolved plan behind it; or plan regions missing · 3: every plan region present and traceable, no primitive chosen from vocabulary alone · 5: + the empty/error/loading states the plan named are all present, and each non-obvious composition choice is argued, not pattern-matched |
| D2 | Primitive correctness | [review] | Right primitive/slot for each need; **no reinvention** (no bare `<div>` where a layout primitive exists, no native form control where a `*-ui` exists); slot/field/a11y contracts honoured (wide controls inside `<field-ui>`, self-labelling widgets NOT; `<table-ui>` paired with `<table-toolbar-ui>`). | 1: bare `<div>` / native `<input>`/`<button>` where a primitive exists, a reinvented control, or a violated field/table-toolbar contract · 3: `adia-lint` NATIVE-PRIMITIVE clean; every need maps to an existing primitive in its correct slot · 5: + the primitive chosen is the best fit over its near-neighbours, slot vocabulary + ARIA exactly honoured, no primitive bent to another's job |
| D3 | Pattern / shell fidelity | [review] | Where a proven pattern or shell applies, the surface matches its anatomy — the right shell for the posture, chrome/nav/content in the documented slots (see the chrome-tier vs content-tier allocation in `references/shell-admin.md`); named patterns composed from canonical parts, not hand-rolled. | 1: wrong shell for the posture, chrome placed outside its slots, or a known pattern hand-rolled · 3: shell matches the posture, `adia-lint` LEGACY-SHELL clean, nav/chrome/content in the documented slots · 5: + bespoke deviations are deliberate and justified, and embedded composites were read so the surface composes against their real anatomy |
| D4 | State & data wiring | [review] | Reactive bindings correct and complete — every dynamic value wired; **no dead binding** (a value slot with no source) and **no illegal wiring** (a binding to a non-existent field, a write where the contract is read-only); loading/empty/error states wired, not only the happy path. | 1: dead or illegal bindings, or only the happy path wired · 3: every dynamic slot wired to a resolvable source; loading/empty/error present and bound (cite `data-wiring`) · 5: + the reactive graph is minimal and correct, and every named state renders against real/mock data |

## REALIZE axis (D5–D7) — real, disciplined, verified, trusted — not asserted

| # | Dimension | Type | What it checks | 1 → 3 → 5 |
|---|---|---|---|---|
| D5 | Token / theme discipline | [gate] | Visuals come from `--a-*` semantic tokens (or the `--md-sys-color-*` bridge; component-leaf `--<component>-*` for one-offs), never raw values where a token exists; the CSS policy holds — **no side-effect CSS, no `<style>`, no inline `style=`, no `<script>`**; every `--a-*` name *resolves* (no fallback-masked typo). | 1: raw hex/px where a token exists, a `<style>`/inline `style=`/`<script>`, a side-effect CSS import, or a broken `--a-*` name at browser default · 3: `adia-lint` RAW-COLOR/RAW-PX clean, zero style/script hits, the import policy followed · 5: + every `--a-*` name verified to resolve, component-leaf overrides use the `-default` seam, the surface re-themes by swapping the token scope with no hardcoded escape hatch |
| D6 | Verify closure | [gate] | Verified against the **rendered** surface, not self-asserted — the verify-target is the real rendered DOM (`adia-probe` / a Playwright snapshot / an audit exit-code), at a declared, honest tier (`structurally-verified-only` vs `nailed`). A green self-audit of the builder's own checks is **not** verify. | 1: verify is a self-assertion ("looks right" / "should render") with no rendered-surface artifact, or `nailed` claimed with no visual ack · 3: `structurally-verified-only` — renders in the browser without errors AND the named audits actually ran and passed (exit codes confirmed, not assumed) · 5: `nailed` — verified against the rendered surface with a human/visual ack (screenshot + notes) tied to the output; the verify exercised the named behaviour, not just presence |
| D7 | Content-trust | [gate] | Any LLM-generated or untrusted content baked in passed the trust gate **before write** — treated as *data, not instructions*; model/user strings escaped and bound as content (no prompt-injection passthrough, no executable markup injected). | 1: LLM/untrusted content concatenated into markup or a handler unescaped, or no trust gate before write · 3: untrusted content bound as escaped data through content slots; no executable markup from a model string; the trust gate ran before write · 5: + every untrusted source identified and bound as data, the gate's pass recorded with the surface, no path by which model output becomes instruction |

## Gate to ship

The two axes gate **independently** — the defect-quadrant rule:

- **COMPOSE:** D1, D2, D3, D4 each ≥ 3 (realises intent with correct primitives, pattern, wiring).
- **REALIZE:** D5, D6, D7 each ≥ 3 (disciplined, verified, trusted).
- **No cross-axis compensation** — a strong COMPOSE cannot offset a sub-3 REALIZE dimension, or vice-versa.
- **Every [gate] dimension is hard** — D5, D6, D7 are mechanically checkable; a check that fails (raw value present, no rendered-surface artifact, unsanitized content) blocks the ship regardless of the [review] scores. No judgment override.

Shippable = both axes clear ≥ 3 **and** zero [gate] fails.

**Top failure to look for first:** the **self-asserted verify closure** (D6 = 1) — a surface
whose "verified" claim is the builder's confidence, not evidence against the rendered surface; it
lets every other defect ship silently, which is the entire reason this rubric exists. The most
common COMPOSE-axis failure is the **keyword-driven premature render** (D1 = 1) — primitives
pattern-matched from prompt vocabulary instead of the resolved plan.

## Output contract (the reviewer fills this)

```
Artifact: <surface/route>   ·   Rubric: composed-surface
| Dim | Type | Score | Finding | Evidence |
|-----|------|-------|---------|----------|
| D1 Intent fidelity       | [review] |  /5 | … | plan-region ↔ surface trace |
| D2 Primitive correctness | [review] |  /5 | … | adia-lint NATIVE-PRIMITIVE |
| D3 Pattern/shell         | [review] |  /5 | … | adia-lint LEGACY-SHELL + shell-admin.md |
| D4 State & data wiring   | [review] |  /5 | … | binding trace vs data-wiring |
| D5 Token/theme           | [gate]   |  /5 | … | adia-lint RAW-COLOR/RAW-PX + token-name grep |
| D6 Verify closure        | [gate]   |  /5 | … | adia-probe artifact + tier |
| D7 Content-trust         | [gate]   |  /5 | … | trust-gate pass before write |
COMPOSE (D1–D4 ≥3): <pass/fail>   REALIZE (D5–D7 ≥3): <pass/fail>   [gate] D5,D6,D7: <all ≥3?>
Verdict: <shippable | blocked: D#…>
```

_Provenance: the two-axis COMPOSE/REALIZE defect-quadrant model is adapted from a component-grading
rubric; the audit tool names and slot contracts here are the factory plugin's own (`adia-lint`,
`adia-probe`, the current primitive catalog). Verify a named audit rule against `scripts/adia-lint.mjs`'s
actual rule set before citing it._
