# Loop Protocol, one full review cycle

Five phases per prompt; human QA gate at cycle close; Phase 5 runs for FAILING
prompts only. `<plugin-root>` below is `$CLAUDE_PLUGIN_ROOT` in Claude Code; the plugin's
installed directory in Codex. Scripts ship in this skill at
`<plugin-root>/skills/gen-ui-review/scripts/` and are run from the
monorepo root (they read `apps/genui/…/gallery-latest.json` and write the
`review/` tree there). Corpus-pattern doctrine consumed by Phase 5:
[corpus-html-patterns.md](corpus-html-patterns.md).

---

## Data model

```text
apps/genui/app/gen-ui-gallery/review/
├── cycle-ledger.json          ← aggregate, schema-gated; read by gen-review-status.mjs
├── cycle-{N}.lock             ← sentinel during an active cycle
└── cycle-1/ … cycle-N/        ← unpadded numbering
    ├── scores.json            ← validates against scores.schema.json
    ├── review-report.md       ← append-only narrative
    ├── cycle-manifest.json    ← provenance (gallery version, decompose timestamp)
    └── screenshots/ raw-dom/ decomposed/   ← per-cycle scratch (gitignored)
```

Durable records are the four committed files; scratch dirs are written and
read within the same run.

## §Setup (before the first prompt)

1. Read `apps/genui/app/gen-ui-gallery/outputs/gallery-latest.json`. Validate
   structure (generated artifact, untrusted): must have `version` (integer),
   `generatedAt` (ISO string), `engines` (array), `groups` (array). Any key
   missing → stop and report. Do NOT validate it against scores.schema.json
   (that is the review OUTPUT contract). Extract only group slugs, prompt
   slugs, engine output component arrays.
2. Determine the cycle number with read-then-lock:
   - `N = max(cycle numbers in review/cycle-ledger.json ∪ on-disk
     review/cycle-* directory numbers) + 1`, or 1 if neither exists.
     (Dirs can outrun the ledger, decompose runs create `cycle-N/` before a
     scoring pass records it. Ledger-only numbering collides with those
     orphan dirs.)
   - Write a `cycle-{N}.lock` sentinel BEFORE writing any cycle data. If the
     lock already exists: another agent is running, stop.
3. Create `review/cycle-{N}/` and a `scores.json` skeleton conforming to
   [scores.schema.json](scores.schema.json), all prompt scores null,
   status `OPEN`.
4. Open `review/cycle-{N}/review-report.md` for append-only writing; header:
   cycle number, timestamp, prompt count, engine list.
5. Confirm the dev server responds (Playwright needs it). Do not boot a
   background server, the operator runs it (see §ManualHandoff).

---

## §Phase 1, Ideal-Output Specification (A data)

Delegate to **`primitive-authoring`** per prompt:

> "Compose the ideal AdiaUI UI for: '{prompt text}'. Output: (1) user intent +
> primary task + states, (2) ASCII DOM wireframe, (3) slot vocabulary table,
> (4) key prop/attr table."

Binary success check: the spec itself is not scored:

- Structured spec returned → `specProduced: true`, proceed.
- Error/empty/off-topic → `specProduced: false`, `status: FAILING` with cause
  `FREE_FORM_HALLUC`, skip Phases 2–5 for this prompt, log in review-report.md.

Capture as A-data: `spec.rootComponent`, `spec.layoutPrimitive`,
`spec.keyComponents` (DOM order), `spec.slotVocabulary`
(`{parent, slot, child}` triples), `spec.states`, `spec.wireframe` (stored as
string, not interpreted).

---

## §Phase 2, Canvas Decomposition (B data)

Mechanized end-to-end by the decompose script, screenshots, DOM walk,
primitive lookup (`TAG_TO_COMPONENT`, the authoritative table), attr
sanitization, overflow gate:

```text
node <plugin-root>/skills/gen-ui-review/scripts/gen-review-decompose.mjs
  --cycle N  [--group <slug>] [--prompt <slug>] [--port 5300] [--settle 2500] [--dry-run]
```

Outputs per prompt under `review/cycle-N/`: `screenshots/<slug>.png`,
`raw-dom/<slug>.json` (internal only), `decomposed/<slug>.json`, the
trust-boundary file:

```json
{
  "promptSlug": "auth-login-form",
  "renderFailure": false,
  "rootComponent": "Card",
  "layoutPrimitive": "Column",
  "components": ["Card", "NativeHeader", "NativeSection", "Column", "Field", "Input", "Button"],
  "slotPositions": [{ "parent": "Card", "slot": "header", "child": "NativeHeader" }],
  "attrs": { "Button": { "variant": "primary", "text": "Sign in" } },
  "unknownElements": [],
  "overflowElements": [],
  "screenshotPath": "review/cycle-N/screenshots/auth-login-form.png"
}
```

Only allowlisted attrs survive (`ATTR_ALLOWLIST` in the script); `data-*`,
`aria-*`, and raw text nodes are discarded. Assess decomposition quality with
[rubric-decompose.md](rubric-decompose.md).

**Overflow gate (mechanical, visual lane).** The script checks text-bearing
elements for `scrollWidth/Height > clientWidth/Height` under computed
`overflow: hidden`; hits land in `overflowElements`. Each entry auto-promotes
to a P1 in Phase 4. This lane is independent of Phase 3: score 93 with
overflow is still FAILING.

**RENDER_FAILURE protocol.** Canvas height < 50px OR empty `components`:
`renderFailure: true`, `status: RENDER_FAILURE` in scores.json, skip Phases
3–5, log in review-report.md, count toward `aggregate.renderFailureCount`.
RENDER_FAILURE does not block cycle close; the same prompt failing 3+ cycles
escalates to the operator as a pipeline bug.

---

## §Phase 3, A-vs-B Gap Scoring

Rubric: [rubric-score.md](rubric-score.md). Inputs: `spec.*` (A) + the
decomposed file (B), never the raw DOM. Before scoring, run the rubric's
§DomainMismatchCheck. Score D1–D5 (0–20 each) + D6 (mechanical root-component
match, 0 or +5; max 105). Record `rubricScore.score`, `rubricScore.delta` vs
prior cycle, dimension breakdown, and cause codes per gap.

**Regression block**: any prompt with `rubricScore.delta < -10` marks the
cycle `BLOCKED`, no cycle close until the regression is explained and the
prior cycle's fix plan audited.

---

## §Phase 4, Cosmetic Audit

Rubric: [rubric-cosmetic.md](rubric-cosmetic.md). Input: the Phase 2
screenshot. Runs for ALL prompts (a structurally passing prompt can still
carry a P1). Record P1/P2/P3 counts + issue list in scores.json.

---

## §Phase 5, Root Cause + Fix Plan

Runs ONLY for `status: FAILING` prompts, structural gate: read
`decomposed/<slug>.json`, confirm `renderFailure: false` and
`status: FAILING` in scores.json first. PASSING and RENDER_FAILURE prompts
never get fix plans. Reads ONLY the decomposed file.

0. **SoT HTML lookup (required).** Identify the canonical HTML page for the
   prompt's domain (corpus-html-patterns.md §SourceOfTruth map). If it has a
   `data-chunk` marker for this pattern: run `npm run verify:corpus`; stale →
   plan a re-harvest. If not: the plan is to ADD the marker to the canonical
   HTML then harvest, the plan's `file:` points at the HTML source, never
   chunk JSON. No canonical HTML for the domain → new authoring task
   (`primitive-authoring`), do not invent a pattern.
1. **Trace causes** using the 9 codes in rubric-score.md §Root-Cause
   Classification. Run the diagnostic confirmation for the suspected code
   before recording it (RETRIEVAL_SCORE → actually run the retrieval search;
   EMPTY_CHUNK → inspect the chunk JSON).
2. **Write the plan.** Each entry: `rank`, `action`, `file`, `impact`, `skill`
   (schema-required). `file` must be inside `apps/`, `catalog/`,
   `packages/gen-ui/engine/corpus/`, or `packages/gen-ui/a2ui/`, anything else is
   flagged for operator review. Corpus-class causes route to `a2ui-maintenance`;
   TRANSPILER_GAP / FREE_FORM_HALLUC route to `primitive-authoring`.
3. Append the ranked plan to `review/cycle-N/review-report.md`.

---

## §CycleClose

1. **Regression block check**, a `BLOCKED` flag from Phase 3 stops here.
2. **Apply fix plans** (FAILING prompts only) via the routed peer skills:
   edit SoT HTML → `npm run harvest:chunks` → `npm run verify:corpus`. This
   skill does not perform the edits.
3. **Regenerate**: `npm run gallery:generate`; confirm 0 console
   errors/warnings in the canvas output.
4. **Human QA sample, per-sweep, not per-cycle** (retired from this
   cycle's own COMPLETE gate 2026-08-12, spec-factory-dx-ws6-measurement.md
   REQ-11, gh#1137). Operator reviews 5 random PASSING prompts against:
   (a) serves the user's task? (b) right primary primitive? (c) would ship
   unchanged? Record the result in the next `qa/dx/` sweep record's
   `humanQA.{sampledPrompts,pass,fail}` (`node scripts/qa/dx-status.mjs`),
   not on this cycle's ledger row. `failCount ≥ 2` still means the
   thresholds are miscalibrated, recalibrate rubric-score.md §Thresholds
   against the human judgments, it just no longer blocks THIS cycle's own
   `status: COMPLETE`.
5. **Schema gate** (must exit 0 before touching the ledger):

   ```bash
   node <plugin-root>/skills/gen-ui-review/scripts/validate-cycle-scores.mjs --cycle N --strict
   ```

6. **Update ledger** (`review/cycle-ledger.json`): `cycleNumber`,
   `completedAt`, `engine`, `status`, `aggregate`
   (passingCount/failingCount/renderFailureCount/meanScore/Δ). `humanQA` is
   RETIRED from this per-cycle row (step 4), do not populate it here; a
   stray value is harmless (ignored) but the field's home is now the
   per-sweep `qa/dx/sweeps/*.json` record. Remove the `cycle-{N}.lock`
   sentinel.
7. **Exit condition**:

   ```bash
   node <plugin-root>/skills/gen-ui-review/scripts/gen-review-status.mjs --check-exit
   ```

   Exit 0 → `status: COMPLETE`; exit 1 → `status: OPEN` (the script lists the
   blockers). Δ = 0 for every prompt across a full cycle → escalate to the
   operator: the causes need substrate changes beyond corpus patching.

---

## §ManualHandoff, human-executed steps per cycle

`gallery:generate` and the decompose script need a running dev server, which
the agent must not boot in the background. Per cycle the operator runs:

```text
[Agent: Phases 1–5, fix plans] → HUMAN: apply data-chunk edits (via a2ui-maintenance plans)
  → HUMAN: npm run harvest:chunks
  → HUMAN: npm run gallery:generate
  → HUMAN: node <plugin>/skills/gen-ui-review/scripts/gen-review-decompose.mjs --cycle N
  → [Agent: re-score Phases 3–5]
  → HUMAN: npm run eval:diff -- --engine <engine>   ← corpus changes must hold the eval floors
  → [Agent: scores.json + ledger + schema gate + exit check]
```

---

## §Modes

**Single prompt** (diagnostic): same phases for the named prompt only; skip
cycle-close regeneration; write `review/single-{slug}-{timestamp}.json`; no
human QA gate.

**Root-cause only**: skip Phases 1–3; requires a prior cycle's decomposed file
+ scores.json as input; run Phase 5 directly.
