# Eval diagnostics, gap diagnosis + regression triage

**Diagnostics run before code changes, always.** The #1 anti-pattern is
"tweak and hope": changing a prompt or adding a chunk without knowing why the
score is low burns real-LLM eval runs and teaches nothing.

## Phase 0, author a stub-mode diagnostic (zero LLM cost)

For each failing intent, capture four signals:

1. **Search ranking**, what `searchChunks()` returns for the intent query.
2. **Composition output**, what the pipeline emitted (HTML or plan).
3. **Tag inventory**, every custom-element-like tag in the emitted HTML.
4. **Coverage delta**, expected_components vs found_components.

Template (runnable with `llmAdapter: null`):

```js
async function diagnose(intent) {
  const search = searchChunks(intent.intent, { limit: 10 });
  const comp = await composeFromIntent({ intent: intent.intent, llmAdapter: null });
  const tags = [...comp.html.matchAll(/<([a-z]+-[a-z-]+)[\s>]/gi)]
    .map(m => m[1]).filter((v, i, a) => a.indexOf(v) === i);
  const found = intent.expected_components.filter(tag =>
    new RegExp(`<${kebab(tag)}-ui[\\s>]`).test(comp.html));
  return { search, tags, found,
    missing: intent.expected_components.filter(t => !found.includes(t)) };
}
```

## Phase 1, classify failures

| Bucket | Symptom | Root cause | Cost |
| --- | --- | --- | --- |
| A. Holdout misalignment | Top-1 retrieved ≠ expected_chunk | Holdout drifted from corpus | minutes |
| B. Coverage gap | Chunk retrieved but HTML lacks expected tags | Chunk HTML doesn't contain those components | hours (new chunks) |
| C. Wrong shell | Retrieval OK, LLM picks wrong page shell | Prompt ambiguity / missing domain→shell mapping | hours (prompt tuning) |
| D. Broken render | HTML emitted but console errors / blank | Missing registrations or bad markup | hours (harvester bug) |
| E. Measurement bug | Output looks right, score low | Scorer regex/casing/substring bug | minutes |
| F. Embedding drift | Async top-1 ≠ sync top-1 | Cosine boosts flip rankings non-deterministically | minutes |

**Fix order: A → E → F → C → B → D.** Fix measurement before content (else
scores can't be trusted); fix determinism before adding content (else evals
fluctuate).

## Phase 2, fixes per bucket

- **A. Holdout alignment**, map each intent's `expected_chunk` to the actual
  top-1; update `packages/gen-ui/engine/evals/corpus/holdout-compose-from-chunks.jsonl`.
- **E. Measurement traps**, PascalCase→kebab (`AgentTrace` → `agent-trace`,
  not `agenttrace`); substring false positives (`pane` vs `panel`,
  `textarea-ui` contains `text-ui`, word boundaries); case sensitivity (`/i`).
- **F. Embedding drift**, prefer sync keyword search for deterministic
  fast-path tiers; keep async cosine for the synthesis tier only (embeddings
  are a tie-breaker by design, see zettel-calibration).
- **C. Wrong shell**, check the `SYSTEM_PROMPT` domain→shell mapping in
  `chunk-synthesizer.js`; add explicit examples for the failing domain and a
  negative constraint ("NEVER default to dashboard-admin-page for
  non-dashboard intents").
- **B. Coverage gap**, author a block chunk the HTML-first way
  ([chunk-authoring](chunk-authoring.md)): demo page in a harvest root (e.g.
  `catalog/ui-patterns/app/<name>/`), `data-chunk` + `data-chunk-kind="block"`
  markers, real component tags so coverage scoring matches, then
  `npm run harvest:chunks`.
- **D. Broken render**, `packages/gen-ui/mcp/gen-ui/scripts/render-fidelity.mjs`
  output (console errors, blank viewport, undefined elements); verify
  registrations in `packages/web-components/index.js`; check the harvester
  didn't strip `data-chunk-slot` from page shells.

## Verification

```bash
npm run eval:compose-from-chunks                       # stub first: fast, free
npm run eval:compose-from-chunks -- --real-llm --report-file   # then real LLM
```

Stop only when all intents pass and the average is stable across 3 runs, and
the SKILL.md floors hold.

## Floor sources, read before quoting a number

The two `check:*-eval-regression` scripts own the floor numbers, read the source
before quoting a number elsewhere; SKILL.md only mirrors them and can drift (it
once silently regressed to `cov≥40` before the mechanical gate existed). Zettel's
floors are a committed file, `packages/gen-ui/engine/evals/health/zettel-floor.json` (gh#1391);
`scripts/release/check-zettel-eval-regression.mjs` loads it at runtime and refuses
to run without it, so re-baselining is a JSON diff, not a source edit. Free-form's
floors are still `ALERT_FLOOR`/`HARD_FLOOR` constants in
`scripts/release/check-free-form-eval-regression.mjs`.

Monolithic floor: cov=100, avg≥95. Dogfood set: 20/20, avg≥95. No mechanical
regression gate exists for monolithic yet: this floor is convention-only, same
failure mode the zettel/free-form gates were built to close. A failing gate
is the artifact, fix at the source (chunk HTML, engine code, tool schema),
re-run the narrowest gate, then the full sequence. A threshold tweak that
papers over a failing gate is a regression, not a fix.

## The eval suite's dimensions

`packages/gen-ui/mcp/gen-ui/scripts/test-evals.mjs` scores 5 weighted dimensions:
structural_validity 30% · intent_alignment 25% · component_coverage 20% ·
card_model_compliance 15% · anti_pattern_count 10%. `--save-baseline`
(`npm run test:evals:baseline`) stores scores; later runs flag any dimension
dropping >5 points or aggregate >3 (exit code 2 = regression).
