# Semantic fail lifting, sub-60 triage procedure

Use when `npm run eval:diff -- --engine zettel --semantic` reports intents
with `semanticScore < 60` (or <70 for the watch list): the judge says the
emitted UI doesn't match what was asked for.

## The judge's three axes (`packages/gen-ui/engine/validate/semantic/index.js:32`)

- **dominantPattern** (weight 0.5), does the root/primary component match the
  intent type (chat, form, calendar, data-display, nav…)?
- **requiredCapabilities** (weight 0.35), are the specific controls the
  intent requires present?
- **forbiddenNoise** (weight 0.15), are off-topic components prominent?

A sub-60 score almost always means dominantPattern scored < 40.

## Triage

1. Read the latest eval run's zettel report, rows with `semanticScore < 70`,
   ascending.
2. Per row, read `semanticAxes.dominantPattern.{expected, observed}` and
   `semanticAxes.requiredCapabilities.missing`.
3. Bucket each failure:
   - **Thin composition**: the retrieved entry exists but is too sparse
     (calendar with nav + weekday labels but no grid).
   - **Wrong composition winning**, retrieval collision; another entry steals
     the intent via keyword overlap.
   - **No matching composition exists**, coverage gap; author a new one.

## Fix strategies, in order of preference

### A. Use a domain-specific primitive as the root

The judge weights root identity heavily: a Card wrapping 38 Badges + 43 Texts
reads as "data-display" even with a 7-column grid inside. Swapping to the
semantically correct primitive fixes it instantly. Historical lifts:
calendar → `CalendarPicker` root (32→65); chat → `Chat` with
`Text[role=user|assistant]` children (52→70+); command palette → `Command` >
`ActionItem` (42→70+). Check the live catalog before authoring children
(`lookup_component` / `get_component_map` MCP tools), invented child
components (`ChatMessage`, `CommandItem`) score zero and trip
`noInventedComponents`.

### B. Resolve retrieval collisions by keyword surgery, both sides

When an unrelated entry wins retrieval, don't just enrich the correct entry, **strip the overlapping keywords from the losing entry too**. Historical
example: `empty-state` kept winning "error state with retry" (sem=32) even
after a dedicated error-state entry existed, until "error state" and "retry"
were removed from empty-state's keywords; then the right entry won at sem=91.

### C. Author the missing coverage

Indicators: multiple intents fail pointing at the same wrong candidate, and no
existing entry's purpose matches `dominantPattern.expected`. Author it the
HTML-first way, a demo page with `data-chunk` markers, then
`npm run harvest:chunks` (see [chunk-authoring](chunk-authoring.md)). Put the
pattern's signature affordance as the dominant child (steps-timeline for a
wizard, textarea+richtext for an editor, accordion for settings) and give it a
rich keyword set (15–20 terms) with exact intent phrases and domain synonyms.

## Verify loop

```bash
node scripts/build/components.mjs --verify
npm run smoke:engines && npm run smoke:register-engine && npm run test:a2ui
node packages/gen-ui/mcp/gen-ui/scripts/eval-diff.mjs --engine zettel --semantic
```

The semantic judge is cached, content-hashed on
(rubricVersion, intent, a2ui-messages), only changed generations re-judge.
Compare `avgSem` and the sub-60 list row by row; hold the zettel floors
(cov≥87, avg≥85, MRR≥0.94) and require `avgSem` ≥ baseline before calling a
lift done.
