# Corpus discipline, retrieval quality, harvest hygiene, drift traps

## Pipeline-side pitfalls (each one has burned a session)

- **"Thinking mode returns garbage"** → API keys must be in `.env` AND
  `scripts/load-env.mjs` imported; otherwise the stub adapter answers with a
  canned 6-component card.
- **"Search misses obvious patterns"** → inspect chunk metadata first
  (`npm run audit:corpus-stats`); bad descriptions/keywords = invisible chunks.
- **"Gate rejects valid intents"** → intent words are filtered by the
  `GATE_STOPS` set + 3-char minimum
  (`compose/strategies/monolithic/generate-instant.js`). Check whether the key
  word is a stop word before touching thresholds.
- **`npm run smoke:chunks` re-harvests chunks as a side effect**, touching
  `corpus/chunks/*.json` mtimes. After smokes, stage only files you actually
  changed, never a blanket `git add -A`.
- **The chunk-synthesizer fast-path threshold is 8** on the blended
  keyword+cosine score. Below it, synthesis fires. Tune in
  `chunk-synthesizer.js` (`STRONG_RETRIEVAL_SCORE`), not by inflating
  `keywordScore()` weights in `chunk-library.js`.
- **`state_id` is opaque to callers**, engine-generated on compose. Never
  construct or parse it; refinements pass back the prior `state_id` so the
  cache chains via `parent_state_id`.
- **After any `@bp` / layout-attribute change to `data-chunk`-annotated HTML,
  run `npm run harvest:chunks` in the same session**, otherwise training
  chunks silently hold stale values.
- **The corpus's inline `style=` is almost entirely structural page-frame
  layout** (no primitive exists for it), don't propose a mass
  "convert to primitives" campaign; it's a known, accepted shape.
- **Harvest roots are a hard-coded list** in `scripts/build/harvest-chunks.mjs`.
  A directory rename that adds new siblings silently drops chunks while the
  harvest "runs clean" (a rename once dropped 26 chunks for ~24h). After any
  root-level rename: check the roots list AND diff chunk counts against the
  previous `_index.json` tally.
- **Subagent-authored demo pages misuse component APIs** (nested content
  instead of label/description attrs, native `<table>` inside `table-ui`
  causing a "No data" overlay). Visual QA every subagent-authored page before
  harvesting it into the corpus.

## Transpile-pass parser lessons (HTML → A2UI)

The transpile pass (`compose/transpiler/`) inherited these hard-won rules, preserve them in any rewrite:

1. **Regex `([\s\S]*?)` can't handle nested same-name tags**, depth-tracking
   tag counting is required (`<card-ui><card-ui>…` matches the inner close).
2. **Only DIRECT children belong in `comp.children`**, not the flattened
   subtree; identify direct children as components not claimed by any other
   component.
3. **Re-ID by array index, not original ID**, multiple components can share
   an original ID; index-mapping guarantees uniqueness.
4. **Subtree walks need a visited set**, shared children reached via multiple
   parents otherwise duplicate.
5. **Boolean-attribute trap**: `text=""` parses as boolean `true`; any code
   doing `c.text.toLowerCase()` needs a `typeof c.text === 'string'` guard.

## Improving search quality without new infrastructure

Three interventions compose multiplicatively; apply in order:

1. **Enrich metadata**, derive descriptions from structural signals
   (Input("Email") + Input("Password") + Button("Sign In") → "Login form with
   email, password fields"). Historically 40% → 95% meaningful descriptions.
2. **Trace which search path each mode actually invokes**, semantic search
   has been built-but-unwired before; one-line wiring changes beat new infra.
3. **Expand synonym/keyword surfaces**, bridge user vocabulary to chunk
   vocabulary ("inbox" → notification) via `data-chunk-keywords`, remembering
   name-token hits outrank keyword hits.
