# Chunk authoring, HTML-first synthesis + harvest

The corpus is **one-format and harvester-driven**. Hand-authored pattern /
composition JSON dirs (`compose/patterns/`, `compose/{fragments,compositions}/`,
`corpus/patterns/`) were retired; the only retrieval substrate is
`packages/gen-ui/engine/corpus/chunks/` (526 chunks + `_index.json`), produced by
`npm run harvest:chunks`. In-tree SoT for this workflow:
`packages/gen-ui/engine/corpus/data-flow.md`.

## The authoring loop

1. **Author a live demo** in a harvest root: `apps/<name>/app/<demo>/`,
   `playgrounds/<name>/`, `catalog/<lib>/app/<demo>/`, or `site/pages/<route>/`.
   Demo pages are live-rendering and human-verified, they are the canonical
   source; chunk JSON is derived output.
2. **Tag retrievable regions** on the bounding element:
   `data-chunk="<slug>"` + `data-chunk-kind="block|page|panel|field"` +
   `data-chunk-domain` + `data-chunk-description` + `data-chunk-keywords`.
   Spec: `.claude/docs/specs/genui-chunk-marker.md`; dev tooling:
   `site/dev-chunks.{js,css}` (the `?chunks` overlay).
3. **Harvest**, `npm run harvest:chunks` (dry-run: `harvest:chunks:dry`)
   walks the source roots, writes `chunks/<slug>.json` + `_index.json`, and
   runs the transpile pass (`compose/transpiler/`) to produce the A2UI
   `template` for annotated chunks.
4. **Embed** (optional but expected for retrieval parity), `npm run build:embeddings:chunks` regenerates `chunk-embeddings.json`.
   Freshness gates: `npm run check:chunks-fresh` + `check:embeddings-fresh`.
5. **Verify**, `npm run smoke:chunks` (stub-LLM, offline) and a rendered
   check of the demo page; then the eval floor:
   `npm run eval:diff -- --engine zettel`.

Never hand-author or hand-edit `corpus/chunks/*.json`, they're build outputs;
the harvester wins on the next run.

## Metadata is the search index

`data-chunk-description` + `data-chunk-keywords` + the slug are what
`keywordScore()` and `searchAll()` match (name-token hits dominate, see
[zettel-calibration](zettel-calibration.md)). A chunk whose name lacks its
entity words is invisible to short queries. Write keyword-rich descriptions
derived from what's IN the chunk (headings, labels, button text), not
generic ("content card") prose.

## Component-catalog examples (the other training signal)

Per-component variant coverage lives in the component **yaml** (SoT) at
`packages/web-components/components/<name>/<name>.yaml`; `<name>.a2ui.json`
sidecars are **generated** by `node scripts/build/components.mjs` and
hook-guarded, edit the yaml and rebuild, never the sidecar. Templates use
PascalCase component names (`Chat`, not `chat-ui`); the registry maps class →
tag. Demo variants in `<name>.examples.html` are the canonical variant list;
yaml examples mirror them 1:1. Primitive/demo authoring itself belongs to the
authoring sibling skill.

## Pitfalls specific to authoring time

- **`?chunks` dev overlay prepends a `<span data-chunk-marker>`** into every
  `[data-chunk]` element, `:first-child` / `:nth-child` rules on those
  children break during dev only; don't chase it as a corpus bug.
- **Page-kind chunks declare slot regions via `data-chunk-slot="X"`**, slot
  regions are NOT themselves chunks. A new page-kind chunk needs a matching
  entry in the slot-validation map (`chunk-composer.js`) or compose-time
  validation rejects its plans.
- **Chunk-kind matters for retrieval budget**, pages and panels ride
  unconditionally into the LLM prompt; blocks compete for the remaining
  `PRE_SEARCH_LIMIT` budget. Mis-kinding a block as a panel inflates every
  prompt.

After harvest, continue with [corpus-discipline](corpus-discipline.md) for the
cross-cutting corpus pitfalls, and [leverage-rules](leverage-rules.md) before
splitting a repeated subtree into its own chunk.
