# Zettel calibration, constants + history + locator/modifier two-pass

**Calibration history is the substrate.** Each tweak left a trail; the same
value may have been tried and rejected before. Recover any decision's context
with `git log -S <CONSTANT> -- packages/gen-ui/a2ui` and the linked PR description.
**Don't change any constant without running `npm run eval:diff` first.** Each
is calibrated against the held-out intent set or production telemetry.

## `STRONG_MATCH_THRESHOLD = 40`

- **File**: `packages/gen-ui/a2ui/compose/strategies/zettel/generator-adapter.js` (`grep -n STRONG_MATCH_THRESHOLD`)
- **Raised**: 22 → 40 post-incident.
- **Reason**: at 22, login-form / signup-form played verbatim too often →
  repetitive output. At 40, only near-perfect retrievals (chart-dashboard=48,
  pricing-tiers=54) match; merely-good falls through to LLM synthesis for
  compositional variety.
- **Scale**: **absolute, not corpus-relative.** `searchAll()`
  (`composition-library.js`) scores per query token: name-token hit **12**,
  keyword hit **8**, tag hit **3**, description hit **2** (stop-words like
  "form"/"page" score 1/1/0/0), plus **+3 per distinct content token matched**;
  a candidate needs ≥2 content-token hits OR a direct name-token hit to score
  at all. Adding/removing corpus entries does not shift any specific
  (query, composition) score: the sum is deterministic per pair. **Don't
  recalibrate as a function of corpus size**, that's a misdiagnosis.
  Practical consequence: short queries need their entity words IN the chunk
  name to clear 40.
- **Tradeoff**: more LLM calls (slower, costlier) ↔ compositional variety.
- **Re-eval before changing**: `npm run eval:diff -- --engine zettel`.

## `STRONG_RETRIEVAL_SCORE = 8`

- **File**: `chunk-synthesizer.js` (`grep -n STRONG_RETRIEVAL_SCORE`)
- **Scale**: corpus-size-**independent** absolute keyword score from
  `chunk-library.js#keywordScore()`: first name-word matches a token **+10**,
  full-query substring **+5**, whole-word token in name **+3**, substring /
  haystack hits **+1** each. Anything below 8 is a "retrieval too weak, synthesize" signal.
- **Async path**: `searchChunksAsync` blends `kw + cos*5`. Cosine ranges 0..1,
  so embeddings contribute 0..5, a pure-cosine match maxes at 5 and can never
  clear 8 alone. **Embeddings are a tie-breaker, not the primary signal. This
  is intentional** (WONTFIX): letting cosine clear the fast-path gate made
  retrieval non-deterministic and flipped top-1 rankings unpredictably.
- **Different scale** than `STRONG_MATCH_THRESHOLD=40` (which scores
  name/keyword/tag/description sums). Never normalize the two, they measure
  different things.

## `PRE_SEARCH_LIMIT = 30`

- **Files**: `chunk-synthesizer.js` + `chunk-refiner.js` (`grep -n PRE_SEARCH_LIMIT`)
- **Reason**: token-budget mitigation, pre-filter the catalog before the LLM
  sees it; the full catalog per prompt would burn tens of thousands of tokens.
- **Synthesizer**: kind-aware allocation, `limit: PRE_SEARCH_LIMIT - pageChunks.length - panelChunks.length`. All pages
  and panels ride unconditionally; blocks fill the remainder. Self-tuning by
  structure.
- **Refiner**: block-only `limit: PRE_SEARCH_LIMIT`, plus all pages/panels on
  top, intentionally more generous because refinement does targeted edits and
  the LLM needs options.
- **Don't naively divide by corpus size** to assess over-permissiveness, the
  kind-aware allocation makes the math non-linear.

## `SCOPE_DRIFT_RATIO = 1.5` + `SCOPE_DRIFT_MIN_ACTUAL = 20`

- **File**: `chunk-synthesizer.js` (exported; `grep -n SCOPE_DRIFT`)
- Composed envelope's component count > 1.5× the sum of bound chunks' counts
  auto-fires a `scope-drift` issue, catches LLM creative expansion that
  hallucinates components beyond the bound chunks.
- The MIN_ACTUAL=20 floor kills false positives on small UIs where
  slot-wrapper noise dominates: <20 components never trips the gate.

## `DEFAULT_MAX_ATTEMPTS = 2`

- **Files**: `chunk-refiner.js` + `chunk-synthesizer.js` (`grep -n DEFAULT_MAX_ATTEMPTS`)
- Validator-driven retry budget. After 2 failed validations the engine emits
  `synthesis-failed`.

## `DEFAULT_MAX_SIZE = 64` (state-cache)

- **File**: `state-cache.js:27`; override via `A2UI_STATE_CACHE_SIZE` env var.
- Per-process, in-memory, survives only as long as the MCP server; multi-turn
  refinement breaks across restarts.
- **Eviction**: LRU on `set` at capacity; `get` and overwriting `set` touch
  recency; `peek` reads without touching.

## `TRACE_INLINE_THRESHOLD_BYTES = 200 * 1024` (issue-reporter)

- **File**: `issue-reporter.js:27`. Above this, traces spill to a sidecar
  `.trace.json` to keep the issue record readable.

## The locator → modifier two-pass (`chunk-refiner.js`)

Multi-turn refinements use two LLM passes:

1. **Locator**, given the intent + a component map of slots and their bound
   chunks, classifies the intent as `targeted` (specific slot/element named or
   verb implies a localized change) vs `untargeted` (broad: "more compact",
   "use teal").
2. **Modifier**, emits ops from a fixed vocabulary:
   `{ rebindSlot, appendToSlot, removeFromSlot, replacePage }`, translated to
   A2UI `updateComponents` messages via `opsToA2UI()`.

Refinements operate on the chunk binding plan only (Phase A simplification);
component-tree refinement is a possible later phase. The wire format holds
either way.
