# Leverage rules, when a repeated subtree earns its own corpus entry

Historical note first: the zettel **fragment mechanism is retired**
(`$fragment` refs now render as visible drift placeholders, see
`compose/strategies/zettel/composer.js`). The leverage discipline below
outlived the mechanism: it now governs whether a repeated subtree gets split
into its own `data-chunk` block vs staying inline in its parent chunk, and it
is the calibration for any future extraction tier.

## The leverage rule

**Do not extract a reusable unit unless it has leverage ≥ 3**, at least 3
consumers would use it. Two exceptions:

1. **Singleton closing a semantic gap**, a distinct domain primitive (e.g. a
   keyboard-shortcut row) justifies extraction at leverage 1 because the
   concept itself needs to be retrievable.
2. **Intra-composition multi-use**, one composition instantiating the same
   subtree N ≥ ~10 times justifies extraction; reuse is measured per instance,
   not per referencing composition. Historical evidence: `calendar-day-cell`
   used 35× inside one calendar composition lifted corpus reuse 26.4% → 33.5%
   on its own.

Sub-leverage extraction bloats the corpus, slows retrieval, and adds
maintenance without improving reuse. Common candidates that DO clear the bar:
card headers, key-value rows, icon+text rows, labeled progress bars, stat
tiles, notification rows.

## Keyword preservation (the extraction-drift lesson)

Extraction SHRINKS the parent's keyword surface: before, the parent contained
the subtree's text; after, it only references the child. Historically,
extracting a card-header from the login-form pattern dropped its retrieval
score 28 → 24 because "login / sign in / email password" tokens left with the
extract; explicitly re-adding those keywords to the parent restored 28.

**Rule: every parent that loses a subtree must be re-enriched with the
semantic tokens that left.** In today's corpus this means the parent chunk's
`data-chunk-keywords` (and re-harvest). Without it, retrieval degrades
silently.

## Threshold discipline after extraction

If a known-good intent starts scoring below a retrieval threshold after a
corpus split, DO NOT lower the threshold. In order:

1. Verify the parent's keywords were preserved (above).
2. Check the new entry's name/description doesn't cannibalize the parent's
   semantic space (retrieval collision, see
   [semantic-fail-lifting](semantic-fail-lifting.md) Strategy B).
3. Only then consider threshold work, with
   [zettel-calibration](zettel-calibration.md) history read first.
