# v1.11.2 — The Numbers Audit (RELEASE_PROTOCOL §4.8)

Six questions, in writing, for every number this release adds or changes.

This release adds **no new metric**. What it does is rarer and worse: it changes the
**population** of two existing ones. `calibration` (Brier, bias) and `overconfident_lapses` were
being computed over a set that silently lost rows — and the rows it lost were the ones that made
the learner look worst. So the figure under audit is not a formula, it is a *denominator that had
a hole in it*, which is the exact shape §4.8 was written for.

| figure | where | what changed |
|---|---|---|
| `calibration.n` / `brier` / `bias` | `stats`, `report`, `/coach` | **population**, not formula — rows that were dropped now arrive |
| `overconfident_lapses` | `stats`, `report` | same population fix |
| `PRODUCTION_MAX` = 2400 | engine constant, both assessor specs | 800 → 2400 |
| `307/307` | README badge, README CLI table, INSTALL-CODEX, INSTALL-PI | selftest count, 302 → 307 |
| `1.11.2` | five version locations + badge | the release version |
| `0 crashes / 600 states / 21,000 calls` | CHANGELOG | fuzz gate result |

## Q1 · Cross-consistent with every other number?

The population fix is the *source* of a cross-inconsistency, not a new one. Before it, two
engine-computed facts about the same session disagreed: the stash recorded that the learner
picked 90 and then lapsed, while `calibration` counted that session as having produced no
confidence datum at all. `stats.self_grading` and the receipt log would show the item; the
calibration `n` would not. **Nothing reconciled those, and nothing would have.** After the fix
both read from the same recorded pick.

Checked side by side rather than reasoned about: a settle whose assessor omits `confidence`
now yields a receipt whose `confidence` equals the stashed pick (selftest
*"the learner's confidence pick survives an assessor that drops or invents one"*), and the
mirror — a declined pick stays `null` even when the assessor supplies a number.

Selftest re-run on the release branch after the last edit: `307/307`, equal to the badge.
Fuzz re-run after the last commit: `0 / 600`.

## Q2 · Which direction does each fail in?

**This is the important row, and it is the reason the bug survived.** The dropped rows were
high-confidence **lapses** — a `(0.9, 0)` pair is the single worst prediction a learner can make,
and deleting it *lowers* the Brier score and *lowers* `overconfident_lapses`. The old behavior
therefore failed in the **flattering** direction, silently, on exactly the learner the metric
exists to warn. That is bug class #1, and it was produced by bug class #5 (evidence dropped
without saying so).

The fix moves both numbers **against** the learner's vanity, which is the correct direction and
the one nobody complains about being wrong. `PRODUCTION_MAX` fails *inconvenient* (a longer
assessor prompt) rather than flattering — and its old value failed flattering in the sharpest
possible way, by making a grade look like a verdict on the answer when it was a verdict on a
truncation. `307/307` and `1.11.2` are pinned mechanically (badge equality; a selftest compares
`ENGRAM_VERSION` to plugin.json).

## Q3 · Denominator, and does it say so?

`calibration.n` **is** the denominator under repair: it counts `(confidence, outcome)` pairs
where both are non-null, and it is printed alongside the Brier score, so a reader can see the
population size. The defect was never that the denominator was unlabelled — it was that rows
left the numerator and denominator *together*, which is invisible in a labelled ratio. Nothing
in the output distinguished "the learner declined to state confidence" from "the grader forgot
to type it back." **After this release the engine can still not tell those apart at read
time — but it no longer creates the second case**, because the pick comes from its own record.
The honest residue: receipts written before v1.11.2 that lost a pick are unrecoverable, the
stash entries behind them are long gone, and this release does not backfill them. Stated here
rather than quietly fixed forward.

`CAL_MIN_N = 10` still gates the verdict to `insufficient-data` below ten pairs; unchanged.

## Q4 · Does anything READ it — and does EVERY SURFACE read it?

Followed to a string a human sees, not just to a field:

- `calibration` → `compute_stats` → `stats` JSON → `/coach` prose → `report`'s dashboard HTML.
- `overconfident_lapses` → same path.

Both surfaces read the same computed dict, so neither can drift from the other. No new field was
added, so there is no new place for a value to be computed and then dropped by one renderer —
the v0.7 failure where the teeth reached the CLI and not the dashboard is not reachable here.

`production_truncated` **is** a field whose reader matters, and it is honest about its limits: it
rides the receipt and is visible in the receipt log and to an appeal, and it is *not* surfaced in
`stats` or the dashboard — it is a per-item provenance mark, not a metric. It is deliberately not
in `EXPORT_RECEIPT_KEYS`; nothing about a truncation leaves the machine.

## Q5 · Can it be reached from the CLI in a way the skills never take?

Yes, and both paths were checked. `rate --production-file` mints a receipt with **no `sid` and no
stash entry** — the skills' `/review` path. There `_stashed_item` returns `None` and every value
falls back to the item, which is correct: on that path the caller *is* the recorder. The
CLI-only hazard is a hand-edited `pending-verify.jsonl`, which is now fuzzed and gated
(string-or-nothing), after it was found tearing a settle mid-batch.

`stash add` remains the only writer of the truncation flag, and `receipt` the only reader.

## Q5.5 · If the number is an INSTRUMENT, does a wrong subject score worse?

`calibration` is an instrument pointed at the learner, so the test is whether a *worse-calibrated*
learner scores worse. **Run, not reasoned about.** Identical input to both engines: 16 reviews,
every one at confidence 90, the learner lapsed 5 of them; the grader omits `confidence` on the
lapses. Read off the rendered dashboard, which is the surface a human actually looks at:

```
v1.11.1   Brier 0.010 · bias -0.100 → underconfident · n=11
v1.11.2   Brier 0.260 · bias +0.212 → overconfident  · n=16
```

**The verdict was inverted, not merely understated.** A learner who confidently failed five
reviews was told they were *under*confident — "you know more than you think", the single most
flattering thing this engine can say, and the exact opposite of the truth. Brier was understated
26×, and the bias **flipped sign**, because every dropped row was a `(0.9, 0)` pair — the
maximally-wrong prediction, and the only kind of row whose removal can reverse the sign.

The caption printed beside those numbers is *"(only answers where you actually stated a
confidence count)"*. The learner stated a confidence on all sixteen. So the sentence containing
the wrong number also contained a false account of why it was trustworthy: bug class #7 wrapped
around bug class #1, produced by bug class #5. Three of the seven, in one line of the dashboard.

The instrument therefore has teeth in the required direction, and this is not a hypothetical
severity: it is what `main` renders today.

### Is it live, or only reachable?

Live, at a low rate, measured on the author's real store rather than a fixture: of **29**
assessor-graded receipts, **1** carries a null confidence — `transformers-ffn /
nexttoken-sufficient`, graded **`lapsed`**. That is the predicted class, hit on the first
instance. Whether the learner declined to state a confidence there or the grader dropped it is
**unrecoverable** — the stash entry is long gone — and that indistinguishability is precisely the
defect: the old design gave those two very different facts the same representation. At this
store's rate the aggregate barely moves (its calibration reads `overconfident` before and after,
so no verdict is inverted for this learner); the inversion above is what the same mechanism does
at a higher drop rate. Both statements are true and neither is the other's excuse.

## Q6 · Does its LABEL survive contact with a reader?

`307/307` is labelled "checks", never "everything works" — the denominator is the suite's own
count, printed by the run.

The label that needed the most care is **`production` on a receipt**. Before this release it was
labelled as the learner's production and was in fact *the assessor's ≤600-char retyping of it* —
a field whose name was true and whose contents were not, which is bug class #7 exactly. It is now
what the label says. `production_truncated` likewise means "the text on this receipt is not the
whole answer", and after this release it can no longer be absent from a receipt that was clipped,
which is the only way that label could have lied.

One label deliberately left alone: the assessor spec still instructs a verbatim echo even though
the engine now overrides it on the stash path. Telling a grader "we will fix this for you"
invites the sloppiness the override exists to absorb, and the echo is still the value used on the
sid-less path.

---

**Verdict:** the two population numbers move, they move against the learner, and they move for a
reason that is written down and tested. No number in this release moves in the flattering
direction.
