# Code + ledger review — 2026-09-05

> **Status (same day, follow-up session):** the "Suggested order" below has been
> worked through. See the *Applied* section at the end for what shipped, what
> is config-only, what the replays showed, and what still needs a restart.

Full read of `src/` and `omp-extension/` plus an analysis of the live ledger
(`~/.auto-model-router/router.db`, 17,184 dispatches, 2026-08-21 → 2026-09-05,
$180.22). Baseline: `bun run typecheck` clean, `bun test` 513/513.

Numbers below are from the ledger unless marked *replay* (`tools/replay.ts`).
"7d" = the seven days ending 2026-09-05T14:43Z ($101.18, 9,336 dispatches).

## Where the money goes (7d)

| Slice | $ | Share | Note |
| --- | --- | --- | --- |
| Total | 101.18 | | ~$14.5/day |
| Sep 3 alone | 38.46 | 38% | key-scoped catalog was 8 models, glm absent 00:14–23:24 UTC; gemini-3.7/3.8-flash ($0.75/M) served everything |
| `hard` tier | 40.76 | 40% | heuristic 30.47 (mostly pre-v0.2.33 failed-tool rows), escalation 5.60, sticky 4.69 |
| model-switch turns (4% of dispatches) | 33.42 | 33% | cache hit 44% vs 91% on same-model turns; moderate→hard alone 25.61 |
| fresh prompt tokens | 49.40 | 56% | cache read 31.70, cache write 4.80, completion 1.99 |
| last 2 days, `hard` | 7.52 of 35.46 | 21% | escalation 4.42 now dominates; heuristic hard is down to 2.18 |

v0.2.33 (failed-tool damping) has been live since the 2026-09-03 restart
(damped rows appear from Sep 3). **v0.2.34 is not live**: both running
`omp.exe` processes started 2026-09-03T00:27–00:29Z, before the commit. All omp
windows need a restart for the circular-loop damping to take effect.

## Findings, ranked by expected $ impact

### 1. Adaptive tier floors now RELAX every tier below its configured floor

`tierPlanFor` takes `min(configured, quantile-band)`. The design comment assumes
a healthy catalog's bands sit above the configured floors; the 347-model catalog
the key admits since 2026-09-05 has a long tail of weak scored models, so the
bands sit far *below* them:

| Tier | Configured coding floor | Effective now |
| --- | --- | --- |
| simple | 40 | 26.6 |
| moderate | 60 | 45.8 |
| hard | 72 | 59.9 |

Consequences measured on Sep 5: `inclusionai/ling-3.0-flash` (coding 50.6) wins
`moderate` outright (433 dispatches, $0.27), and *replay* shows
`google/gemma-4-26b-a4b-it` (coding 39.3) taking 1,175 of 4,000 recent
dispatches at `simple`/`trivial`. ling's probe-escalation rate is 3.6% (16/447)
against glm-5.3-flash's 0.26% (22/8,346), and each escalation lands on
`claude-opus-5` via the capability floor: 7 escalated turns cost $3.02 in two
days, eleven times the cost of all 433 ling turns combined.

*Replay* of the 717 Sep-5 dispatches with `adaptiveTierFloors=false`: moderate
returns to glm-5.3-flash (+347 dispatches), direct cost +$0.28 (+28%) — before
counting the escalations replay cannot model, which cost more than that.

**Fix (code):** relax only when the configured floor leaves the tier thin. In
`buildCandidates`, count rankable models passing the configured floor on the
request axis; use the adaptive band only when that count is below a small
minimum (e.g. 3). That preserves the guardrail-narrowed case the feature was
built for and makes a wide catalog behave exactly like the config says.
Interim (config, hot-reloads): `adaptiveTierFloors: false`.

### 2. Trust prices a failure as a same-model retry; the real price is the next tier

`effectiveUsd = expectedUsd / successRate` treats a 4% failure rate as a 4%
surcharge. The observed cost of a failure is a re-dispatch of the whole prompt
on the escalation target: for ling → opus that is ~700× the turn's own cost, so
the true surcharge is ~25×, not 4%. Every model's trust sits at 0.93–0.997, so
the divisor is effectively inert and `minTrust: 0.7` never trips.

**Fix (code):** `effectiveUsd = expectedUsd + escalationRate × forecast(escalation target)`,
where the target is the next tier's current winner. This is what makes cheap
flaky models lose to cheap reliable ones. Keep the divisor only for hard errors.

### 3. Escalation re-dispatches the model that just failed (49 of 71 in 7d)

Probe escalation sets `escalateFrom` but never adds the failing slug to
`excludeSlugs` (only HTTP-error failover does). 7d: `gemini-3.8-flash moderate →
gemini-3.8-flash hard` ×11, `gemini-3.7-flash moderate → gemini-3.7-flash hard`
×12, `gemini-3.7-flash simple → gemini-3.7-flash moderate` ×8, and so on — the
same model one tier up, paying the tier premium for a provider hiccup.
Escalated attempts cost $5.50 in 7d.

**Fix (code):** on a probe escalation, push `decision.slug` onto `failedSlugs`.
For `empty_completion`, `refusal` and `upstream_error` — signals that indict
the provider rather than the tier — try a same-tier sibling first, exactly as
`MAX_SAME_TIER_FAILOVERS` already does for 5xx/429, and only then step up.

### 4. Cross-tier switches forfeit the cache with no economic check

The stay/switch arithmetic in `select.ts` step 4 runs only when the warm model
is a candidate in the *new* tier. A tier change therefore always switches
cold: 7d moderate→hard 117 switches, $25.61, 34% cache hit. 84 of those 215
uncaptured switches carried `lastToolFailed`, which v0.2.33 has since damped
(2d hard-heuristic spend is $2.18). What remains is structural:

- **Myopic stay/switch keeps expensive models warm.** After an escalation put
  `kimi-k3` in a conversation, "stay $0.0589 ≤ switch $0.1131 × 1.3" held it at
  `moderate` for 33 dispatches ($4.12, $0.125/turn) where glm would have been
  $0.003/turn warm. The comparison is correct for one turn and wrong for the
  12-turn run that followed. Amortise the switch cost over an expected horizon
  (deep loops average 25 dispatches per user turn): switch when
  `H × (stayWarm − newWarm) > switchCold − stayWarm`, with H ≈ 5–10.
  *Replay* with `switchMargin=1.0` changes 0 decisions — the margin is not the
  lever, the horizon is.
- When a tier change is driven by a single soft signal on a mechanical
  continuation, consider requiring the signal to persist for two turns before
  paying a cold switch to a 20–50× model.

### 5. Sep 3: a catalog collapse cost ~$32 in one day, silently

`no candidates in trivial (8 rejected)` on every Sep-3 row: the key-scoped
catalog had 8 models and glm was absent for 23 hours. `gemini-3.7-flash` served
833 dispatches ($19.05) and `gemini-3.8-flash` 666 ($13.19) at 10× glm's rate.
Whether that was a guardrail edit or an upstream blip, the router accepted a
352→8 shrink with no warning (`doRefresh` only rejects an *empty* payload).

**Fix (code):** log at `warn` and expose on `/health` when a refresh shrinks
the catalog by more than ~50%; optionally keep the previous snapshot for one
refresh interval before adopting the shrink, so a transient blip does not
reroute a whole day.

### 6. Client abort after the generation finished is recorded as an error

2,068 rows (12%) carry `error = "request aborted"`, and 1,842 of them have a
`finish_reason` and full usage — the upstream generation completed and the
client closed before `[DONE]` was read. On those turns `runTurn` returns from
`onUpstreamError` without saving state: `turn` is not incremented (the next row
reuses the same turn number in 1,087 of 1,107 cases), `currentSlug`,
`cacheWarmSlug`, `lastPromptTokens` and `compactionPlan` are not updated, the
agentdox transcript is skipped, and the row is excluded from latency stats.
The rate tracks the omp binary: 16–19%/day before the 2026-09-02 omp update,
2–7% since — so the client-side cause is mostly gone, but the router should
not depend on it.

**Fix (code, `turn.ts`):** in the stream `catch`, an abort with
`finishReason !== null` is a completed generation: fall through to the normal
commit path (the dead sink no-ops), record `error: null`, save state.

### 7. Compaction re-plans every turn because the budget is unreachable

`overBudget = compactedTokens > budgetTokens` is always true: post-compaction
prompts are 100–160k against a 40k budget, so `floorRatio` never rations and
a new edit is added the moment a tool result ages past `protectRecentTurns`.
7d, same-model turns: plan changed 1,031× at 79.5% cache hit and $0.0120/turn
vs 92.6% and $0.0067 when the plan held. That churn is worth ~$3–5/week.

**Fix (code):** ration new edits — only extend the plan when the compacted
prompt has grown ≥ X% (say 10%) since the last re-plan, or every K turns;
carried edits still apply verbatim in between.

### 8. Token calibration pairs the wrong bytes with the wrong tokens

`estimatePromptTokens` records *pre-compaction* bytes; `ledger.record` pairs
them with *post-compaction, post-context-block* billed tokens. With ~43k tokens
compacted per turn the learned ratio absorbs compaction: `actual / rawEstimate`
≈ 1.10 but `actual / compactedEstimate` ≈ 1.50 — the size selection actually
uses is 33% low (forecasts, `minContext`, price tiers, switch arithmetic).

Separately, `nex-agi/nex-n2-mini` reports prompt tokens **8.4×** the estimate
(same bytes glm bills at 1.1×) and has poisoned the `qwen3` family to
1.61 bytes/token; 27 catalog models carry that tokenizer and will be
over-estimated ~2×.

**Fix (code):** calibrate on dispatched bytes (`promptBytes − savedBytes +
contextBlock.length`), and reject calibration samples whose `actual/estimate`
is outside ~[0.4, 2.5].

### 9. Hold-length experiment is diluted; exploration verdicts are blind

- `breakHoldOnMechanical` (on since Aug 30) breaks the hold on 96% of turns:
  7d 492 holds broken vs 233 held; sticky rows in the 4 turns after an
  escalation total $0.97. Per-turn cost by arm (2/3/4) is 0.0103/0.0110/0.0086
  with wildly different conversation lengths — no separable signal. Close it.
- Exploration's "cheaper tier sufficed" verdict is the probe, which rejects
  ~1% of turns at every tier. Explored turns: hard→moderate 2/160 rejected,
  moderate→simple 3/855 — indistinguishable from baseline. The probe only sees
  structural failure, so exploration cannot learn quality. (It did save money:
  160 hard→moderate turns cost $1.26 against ~$24 at hard.) Either stop it, or
  replace the verdict with an outcome proxy — loop length after the turn, or
  whether the next tool call succeeded.

### 10. Smaller classifier and scoring items

- **Stale-image weight.** `W_IMAGES` fires on `hasImages` (any image in
  history); 5,603 of 7d heuristic rows carry one. 339 rows ($3.61) sit one tier
  higher only because of that +0.04. Switch to `hasNewImage`, the same
  principle `classifyTask` already applies.
- **Latency scoring is inert.** The reference wait uses 1,024 expected
  completion tokens (34s at 30 tok/s) while actual completions average ~200;
  glm's 9.1s mean TTFT and 152 turns >30s TTFT in 7d never earn a penalty, and
  `maxExpectedWaitMs: 160000` never fires. Use the ledger's measured mean
  completion tokens as the expected completion.
- **Reasoning weight** `medium: 0.07` rode on 662 dispatches ($5.80) in 7d;
  still the single largest constant offset after the loop-depth ramp.

## Code defects (not ledger-visible)

| Where | Defect | Effect |
| --- | --- | --- |
| `src/server/http.ts` `handleChatCompletions` | `acquireTurn()` succeeds, then a JSON/parse failure returns before `runTurn`, so `releaseTurn` never runs | after 24 malformed bodies the router answers 429 forever until restart |
| `src/router/escalate.ts` + `turn.ts` | `maxHoldMs` is checked only inside `observe()`, i.e. when a chunk arrives; keep-alive comments are dropped by `parseSse` | a stream that emits nothing holds the client until the upstream gives up: 47 `empty_completion` rows waited 9–150s (gemini-3.8-flash 14×, avg 40s). **Deliberately left alone**: enforcing the 8s ceiling on a timer would escalate every glm-5.3-flash turn whose first token arrives after 8s — its mean TTFT is 9.1s — so the accidental leniency is protective. A timer-driven check only makes sense with `maxHoldMs` raised to ~2× typical TTFT (30s). |
| `src/server/turn.ts` | probe escalation does not exclude the failing slug (finding 3) | 69% of escalations retry the same model |
| `src/server/turn.ts` | abort-after-finish path skips state save (finding 6) | stale hysteresis/cache/compaction state on 2–19% of turns |
| `src/router/tier-plan.ts` | `min(configured, adaptive)` with no thinness test (finding 1) | floors relaxed on a wide catalog |
| `src/tokens/estimate.ts` + `src/cost/ledger.ts` | calibration bytes/tokens mismatch, no outlier rejection (finding 8) | biased estimates, poisoned tokenizer family |
| `src/router/candidates.ts` | trust divisor (finding 2) | flaky cheap models never lose |
| `src/router/tier-plan.ts` `tierPlanFor` | memo keyed on the config object, which hot reload mutates in place | a `filters.includeFree` edit is ignored until the next catalog refresh (≤5 min); cosmetic |

## Verified fine

- Cache-breakpoint placement, compaction carry-forward and validation, cost
  breakdown arithmetic, unattributable-error trust exclusion, hot reload, the
  per-slug newest-first index. Same-model cache hit is 91–92% inside the 5-min
  warm window and 44% beyond it; the `cacheWarmTtlMs` model matches reality.
- Per-turn routing cost: `buildCandidates` runs in 20–28ms per tier on the
  347-model catalog including trust/latency queries (31ms for 31 slugs). Not a
  bottleneck; the `trustWindowDays` note about growth still applies at 75k rows.
- The LLM adjudicator remains dead (0 `llm` rows); `ambiguityThreshold: 0` is
  correct.
- `empty_completion` escalations are genuine empty streams, not the 8s hold
  ceiling firing early (latency 2–150s, `ttft` null, `finish` stop/null).

## Suggested order

1. Config now (hot-reloads, routing-neutral for the narrowed-catalog case):
   `adaptiveTierFloors: false`; restart all omp windows for v0.2.34.
2. Code: findings 3 + 6 + the `http.ts` slot leak (small, low risk).
3. Code: finding 1 (thinness-gated relaxation) so adaptive floors can go back on.
4. Code: findings 2 and 4 (escalation-aware effective cost, horizon-amortised
   switch). Validate each with `tools/replay.ts` on the post-Sep-5 population.
5. Code: findings 7 and 8; catalog shrink warning (5).
6. Close the hold-length experiment; decide what exploration should measure.

## Applied (2026-09-05, follow-up session)

Typecheck clean, 539 tests pass. Schema is now **v16**. Everything below is in
the working tree (and, for `turn.ts`/`http.ts`, already in commit `332bfbe`,
which another session swept up while committing its own batched-signals
change — the message on that commit does not mention them).

| # | Finding | What shipped | Live effect |
| --- | --- | --- | --- |
| 1 | Adaptive floors relaxed on the wide catalog | `tier-plan.ts`: a configured floor stands whenever ≥3 rankable models meet it (`MIN_FLOOR_ADMITS`); only a thin tier relaxes to the band | config `adaptiveTierFloors: false` (hot-reloaded, in effect now); delete the key once every omp window runs this build |
| 2 | Trust divisor misprices failure | `filters.escalationCostWeight` (default 0) + `Ledger.escalationCost()`: a model's measured escalation rate × the ledger's measured $/prompt-token of escalated retries | **left off.** Replay on 5,000 Sep 2–5 dispatches at weight 1: 586 decisions move off ling-3.0-flash onto solar-pro4/mistral-nemo, direct cost +1.3%; the avoided-escalation benefit is not something replay can model. With the floors fixed the expensive case (ling at moderate → opus) no longer arises, so the term is a hedge, not a fix. Enable if ling-class models keep winning `simple`. |
| 3 | Escalation re-dispatches the same slug | `turn.ts`: the failing slug joins `excludeSlugs` on every probe escalation; `empty_completion`/`refusal`/`upstream_error` try a same-tier sibling (bounded by `MAX_SAME_TIER_FAILOVERS`) before stepping up | on restart |
| 4 | Cross-tier switch / myopic stay-switch | `hysteresis.switchHorizonTurns` (default 1): `H × stayWarm` vs `switchCold + (H−1) × newWarm` | replay at 8: **0 of 5,000 decisions change** while glm is in the catalog (the one-turn rule already switches away from a dear warm model when the winner is cheap cold). It only bites when the winner is itself dear cold — the Sep 3 kimi case. Add `switchHorizonTurns: 4` to config **after** the restart (the old schema rejects unknown keys); a commented block is in place. |
| 5 | Catalog collapse was silent | `openrouter-catalog.ts`: a refresh keeping <50% of ≥20 models logs at `warn` and is exposed as `catalog.shrink` on `/health`; cleared on recovery | on restart |
| 6 | Abort-after-finish recorded as error | `turn.ts`: an abort with `finishReason` set is a settled generation — commit path, state saved, `error: null` | on restart |
| 7 | Compaction re-plans every turn | `compaction.replanGrowthRatio` (default 1) + `conversations.compaction_plan_tokens` (v16): extend a plan only once the compacted prompt has grown by the ratio; fit-to-window never rationed | add `replanGrowthRatio: 1.1` **after** the restart; commented block in place |
| 8 | Calibration bias / poisoned qwen3 | `adjustPendingEstimate()` pairs billed tokens with the dispatched bytes (post-compaction, plus context block); samples outside 1.5–8 bytes/token are rejected; **v16 deletes every `token_calibration` row** so families relearn on correctly paired data | on restart (first open runs the migration) |
| 9 | Hold-length experiment diluted | config `exploration.holdTurns.enabled: false` | hot-reloaded |
| 10 | Stale-image weight | `classify.ts`: `W_IMAGES` fires on `hasNewImage` | on restart |
| defect | `http.ts` slot leak | released on parse failure; regression test | on restart |
| reviewer | unreachable "user-visible failure" branch | `features.ts` scans the tool run behind the newest user turn, so the full failed-tool weight is reachable; damping helper deduplicated; stale header comments and README row fixed | on restart |

Not done, by choice: the timer-driven `maxHoldMs` (see the defects table — it
would mis-escalate glm), and the exploration verdict redesign (a product
decision: what should "the cheaper tier sufficed" measure?).

**Deploy order.** Restart every omp window (two `omp.exe` processes started
2026-09-03T00:27–00:29Z are still on v0.2.33 code). Then add the two
commented config keys and delete `adaptiveTierFloors: false`. Watch the first
ledger rows for `[re-plan rationed]`, `failover: … empty_completion`, and
`compaction_plan_tokens` populating; `token_calibration` will be empty until
~20 samples per family have landed.
