# V16.4 — Measured Adaptive Reasoning Runtime

## Goal

Stronger on hard tasks, faster on easy ones, less wasted tokens/context/tool
calls, DeepSeek Web consulted exactly when measured evidence justifies it.
No verifier weakening, no authority change, no heavy orchestration.

## What changed

- **Lazy runtime hydration** (`lib/lazy-runtime.mjs`). Core boot (safety, task
  policy, minimum router, runtime config, verification contracts) stays eager.
  The browser lane, the managed browser-worker client, the DeepSeek adapter, the
  web-reasoning lane, the code-intelligence surface and repo-map hydrate once
  per process via a cached dynamic import: concurrent callers join one promise,
  transient failure never poisons the cache, and every hydration reports
  `lazyModulesLoaded / Hits / Misses / JoinCount / LatencyMs / Failures`.
  These seven modules are no longer reachable from any static import in
  `pi/extensions/ues.ts`; the measured static boot graph drops from 127 to 106
  files and from 1,576,012 to 1,005,013 bytes. The 106 figure includes two
  small eager measurement modules (`lib/verified-task-cost.mjs`,
  `lib/repo-map-measurements.mjs`, 111 lines together, no heavy dependencies)
  that the controller's finalize path reports on every run.
  The LSP pool/provider subtree stays EAGER on purpose: verifier policy
  (`lib/fast-static-verification.mjs`) statically imports it, and a verifier
  that hydrates after the tool would already have executed.
- **Structural escalation V2** (`lib/web-reasoning-structural.mjs`, called by
  `lib/web-reasoning-escalation.mjs`). Priority: runtime evidence > repository
  structure > verifier history > task text. Vietnamese signals are
  fallback-only and never override grounded structural evidence. AUTO still
  skips version bumps, doc edits, typos and trivial deterministic fixes;
  FORCE/OFF semantics unchanged. The shared signal vocabulary lives in
  `lib/web-reasoning-signals.mjs` so the structural module can use it without
  an import cycle back into the router.
- **Fresh-evidence follow-up** (`lib/fresh-evidence.mjs`). The follow-up path
  splits `providerSeenEvidence` from `currentRepositoryEvidence`, refreshes
  local state, and sends only changed sections with prior/current fingerprint
  provenance. Unrefreshable state fails closed to local-only.
- **Adaptive Decision Packet tiers** (`lib/decision-packet-tiers.mjs`).
  SMALL ≤10k, MEDIUM ≤20k, LARGE ≤36k chars; the 48k ceiling is now an
  emergency upper bound only. First send carries the smallest sufficient tier;
  follow-ups expand only missing sections, never the whole packet.
- **Adaptive follow-up budget** (`lib/followup-budget.mjs`). Canonical default
  `maxFollowUps = 1` (hard max 2). A second follow-up requires fresh verifier
  evidence with a changed fingerprint, an unresolved first follow-up, benefit
  above cost, submit budget, and a healthy session. All external submits keep
  zero automatic retry with explicit accounting.
- **Advisor benefit learner** (`lib/advisor-benefit-learner.mjs`). Bounded
  observational learner (sample floor 8, hysteresis, 500-key GC). It only
  nudges the AUTO consultation weight; it cannot touch correctness policy,
  cannot disable verifier/security gates, and cannot force external web use.
- **VerifiedTaskCost** (`lib/verified-task-cost.mjs`). Unified cost metric
  meaningful only when the local verifier PASSED, with per-component
  MEASURED / DERIVED_FROM_MEASURED / NOT_MEASURED provenance. Missing token
  fields report null/partial, never fabricated zero.
- **Real-task A/B corpus** (`evals/web-reasoning-corpus.json`). Seed set across
  bug diagnosis, multi-file fix, architecture, API/backend, frontend, monorepo,
  refactor, security, ambiguous root cause, failing tests and trivial-local.
  Reuses the existing `--executor=pi` receipt import in
  `scripts/bench-web-reasoning-ab.mjs`; deterministic fixtures stay the
  regression gate. Promotion requires false PASS = 0 and measured token data
  for any efficiency claim.
- **Parallel read-only consult prep** (`lib/consult-prep.mjs`). Lane A
  (capability/session readiness) and Lane B (retrieval/packet/evidence/
  redaction/fingerprint) join before any submit. Fill/click/Send are never
  parallelized. AUTO lane failure falls back local; FORCE reports
  `WEB_REASONING_UNAVAILABLE`.
- **Release test coordinator** (`scripts/release-test-coordinator.mjs`,
  `npm run release:coordinator`). Builds the UNION of all coordinated eval
  file lists, runs each file once, saves per-file receipts, and derives every
  eval verdict. Focused eval commands are unchanged standalone; uncertain
  mappings fail closed to the legacy chain.
- **Repo-map measurements** (`lib/repo-map-measurements.mjs`). Top-K recall,
  edit-site recall, first-correct-file rank, bytes before first correct edit,
  repeated-retrieval rate, query latency. Weights change only with benchmark
  proof.
- **CI alignment** (`.github/workflows/ci.yml`). Ubuntu Node 22, Ubuntu Node
  24, Windows Node 22, Windows Node 24 (Windows Node 22 is the floor check
  for the declared `node >=22.19` minimum). Publish (Node 24, trusted
  publishing) untouched.

## Invariants kept

Independent/integration/visual verifiers, Evidence Store, dirty-work and
`.env` guards, containment, destructive-command guard, ownership fencing,
process-tree cleanup, Windows-safe cleanup, LSP fail-closed diagnostics,
static completeness, browser action taxonomy, zero replay of external side
effects, DeepSeek consultant-only boundary (`canProducePass` stays false for
advice), packet secret redaction, untrusted-output boundary, zero automatic
external retry.

## Monolith note (Phase 13)

`pi/extensions/ues.ts` had 77 eager `lib/` imports before this wiring and 74
after. What remains eager is what MUST be: safety, task policy, permission
policy, MCP tool policy, execution ownership, workspace root/hygiene,
fast static verification and the LSP provider/pool it statically depends on.
Everything else on the hot-but-heavy paths now hydrates through
`lib/lazy-runtime.mjs`, and `/ues-status` reports the measured hydration set
(`details.lazyRuntime.hydrated`) so the claim is observable in a real run
rather than asserted in a fixture.

## Non-goals

No claim of model improvement from deterministic fixtures. No token-saving
number without provider token telemetry. No live DeepSeek in CI.
