# V16.10 — Context Intelligence & Measured Economy

Release: **16.10.0** · supersedes 16.9.0 · Node >= 22.19 · Windows-safe

V16.10 does not add a new executor, a swarm, a vector database, or an always-on browser. It
adds **six single-owner capabilities** that each make a run cheaper or safer *without* weakening
correctness. The local verifier is still the only thing that can produce PASS, DeepSeek is still
an untrusted stateful advisor, and Pi is still the only executor.

Every capability is bound by the same laws:

- **ONE BEHAVIOR, ONE OWNER** — no module re-implements a behavior another module owns.
- **Deterministic first, model last** — a model summarizer is consulted only after the
  deterministic ladder had its chance, and only if a caller explicitly supplies one.
- **Honest provenance** — a value is `MEASURED`, `DERIVED`, `ESTIMATED` or `NOT_MEASURED`, and
  nothing is ever promoted from an estimate to a measurement.

---

## 1. The six capabilities

| Capability | Module | Owns | Never does |
| --- | --- | --- | --- |
| A. Context Kernel V2 | `lib/context-kernel-v16-10.mjs` | the context **allocation decision** | call a model on its own, drop a pinned segment |
| B. Tool Output Budgeter | `lib/tool-output-budgeter-v16-10.mjs` | tool-output **shaping** | own a store, do IO, call a model |
| C. Repo Intelligence V2 | `lib/repo-intelligence-v16-10.mjs` + `lib/repo-intelligence-cache-v16-10.mjs` | one **composed** repo brief + its cache | recompute the graph/index/map it composes |
| D. Semantic Tool Router | `lib/semantic-tool-router-v16-10.mjs` | an **advisory** ordering over the caller's universe | widen the surface, advertise a denied tool |
| E. Verification Ladder | `lib/verification-ladder-v16-10.mjs` | the **cheapest-sufficient rung** order | run tests itself, promote weak evidence to PASS |
| F. Metrics V2 | `lib/efficiency-metrics-v16-10.mjs` | **aggregation** of capability receipts | write a ledger, run a task, infer quality |

## 2. Context Kernel V2

The kernel is a pure planner (`planContextKernel`) plus a small async applier
(`applyContextKernel`). It orchestrates the existing owners — `seen-context-ledger`,
`context-pruning`, `reversible-context` — and owns only the allocation and its ledger.

Segments carry a tier: `pinned` → `task` → `evidence` → `memory` → `history`. Each tier has a
budget share, and a tier may borrow unused budget from later tiers but **never** from `pinned`.
If pinned content alone exceeds the budget the plan reports `overBudget: true` rather than
silently dropping a system rule.

Deterministic compaction is a strict ladder: collapse whitespace → drop duplicate consecutive
lines → structural head/tail. **A candidate that is not strictly smaller is refused**
(`applied: false`, reason `compaction-not-beneficial`). A no-op is always preferred to a
regression. The ordering law is observable as `orderingLaw: "deterministic-first-llm-last"`.

## 3. Tool Output Budgeter

`shapeToolOutput(toolName, text, options)` picks a strategy deterministically — `read-file`,
`search`, `test`, `diff` or `generic` — and shapes accordingly:

- a read keeps head **and** tail (the API lives at the top, the closing logic at the bottom);
- a search routes rows and never dumps every match body;
- a passing suite collapses to a count summary;
- a failing suite exposes the failing command, failed names, first assertion and bounded stack
  frames instead of thousands of passing lines;
- a diff shows the changed files and hunks.

Every shaped output carries a receipt with `originalChars`, `visibleChars`, `omittedChars`,
`noticeOverheadChars`, `truncated`, `retrievalHandle` and `expansionAvailable`. When a body is
truncated the model-visible text always names original/visible/omitted and the handle — there is
no silent cut. Token counts are `NOT_MEASURED`; only a provider response can produce a measured
token count.

`lib/tool-output-governor.mjs` remains the single governor; it delegates shaping to this module
and preserves the exact raw bytes in the evidence store before shaping.

## 4. Repo Intelligence V2

One orchestrator composes `repo-graph`, `semantic-index`, `repo-map` and `affected-tests` through
`getOrComputeRepoIntel`, a fingerprint-scoped persistent cache (`.ues-cache/repo-intel-v16-10`).

Cache properties:

- **atomic writes** (temp file + rename, with a Windows retry on transient locks);
- **in-flight coalescing** per key, so two concurrent callers share one computation;
- **bounded** by both entry count and total bytes, with LRU eviction;
- **schema-version validated**; a corrupted entry degrades to a MISS, never to a wrong answer;
- TTL expiry, and an L1 memory layer in front of the L2 disk layer.

## 5. Semantic Tool Router

The router classifies a task into deterministic intents (`INSPECT_FILE`, `LOCATE_CODE`,
`EDIT_CODE`, `CREATE_FILE`, `RUN_VERIFY`, `FETCH_EVIDENCE`, `MANAGE_SERVICE`, `BROWSE_WEB`,
`LIST_STRUCTURE`) and ranks the caller's `universe` against them. "Semantic" here means
capability + purpose keyword + synonym matching — **never** an embedding, a model, or a network
call.

Invariants, each provable by `assertRouteRespectsDenied`:

- the router can only order tools the caller already owns (`widened: false`);
- a denied tool never appears;
- a writer-only tool is gated on `writer === true`;
- the discovery dispatcher (`ues_tool_search`) is capped below any concrete match.

The router is **advisory**: `compileToolSurface` still decides the advertised set. The router
only feeds an ordering into `coreToolPriorities` via `mergeRouteIntoPriorities`.

## 6. Verification Ladder

The ladder orders rungs cheapest-first: `reuse` → `static` → `affected-tests` → `full-suite` →
`independent-verifier`. Each rung proves a strength: `syntax` → `unit` → `behavior` →
`independent`. A claim states the minimum strength it requires; the ladder picks the cheapest
rung that meets or exceeds it.

Two laws make the ladder trustworthy:

1. **A rung that passes but is too weak does not satisfy the claim.** It is recorded as
   `passed-insufficient` and the ladder keeps escalating. It never stops early on evidence that
   does not actually prove the claim.
2. **Absence of evidence is `UNVERIFIED`, never `PASS`.** Only a rung that actually ran and
   failed yields `FAIL`.

The ladder never spawns a process itself; it orchestrates `findReusableVerification`,
`collectFastStaticEvidence` and `resolveAffectedTests`, and a caller injects the suite runner.

## 7. Metrics V2

Metrics V2 is the single aggregation owner for the capability receipts. It:

- sums only over rows that actually **measured** a field, returning the honest triple
  `(value, measuredRows, unmeasuredRows)` so "no data" is never mistaken for "zero";
- refuses a ratio unless its denominator is positive **and** fully measured;
- labels char→token savings `ESTIMATED` and keeps headline token savings explicitly
  `NOT_MEASURED` with a reason;
- counts success only from a PASS verdict;
- is pure-read — it never writes a ledger and never runs a task.

It reports `qualityClaim: "NOT_INFERRED_FROM_EFFICIENCY"`: efficiency numbers never imply a
correctness claim.

## 8. Wiring

`pi/extensions/ues.ts`:

- builds the runtime context pack through Context Kernel V2 (`compactContextPack` delegates the
  budget decision to `planContextKernel` + `compactDeterministically`, with a proven per-section
  cap as the fail-safe);
- orders candidate tools through the Semantic Tool Router (`mergeRouteIntoPriorities`,
  ordering-only) at both tool-surface call sites;
- surfaces Metrics V2 in the `/ues-status` digest;
- exposes two new read-only `ues_code` actions: `repo-intelligence` and `verification-plan`.

`pi/extensions/ues-child-runtime.ts` governs built-in tool output through
`lib/tool-output-governor.mjs`, which now shapes through the Tool Output Budgeter.

## 9. Measurement

`npm run bench:v16.10` drives each capability owner directly with a deterministic fixture. On a
175 KB repetitive test log:

| Capability | Before | After | Provenance |
| --- | --- | --- | --- |
| Tool Output Budgeter | 175,013 chars | 2,146 chars | MEASURED |
| Context Kernel V2 | 175,013 chars | 4,100 chars | MEASURED |
| Provider tokens | — | — | NOT_MEASURED |

Wall-clock, char counts and call counts are `MEASURED`. Provider token usage is
`NOT_MEASURED` and nothing is extrapolated into a total task speedup.

## 10. Deferred to V16.11

Browser Transport V2, the MutationObserver lane, and the `advisor-session-manager` ownership
migration are explicitly deferred. They are out of scope for 16.10.0.
