# V16.6 — Unified Adaptive Orchestration & DeepSeek Reasoning Partner

Release: **16.6.1** (DeepSeek Runtime Integrity & Economy Hardening) · supersedes 16.6.0 ·
Node >= 22.19 · Windows-safe

V16.6 replaces V16.5's "one task policy, one advisor question, one progress line" with
**one budget per run** and a DeepSeek session that behaves like a reasoning partner instead of
a single-shot consultant.

Nothing about correctness moved. Every V16.3/V16.4/V16.5 safety bound is still in force, the
local verifier is still the only thing that can produce PASS, and a false PASS is still zero.

---

## 1. What is new

| Area | V16.5 | V16.6 |
| --- | --- | --- |
| Spend decisions | task-policy profile picked each number | `lib/orchestration-budget-v16-6.mjs` decides **all** numbers once per run |
| DeepSeek turns | `maxConsultations = 1` | canonical **turn budget**: 0 / 2 / 4 / 6 by complexity, lane-clamped |
| DeepSeek modes | on/off/auto | `UES_REASONING_MODE` = `economy` \| `balanced` \| `deepseek-first` |
| Advisor roles | 5 question types | 9 roles, one primary thread + at most one secondary |
| Conversation | stateless submits | one conversation session per run, with rotation, capsule and cache |
| Tool descriptions | full text | `full` / `compact` / `minimal` profiles, safety-audited |
| Progress | one observer | Progress Observer V2 (`compact` default, `detailed`, `off`) |
| Telemetry | ad-hoc fields | every number carries a **provenance** label |

## 1a. What 16.6.1 fixed

16.6.0 shipped the architecture above. The 16.6.1 audit found that several of its wiring
invariants did not hold on the production path. Each row is a confirmed defect reproduced by
`test/v16-6-1-regressions.test.mjs`; every one of those tests fails on 16.6.0.

| Area | Defect on 16.6.0 | 16.6.1 |
| --- | --- | --- |
| Escalation | `refineOrchestrationBudget` changed the profile but left `deepSeekMode`, `deepSeekTurnBudget`, `deepSeekAdvisorRole`, `deepSeekPacketTier`, `parallelReasoning` and `maxDelegationDepth` at their pre-escalation values, so a real DEEP run advertised `deepSeekMode:"off"` with a zero-turn advisor budget | refinement recomputes the whole decision from the stored inputs: one input, one decision |
| Context budget | the policy applied `PROFILE_SPEND[profile].contextBudget` while telemetry reported the pressure-adjusted value (18k reported, 20k spent) | the applied policy uses the budget's own value, floored at 8,000 chars |
| Consult cache | the key hashed a `status + path` change list and a 128-character prefix of the advisor packet, so a re-edit of the same path and a new verifier failure both replayed stale advice | the key hashes the real diff, per-file content, the whole packet, the evidence fingerprint, and workspace/provider/model/role/phase |
| Session rotation | "rotation" zeroed counters on the same session object; the web conversation kept its full context while local accounting claimed a fresh one | `rotateConversationSession()` closes conversation A and opens a distinct conversation B; the browser profile is never touched |
| Evidence protocol | the adapter asked for a fixed JSON schema while the parser only read `EVIDENCE: <kind>` lines | one canonical JSON `evidenceRequests` array; the line form is a compatibility parser over the same normalized shape |
| Evidence budget | per-request char limits only, so 8 requests x 16,000 chars each passed individually and all shipped | a cumulative per-exchange and per-run character budget (24,000 / 60,000) with an honest refusal |
| Path containment | a relative evidence target was resolved against `process.cwd()`, not the workspace root | resolved against the declared root; a sibling whose name shares a prefix is rejected |
| Provider states | a missing capability state defaulted to `ready`, and `startSession` mapped every non-`needs-auth` state to `ready` | only an explicit recognized READY is READY; everything else fails closed with its own state |
| Session reuse | reuse navigated to the entry URL while its own comment claimed a read-only probe, silently discarding the conversation | reuse takes a read-only health probe; auth wall, selector drift and a lost session stay distinct |
| Resume capsule | only the rendered `content` was redacted, the structured fields kept raw secrets, and the truncation marker could exceed `maxChars` | sanitization happens before the object exists, over every field; `maxChars` is exact |
| Parallel reasoning | `Math.max(1, maxParallel - 1)` manufactured a reader lane out of a `maxParallel: 1` budget, and the check was a denylist | the DeepSeek writer occupies a lane, `maxParallel <= 1` means no overlap, and only allowlisted read-only operations may overlap |
| Session pool | `activeWriters() + 1` double-counted the lease it had just registered, granting one writer fewer than its own bound and reporting a phantom one | writers are counted once and the high-water mark only rises on a granted lease |
| Decision packet | the budget was judged against `JSON.stringify(sections).length` while the outbound payload is the rendered packet | the budget is judged on the rendered payload; essential constraints are compacted by dropping whole rows, never by clipping one mid-sentence |
| Follow-up delta | `changedSections` listed every section that differed, including the ones the budget dropped | `changedSections` is exactly what was sent, `omittedChangedSections` names the rest, and a partial delta is flagged `misleading` |
| Follow-up bounds | four different defaults (lane 1/1, adapter 4, wrapper 4/16, turn policy 3/2) | one `WEB_REASONING_BOUNDS` table; the second follow-up is gated by `lib/followup-budget.mjs` |
| Advisor learner | every consultation was recorded with `finalVerifiedResult: null`, so benefit and harm stayed at 0 forever, and `taskClass` was read off a field the task policy never sets | samples stay pending until the local verifier resolves them; benefit needs acceptance and a favorable verifier delta |
| Advice grounding | advice naming no existing file was rejected outright, discarding architecture reasoning | grounding is wider than a file (symbol, topology, runtime, verifier, constraint, dependency graph) and is always local |
| Output economy | consecutive distinct `npm warn deprecated <package>` lines were merged into one, losing distinct package names | unique deprecation evidence is preserved verbatim; `auto` is the default and refuses with an explicit reason |
| Telemetry | `deltaTokens: derived(0)` and `savedTokens: derived(chars/4)` overstated measurements | `ESTIMATED` for chars to tokens, `NOT_MEASURED` when unobservable |
| Progress observer | `secretsEmitted: measured(0)` asserted a scan that never ran | a real post-render secret scan runs; chain-of-thought is reported as a policy invariant |
| Prefix drift | the baseline key was `provider + model`, so two projects in one process compared against each other | the key is a hashed workspace identity plus provider and model |
| Lazy graph | `pi/extensions/ues.ts` statically imported the consult cache and the parallel planner while they were advertised as lazy | removed; the `DEEPSEEK_SESSION` lazy stack hydrates only on a real consult |
| Tool routing | `model` was a keyword for both code intelligence and database access | routing is evidence-first (runtime, then repo/symbol structure, verifier, phase, text); `model` alone never means a database model |
| Description learner | keyed on `model + risk`, so a profile proven for a 6-tool surface was reused for a 20-tool one | keyed on model family, risk, phase, surface fingerprint and description schema version |

---

## 2. The unified budget

`lib/orchestration-budget-v16-6.mjs` produces one record per run:

```js
{
  executionProfile: "FAST" | "BALANCED" | "DEEP",
  taskPolicyExecutionProfile: "fast" | "standard" | "deep",   // integration vocabulary
  contextBudget, skillBudget { maxSkills, capsuleChars },
  maxAdvertisedTools, toolDescriptionProfile,
  deepSeekMode, deepSeekTurnBudget, deepSeekAdvisorRole, deepSeekPacketTier,
  delegationMode, maxChildren, maxParallel,
  verificationStrategy,
  reasons: [{ signal, basis, impact }], measurements, provenance, fingerprint
}
```

### Evidence priority

`runtime evidence > repository structure > verifier evidence > task text` (cap: **+1**).

Task wording alone can never reach DEEP, no matter how dramatic it is. Every reason records
which basis produced it, so a reader can see *why* the profile moved.

### Profile spend (the only place these numbers live)

| Profile | context chars | skills | capsule chars | advertised tools | description profile | delegation | verification |
| --- | --- | --- | --- | --- | --- | --- | --- |
| FAST | 8,000 | 1 | 1,200 | 6 | minimal | parent-direct | targeted |
| BALANCED | 20,000 | 3 | 2,600 | 12 | compact | 1 child / 2 parallel | targeted + affected |
| DEEP | 48,000 | 5 | 4,200 | 20 | full | 1 child / 2 parallel | targeted + integration |

Delegation hard bounds (never relaxed): **max 3 children, max 3 parallel, depth max 2**.

### Integration rule

* Spending **less** than the task policy allowed is always allowed — that is the point.
* Spending **more** requires an evidence *floor*: `risk=high|critical`, `mode=long-horizon`,
  `evidence-score>=3`, `hard-task-cannot-be-fast`, `repeated-verifier-failure>=2`, or a proven
  single-file bound.
* An existing DEEP floor is always kept.

So V16.6 never inflates a run just because a heuristic looked at it, and it never shrinks one
that risk already escalated.

### The task-signal bridge

`lib/task-signal-bridge-v16-6.mjs` is the single place that translates the shape
`lib/task-policy.mjs` actually returns (`risk`, `score`, `signals[]`, `domains[]`,
`executionProfile`, `singleFileBounded`) into the decision vocabulary the budget and the advisor
selector consume (`changeKind`, `taskClass`, `intent`, `deterministic`, `risk`).

It exists because the two vocabularies did not match, and three real defects followed:

1. **Risk precedence inversion.** The budget read `taskPolicy.risk || input.risk`. Because
   `classifyEngineeringTask` always returns a `risk` *string* (defaulting to `"low"`), an explicit
   caller escalation of `risk: "high"` was silently discarded — high-risk analysis could never
   escalate. The bridge now takes the **maximum** severity of the caller signal and the
   repository-detected domain, in both directions: a caller can escalate, and can never downgrade.
2. **Mode precedence inversion.** The same shape discarded a declared `mode: "long-horizon"`.
3. **A decorative role table.** `selectAdvisorRolesV2` reads `intent`/`taskClass`, which the task
   policy never produces, so *every* task fell through to `fallback:root-cause`. The nine-role
   table never chose a role in production.

Determinism is asserted **positively** (a `docs`/`version`/`config`/`test` change kind with no
risk signal). An earlier version treated "no risk signal + task-policy score 0" as deterministic,
which is what the classifier returns for any short, keyword-free sentence — it downgraded real
bugs and UI tasks to FAST with zero DeepSeek turns. Absence of evidence is not evidence.

A bounded, closed-list **deterministic-shape detector** (version bump, typo, docs/comment edit)
supplies the change kind when the caller does not. It is deliberately asymmetric: it can only
ever contribute a *spending ceiling* (FAST, no DeepSeek), never an escalation, never an authority.
That is the only keyword-shaped component in the decision path.

`nextjs` is deliberately **not** mapped to the `ui` change kind: Next.js is full-stack, and
"route handler returns 500" is a backend bug hunt that was being routed to the UI advisor.

### Determinism

`fingerprint` is a SHA-256 over the decision fields only — no timestamps — so the same inputs
always produce the same budget, and two different decisions can never share a fingerprint.

---

## 3. DeepSeek as a reasoning partner

DeepSeek remains a **consultant**. It has no filesystem, git, terminal or permission authority,
it never decides PASS, and it never sees secrets. Its output is untrusted external evidence
that is recorded, bounded, and handed to the executor.

`lib/deepseek-turn-policy-v16-6.mjs` is the single turn policy:

| Complexity | economy | balanced | deepseek-first |
| --- | --- | --- | --- |
| easy | 0 | 0 | 0 |
| normal | 0 | 2 | 3 |
| hard | 2 | 4 | 4 |
| very-hard | 3 | 6 → 5 (clamped) | 6 → 5 (clamped) |

`economy` never spends more turns than `balanced`. The lane clamps consultations to 3 and
follow-ups to 2 (`LANE_SAFETY_BOUNDS`), so the effective operational maximum is 5 turns.
`UES_WEB_REASONING_MODE=off` yields 0 turns; `force` yields a floor of 1.

`UES_REASONING_MODE` controls **degree of participation only**. It can never enable the browser
lane and can never bypass `lib/web-reasoning-escalation.mjs`, which stays the only component
allowed to decide "ask DeepSeek". An invalid value falls back to `balanced` with
`normalized: false` — it never silently becomes `off`.

---

## 4. Session intelligence

Browser/profile session and conversation session are separate concerns.

* `lib/deepseek-session-budget.mjs` — one conversation per run: turns, chars, errors, idle
  expiry, rotation, telemetry.
* `lib/deepseek-session-pool.mjs` — bounded leases: **one writer**, read-only overlap allowed,
  hard concurrency 3.
* `lib/deepseek-resume-capsule.mjs` — bounded (≤ 12,000 chars), deterministic, secret-scanned
  continuity for the next delta.
* `lib/deepseek-consult-cache.mjs` — bounded in-memory LRU (64 / hard 128, 30 min TTL) keyed on
  HEAD + diff + evidence + role + phase + mode + question + constraints.
* `lib/deepseek-evidence-requests.mjs` — 8 allowlisted evidence kinds, ≤ 4 per exchange and
  ≤ 8 per run, workspace-contained, `.env` / key material / `.git/` / `node_modules/` / traversal
  denied, redact → bound → re-scan before anything can leave.

There is **no second persistent store**: the cache is in-memory and bounded, and the Evidence
Store remains the only durable artifact store.

### Context window

Never hard-coded 64K. `resolveContextWindowTokens()` resolves detected → override →
conservative fallback (32,768, labeled `ESTIMATED`), clamped to [4,096, 262,144].
`UES_DEEPSEEK_CONTEXT_TOKENS=auto` is the default.

---

## 5. Tool surface

* **Advertised tools** are capped per profile (6 / 12 / 20). Deferred tools stay reachable
  through the existing `ues_tool_search` hydration dispatcher, and every safety capability stays
  runtime-enforced — the cap narrows the *advertised* list, never the capability set.
* **Description profiles** (`full` / `compact` / `minimal`) shrink text only. A compression that
  would drop a protective, permission, containment, failure or re-read line is rejected and the
  original description is restored (`fellBack: true`). The `ues_tool_search` dispatcher is never
  compressed. Invalid env values resolve to `full` with `normalized: false`.
* **Stable-prefix guard** (`CACHE` / `BALANCED` / `TOKEN`) is report-only: it never reorders the
  surface, and an env budget may only tighten.
* **Tool-output economy** collapses only explicit noise families (dependency-install runs,
  progress frames, repeated identical lines) and is **off by default**
  (`UES_TOOL_OUTPUT_ECONOMY=off`). `assertNoLossyTransform()` is the executable proof that
  preserved evidence (errors, `path:line`, diffs, exit codes, security warnings) survives; when
  nothing qualifies the output is byte-identical.

---

## 6. Progress Observer V2

Header format: `UES 16.6 · <REASONING MODE> · <PROFILE>`, e.g.
`UES 16.6 · DEEPSEEK-FIRST · BALANCED`.

Modes: `compact` (default), `detailed`, `off`. Notes are redacted and bounded to 160 chars,
lanes are capped at 8, and telemetry reports `chainOfThoughtEmitted: 0` and
`secretsEmitted: 0` as measured facts.

---

## 7. Telemetry and provenance

`lib/measurement-provenance.mjs` defines the only labels allowed:

`MEASURED` · `DERIVED` · `ESTIMATED` · `NOT_MEASURED`

A number without a label is invalid. Characters→tokens is always `ESTIMATED`. Anything not
observed is `NOT_MEASURED`, never a guess.

`lib/run-telemetry.mjs` gains an additive `v16_6` block (the telemetry `schemaVersion` stays 2)
carrying the budget, the turn budget, the DeepSeek counters, the observer summary, and explicit
`NOT_MEASURED` token/latency savings until a real A/B run measures them.

---

## 8. Evaluation

`npm run eval:v16.6` runs the deterministic 20-scenario corpus A/B against V16.5 behavior and
asserts ten invariants. `test/v16-6-1-regressions.test.mjs` adds 47 production-wiring
regressions for the defects listed in section 1a; every one of them fails on 16.6.0. Six
invariants below were added after measurement showed the original set could not fail:

| Invariant | Why it exists |
| --- | --- |
| `fingerprintsDeterministic` | the same input must produce the same decision |
| `turnBudgetWithinPolicyCeiling` | turns never exceed the frozen cap of 6 |
| `laneFollowUpsBounded` | consultations ≤ 3 and follow-ups ≤ 2 (the V16.3 lane hard max) |
| `packetNeverLarger` | a V16.6 packet is never bigger than V16.5's |
| `noEscalationWithoutEvidence` | every DEEP profile has a recorded evidence floor |
| `noDeepSeekOnDeterministicWork` | deterministic work spends **zero** DeepSeek turns |
| `riskEscalationReachesDeep` | a high-risk task actually reaches DEEP (catches **under**-escalation) |
| `advisorRolesDiscriminate` | ≥ 3 distinct advisor roles across the corpus (catches a decorative role table) |
| `noAdvisorRoleWhenDeepSeekOff` | an off run reports `none`, not a specialist it never asks |
| `delegationWithinHardMax` | children and parallelism never exceed 3 |

> The first implementation shipped `laneFollowUpsBounded` as `(x <= 1 ? x : x) <= 5` - identical
> ternary branches, which can never fail and reported a vacuous PASS. `noEscalationWithoutEvidence`
> returned `true` for every non-DEEP scenario, so it could not detect an under-escalation, which is
> the more dangerous direction. Both are replaced above, and
> `test/v16-6-measured-regressions.test.mjs` (test `M1`) proves each predicate rejects a
> deliberately broken row.

Current measured deltas (deterministic fixture run, **not** a live-model result):

| Metric | V16.5 | V16.6 | Note |
| --- | --- | --- | --- |
| skill capsule chars | 40,000 | 44,800 | larger on purpose: DEEP evidence tasks |
| advertised tools | 158 | 216 | profile-driven, capped at 6/12/20 |
| context chars | 248,000 | 424,000 | same trade: evidence-driven DEEP budget |
| DeepSeek turn budget (corpus total) | 20 | 45 | V16.5 was 1 turn per task |
| advisor packet chars | 12,476 | 12,476 | unchanged |
| DeepSeek turns on FAST scenarios | 1 | 0 | the release's main efficiency claim |

V16.6 deliberately spends **more** context and skill characters on evidence-heavy scenarios than
V16.5 did. The efficiency claim is not "smaller budgets everywhere"; it is that a trivial task
spends nothing, a normal task spends ~2 bounded turns, and a DEEP task spends more than V16.5
was even capable of requesting. Context and tool totals rising while FAST work costs zero is the
intended shape, and the corpus is the evidence for it.

### Tool-description economy (measured, not claimed)

Measured against the real Pi tool surface (5 UES tools, 1,383 chars):

| Profile | Chars | Saved | Protective units dropped |
| --- | --- | --- | --- |
| `full` | 1,383 | baseline | 0 |
| `compact` | 1,166 | 16% | 0 |
| `minimal` | 552 | 60% | 0 |

**The 40-70% design target is met by `minimal`, not by `compact`.** `compact` measures **16%**
on this surface. It is reported as measured rather than restated as 40%. Tool-selection quality
under either profile remains **NOT_MEASURED**: these are byte measurements, not evidence that a
model still picks the right tool. 16.6.1 additionally scoped the description learner to the tool
surface it was measured against, so a ratio proven for one surface can no longer be transferred
to another.

The first implementation measured **0% on `compact` and 1% on `minimal`** against this same
surface, because every UES tool description is a single prose paragraph and the compressor only
split on line boundaries. Two defects were fixed: prose descriptions are now segmented into
sentences, and the safety classification for structural lines uses the broad protective pattern
rather than the narrow sentence set (a bullet like `- Writes are workspace-contained and never
touch .env files.` was previously classed as ordinary `bullet` prose and silently dropped by
`minimal`).

### Not measured

Provider tokens, wall-clock latency, cost, model quality and tool-selection quality are
`NOT_MEASURED`. This corpus runs no model and no provider. In particular:

- **Tool-selection quality under `compact`/`minimal` is unverified.** The compression ratios above
  are byte measurements; whether a model still selects the right tool at a 60% shorter
  description is not established here. The description-profile learner only promotes a profile
  after ≥ 8 samples with a healthy verified-pass rate and a low selection-error rate, and
  widens back to `full` on any observed drop.
- **Latency overlap savings are unmeasured.** `parallel-reasoning-v16-6.mjs` measures overlap
  windows in ms but reports `overlapWallMs` / `serialBaselineMs` / `wallClockSaved` as
  `NOT_MEASURED`, because no serial baseline was captured.
- **DeepSeek Web token counts are estimates.** DeepSeek's UI exposes no token counter, so every
  DeepSeek token figure is `ESTIMATED` from measured characters and never presented as exact.
- **Provider token counters are unavailable.** `providerEvidenceTokens` and `providerTokensSaved`
  are `NOT_MEASURED`, never a derived zero.
- **Tool-output economy savings are byte measurements.** `savedChars` is `DERIVED` from measured
  lengths and `estimatedSavedTokens` is `ESTIMATED` from `chars / 4`.
- **The evidence budget's refusal is bounded, not strategic.** When the cumulative character
  budget is exhausted the run continues with local reasoning; no claim is made that the omitted
  evidence would have changed the outcome.
- **Wall-clock savings from parallel reasoning are still unmeasured.** 16.6.1 tightened the rule
  (the DeepSeek writer counts against the lane budget, unknown operations are refused) but
  captured no serial baseline, so `wallClockSaved` stays `NOT_MEASURED`.

---

## 9. Environment

| Variable | Values | Default |
| --- | --- | --- |
| `UES_REASONING_MODE` | `economy` \| `balanced` \| `deepseek-first` | `balanced` |
| `UES_TOOL_DESCRIPTION_PROFILE` | `full` \| `compact` \| `minimal` \| `auto` | `full` |
| `UES_TOOL_OUTPUT_ECONOMY` | `off` \| `auto` \| `on` | `auto` |
| `UES_PROGRESS_OBSERVER_V2` | `compact` \| `detailed` \| `off` | `compact` |
| `UES_DEEPSEEK_CONTEXT_TOKENS` | `auto` or a token count | `auto` |
| `UES_PREFIX_DRIFT_BUDGET` | 0..2 (tighten only) | per mode |

---

## 10. What did not change

Correctness gates, verifier strictness, permission lattice, workspace containment, Evidence
Store integrity, process cleanup, browser safety, Windows support and LSP fail-closed behaviour
are untouched. `lib/web-reasoning-escalation.mjs` is still the only component that may decide to
ask DeepSeek. V16.5's observer module is kept intact next to Observer V2.