# V16.5 — Agent Skill & Delegation Intelligence

Architecture, contracts, bounds, and measurements for the V16.5 runtime slice.

Scope: **skill selection, tool surface, specialist delegation, handoff, and the web advisor**.
V16.5 does not rewrite the repo map, the semantic index, LSP, the browser taxonomy, the process
supervisor, the Evidence Store, or the verification architecture.

---

## 1. Architecture

```
task text
  │
  ├─ lib/skill-router.mjs      detectIntent → rankSkills → routeSkills
  │     └─ lib/skill-registry.mjs   48 machine-readable contracts (no skill bodies)
  │
  ├─ lib/skill-capsule.mjs     compileSkillCapsule → bounded capsule + provenance
  │                              expandSkillCapsule → explicit full-text escape hatch
  │
  ├─ lib/tool-surface-v3.mjs   predictCapabilities → compilePhaseToolSurface
  │                              hydrateCapability → deterministic receipt
  │     └─ lib/tool-surface-economy.mjs (V16.2) remains the final advertised-surface authority
  │
  ├─ lib/subagent-fabric.mjs   decideDelegation → planDelegation → session/child lifecycle
  │     └─ lib/delegation-safety.mjs  classifyScope → assessParallelSafety → buildDelegationWaves
  │     └─ lib/delegation-fleet.mjs   runDelegationWave (bounded concurrent wave EXECUTOR)
  │           ├─ reuses lib/task-graph.mjs computeSafeWaves for wave ORDER
  │           └─ reuses the existing Pi child spawn + lib/process-supervisor.mjs for PROCESSES
  │
  ├─ lib/verified-handoff.mjs  createHandoffCapsule → Evidence Store raw + bounded capsule
  │
  ├─ lib/deepseek-advisor-roles.mjs  specialist question types + bounded packets
  │     └─ lib/web-reasoning-lane.mjs (V16.3/16.4) remains the transport
  │
  ├─ lib/advisor-benefit-learner-v2.mjs  bounded AUTO consult weight
  ├─ lib/reasoning-doctor.mjs   read-only readiness (`ues doctor --reasoning`)
  └─ lib/agent-progress-observer.mjs   observer-only fleet view
```

Single production importer: `lib/v16-5-runtime.mjs`, imported by `pi/extensions/ues.ts`.

---

## 2. Skill Registry

`lib/skill-registry.mjs` compiles one contract per skill from a single bounded spec table plus
the `SKILL.md` frontmatter. No skill body is duplicated and none is needed to route.

```jsonc
{
  "id": "bug-diagnosis",
  "intents": ["ambiguous-failure", "regression", "root-cause", "test-failure"],
  "taskClasses": ["bugfix", "verification"],
  "requiredCapabilities": ["read-file", "run-shell", "search-text"],
  "optionalCapabilities": ["code-intelligence", "deepseek-advisor", "diagnostics"],
  "requiredTools": ["bash", "grep", "read"],
  "optionalTools": ["ues_code", "ues_web_reasoning"],
  "forbiddenActions": ["deploy", "force-push", "publish"],
  "contextClass": "procedure",
  "sideEffectClass": "none",
  "outputContract": "bug-diagnosis-report-v1",
  "verificationRequirements": { "required": false, "independent": false },
  "composableWith": ["change-impact-analysis", "code-review", "implementation-engineer"],
  "hostSupport": ["pi"],
  "version": 2
}
```

Bounds and invariants:

- `MAX_CONTRACT_CHARS = 1600` per contract; the largest shipped contract is ~1.1k chars.
- `MAX_REGISTRY_CHARS = 120_000` for the whole metadata surface (~25k chars measured).
- `sideEffectClass` must be `"none"` — skills are advisory, the runtime owns side effects.
- `composableWith` is **derived** from shared task classes and capped at 10; it is never
  hand-maintained.
- `validateSkillContract()` fails closed on a missing field, unknown id, version drift, a
  non-`none` side-effect class, or an over-budget contract.
- `skillRegistryDrift()` reports any shipped `SKILL.md` without a contract, and any contract
  without a shipped skill.

---

## 3. Skill Router

Pipeline: `task → detectIntent → evidence → rankSkills → routeSkills → minimal activation`.

Signals, all deterministic and all recorded per candidate:

| Signal | Meaning |
| --- | --- |
| `exact-intent` | Intent id detected in the normalized task |
| `stack-match` | Framework token (`next.js`, `nestjs`, `react native`, …) |
| `repo-evidence` | Changed-file paths (routes, schema/migrations, tests) |
| `task-class` | Detected class overlaps the contract's task classes |
| `required-capability` | Task requires a capability the skill requires |
| `risk-domain` | Policy domain (`security`, `payment`, `data`) |
| `intent-domain-escalation` | A narrow intent inside a broader risk domain |
| `learned-utility` | Measured SkillUtility above/below hysteresis (neutral below the floor) |
| `context-cost` | Tie-breaker: `gate` > `procedure` > `reference` load cost |
| `negative-guard` | Suppresses framework/reference skills on a docs-only or docs-intent task |

Activation is a **relative cutoff**, not "top N unconditionally": `cutoff = max(8, leaderScore ×
0.4)`. Unconditional top-N is exactly what pulls irrelevant skills into context.

Two documented ways past the default target of 3, both always reported in `expandedReason`:

- `above-default-cutoff` — more skills cleared the evidence cutoff than the default allows;
  exactly one extra slot is granted.
- `composition-with-independent-evidence` — a strongly matched skill pulls in one direct peer
  that has its own independent evidence and a non-redundant task class.

Language handling is diacritic-insensitive (`normalizeTask`) and distinguishes `en`, `vi`, and
`mixed` by English-technical-token ratio, so a Vietnamese task that embeds English identifiers is
routed as mixed rather than mis-parsed.

### Skill Utility

```
SkillUtility = verifiedUsefulness
             - contextCost × 0.1
             - unnecessaryToolCost × 0.1
             - falseActivationRate × 0.5
             + verificationContribution × 0.05
```

- `minSamples = 8`. Below that the verdict is `NEUTRAL` and `evidence = NOT_MEASURED`.
- Hysteresis margin `0.12`; bounded store with insertion-order eviction.
- The learner **cannot delete a skill** (`deletionAllowed: false`) and **cannot change safety
  policy** (`safetyPolicyMutable: false`).

---

## 4. Skill Capsule

Concatenating four `SKILL.md` bodies to answer a payment+auth+Next.js+bug task is how skill noise
is created. A capsule compiles one bounded block instead.

Budget allocation order is fixed:

1. Task contract
2. **Required constraints** — allocated first, never truncated, never dropped
3. Procedure
4. Applicable rules
5. Reference notes

Constraint detection matches `must`, `must not`, `never`, `always`, `require*`, `forbidden`,
`do not`, `rollback`, `idempot*`, `verify`, `security`, `integrity`, `transaction`,
`fail-closed`, `secret`, `permission`, `authoriz*`.

Every section records provenance (`skillId`, `heading`, `lineStart`, `entryCount`). Ordering is
`skillId → source line → text`, so identical inputs give identical output and an identical cache
fingerprint. `expandSkillCapsule()` is the explicit escape hatch back to full section text.

Telemetry: `skillsConsidered`, `skillsActivated`, `skillCapsuleChars`, `rawSkillCharsAvoided`,
`skillCacheHit`, `skillExpansionCount`. No provider-token figure is derived from characters.

---

## 5. Tool Surface

`predictCapabilities()` maps normalized task text plus an optional phase to a capability set;
`compilePhaseToolSurface()` turns that into a model-facing advertised list with a bounded profile
limit. Host-provided capability tools (browser/MCP) are resolved through the same capability
vocabulary as hydration, so they do not need a fixed name list.

Safety is separated, not merged:

- `SAFETY_CAPABILITIES` (11 entries: permission lattice, workspace containment, destructive-shell
  policy, dirty-work guard, local-env write guard, secret redaction, execution ownership,
  Evidence Store, verification gate, external-side-effect no-replay, thinking-level preservation)
  are runtime-enforced and reported separately from `advertised`.
- `hiddenToolsRemainPolicyChecked = true` — hiding a tool is not removing a permission.
- `denied` tools are never advertised and never hydratable; they are reported in
  `deniedCapabilities` so a refusal has an explicit reason.

Hydration is per-surface and once-only:

| Deny reason | Condition |
| --- | --- |
| `NO_EVIDENCE` | The request carries no evidence string |
| `ALREADY_HYDRATED` | The capability was already granted on this surface |
| `POLICY_DENIED` | The underlying tool is policy-denied |
| `ALREADY_ADVERTISED` | The capability is already model-facing |
| `UNKNOWN_CAPABILITY` | Not in the deferred set |
| `SIDE_EFFECT_REQUIRES_POLICY` | Writer tool without `allowSideEffectHydration` |
| `BUDGET_EXHAUSTED` | Per-surface hydration budget reached |

Every request returns a hash-bound receipt with `granted`, `reason`, `policyChecked`,
`permanentToolLoss: false`, `permissionRemoved: false`, and the enforced safety set.

V16.2 `compileToolSurface()` remains the final advertised-surface authority and stable-prefix
owner. V16.5 supplies task-phase **priority** only.

---

## 6. Subagent Fabric

Roles reuse the existing 12-agent catalog; no new role was added.

| Role | Agent | Read-only |
| --- | --- | --- |
| `explore` | `codebase-mapper` | yes |
| `diagnose` | `debugger` | yes |
| `implement` | `executor` | no |
| `review` | `reviewer` | yes |
| `test-analysis` | `integration-verifier` | yes |
| `architecture` | `architect` | yes |

`decideDelegation()` returns `delegate` or `parent-direct` with the evidence that decided it:
independence-required, context-separation, decomposable, high-context-pressure, versus
single-file-target, deterministic-change, low-context-pressure, no-specialist-role-requested.

Bounds:

| Bound | Value |
| --- | --- |
| Active children | default 2, hard max 3 |
| Delegation depth | default 1, hard max 2 |
| Cycle guard | an agent may not reappear in its own stack |
| Child task text | 3,000 chars |
| Child context | 6,000 chars (+ task) |
| Timeout | 180s default, per-child override |
| Inactivity | heartbeat interval × 8 |

`buildChildContext()` returns `notCopied: [parent-conversation, full-skill-bodies, all-tools,
complete-repository, unrelated-prior-attempts]` so the omission is auditable rather than implicit.

Lifecycle guarantees: `finalizeChild()` emits a receipt with `canProduceVerdict: false` and
`permissionGrant: false`; partial output on failure is preserved by reference; `cancelAllChildren()`
plus `sweepChildren()` guarantee `orphansRemaining === 0`.

---

## 7. Handoff Capsules

Raw child output → Evidence Store. Parent context → bounded capsule.

```
{ childId, parentId, role, agent, task, findings, relevantFiles, symbols,
  evidenceRefs, proposedActions, unresolvedQuestions, risks,
  verificationStatus, childVerificationClaim, confidence,
  canProducePass: false, isTaskVerdict: false, canGrantPermission: false,
  rawEvidenceAvailable, rawEvidenceRef, measurements, redaction, source }
```

Budget 2,000–6,000 chars (default 4,000). `verificationStatus` is always `not-verified`; a child's
own claim is preserved separately as `childVerificationClaim`. `assertHandoffAuthority()` fails
closed if any authority field is tampered.

Measured: `rawChildChars`, `handoffChars`, `handoffRatio`, `handoffRecallCount`,
`rawEvidenceRehydrations`, `parentContextGrowthChars`. `recordRawRehydration()` rejects a ref that
does not match the stored raw ref.

Every text field passes through `lib/secret-redaction.mjs`.

---

## 8. Parallel Delegation

`classifyScope()` derives `parallelClass`:

- `read-only` — may run concurrently
- `disjoint-writer` — one writer with a single root, no shell/service/external side effect
- `serial-only` — destructive shell, external side effect, or a mutable service a writer touches

Blocked pairs: writer-overlap on shared files or roots, overlapping files generally,
destructive-shell, external-side-effect, mutable-service, and capacity above the parallel max.
`buildDelegationWaves()` greedily emits waves that are individually safe and bounded.

### 8.1 Bounded concurrent execution

`classifyScope()`/`buildDelegationWaves()` decide; `lib/delegation-fleet.mjs` **executes**.
It is deliberately not a scheduler and not a swarm:

- It never spawns a process. The caller supplies `execute`, which is the existing Pi child
  runtime; `lib/process-supervisor.mjs` remains the only process framework.
- It never invents a second wave order. `computeSafeWaves()` (task graph) owns ordering,
  `buildDelegationWaves()` (delegation-safety) owns safety.
- It never grants authority. Every returned outcome carries `canProduceVerdict: false`.

One wave runs through a bounded worker pool of `min(maxParallel, waveSize)` workers. Workers
pull children in deterministic key order, and `Promise.allSettled` over the workers means one
child throwing, timing out, or being reaped can never cancel an unrelated sibling.

| Bound | Value | Source |
| --- | --- | --- |
| Active children | default 2, hard max 3 | `FABRIC_LIMITS` / `UES_MAX_ACTIVE_CHILDREN` |
| Delegation depth | default 1, hard max 2 | `FABRIC_LIMITS.hardMaxDepth` |
| Watchdog tick | 250 ms | `FLEET_LIMITS` |

`resolveFleetConcurrency()` clamps any request into `[1, 3]`; a request above the hard max is
capped, never honoured.

Preserved invariants on the parallel path:

- **Per-child timeout** — `registerChild()` stores the timeout; the fleet watchdog calls the
  same `sweepChildren()` the single-child path uses.
- **Inactivity watchdog** — identical reaper; on reap the fleet aborts *that child's* signal
  so the process supervisor terminates the tree.
- **Cancellation** — the parent signal is linked into a per-child `AbortController`; a child
  aborted this way keeps its real stop reason (`cancelChild`, not `finalizeChild`).
- **Execution ownership / cleanup** — unchanged. The fleet has no filesystem, git, or
  credential path at all.
- **No orphan** — the fleet awaits every worker before returning, and the teardown cancels any
  non-terminal child and records it.
- **Evidence binding** — each child keeps its own `outputRef`/`handoffRef` on its own receipt.

Failure posture: a child that throws, is aborted, or returns a non-zero `exitCode` is a
**failure**. Its partial payload is dropped from `result`, the wave row becomes an honest
failure row, and `runDelegationWave()` always returns `passed: false`. Only the local verifier
can produce a verdict.

Writer conflicts are not co-scheduled: `computeSafeWaves()` already defers a task whose declared
write scope overlaps an already-selected task to a later wave, and each writer task already runs
inside its own Git worktree sandbox. `delegation-safety` adds the fail-closed classification on
top (overlapping roots, mutable services, external side effects, destructive shell).

### 8.2 Measurements

Measured on the deterministic fleet fixtures in `npm run eval:v16.5:measure` (child bodies are
deterministic in-process stand-ins, the executor is the production one):

| Measurement | Provenance |
| --- | --- |
| `safeWaveCount`, `parallelDelegations`, `serializedDelegations` | MEASURED |
| `maxObservedChildConcurrency` | MEASURED |
| `childQueueMs`, `childExecutionMs`, `parallelWallMs`, `sequentialEquivalentMs` | MEASURED |
| `overlapSavingsMs` | MEASURED (dispatch overlap of child execution windows only) |
| model latency / provider tokens / end-to-end task duration | `NOT_MEASURED` |

`delegationFleetTelemetry()` sets `speedupClaim: null` unconditionally and records
`measurementScope: "child execution windows inside this process only"`. `overlapSavingsMs` is
not called a speedup anywhere in this repository.

---

## 9. DeepSeek Advisors

Five specialist question types, each with required inputs, explicit forbids, and expected
outputs:

| Role | Question |
| --- | --- |
| `root-cause` | Rank the candidate causes and give the cheapest falsification test for the top one |
| `architecture` | Compare viable architectures under these constraints; name the smallest viable one |
| `alternative-fix` | Propose the smallest correct alternative fix and state what it would break |
| `adversarial-review` | Do not redesign. Find the strongest reason this patch could still be wrong |
| `verifier-failure` | Given this exact verifier failure, what is the next discriminating check? |

`selectAdvisorRole()` routes deterministically: a failed command wins, then review-with-patch,
then planning, then ambiguous symptoms, then patch-without-review.

Packets are bounded per section and overall. `ADVISOR_AUTHORITY` is consultant-only:
`canProducePass: false`, `isTaskVerdict: false`, no filesystem, no git, no terminal, no secrets,
`instructionAuthority: "none"`. Follow-ups stay bounded by `lib/followup-budget.mjs` (V16.4); the
role only declares how many would be structurally useful (1–2).

---

## 10. Learning

`lib/advisor-benefit-learner-v2.mjs` observes `consulted`, `adviceAccepted`,
`verifierAttemptsBefore/After`, `finalVerifiedResult`, `wallTimeDeltaMs`, `toolCallDelta`,
`providerTokens`, `browserLatencyMs`, `followUps`, `fallbacks`.

Keys are scoped by `taskClass | subsystemBucket | ambiguityClass | failureClass | provider | model |
advisorRole`. `minSamples = 8`, hysteresis `0.15`, `maxKeys = 500`, bounded GC, persisted under
`.ues-work/.advisor-learner-v2.json` with a size cap.

Allowed effect: **move the AUTO consultation weight only** (`effect: "auto-consult-weight-only"`).
Forbidden and asserted: force-web, disable-verifier, change-permissions, produce-pass,
bypass-security, override-user-mode.

---

## 11. Reasoning Doctor

`ues doctor --reasoning` (and `--json`) is read-only. It reports provider, mode, adapter
availability, configured profile, session readiness *only when safely observable*, last measured
latency, recent consultation count, recent accept/reject, benefit-learner state, packet-tier stats,
follow-up health, the five advisor roles, and an authenticated-profile boolean.

It submits no prompt, mutates nothing, and prints no credential, cookie, or storage value.
Unavailable values are `NOT_MEASURED` or `unknown`, never guessed.

---

## 12. Progress Observer

Observer-only fleet view:

```
UES
+- phase: delegation
|  +- v explore  - mapped auth surface
|  +- * diagnose  - reproducing failure  2.1s
|  +- - deepseek  - not needed
`  +- o verify
`- observer-only: this view has no runtime authority
```

No chain-of-thought is stored or rendered. Actions are bounded to 120 chars and secret-shaped
substrings are redacted. At most 12 lanes. The view has no runtime, verdict, or permission
authority.

---

## 13. Measurements

`npm run eval:v16.5` runs a 15-task deterministic corpus through both the V16.4 baseline path
(legacy `compileSkillContext` + V16.2 tool-surface economy) and the V16.5 pipeline.

Measured on the current corpus:

| Measurement | V16.4 baseline | V16.5 | Provenance |
| --- | --- | --- | --- |
| Skills considered per task | — | 48 | MEASURED |
| Skills activated per task | 3 | 1.8 avg / 3 max | MEASURED |
| Skill context chars per task | 1,196 avg | 1,220 avg | MEASURED |
| Raw skill chars avoided | — | 5,182 total | MEASURED |
| Advertised tools per task | 8.0 | 5.8 | MEASURED |
| Tool schema chars per task | 12,287 | 7,873 | ESTIMATED |
| Delegations / parent-direct | — | 7 / 8 | MEASURED |
| Handoff raw → capsule chars | — | 88,000 → 423 (0.48%) | MEASURED |
| Advisor packet chars | — | 642 | MEASURED |
| Safe parallel waves | — | 3 | MEASURED |
| Parallel / serialized delegations | — | 6 / 4 | MEASURED |
| Max observed child concurrency | — | 2 (budget default 2, hard max 3) | MEASURED |
| Child queue / execution ms | — | 504 / 1,510 | MEASURED |
| Parallel wall vs sequential equivalent | — | 1,008 / 1,510 ms | MEASURED |
| Dispatch overlap | — | 502 ms | MEASURED (dispatch overlap only, not a speedup) |

Honest reading:

- The **tool surface** is materially smaller: −4,413 estimated schema chars per task (−35.9%).
- The **skill context** is not smaller than V16.4 on this corpus. V16.5's value there is
  provenance, guaranteed constraint preservation, bounded composition, and the ability to
  compose four skills without concatenating four bodies — not a character reduction.
- **Bounded parallel delegation overlaps in dispatch**: on the deterministic fleet fixtures the
  safe read-only pair overlapped (502 ms of measured overlap across 1,510 ms of child execution
  in 1,008 ms of wall clock). That is a dispatch-overlap measurement, **not** a task-level
  speedup.
- **Provider tokens, real-model wall-clock speed, and model quality are `NOT_MEASURED`.** Those
  require a live `ues trial` run. Nothing in this document claims them.

A real-task A/B extension point exists at `scripts/bench-task-ab.mjs` and `ues trial`; the
V16.5 deterministic fixture deliberately makes no real-model claim.

---

## 14. Safety Invariants

V16.5 does not weaken any of the following:

local final verifier · integration verifier · visual verifier when required · Evidence Store ·
static diagnostics completeness · dirty-work guard · `.env` protection · workspace containment ·
destructive-shell policy · execution ownership · process-tree cleanup · Windows cleanup barrier ·
browser action taxonomy · external-side-effect zero automatic replay · DeepSeek consultant-only ·
`canProducePass = false` · untrusted external content boundary · secret redaction · selected
thinking level · no automatic publish/push/deploy · bounded parallel delegation (unsafe scopes
serialized, concurrency never above the hard max, no orphan process).
