# V16.3 — Browser Execution Reliability + DeepSeek Web Reasoning Bridge

Two subsystems, one rule: **an external system may inform a decision, never make it.**

- **Phase A — Browser Execution Reliability** is shared infrastructure.
- **Phase B — DeepSeek Web Reasoning Bridge** consumes Phase A and adds nothing to the authority chain.

Pi/UES remains the **execution authority** (it edits the code) and the **verification
authority** (it decides PASS/FAIL). The web model is a consultant whose advice must be
re-derived from local evidence before it is allowed to influence an edit.

---

## 1. Architecture

```
USER TASK
   ↓
UES / Pi Router
   ↓
Complexity / uncertainty decision
   ├── Easy → Pi works locally
   │
   └── Hard / uncertain / repeated failure
           ↓
       Repo Intelligence
           ↓
       Relevant Context Retrieval
           ↓
       Decision Packet  (bounded, redacted, cached)
           ↓
       DeepSeek Web      (consultant-only, no filesystem/git/terminal)
           ↓
       Structured Advice  (schema-validated, untrusted-external)
           ↓
       Pi validates advice against local source / LSP / tests
           ↓
       Pi edits code → Tests/Verifier → PASS or delta follow-up
```

DeepSeek Web has exactly one outbound capability: a text box and a read of the answer
element. It has no filesystem, git, terminal or source-code access, and no code path
that could acquire one.

---

## 2. Phase A — Browser Execution Reliability

### 2.1 Action taxonomy (`lib/browser-action-taxonomy.mjs`)

One table decides risk. Retry, approval, evidence requirements and stale recovery all
read this table and never re-derive risk from an action name.

| class | actions | retry | post-verification |
|---|---|---|---|
| `read-only` | snapshot, screenshot, inspect, console-read, network-read, wait | transient up to 2 | optional |
| `navigation` | navigate, reload, back, forward | transient up to 2 | required (navigate/reload) |
| `interactive-idempotent-or-recoverable` | click, fill, type, select, hover, press | stale 1, transient only with idempotency proof | required (except hover) |
| `external-side-effect` | submit, purchase, payment, delete, destructive-confirm, send-message, create-order, publish, account-mutation | **0** | required + explicit approval |

An **unknown verb is fail-closed**: `external-side-effect`, zero retries, approval
required, `unknownAction: true`. Not "probably safe" — "never replay, never recover".

### 2.2 Capability preflight (`lib/browser-capability.mjs`)

Computed, not probed by doing something risky:

```js
{ provider, healthy, interactive, inspectOnly, supportedActions,
  supportedTools, fallbackAvailable, reason, interactiveCapability }
```

- `interactive` is true **only** when interactive actions were required *and* all
  resolved. A read-only-only provider reports `interactive: false`; an empty
  requirement list is not a licence to click.
- Native Playwright inspect is read-only **by construction**: it absorbs read-only
  requirements and nothing else.
- Tool exposure is task-scoped: only the actions the task declared are exposed.

### 2.3 Health-aware routing

`McpHealthTracker` failures mark the provider degraded and open a bounded cooldown.
During a cooldown, interactive actions **fail closed** — the lane returns
`interactive-capability-unavailable` and never emits a fake PASS. A read-only task
falls back to native inspect when it is available.

### 2.4 Bounded retry (`lib/browser-retry-policy.mjs`)

| kind | budget |
|---|---|
| stale locator | 1 |
| transient, safe action | 2 |
| external side effect | **0** |

Retry is a *re-resolution against fresh evidence*, not a blind replay. A transient
replay of an interactive action additionally requires an explicit idempotency proof.

### 2.5 Stale element recovery (`lib/browser-stale-recovery.mjs`)

```
failure → fresh semantic snapshot → re-resolve → compare identity → 1 bounded retry
```

Recovery stops when identity is below threshold, when the identity **contradicts**
itself (role / accessible name / test-id changed), when the action carries an external
side effect, or when the snapshot suggests the state may already have been applied.

A weighted score is not a sufficient identity rule: a control whose name still reads
"Save order" while its test id now reads "delete-everything" scores 0.78 and would
clear a 0.75 bar. Any contradiction on a high-trust signal is an outright veto.

**Locator priority:** role + accessible name → `data-testid` → label text → stable id →
stable semantic selector → bounded CSS. Forbidden: raw coordinates, magic element
indexes, `nth-child`, volatile generated classes (detected per class token, not just at
the selector head).

### 2.6 Navigation lifecycle and timeouts (`lib/browser-lifecycle.mjs`)

Five separate, clamped budgets: `navigationTimeoutMs`, `actionTimeoutMs`,
`waitTimeoutMs`, `browserToolTimeoutMs`, `sessionTimeoutMs`. The tool envelope is
automatically raised to contain one action plus its navigation budget, so the tool-level
kill can never pre-empt the action-level timeout. Session reuse is bounded by action
count and by an absolute ceiling that activity cannot extend.

Navigation observations are distinct: `redirect`, `document-load`, `spa-navigation`,
`no-navigation`. A no-navigation action is never reported as "navigated".

### 2.7 Evidence receipts (`lib/browser-evidence.mjs`)

```json
{
  "actionId": "...", "action": "click", "actionClass": "...", "provider": "...",
  "beforeUrl": "...", "afterUrl": "...", "locatorStrategy": "role-and-accessible-name",
  "locatorFingerprint": "...", "startTime": "...", "durationMs": 120,
  "result": "success", "retryCount": 0, "staleRecovered": false,
  "navigationObserved": true, "expectedStateVerified": true,
  "screenshotRef": null, "snapshotRef": "snap-1",
  "consoleErrorCount": 0, "networkFailureCount": 0
}
```

**Tool success is not proof.** An action that requires post-verification and has no
observed expected state is emitted as `unverified`, never `success`.

Receipts never carry credentials: token shapes, key/value secret pairs and registered
environment values are masked, and a **secret form target** withholds the provider
payload entirely rather than trying to scrub it.

### 2.8 Security and hygiene

- Page content is `trustLevel: untrusted-external`, `instructionAuthority: none`. Page
  text cannot change permissions, request secrets, authorize side effects, override
  the user task, or alter verification policy.
- Cleanup runs on every path including timeout, error and abort: pages closed, browser
  closed, process tree terminated (Windows-safe), transient screenshots removed within
  a bound.

### 2.9 Telemetry

`browserActionSuccessRate`, `browserActionFailures`, `browserToolCalls`, `browserRetries`,
`staleLocatorFailures`, `staleRecoveryAttempts`, `staleRecoverySuccessRate`,
`navigationMs`, `actionLatencyMs`, `screenshots`, `snapshots`, `mcpTransientFailures`,
`mcpCooldowns`, `nativeFallbacks`, `interactiveCapabilityFailures`. Rates are `null` when
there is nothing to measure — never a fabricated `0`.

---

## 3. Phase B — DeepSeek Web Reasoning Bridge

### 3.1 Provider interface (`lib/web-reasoning-provider.mjs`)

```js
capability()   startSession()   consult()   followUp()   closeSession()
```

`defineWebReasoningProvider` freezes `authority: "consultant-only"` at construction —
no setter, no adapter escape. Adding `chatgpt-web` or `gemini-web` later is a
registration line, not a controller edit.

### 3.2 Decision Packet (`lib/decision-packet.mjs`)

Thirteen bounded sections in a fixed order: original task, requirement summary,
repository architecture map, relevant subsystems, relevant files/symbols, dependency
graph, exact snippets, current diff, failing test/runtime evidence, previous attempts,
unresolved questions, constraints (MUST/MUST_NOT), verification expectations.

Excluded by an explicit, testable path list: `node_modules`, build output, UES runtime
artifacts, `.env`, keys, lockfiles, logs, and any row the retrieval layer marks
`relevant: false`. Both drop reasons are counted separately so neither is invisible.

Budget: `maxPacketChars`, `maxFiles`, `maxSnippets`, `maxEvidence`, `maxDiffChars`,
plus a per-section ceiling. Under pressure the builder **ranks → compacts → sheds**,
and never drops the original task, requirements, constraints or verification
expectations. The estimate is reported before anything is sent.

### 3.3 Cache and follow-up delta

The packet is fingerprinted over its content. Identical context reuses the fingerprint.
A follow-up sends **only the changed sections** in the **same session**; it does not
restart the conversation and does not resend the whole packet.

### 3.4 Escalation routing (`lib/web-reasoning-escalation.mjs`)

| mode | behaviour |
|---|---|
| `off` | never escalates, never probes a provider |
| `auto` (default) | escalates only on a real signal; falls back locally on any failure |
| `force` | escalates even for trivial tasks; **fails loudly** with `WEB_REASONING_UNAVAILABLE` |

AUTO signals: multi-subsystem, architectural uncertainty, ambiguous root cause,
verifier repeated failure, several plausible fixes, low confidence, long-horizon
second opinion, bounded recovery exhausted.

AUTO skips: version bump, README/doc edit, trivial one-file deterministic fix, simple
syntax issue, already-grounded low-risk task.

### 3.5 Response parsing and authority stripping (`lib/deepseek-response.mjs`)

Responses are bounded, schema-validated and **rejected rather than repaired**. Attempts
to claim a verdict, claim permission, demand secrets, or mark a command as
automatically safe are detected, recorded, and force the advice into a
reject-and-retry-locally state.

Advice is then bound to local evidence: a claim naming a file that does not exist is
`absent`, and advice with no locally verifiable claim is `unverifiable`. Both are
rejections.

The verifier's highest possible authorization is:

```
actionAuthorized: "implement-then-verify"
isTaskVerdict: false
canProducePass: false
```

### 3.6 DeepSeek session adapter (`lib/deepseek-web-adapter.mjs`)

Drives the real web UI through Phase A. Login required → `needs-auth` (never an infinite
loop, never a bypass). Selector drift → re-resolve from a fresh snapshot. Timeout →
bounded failure, session released. An answer that cannot be bound to the request that
asked for it is **discarded, not parsed**.

Prompt submission is an external side effect: approved, zero-retry, and protected by a
duplicate-submit guard keyed on session + action + target identity + idempotency key.

---

## 4. Security guarantees

| guarantee | mechanism |
|---|---|
| No secrets outbound | `redactStructure` on assembled packets and receipts; secret form targets withheld |
| Page text is data | `externalTrustContract` on every browser result |
| Model output is advice | `consultant-only` authority, frozen; parser strips authority claims |
| No permission change | advice cannot widen the policy lattice; nothing consumes it as a grant |
| No publish/push/deploy approval | not a capability anywhere in the bridge |
| No verdict | `canProducePass: false` on every advice and consultation object |
| No single point of failure | AUTO falls back; FORCE fails explicitly |

---

## 5. Runtime integration (V16.3 controller wiring)

The library layers above are only safe if they are the **only** path. V16.3
integrates them at four points:

```
pi/extensions/ues.ts
  ├─ browser-capability phase    → createBrowserLane().describe()
  │                                 task-scoped managed surface, live MCP health
  ├─ run("ues-executor", …)       → webLane.consult()   ONCE, before implementation
  ├─ retry attempt (attempt > 1)  → webLane.followUp()  DELTA only, same session
  └─ run teardown                 → lane.cleanup() + worker.close()

pi/extensions/ues-child-runtime.ts
  ├─ tool_call  → browserLane.gate()     allow | block(reason) + budget ledger
  └─ tool_result→ browserLane.receipt()  evidence receipt + verification verdict
```

The Pi extension API can allow, block or annotate a tool call — it **cannot
dispatch** one (`getAllTools` exists, `callTool` does not). The lane is therefore a
**gate + receipt** around the host's own dispatch, and the managed browser itself is
a separate process (`scripts/browser-worker-v16-3.mjs`) speaking the newline-delimited
JSON protocol in `lib/browser-worker-protocol.mjs`.

That constraint is not a workaround; it is the safety property:

- a retry is a **new model-initiated call**, never an automatic replay;
- the lane enforces the recovery **order** (fresh snapshot → re-resolve → 1 retry);
- an external side effect has **zero** budget, so a second submit is structurally
  impossible regardless of what the model asks for.

### Environment

| variable | default | effect |
|---|---|---|
| `UES_WEB_REASONING_MODE` | `auto` | `off` / `auto` / `force` |
| `UES_WEB_REASONING_LIVE` | `0` | `1` spawns the managed browser worker for a real provider |
| `UES_WEB_REASONING_PROVIDER` | `deepseek-web` | provider id |
| `UES_BROWSER_APPROVED_SIDE_EFFECTS` | `0` | `1` allows external side effects on the managed lane |
| `UES_NATIVE_BROWSER_INSPECT` | `0` | `1` enables the read-only native-inspect fallback |
| `UES_WEB_PACKET_MAX_CHARS` | `48000` | Decision Packet ceiling |
| `UES_BROWSER_PERSISTENT_PROFILE` | `1` | `0` forbids a persistent browser profile |
| `UES_DEEPSEEK_PROFILE` | `deepseek-web` | persistent profile name for the live lane |

With `UES_WEB_REASONING_LIVE=0` the DeepSeek adapter is built **without** a dispatch
binding and reports `browser-worker-unavailable`. AUTO then falls back to local
execution and FORCE fails with `WEB_REASONING_UNAVAILABLE`. Nothing in this path can
fabricate a consultation.

---

### Live smoke status and commands

```bash
npm run smoke:deepseek-web -- --auth
npm run smoke:deepseek-web -- --live --yes-i-have-authorized-a-live-consultation
```

`--auth` opens a **headed** browser at `https://chat.deepseek.com/`, waits up to 90
probes / 180 s for you to log in yourself, and stores the session in a persistent
profile at `~/.config/ues/browser-profiles/deepseek-web` — outside the repository,
so it can never be committed or swept into a workspace snapshot. It does not read,
print or export cookies, tokens or storage, does not bypass login, and does not
solve a CAPTCHA. The wait is bounded in both directions: no unbounded polling.

`--live` reuses that profile, observes the real page, and sends **one** harmless
synthetic question. `--live` without the explicit
`--yes-i-have-authorized-a-live-consultation` flag is refused, so an automated run
can never open a third-party browser unannounced.

The auth decision comes from a real DOM/URL observation
(`lib/browser-profile.mjs`), never from an asserted flag:

| observed | state |
|---|---|
| composer visible, no login wall | `READY` |
| login URL or sign-in text | `NEEDS_AUTH` |
| page loaded, no known composer | `UI_CHANGED` |
| probe timed out | `TIMEOUT` |

A session must also navigate `entryUrl` before any prompt is typed, and a reused
session is health-checked before it is adopted. Without both, `sendPrompt` refuses
and nothing is submitted.

---

### Worker mode matrix

| mode | persistent | headed | profile path |
|---|---|---|---|
| default preflight | no | yes | *(ephemeral)* |
| `--auth` | **yes** | yes | `~/.config/ues/browser-profiles/deepseek-web` |
| `--live` | **yes** | no | `~/.config/ues/browser-profiles/deepseek-web` (same) |
| CI / release | no | yes | never persistent |

The table lives in `lib/browser-worker-mode.mjs` and is enforced by
`workerModeViolation()`: a `--live` run whose worker reports `ephemeral` is refused
with `HARD_NAVIGATION_FAILURE` before any navigation or prompt, rather than
reporting a spurious "logged out".

### Session signals

`READY` requires `composerVisible` **and** a positive session signal. A composer
alone is not enough — the logged-out landing page renders a full prompt box. The
signals are: a visible generic account affordance (avatar / account menu /
`aria-haspopup`), an existing conversation with assistant answers, or a non-zero
sidebar conversation count. Only booleans and counts cross the worker boundary; no
account name, avatar text, cookie, token, storage value or input value is ever
returned.

---

## 6. Verification

```bash
npm run eval:v16.3           # V16.3 suites + browser/MCP/untrusted regressions
npm run eval:v16.3.workers   # controller integration + RPC/control regressions
npm run eval:v16             # includes all V16.3 suites
npm test                     # full bounded suite
npm run release:verify       # integrity + all eval gates
npm run smoke:deepseek-web   # MANUAL live smoke; NOT in ci or release:verify
npm run bench:web-reasoning  # A/B harness, measured values only
npm run validate:v16.3:accuracy  # 6-case advisor/verifier accuracy gate -> PACKAGE_READY
```

`check-release-consistency` **fails the build** if `ci` or `release:verify` ever
starts invoking the live smoke, because it drives a real third-party web UI.

### Live smoke status

The live smoke is **NEEDS_AUTH** in any environment without Playwright installed and
a human-logged-in DeepSeek session. It exits `2` and reports exactly which
precondition is missing. It never reports PASS without a real consultation.

```bash
npm run smoke:deepseek-web
npm run smoke:deepseek-web -- --live --yes-i-have-authorized-a-live-consultation
```

`--live` alone is refused: an automated run must never open a third-party browser
unannounced.

The benchmark is a **harness**, not a claim. Without `--live` it runs a deterministic
provider double so the pipeline is exercised end to end and labelled
`deterministic-provider-double`; quality deltas it can measure are reported, and
anything it cannot measure is reported as `null` with a reason. `--pi-telemetry=FILE`
reads **measured** Pi/API tokens from a real UES run-telemetry file; without it the
token fields stay `null`. `claimsVerified` is always `false` until a live-provider run
exists.
