# 001 — Terminal Gate Round 3: Oracle vs agbrowse

Date: 2026-07-24
Status: P (research)
Theme: T1 + T9 — terminal completion gate and stop-control scoping
Oracle delta anchor: `114610762975ca96136e92387c31fd875c6b03a2..HEAD`
Oracle files read: `src/browser/actions/assistantResponse.ts`, `src/browser/actions/thinkingStatus.ts`, and (for T9's selector definition) `src/browser/constants.ts`
agbrowse files read: `web-ai/chatgpt.mjs`, `web-ai/chatgpt-response-observer.mjs`, `web-ai/chatgpt-response-dom.mjs`, `web-ai/tab-finalizer.mjs`

## Upstream mechanisms

The rows below describe the actual predicate/state/scoping changes in each `git show`, not merely the commit subjects. Some commits are intermediate steps later tightened or removed; that lineage matters because `67da293a` deliberately replaces quiet-window recovery, including `a84f52e3` weak-evidence aging, with fail-closed scoped completion proof.

| Commit | Mechanism from actual diff | Bug class closed |
| --- | --- | --- |
| `9f6703bf` | Replaces `progress, [role=progressbar], [aria-valuenow], [data-testid*=progress]` with genuine `progress, [role=progressbar]`; determinate HTML/ARIA bars count as live only while `value < max` (ARIA max defaults to 100), while valueless bars remain indeterminate/live. | Generic sliders/spinbuttons and completed progress bars permanently vetoing completion. |
| `1e2f71a0` | Recognizes short heading-prefixed completed summaries by changing `startsWith("thought for ")` to bounded `includes("thought for ")`; sidecar completion checks use visible text rather than joined test-id metadata; image-only chrome accepts optional `Reasoning`/`Pro thinking` prefixes. | A mounted `Reasoning Thought for 12s` summary being mistaken for active reasoning forever, and prefixed image chrome leaking into the answer. |
| `0071c547` | Moves the standalone live-progress veto from `document` to the latest conversation turn; verified sidecar panels remain a separate path. | Unrelated persistent page progress UI blocking a completed response. |
| `ded58d44` | Splits activity into strong (stop, shimmer, busy, status, progress) and weak (text-only sidecar). Strong activity resets/vetoes finished-action proof; weak stale sidecar evidence may be overridden by stable, debounced completion controls. | Either letting a transient action bar beat real live work, or allowing a stale mounted reasoning sidecar to hang completion. |
| `93ccb79d` | Prefix status matching is allowed only in verified thinking/reasoning test-id chrome; generic status/live nodes require exact active labels. Sidecar progress is strong only when the panel metadata identifies thinking/reasoning/sidecar. | Unrelated status copy or progress inside a generic dialog/aside becoming strong model-activity evidence. |
| `86d1fb2b` | Scopes shimmer and `aria-busy=true` to the latest turn or a verified thinking panel. Replaces the broad short-text completed-summary test with a whole-string anchored duration/worded-duration grammar, so `Thought for 2s: Searching…` remains live. | Global busy UI causing hangs, and a growing live trace being prematurely classified as a completed reasoning summary. |
| `9454ef4d` | Passes sampled response `messageId`/`turnId` and `minTurnIndex` into completion detection. The finished action must belong to the same latest assistant turn by identity, or (without identity) the turn must be at/after the new-turn baseline. | A persistent action bar from an older response proving completion of the current sampled text. |
| `a84f52e3` | Intermediate fallback design: weak sidecar activity carries a normalized text key; changing weak evidence resets quiet time, while static weak evidence may be ignored only after at least `max(quietMs*5, 60s)`. Generic role/status live regions are removed from strong status matching. | A live-changing sidecar being treated as stale, while a truly stale mounted sidecar blocks selector-drift recovery forever. Superseded by `67da293a`. |
| `67da293a` | Removes quiet-window proof entirely. Terminal success now requires a stable, debounced finished-action bar; completion evidence can be carried directly by a fallback snapshot as `completionVisible`, and otherwise must be correlated by identity/baseline. Weak-evidence key/aging state is deleted. | Stable preamble + absent stop being treated as completion when no positive, response-scoped terminal evidence exists. |
| `7b107769` | In wrapperless markdown fallback, when no normal turn wrappers exist, requires the candidate node to be DOM-following the latest user node; it no longer treats `hasTurns=false` as automatically after `minTurnIndex`. | Old/user-echo/wrapperless markdown being accepted because turn-index correlation was unavailable. |
| `57d4a7af` | Extends the anchored completed-reasoning grammar with optional trailing `Edit`. | Completed `Thought for … Edit` action chrome being treated as active reasoning indefinitely. |
| `99b30cfa` | Keeps exact stop test IDs, but scopes the broad aria-label fallback to `form button` and excludes labels containing dictation, voice, or read. | Read-aloud/voice/dictation “stop” controls elsewhere (or in the composer) falsely reporting model generation. |

## agbrowse current state

- The authoritative loop slices assistant texts after `baseline.assistantCount`, filters placeholders, then reads document-level streaming state (`web-ai/chatgpt.mjs:607-617`). It verifies the latest assistant follows the latest user (`web-ai/chatgpt.mjs:505-529`, `web-ai/chatgpt.mjs:618-623`).
- Completion is still permitted from text equality plus a length-dependent quiet window: 1s with a finished action, otherwise 2–8s depending on length (`web-ai/chatgpt.mjs:624-634`). Thus finished controls accelerate completion but are not required positive proof.
- `isStreaming()` searches document-wide stop controls (`web-ai/chatgpt.mjs:909-917`) and document-wide `progress`/`[role=progressbar]` without checking determinate completion (`web-ai/chatgpt.mjs:918-927`). Its sidecar heuristic is binary, uses joined visible text/aria/test-id, and excludes any label containing `thought for` (`web-ai/chatgpt.mjs:928-963`).
- `isResponseFinished()` selects the latest recognizable assistant turn and finds a finished action inside it, but receives neither the sampled text/identity nor the baseline index (`web-ai/chatgpt.mjs:970-993`).
- Placeholder filtering includes exact `reasoning`, `pro thinking`, `finalizing answer`, and prior-round `ChatGPT said: Answer now`, but has no completed-reasoning-action grammar (`web-ai/chatgpt.mjs:71-88`, `web-ai/chatgpt.mjs:1344-1347`). `cleanAssistantText` strips only a narrow leading `Thought for ...` shape (`web-ai/chatgpt.mjs:1352-1355`).
- The MutationObserver is expressly non-authoritative (`web-ai/chatgpt-response-observer.mjs:3-11`), but its wake predicate also uses a document-wide stop selector (`web-ai/chatgpt-response-observer.mjs:34-65`). Timeout recovery accepts either finished controls or a stable re-read as sufficient (`web-ai/chatgpt-response-observer.mjs:117-145`; caller gate at `web-ai/chatgpt.mjs:765-810`).
- DOM extraction is selector/count based and returns text only—no message/turn identity or scoped completion bit (`web-ai/chatgpt-response-dom.mjs:3-12`, `web-ai/chatgpt-response-dom.mjs:20-40`, `web-ai/chatgpt-response-dom.mjs:50-72`). It has no wrapperless markdown fallback equivalent.
- Finalization trusts the caller's `complete`/`completed` status and does not independently validate terminal evidence (`web-ai/tab-finalizer.mjs:7-8`, `web-ai/tab-finalizer.mjs:50-72`). That separation is reasonable, but means the poll/recovery gate must be correct before calling it.

Prior-round G1–G6 are real protections but do not cover the new scope/correlation rules. Exact-string stability handles rewrites; progress/sidecar vetoes exist; turn ordering rejects a recognizable stale assistant turn; placeholders are broader. Round 3 goes beyond those protections by requiring current-turn activity scope, distinguishing strong from weak evidence, binding completion controls to the sampled response, correlating wrapperless candidates, and ultimately refusing quiet-only completion.

## Per-mechanism classification

| Commit / mechanism | agbrowse classification | Priority | Evidence and required change |
| --- | --- | --- | --- |
| `9f6703bf` genuine/live progress predicate | **Gap** | **P2 hardening** | Selectors are already restricted to `progress` and `[role=progressbar]`, matching half the fix, but every visible determinate bar—including `value == max`—is considered streaming (`web-ai/chatgpt.mjs:918-927`). Evaluate each bar and return live only for indeterminate or `value < max`. |
| `1e2f71a0` heading-prefixed completed reasoning | **Gap** | **P2 hardening** | The sidecar code excludes any joined label containing `thought for`, so it recognizes prefixed summaries but also suppresses live traces containing that phrase (`web-ai/chatgpt.mjs:939-958`). Add a visible-text-only, anchored completed-summary predicate and corresponding response/image chrome normalization rather than broad `includes`. |
| `0071c547` live progress scoped to current turn | **Gap** | **P1 real risk** | Progress locators are page-global (`web-ai/chatgpt.mjs:918-927`). Move progress evaluation into the latest assistant/current response turn; allow a separate path only for metadata-verified thinking sidecars. |
| `ded58d44` strong versus weak activity | **Gap** | **P2 hardening** | `isStreaming()` returns one boolean; text-only sidecar and stop/progress have identical veto strength (`web-ai/chatgpt.mjs:909-965`). Return structured activity strength and let only response-scoped strong activity veto positive completion proof. |
| `93ccb79d` verified thinking chrome scope | **Gap** | **P1 real risk** | Global progress is unscoped and any geometrically right-side thinking-looking panel is a veto; no verified panel metadata is required (`web-ai/chatgpt.mjs:918-963`). Restrict strong panel evidence to thinking/reasoning/sidecar metadata and current-turn indicators. |
| `86d1fb2b` scoped busy + anchored summary grammar | **Gap** | **P2 hardening** | agbrowse does not currently use shimmer/`aria-busy` in the general gate, so that subpart creates no global-busy false positive; however its completed-summary exclusion is the over-broad `label.includes('thought for')` (`web-ai/chatgpt.mjs:939-946`). Implement the anchored whole-label duration grammar before adding any busy signals, and scope future busy checks to current turn/verified panels. |
| `9454ef4d` completion controls bound to sampled response | **Gap** | **P1 real risk** | `isResponseFinished(page)` has no sampled identity or baseline argument; it merely checks the latest recognizable assistant turn (`web-ai/chatgpt.mjs:970-993`), while extracted rows carry only text (`web-ai/chatgpt-response-dom.mjs:20-40`). Return turn/message identity (or a scoped completion bit) and require identity/baseline correlation. |
| `a84f52e3` weak-evidence aging | **Gap** | **P3 defer** | No weak evidence key/age exists; sidecar evidence is binary (`web-ai/chatgpt.mjs:928-963`). This exact upstream mechanism was superseded by `67da293a`; do not port it independently unless agbrowse intentionally retains quiet fallback. The applicable risk should instead close through scoped positive proof. |
| `67da293a` scoped terminal evidence only | **Gap** | **P1 real risk** | agbrowse completes after stable text even when `finished=false` (`web-ai/chatgpt.mjs:624-634`) and timeout recovery likewise permits stable re-read (`web-ai/chatgpt-response-observer.mjs:122-145`, `web-ai/chatgpt.mjs:785-810`). Require response-scoped completion proof for normal/recovery text; fail closed on selector drift, retaining separately proven generated-image completion. |
| `7b107769` wrapperless completion correlation | **Gap** | **P2 hardening** | DOM readers only recognize configured assistant wrappers and return flat strings (`web-ai/chatgpt-response-dom.mjs:3-12`, `web-ai/chatgpt-response-dom.mjs:20-40`); ordering evaluation fails open on unrecognized structure (`web-ai/chatgpt.mjs:510-529`). Add a wrapperless markdown fallback that rejects user echoes and requires DOM-following the latest user when turn indexing is unavailable. |
| `57d4a7af` completed reasoning action with `Edit` | **Gap** | **P2 hardening** | Placeholder/final-answer logic has no anchored `Thought for <duration> [Edit]` action-chrome predicate (`web-ai/chatgpt.mjs:71-88`, `web-ai/chatgpt.mjs:1344-1355`). Add the optional-`Edit` completed-summary grammar to activity and answer/image normalization. |
| `99b30cfa` composer-scoped stop aria fallback | **Gap** | **P1 real risk** | Both the authoritative poller and shared DOM constants use document-wide `button[aria-label*="Stop" i]` (`web-ai/chatgpt.mjs:909-917`, `web-ai/chatgpt-response-dom.mjs:9-12`), inherited by observer/recovery (`web-ai/chatgpt-response-observer.mjs:34-65`, `web-ai/chatgpt-response-observer.mjs:151-167`). Scope fallback to composer form and exclude dictation/voice/read controls while retaining exact test IDs. |

## Proposed gap rows

- G-T1-R3-01 | genuine/live determinate progress predicate | Gap | `web-ai/chatgpt.mjs` | `chatgpt.mjs:918-927` treats every visible determinate bar as live
- G-T1-R3-02 | completed-reasoning summary/action grammar | Gap | `web-ai/chatgpt.mjs` | `chatgpt.mjs:939-946,1344-1355` uses broad sidecar exclusion and lacks anchored optional-Edit grammar
- G-T1-R3-03 | current-turn activity and verified-panel scoping | Gap | `web-ai/chatgpt.mjs` | `chatgpt.mjs:909-963` evaluates stop/progress globally and sidecar evidence without strong/weak scope
- G-T1-R3-04 | completion-control binding to sampled response | Gap | `web-ai/chatgpt.mjs`, `web-ai/chatgpt-response-dom.mjs` | `chatgpt.mjs:970-993`; `chatgpt-response-dom.mjs:20-40` has no identity/baseline correlation
- G-T1-R3-05 | response-scoped positive terminal proof | Gap | `web-ai/chatgpt.mjs`, `web-ai/chatgpt-response-observer.mjs` | `chatgpt.mjs:624-634,785-810`; `chatgpt-response-observer.mjs:122-145` still accepts quiet text stability
- G-T1-R3-06 | wrapperless completion correlation | Gap | `web-ai/chatgpt-response-dom.mjs`, `web-ai/chatgpt.mjs` | `chatgpt-response-dom.mjs:3-40`; `chatgpt.mjs:510-529` has no wrapperless candidate path and ordering can fail open
- G-T9-R3-01 | composer-scoped stop aria fallback | Gap | `web-ai/chatgpt-response-dom.mjs`, `web-ai/chatgpt.mjs` | `chatgpt-response-dom.mjs:9-12`; `chatgpt.mjs:909-917` use a document-wide aria-label fallback

