# Bug Fix Log

Purpose: every defect gets durable evidence, root cause, regression coverage, and validation.

## Required entry

- **Date / symptom**
- **Detection source**: unit test, process test, real peer, user report
- **Reproduction**
- **Root cause**
- **Fix commit**
- **Regression test**
- **Validation**
- **Residual risk / follow-up**

## Entries

### 2026-08-03 — Nested peer workflow deadlocked behind parent active ID

- **Detection source:** live three-peer RPS/debate dogfood.
- **Symptom:** `bop` held parent RPS request active; `pib` answer arrived and was marked processed, but no turn fired. Debate request queued indefinitely.
- **Reproduction:** parent request `bip→bop`; while parent active, `bop→pib` question; `pib→bop` normal answer on same thread. Single-active scheduler suppresses answer turn.
- **Root cause:** one `activeId` could neither represent suspended parent work nor make continuation/urgent messages exact valid targets.
- **Fix commit:** none; user requested no commits.
- **Fix:** bounded persisted LIFO `activeStack`; live-known same-thread response and trusted urgent promotion; atomic exact-frame completion; CAS dequeue; bounded overflow; startup completed-frame recovery.
- **Regression tests:** state stack/CAS/recovery; policy trust/promotion; core continuation/urgent/overflow/startup; deterministic reply/resolve completion-vs-promotion races.
- **Validation:** 172/172 tests; TypeScript clean; lifecycle E2E, canonical PTY E2E, and exact three-peer presence E2E pass; real restarted-Pi nested debate/urgent acceptance pending.
- **ADR:** `docs/adr/0005-nested-active-work-stack.md`.
- **Residual risk:** no hard provider-turn cancellation; continuous urgent arrivals can delay lower frames until depth cap blocks further promotion.

### 2026-08-03 — Third peer initialization removed a live peer from presence

- **Detection source:** user live-session report and `docs/screenshots/offline.png`.
- **Symptom:** after connecting new Pi identity `pib` to `bop`, status became `[bop] ↔ bip (offline), pib` while `bip` still owned a watcher and heartbeat.
- **Reproduction:** connect real Pi `bip↔bop`, detach a read-only secondary `bip`, then connect real Pi `pib→bop`; inspect `bop` status and shared AMQ config.
- **Root cause:** each attach called `amq coop init --force --agents <self>,<peers>`, replacing the shared configured roster. `pib→bop` reduced it to `bop,pib`, so public `presence list` omitted live `bip`. Secondary sessions could also emit handle-global detached notices, and detached state overrode newer presence.
- **Fix commit:** pending.
- **Regression tests:** merge-preserving transport test in `tests/mailbox.test.mjs`; ordered presence tests in `tests/peer-status.test.mjs`; owner-gated detach assertions in `tests/pi-extension-peer.test.mjs`; exact real-Pi harness `npm run e2e:peer-presence`.
- **Validation:** 155/155 tests; TypeScript clean; lifecycle E2E passed; canonical PTY E2E passed; exact isolated real-Pi flow reports status `[bop] ↔ bip, pib` and roster `bip,bop,pib`.
- **Postmortem:** `docs/postmortems/2026-08-03-three-peer-roster-clobber.md`.
- **Residual risk:** graceful lifecycle ordering uses wall-clock timestamps; instance-generation identifiers remain a future hardening option for hosts with severely skewed clocks.

### 2026-08-02 — AMQ read header shape broke resolve

- **Detection source:** real peer dogfood; `bi-boss` reported valid resolve failure.
- **Symptom:** `Resolve <id>: read returned no message matching requested id (got id undefined)`.
- **Reproduction:** real CLI returned `{body, header:{id,...}}`; lifecycle expected top-level `id`.
- **Root cause:** transport wrapper assumed normalized read shape without normalizing the real AMQ JSON contract.
- **Fix commit:** `43b7444`.
- **Regression test:** `tests/mailbox.test.mjs` uses header-shaped `amq read --json` output and asserts normalized `id`, `from`, and `body`.
- **Validation:** 131/131 tests; TypeScript clean; real explicit read/resolve later passed.
- **Residual risk:** add contract fixtures for future AMQ CLI schema changes.

### 2026-08-02 — Handshake peer became implicit primary target

- **Detection source:** independent ADR review.
- **Symptom:** handshake-discovered peer could become default send target.
- **Root cause:** extra-peer merge assigned `primaryPeer`; command send also fell back to a sole arbitrary peer.
- **Fix commit:** `9f14e93`.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts handshake peers remain discovery-only and no sole-peer fallback exists.
- **Validation:** 131/131 tests; TypeScript clean.
- **Residual risk:** presence/discovery remains visible without trust promotion by design.

### 2026-08-02 — Restored partial mailbox failed migration

- **Detection source:** user real-session report.
- **Symptom:** `amq list ... --cur failed: mailbox for "bpat" is missing ... inbox/new`.
- **Root cause:** restored attached session entered migration before repairing an incomplete mailbox directory.
- **Fix commit:** `ad37c6b`.
- **Regression test:** `tests/pi-extension-connect.test.mjs` asserts `startWatch` initializes mailbox before migration.
- **Validation:** 132/132 tests; TypeScript clean.
- **Residual risk:** external deletion during an active watcher remains handled by the next bounded recovery cycle.

### 2026-08-02 — New peer connection lacked visible status update

- **Detection source:** user real-session observation.
- **Symptom:** remote peer connected, but status remained `bi-boss` and no notification appeared.
- **Root cause:** housekeeping notice was consumed internally and `setStatus` was never called after extra-peer discovery.
- **Fix commit:** `f65b804`.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts connection notification and status refresh.
- **Validation:** 131/131 tests; TypeScript clean.
- **Residual risk:** discovered peer stays discovery-only and does not become primary automatically.

### 2026-08-02 — Real-Pi e2e wait false negatives

- **Detection source:** real-Pi e2e runs.
- **Symptom:** valid tool output was missed by tmux wait logic.
- **Root cause:** scrollback line baselines, OSC hyperlink escapes, carriage returns, and whitespace formatting.
- **Fix commit:** `26ae469`.
- **Regression test:** shell e2e wait logic uses tool-occurrence deltas and terminal-output normalization.
- **Validation:** cross-CWD e2e passed; real-Pi flow passed once; later provider quota blocked reruns.
- **Residual risk:** provider failures remain external to bridge transport.

### 2026-08-02 — Fresh peer presence could not revive offline status

- **Detection source:** user live-session report plus direct presence inspection.
- **Symptom:** live peers could exchange mail, but status showed `(offline)`; `amq presence list` showed an active peer with fresh `last_seen`.
- **Reproduction:** set local peer status to `disconnected`, provide active presence with `last_seen` inside the 30-second TTL, and reconcile status.
- **Root cause:** `withLivePeerStatus()` preserved `disconnected` status before checking fresh presence, making offline state sticky after one detach notice or stale sample.
- **Fix commit:** pending.
- **Regression test:** `tests/peer-status.test.mjs` asserts fresh presence revives disconnected peers and stale presence marks them offline.
- **Validation:** pending final commit; 137/137 tests and TypeScript clean before final runtime restart.
- **Residual risk:** sessions with no presence heartbeat become offline after 30 seconds even if their transport remains reachable; runtime heartbeat behavior needs separate E2E coverage.

### 2026-08-02 — Live watcher did not refresh presence

- **Detection source:** user live-session report; direct presence inspection showed an active transport with stale `last_seen`.
- **Symptom:** sessions could exchange mail while presence-based status marked them offline after the TTL.
- **Reproduction:** keep watcher transport alive beyond 30 seconds without invoking presence update; inspect `presence list` and status.
- **Root cause:** AMQ watcher owned mailbox polling but never refreshed agent presence; TTL measured stale metadata rather than watcher liveness.
- **Fix commit:** pending.
- **Regression test:** `tests/watch-lifecycle.test.mjs` asserts initial/interval heartbeat and inactive update on stop.
- **Validation:** 138/138 tests; TypeScript clean; lifecycle E2E passed.
- **Residual risk:** abrupt SIGKILL cannot publish inactive presence; TTL remains required fallback.

### 2026-08-02 — Resolve reactivated completed answer

- **Detection source:** live peer dogfood during E2E; resolve output reported an old answer as next active.
- **Symptom:** after reply/resolve, `Next active` pointed to a previously completed answer, causing repeated peer/model turns.
- **Reproduction:** retain completed answer in `inbox/new`, resolve current request, run `dequeueNext`; old answer became active.
- **Root cause:** `processedIds` only prevented injection dedupe. Queue advancement had no completion state and treated every unread actionable answer as pending; receive planning also ignored completion state.
- **Fix commit:** pending.
- **Regression test:** `tests/queue-lifecycle.test.mjs` marks old answer completed and asserts pending question becomes active; `tests/receive-loop.test.mjs` asserts completed answers are not re-injected; reply/resolve lifecycles now mark completion after successful read.
- **Validation:** 140/140 tests; TypeScript clean; lifecycle E2E passed.
- **Residual risk:** completion IDs are bounded to 500; exceptionally old retained mail can re-enter after eviction and needs explicit history filtering.

### 2026-08-02 — Existing-root authentic E2E caught late duplicate answer

- **Detection source:** real-root authentic E2E using `bpat` and `bi-boss` with `openai-codex/gpt-5.4-mini`.
- **Symptom:** after ping reply/resolve and later conversation steps, `bpat` received a second `answer` body `ack` referencing the original ping. Strict actionable-leftover assertion failed.
- **Reproduction:** run `run-authentic-existing-root-e2e.sh`; preserve old mailbox/history; complete ping → reply → resolve, then hold quiet window. Late duplicate answer appears.
- **Root cause:** not yet proven; likely an already-running model turn emitted a stale reply after original message completion. Bridge has no hard cancellation/runaway-send guard.
- **Fix commit:** pending diagnosis.
- **Regression test:** existing-root authentic E2E strict zero-actionable assertion; artifact message retained under the run's unique subject tag.
- **Validation:** attach, bidirectional send/read/reply/resolve passed; run intentionally failed before urgent/reconnect completion due actionable leftover.
- **Residual risk:** urgent intervention cannot claim cancellation until explicit model-turn stop semantics exist.

### 2026-08-03 — Stale active mailbox hijacked natural conversation

- **Detection source:** natural existing-root E2E with `bpat`, `bi-boss`, and `e2e-referee`.
- **Symptom:** harness appeared to show `bi-boss` ignoring a new RPS invitation; inspection showed `bpat` answered an old unresolved RPS invitation instead.
- **Reproduction:** retain old `bpat/inbox/new` RPS request and stale `activeId`, then start fresh natural prompt. `bpat` emits answer with old refs; no fresh invite reaches `bi-boss`.
- **Root cause:** old actionable mail remained pending; `activeId`/processed state had no stale-work boundary. Harness also used broad keyword matching and misclassified old answer as fresh invite.
- **Fix commits:** `fd14d95`, `2483e42`, `d2b92df`, `029c6c3`, plus current-work status diagnostics.
- **Postmortem:** `docs/postmortems/2026-08-03-stale-active-mailbox-natural-e2e.md`.
- **Regression tests:** receive/state takeover preservation, explicit target guards, active-work diagnostics, isolated exact natural E2E.
- **Validation:** 149/149 tests; TypeScript clean; exact natural sequential E2E passed twice; reused-root mode remains diagnostic.
- **Policy:** never auto-clear unfinished active work. Resume, explicitly resolve/abandon, or quarantine.
- **Residual risk:** urgent preemption while another active model turn exists remains post-baseline work.

## Process rule

No future bug fix is complete until this file and an applicable postmortem are updated, with a regression test named or linked.

### 2026-08-02 — Graceful peer shutdown left stale connected status

- **Detection source:** user edge-case review.
- **Symptom:** peer session quit, but remote status bar retained the peer as connected and emitted no disconnect notification.
- **Root cause:** session shutdown released local ownership but sent no peer lifecycle notice; housekeeping handled only `attached` events.
- **Fix commit:** pending.
- **Regression test:** `tests/pi-extension-peer.test.mjs` covers disconnect notice, offline status, and shutdown notice; `tests/pi-bridge-view.test.mjs` covers detached housekeeping classification.
- **Validation:** pending final commit; graceful shutdown covered. SIGKILL remains a separate presence-reconciliation case.
- **Residual risk:** abrupt kill or host crash cannot send a notice; presence TTL/reconciliation remains ADR-0003 follow-up.

### 2026-08-02 — Model context omitted kind/priority policy and send evidence omitted priority

- **Detection source:** user inspection of live Pi context.
- **Symptom:** injected bridge context described routing but not message-kind/priority semantics; sent-message evidence showed kind but not effective priority.
- **Root cause:** policy lived in documentation/tool schemas only; model-facing context and textual evidence were incomplete. Running Pi also retained pre-change extension code until restart.
- **Fix commit:** pending.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts policy instructions in model context; `tests/mailbox.test.mjs` asserts priority in send evidence.
- **Validation:** 133/133 tests; TypeScript clean.
- **Residual risk:** live sessions require restart to load extension changes; future runtime diagnostics should expose loaded extension revision.

### 2026-08-02 — Collaboration stalled behind stale activeId

- **Detection source:** user report plus live mailbox/state inspection.
- **Symptom:** peer received notification but did not start work on next request.
- **Evidence:** `bpat` and `bi-boss` had `activeId` values whose messages were already in `inbox/cur`; both had sent valid replies. Subsequent normal actionable messages therefore had `triggerTurn=false`.
- **Root cause:** active state was not reconciled when a message left `inbox/new` through an alternate reply/send path.
- **Fix commit:** pending.
- **Regression test:** `tests/receive-loop.test.mjs` verifies stale active IDs are cleared before planning the next batch.
- **Validation:** 134/134 tests; TypeScript clean.
- **Residual risk:** agents should still reply and resolve the original ID; model-facing context now states this workflow. Reconciliation is safety net.

### 2026-08-02 — Resolve guidance used wrong tool parameter name

- **Detection source:** live collaboration drive; first resolve call failed validation.
- **Symptom:** model-facing guidance said `messageId`, but `amq_bridge_resolve` schema requires `id`.
- **Root cause:** semantic message identity name differed from tool parameter name.
- **Fix commit:** pending.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts context uses `amq_bridge_resolve({ id: messageId })`.
- **Validation:** pending final commit; corrected retry succeeded and resolved the message.
- **Residual risk:** running Pi sessions retain old context until restart/refresh.

### 2026-08-02 — Explicit inbox reads exposed processed mail and sustained collaboration loop

- **Detection source:** real-peer game drive and user observation.
- **Symptom:** peers reread old RPS messages, reasoned repeatedly, and emitted 59 new messages; urgent stop did not cancel active generation.
- **Evidence:** `bpat` reached 59 messages in `inbox/new`; `bpat.activeId` pointed to completed non-actionable answer `15:25:49`; `bi-boss.activeId` pointed to completed answer `15:35:10`. Stop messages remained pending.
- **Root cause:** processed IDs only guarded watcher reinjection. Inbox tool/command exposed processed `inbox/new` messages; active reconciliation treated any present message as active, including answers/status.
- **Fix commit:** pending.
- **Regression test:** receive-loop active reconciliation; T6 inbox filtering assertions.
- **Validation:** 134/134 tests; TypeScript clean.
- **Residual risk:** urgent priority remains wake semantics, not hard cancellation of an already-running model turn. Manual process stop remains emergency control.

### 2026-08-02 — Detach command gives no immediate feedback and blocks input

- **Detection source:** user real-session observation during manual detach.
- **Symptom:** no `Detaching...` message appears; user input remains blocked while detach cleanup runs.
- **Reproduction:** run `/amq-bridge detach` while watcher/peer cleanup is active.
- **Root cause:** command awaits full detach lifecycle and invokes synchronous cleanup (`spawnSync`) before returning control to Pi UI.
- **Fix commit:** pending.
- **Regression test:** pending; must assert immediate detaching feedback, asynchronous cleanup, completion/failure notification, and serialized reattach.
- **Validation:** not fixed yet.
- **Residual risk:** background cleanup must preserve ownership/re-attach ordering and report failures visibly.

### 2026-08-02 — Disconnect warning appeared but status bar kept peers online

- **Detection source:** user real-session observation after both peers quit.
- **Symptom:** `AMQ peer disconnected: <handle>` appeared, but status bar still displayed peers normally connected.
- **Root cause:** status rendering trusted persisted roster status instead of reconciling peer availability from live AMQ presence.
- **Fix commit:** pending.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts live peer-status reconciliation uses `listPresence`.
- **Validation:** 134/134 tests; TypeScript clean.
- **Residual risk:** presence can lag briefly; explicit detached status remains authoritative until reconnect notice.

### 2026-08-02 — Stale active presence kept quit peers online

- **Detection source:** user restart observation plus live AMQ presence inspection.
- **Symptom:** no peer sessions existed, but status bar still showed `bpat` and `bi-boss` as online.
- **Evidence:** AMQ presence reported both `status=active` with `last_seen` several minutes old; ownership state was null for both peers.
- **Root cause:** status rendering trusted AMQ presence status without applying freshness/TTL validation. Quit peers left stale active presence records.
- **Fix commit:** pending.
- **Regression test:** `tests/pi-extension-peer.test.mjs` asserts presence TTL and freshness checks.
- **Validation:** 134/134 tests; TypeScript clean.
- **Residual risk:** a peer can appear online for the TTL window after abrupt exit; explicit detached notices still provide immediate offline state when loaded.
