# ADR 0005: Bounded nested active-work stack

## Status

Accepted — implemented for Pi + AMQ filesystem baseline.

## Context

The scheduler stored one `activeId`. While `bop` handled a parent request from `bip`, it sent a child question to `pib`. `pib` answered, but the bridge marked that answer processed and suppressed its turn because the parent remained active. The parent waited for a response that had already arrived: nested peer workflows deadlocked.

The same model made trusted urgent messages wake a turn without becoming the exact active target. Reply/resolve therefore rejected the urgent ID, so urgency and exact-target safety contradicted each other.

Weakening exact-target validation would reopen historical/wrong-message completion defects.

## Decision

Persist a bounded LIFO active-work stack while retaining `activeId` as a compatibility mirror of the top frame:

```json
{
  "activeId": "urgent-id",
  "activeStack": ["parent-id", "continuation-id", "urgent-id"]
}
```

Rules:

1. New ordinary actionable work activates only when stack is empty.
2. A same-thread `answer` or `review_response` from a known peer with fresh presence is a continuation. It pushes above current work and triggers a turn.
3. A trusted urgent actionable pushes above current work and triggers a turn.
4. Reply/resolve still accepts only stack top.
5. Completion removes its exact frame atomically. If newer work was pushed concurrently, newer frames remain. Otherwise previous frame resumes.
6. Queue advancement uses compare-and-set and activates only when stack is empty.
7. Stack depth is capped at 8. Overflow never kills watcher: its ID is atomically deferred, current top is woken to free capacity, and deferred work auto-promotes on the next completion.
8. Startup reconciles completed top frames before entering blocking watch and emits a recovery turn for resumed work.
9. `activeId`-only v3 state remains valid. First stack mutation imports it as one frame.
10. Scheduler/core module changes require full Pi process restart. `/reload` may retain cached dependency modules and is not an upgrade boundary.

## Trust

Urgent promotion uses existing trusted-peer policy. Continuation promotion is narrower but distinct: sender must already be in current peer roster and have fresh live presence; kind and thread equality alone are insufficient.

## Failure ordering

- Promotion persists stack frame before processed-id record. Crash may re-notify, but cannot lose promoted work.
- Completion records terminal state before removing frame. Startup removes completed top frames after crash.
- Completion removes requested frame from any stack position to preserve urgent work pushed after target validation.
- Dequeue cannot overwrite a concurrently promoted frame.

## Consequences

Positive:

- nested agent-to-agent workflows continue without manual wake;
- urgent work becomes an exact valid target and resumes interrupted work afterward;
- exact-ID completion guard remains strict;
- completed-parent/preemption races preserve newer work;
- bounded overflow cannot stop watcher.

Costs and limits:

- LIFO scheduling can delay parent work; depth cap bounds interruption growth;
- continuous urgent traffic may still delay lower frames, though new promotions stop at depth 8;
- same-thread continuation requires live known-peer evidence;
- no hard cancellation of an already-running provider turn; urgent work runs at next available turn;
- full process restart required after scheduler upgrades.

## Verification

- state push/pop/remove/CAS/recovery tests;
- continuation and urgent policy tests;
- deterministic completion-vs-promotion race tests for reply and resolve;
- overflow test proving no fatal watcher failure;
- startup recovery test proving resume occurs before blocking watch;
- real Pi nested debate/urgent acceptance required before release claim.
