# Feature: Autopilot Circuit-Breaker

<!-- toc -->
- [Trip conditions](#trip-conditions)
- [Wiring status](#wiring-status)
- [Action on trip](#action-on-trip)
- [Why this is the right autopilot exception](#why-this-is-the-right-autopilot-exception)
- [The runner-level breaker (a different failure, a different layer)](#the-runner-level-breaker-a-different-failure-a-different-layer)
- [Two more things a long-running runner needs](#two-more-things-a-long-running-runner-needs)
<!-- /toc -->

**Pattern**: autopilot runs with zero interaction, which is exactly when a silent failure loop is most expensive  -  an agent can burn a budget re-attempting the same broken fix, or thrash between two phases, with nobody watching. A circuit-breaker converts "keep going no matter what" into "keep going until a defined unsafe condition, then halt and hand back to the user." This is the sanctioned autopilot pause (same class as the Phase 5 channels pause): the run stops, records why, and waits for an explicit `resume`.

**Gated by `prefs.global.autopilotCircuitBreaker`** (`enabled` default true, `identicalFindingCycles` default 2, `maxReworkCycles` default 3; `schemas/prefs.schema.json`). Halting is always safe, so the breaker itself defaults on. Disable per-run only with an explicit override. Complements, does not replace, the existing autopilot safety rules (build-fail max 3 retries, Phase 4 blocking-finding rework, destructive-op confirmations).

## Trip conditions

Any one trips the breaker. All are evaluated from `agent-state.json` + telemetry, no extra model calls.

| # | Trigger | Detection | Default threshold |
|---|---|---|---|
| 1 | **No-progress stall** | two consecutive phase checkpoints with no forward transition (same `phase` + `step` + `reworkCount`, no new artifact/commit) | 2 checkpoints |
| 2 | **Identical repeated failure** | same normalized build-error signature, or the same Phase 4 finding fingerprint, recurs across consecutive Phase 3 rework cycles | 2 cycles |
| 3 | **Rework storm** | Phase 4 -> Phase 3 rework cycles exceed the cap (distinct from the build-retry cap) | `maxReworkCycles` = 3 |
| 4 | **Cost drift** | cumulative spend crosses the `costBudget` ceiling, or the projected next-phase spend would exceed it, after model-fallback has already downgraded | `costBudget` ceiling |
| 5 | **Merge/rebase conflict** | Phase 4 push blocked by a conflict that requires history reconciliation | any conflict |

Trigger 2 is the key addition over the plain build-retry cap: a build can "fail differently" three times (legitimate iteration) or "fail identically" twice (stuck). Only the identical-failure case is a stall; the retry cap catches the rest.

## Wiring status

| Trigger | Evaluated by | Status |
|---|---|---|
| 2, finding half | `review-delta.mjs` exit 3 at Phase 3 Step 3.8: a blocking/important finding whose `fingerprint` (finding-fingerprint.mjs) stays in the accepted set for `identicalFindingCycles` consecutive rounds | **code** (v16.20.0) |
| 3 | Phase 3 re-entry item 6: the `retryCount === 3` hard-kill records the trip | **code** (v16.20.0) |
| 2, build-error half | needs a build-log signature normaliser | documented behaviour, no script yet |
| 1 | needs checkpoint-to-checkpoint artifact diffing | documented behaviour, no script yet |
| 4 | belongs to `cost-budget-check.mjs` | documented behaviour, no script yet |
| 5 | Phase 4 push | documented behaviour, no script yet |

State shape: `state.circuitBreaker = {tripped, trigger, detail, checkpoint: {phase, step, iteration}, trippedAt, counters: {identicalFindingCycles, reworkCycles}}` (`schemas/agent-state.schema.json`). The per-round classification the finding half reads lives in `state.reviewIterations[i].delta` (`new`, `stillPresent`, `resolved`, `downgraded`, `recurrence`, `plateau`). `smoke-autopilot-circuit-breaker.sh` asserts the schema fields, the scripts and the phase wiring, not only this prose.

## Action on trip

1. Set `agent-state.json.circuitBreaker = {tripped: true, trigger: <#>, detail, checkpoint}` and flip `autopilot` handling to paused (the run does not continue unattended).
2. Emit one actionable line per the progress contract: what tripped, the evidence (error signature / cycle count / spend vs ceiling), and the single next action (`resume #N` after a fix, or `kill #N`).
3. Never auto-resolve the underlying cause  -  no force-anything, no conflict auto-merge, no budget self-raise. The breaker hands control back; it does not paper over the problem.
4. `resume #N` clears `circuitBreaker.tripped` (keeping `counters`) and continues from the recorded checkpoint. If the same trigger fires again immediately, the breaker re-trips (no silent bypass).

## Why this is the right autopilot exception

Autopilot's contract is "no interaction on the happy path." The circuit-breaker fires only off the happy path, where continuing unattended is the *less* safe choice: repeating a proven-broken action, or spending past a ceiling the user set, is not autonomy, it is a runaway. Halting with a precise reason is cheaper than the tokens (and trust) a silent loop burns.

Inspired by autonomous-loop operators that gate on explicit stop conditions (no-progress, identical-stack-trace repetition, cost drift, conflict blocking), adapted to this pipeline's phase/rework/state model.

## The runner-level breaker (a different failure, a different layer)

Everything above is the IN-RUN breaker: one run, one state file, triggers that
mean "this task is not converging". It cannot see the failure a server actually
hits, which is not about any one task.

An expired token, a `claude` binary that no longer launches, a full disk: every
item fails identically, a few minutes apart, until the queue is empty. Nothing
above notices, because each individual run failed for its own apparently local
reason. What is left afterwards is a ledger of failures with no indication of
which came first and no queue to resume.

`autopilot-runner.mjs` therefore refuses to take NEW work after
`MA_AP_BREAKER_LIMIT` consecutive attempts produced nothing (default 3, zero
disables). The count is CONSECUTIVE and read from `attempted.jsonl` rather than
kept as a counter, because a counter is a second piece of state that can
disagree with the ledger - and the ledger is what a person reads when they ask
why the runner stopped. The breaker and the rate-limit backoff read the rows of
a trailing window (`MA_AP_ATTEMPT_WINDOW_SEC`, 7 days by default), not a fixed
number of rows, so the `blocked-*` rows a refused tick writes cannot push the
failures out of view. One run counts once, by its `taskId`, even when it wrote
two rows (`timed-out`, then the outcome a later tick retired it with).

What counts is the outcome contract in `_autopilot-outcomes.mjs`, the one
module the runner, intake and this breaker all read: every RETRYABLE outcome
except `died` and `rate-limited` is a failure, and three kinds of row are
deliberately not:

- PARKED outcomes (`needs-input`, `awaiting-answer`, `verification-failed`) are
  runs that did their work and are waiting for a person. That is the system
  behaving correctly, and counting it would stop a machine whose only problem is
  that somebody has not answered yet. They end a failure chain. The exception
  is a publish refusal that describes the machine: a `verification-failed` row
  whose `publishGate` is `remote`, `push` or `pr-create` means the push URL
  could not be resolved or the host refused the push or the PR, which every next
  item meets at the same step. Those count (`ENVIRONMENTAL_PUBLISH_GATES` in the
  runner). The runner writes `publishGate` on every row the publish step
  refused.
- `blocked-*` items never ran at all - arming refused them - so they say nothing
  about whether a run would have worked. They neither count nor end a chain,
  and they do not move the cooldown clock either: arming writes one on every
  tick it refuses, which would otherwise keep the cooldown from running out.
- `died` (a run the supervisor lost, e.g. across a reboot) and `rate-limited`
  (which has its own backoff, doubling per consecutive limit) are skipped the
  same way: neither counts, and neither passes for the success that ends a chain.

When it opens, the reason goes into `queue.json.blockedReason` (so
`autopilot-status` shows it rather than reporting an idle queue), a
`breaker-open` line goes into `ticks.jsonl`, and the log names the outcome that
opened it plus the override.

It is HALF-OPEN rather than latched, and that distinction is load-bearing. The
breaker returns before the only writer of `attempted.jsonl`, so a version that
simply stopped could never record another attempt, the consecutive count could
never fall, and the runner would be off for good - while its own log said
"until an attempt succeeds". After `MA_AP_BREAKER_COOLDOWN_SEC` (30 min by
default) one probe is let through: a recovered machine succeeds and the count
resets on its own, and a still-broken one fails, becomes the newest attempt, and
closes the breaker for another cooldown. One wasted run per cooldown is the
price of not needing a person to notice.

## Two more things a long-running runner needs

**The log.** launchd appends the runner's stdout to `runner.log` forever. On a
laptop that file is never read; on a machine ticking every few minutes for
months it is what fills the disk, and a full disk stops the runs the log existed
to record. Each tick truncates it past `MA_AP_LOG_MAX_BYTES` (default 5MB),
keeping the last 256KB in `runner.log.1`.

Truncating the same inode is the mechanism, not renaming: launchd holds the file
open with `O_APPEND`, so a rename leaves that descriptor writing into the renamed
file while the new one stays empty forever. That is the rotation that looks
correct and silently stops logging.

**Telemetry.** `ticks.jsonl` gets one JSON line per tick - what the tick did,
what it took, what came out, how many consecutive failures preceded it. The
human log answers "what happened just now"; this answers "how has this been
behaving for a week", which is the question a server raises and prose cannot be
asked. Past `MA_AP_TICKS_MAX_BYTES` (2MB) only the newest `MA_AP_TICKS_KEEP`
lines (5000) are kept.

**The attempt log.** `attempted.jsonl` is the queue's memory, so past
`MA_AP_ATTEMPTED_MAX_BYTES` (4MB) it is compacted, not cut. Rows from the last
30 days stay whole. An older row stays when intake or the cleanup report still
reads it: the newest row of its item, and every non-`blocked-*` row of an item
that is not finished (its attempt count, research rounds and the task ids that
tie a parked run to its worktree). The full log before compaction is kept in
`attempted.jsonl.1`.
