# Feature: Autopilot Operations

<!-- toc -->
- [Awake agent](#awake-agent)
- [Sleep lock](#sleep-lock)
- [Credential probe](#credential-probe)
- [Parallel cap](#parallel-cap)
- [Cleanup report](#cleanup-report)
- [Daily digest](#daily-digest)
- [Status fields](#status-fields)
- [Config validation and item identity](#config-validation-and-item-identity)
- [Reference](#reference)
<!-- /toc -->

**Pattern**: the runner does more than run items. A machine that works through a
queue overnight also has to stay awake while it works, find out before it arms
that it can read its own credentials, clean up after itself, and say what it did.
The awake agent belongs to the mode and runs from `autopilot-on` to
`autopilot-off`; the rest are owned by the runner tick and run outside every
child session. The awake agent, the sleep lock and the credential probe are on
by default, since a queue that stops while the Mac sleeps, a run that sleeps
mid-session, or one that cannot read its own credentials fails for a reason
nobody chose; the other three are off, unlimited or read-only unless the config
turns them on.

| Operation | When | Default | Where it shows |
|---|---|---|---|
| Awake agent | from `autopilot-on` to `autopilot-off` | on, display may sleep | `awake` in `autopilot-status --json` |
| Sleep lock | from the launch to the end of the tick | on | `inhibitorHeld` in `autopilot-status --json` |
| Credential probe | `autopilot-on`, and every tick for the item's source | on | the arming refusal, `blockedReason` |
| Parallel cap | before an item is taken | unlimited | `maxParallelAgents` |
| Cleanup report | first tick of each local day | dry run | `gc-report-<date>.json` / `.md`, `lastGcReport` |
| Daily digest | first tick of each local day | off | `digest-<date>.json` / `.md`, `lastDigest` |

Code: `scripts/_autopilot-ops.mjs`, called from `autopilot-runner.mjs`; the awake
agent is `scripts/autopilot-awake.mjs`; the probe lives in `autopilot-arming.mjs`
and `lib/credential-store.sh probe`.

## Awake agent

launchd's `StartInterval` does not wake a sleeping Mac; it fires on the next
wake. With only the sleep lock below, the machine can sleep between ticks, the
next tick never fires, and the queue stops with nothing to say so. On a machine
kept as a server that is the common case, not an edge.

So `autopilot-on` installs a second launchd agent next to the schedule,
`com.multi-agent.autopilot-awake`, rendered from
`templates/multi-agent-autopilot-awake.plist.template`. It runs
`/usr/bin/caffeinate -s -i` with `KeepAlive`, for as long as it is loaded, and
`autopilot-off` boots it out and deletes its plist. Loading it is the same
decision as turning the mode on; it never outlives the mode.

- `-s` prevents system sleep, and macOS honours it **on AC power only**.
- `-i` prevents idle sleep, and that one holds **on battery too**. A laptop left
  unplugged with the mode on stays awake until it is plugged in or the mode is
  turned off; `awake.enabled: false` removes the agent and leaves only the
  in-flight lock.
- `-d` keeps the display on as well, and is added only with
  `awake.display: true`. `autopilot-on` asks; the default is no.

The display option costs energy and screen wear, and autopilot itself does not
need it: builds, `simctl` screenshots and headless browsers all work with the
display asleep. Display sleep also does not lock the keychain. System sleep can
(a login keychain set to lock on sleep), which is what `-s -i` prevents, and why
the display is the part left optional.

`autopilot-awake.mjs install` writes the plist 0644, boots out a job already
loaded so a changed answer applies now, then `launchctl bootstrap gui/<uid>`
with `launchctl load` as the fallback. Without `--display` / `--no-display` it
reads `awake.display` from the config. `autopilot-awake.mjs status --json` is
`{held, display}`: held when `launchctl list com.multi-agent.autopilot-awake`
reports a running PID, or, when launchd cannot say, when a `caffeinate -s -i`
process is running (the in-flight lock carries `-w <pid>` and is never counted);
the display answer is read from the installed plist. `MA_AP_LAUNCH_AGENTS`,
`MA_AP_LAUNCHCTL_BIN` and `MA_AP_PS_BIN` replace the directory and the binaries,
which is how the tests run it against a temp directory.

## Sleep lock

launchd's `StartInterval` does not wake a sleeping Mac, and a Mac that falls
asleep in the middle of a run leaves a session half done. So from the moment the
runner launches a session until the tick ends - research rounds, the resume and
the publish step included - it holds a sleep inhibitor:

`caffeinate -s -w <runner pid>`. `-s` asserts against system sleep and macOS
honours it **only on AC power**; `-w` ties it to the runner. It covers the run in
flight whether or not the awake agent is installed, and holds nothing on battery:
keeping an unplugged laptop awake is the awake agent's `-i`, chosen with the mode,
not something a single run decides. The display is not held on. The platform is
macOS only (ADR-0012); anywhere else the tick holds nothing and logs
`sleep lock: not held on <platform> - macOS only, see ADR-0012`.

The inhibitor is spawned in its own process group and released on every way out
of the tick, a failed launch and an exception included: the runner kills the
process group of the pid it spawned, never a name. It cannot outlive the runner
even when the runner is killed first, because `-w` ends it when the runner pid
goes. `inhibitor.json` in the autopilot root records the pid;
the next tick removes a record whose runner is gone, and stops its process only
when that pid is still running the recorded binary. `MA_AP_INHIBIT_BIN` replaces
the binary, which is how the tests run it without the real one.

A tick with nothing to take holds nothing, and neither does a parked run a person
resumes by hand: that session is theirs.

## Credential probe

`prefs.global.serviceStatus` says what a credential did the last time it was
used. It does not say whether the credential can be read **now, by a background
process**. The probe answers that, without reading the value into any process
that could print it:

| Service | Needed when a repo | Probe |
|---|---|---|
| GitHub | takes GitHub issues, or publishes to GitHub (the default host) | `gh auth status` |
| Jira | takes Jira items | `prefs.global.keychainMapping.jira` and `prefs.global.hosts.jira` are set, then `credential-store.sh probe jira` |

`credential-store.sh probe <key>` resolves the key through `keychainMapping` the
same way `get` does, reads the item with its output sent to `/dev/null`, and
prints one word: `readable` (exit 0), `missing` (exit 1) or `locked` (exit 5).
On macOS the item is really read, not only listed, because an item whose access
control still asks before handing the value over is exactly the one a background
tick cannot use; `errSecInteractionNotAllowed` is what tells a locked keychain
from a missing item. A store that has not answered in 20 seconds
(`MA_AP_PROBE_TIMEOUT_MS`) is reported as waiting on a prompt nobody can answer.

Two places run it:

- **`autopilot-on`** runs `autopilot-arming.mjs probe` after the config is
  written and before the schedule is offered. A missing, unmapped or locked
  credential refuses **Start**, naming the logical key and the fix; "Configure
  only" stays available.
- **Every tick** runs `autopilot-arming.mjs check --source <source> --repo
  <owner/name> --probe` for the item it is about to take. The tick runs in the
  background session, and a login keychain set to lock on sleep or after a
  timeout is readable from the terminal that armed the mode and locked for the
  tick that fires after the Mac slept. `--repo` adds the credential the item's
  repo publishes with: `gh auth status` for a GitHub host, so an expired `gh`
  stops a Jira item before its run is paid for rather than at the push. A
  Bitbucket push goes through the git credential helper, which has no probe
  short of a network call, so nothing is added there. A refusal is recorded as
  `blocked-credential` for that item only; a locked Jira token does not stop
  GitHub items.

On macOS the probe also reads the login keychain's lock settings
(`security show-keychain-info`). Locking on sleep or after a timeout is a
**warning**, not a refusal - the credential is readable at that moment - and it
names the fix: Keychain Access, login, Change Settings, or `security
set-keychain-settings` with no options.

The probe runs in the runner, never in a child: the unattended guard still
refuses any Keychain read inside a session.

## Parallel cap

`maxParallelAgents` in `config.json` caps runner-launched sessions that are
working at once.

What the runner does without it, unchanged: `runner.pid` is exclusive, so the
runner supervises one session at a time; `slots` counts the entries it is
running and does not count parked ones, because a parked run is waiting for a
person, not working. A parked run a person answered can therefore be busy again
alongside the next item.

With `maxParallelAgents: N`, the count before an item is taken is the entries the
runner is running plus every parked entry whose session reports `busy` or
`running` in `claude agents --json`. At N, nothing is taken that tick
(`parallel-cap` in `ticks.jsonl`). An agent list that cannot be read counts as
unknown and takes nothing, since a cap that guesses zero is not a cap. Absent or
null is unlimited and adds no check at all: the cap is a ceiling, never a raise.

The claim itself is an exclusive create of `runner.pid`, so two ticks that both
passed recovery cannot both launch; the loser logs `another tick claimed first`
and leaves. The claim is written whole to a temp file and hard-linked into place
(`link()` fails when the name is taken), so no tick ever reads an empty or
half-written claim, and every other file the runner owns is replaced by temp
file and rename. A `runner.pid` that does not parse is removed with a log line
once it is older than a minute, rather than left to refuse every claim; a
younger one is treated as a claim being written and the tick leaves.

## Cleanup report

On the first tick of each local day (stamped in `daily.json`) the runner runs
`gc-abandoned.sh` over **the worktrees it created**, and nothing else:

- The list is `worktrees.jsonl`, appended when a run's state is matched by the
  session id the runner launched it with - the same proof the runner's own
  retirement uses before it removes anything (`createdWorktree`). A hand-made
  worktree, or an attended run's, never enters it. The proof is checked again
  when the report is built: a path is listed only while the newest run state
  naming it carries the ledger row's `sessionId`, so a worktree an attended run
  later opened at the same path is not the runner's to clean up.
- A worktree the queue still holds is taken off the list, and so is one whose
  item's newest attempt in `attempted.jsonl` is PARKED (`awaiting-answer`,
  `needs-input`, `verification-failed`): a parked run leaves the queue once its
  session ends, and its worktree is what the person who looks at it needs.
- `gc-abandoned.sh --only <list> --report <file>` applies its usual rules to the
  listed paths only: a run waiting on you is kept (an open PR, the Commit or
  Report phase, `awaiting_input`, and any state carrying `verificationFailed` or
  `waitingFor`, whatever its `status`), uncommitted work is stashed and
  its worktree kept, nothing outside `<repo>/.worktrees/` is touched, and an
  unlisted run's state is not rewritten.

`gc-report-<date>.json` (every candidate: task, path, action, reason, idle days,
size) and `gc-report-<date>.md` land in the autopilot root at 0600. The newest 14
are kept.

The report is a dry run. `gc: { autoDelete: true }` adds `--yes`, still limited to
the list. A job that fails is stamped anyway and tried again the next day, not on
every tick.

## Daily digest

Off by default. `digest: { enabled: true }` computes, on the first tick of each
local day, the trailing 24 hours from `attempted.jsonl` and `queue.json`:

- items attempted (distinct items that ran; an arming refusal did not run)
- outcomes by class (`terminal`, `retryable`, `parked`, `blocked`) and by outcome
- PRs opened
- parked items waiting for an answer
- cost, under the spend ceiling's rule: a row's `usd` is the run's cumulative
  cost when it was written, so a run recorded more than once (parked and later
  settled, researched and then resumed) counts once, at its largest figure, keyed
  by the runner's `taskId`; a run no row measured is counted as unknown, never as
  zero (`_cost.mjs` `attemptSpend`, shared with `autopilot-arming.mjs`
  `spendWindow`)

It is written to `digest-<date>.json` and `.md`. It is **sent** only when
`reportChannels` is also true and `channels.mode` is `configured`, through the
same path as the post-PR report (`reportDigest` in `autopilot-publish.mjs`): the
text passes the outbound gate first, goes to the script as a 0600 body file, and
the script runs with `MULTI_AGENT_UNATTENDED` unset. The `jira` target posts one
comment on `digest.jiraIssue`, a standing issue the operator chooses, since a
digest belongs to no single item; without one it is reported as skipped. `pr`,
`confluence` and `wiki` are skipped. What was sent is recorded under `sent` in the
day's JSON.

## Status fields

`autopilot-status --json` (and `status.json`) carry:

| Field | Type | Meaning |
|---|---|---|
| `inhibitorHeld` | boolean | `inhibitor.json` names a pid that is alive |
| `lastGcReport` | string | path of the newest `gc-report-<date>.json`, empty when none |
| `lastDigest` | string | path of the newest `digest-<date>.json`, empty when none |
| `maxParallelAgents` | integer or null | the configured cap, null when unlimited |
| `awake` | object | `{held, display}`, both booleans: the awake agent is running, and whether it keeps the display on |

## Config validation and item identity

`config.json` is checked against `schemas/autopilot-config.schema.json` each
time it is loaded (`_autopilot-config.mjs`), because its numbers are ceilings
and a value that is not a number compares false with everything. A config the
schema refuses stops the queue and names the field: the runner takes and
publishes nothing and writes the reason to `queue.json.blockedReason` (a
`config-invalid` tick), intake queues nothing and exits 3, and
`autopilot-arming.mjs check` answers `blockedBy: config`.

A GitHub item's id is `owner/repo#N`, so two repos that share a short name
never share an item's history. A row recorded under the older short id
(`repo#N`) is still read for that item, erring toward leaving it alone. A Jira
issue that more than one repo's search returns is queued once, for the first
repo in config order; each other match is listed in `queue.json.dropped` with
the repo it was queued for. Give each repo a `jiraJql` that matches only its
own issues when several repos take Jira items.

## Reference

Scripts: `autopilot-runner.mjs` (`dailyJobs`, `recordWorktree`, the claim and the
sleep lock around the run), `_autopilot-ops.mjs`, `autopilot-awake.mjs`, `autopilot-arming.mjs`
(`probe`, `check --probe`), `lib/credential-store.sh probe`, `gc-abandoned.sh
--only --report`, `autopilot-publish.mjs` (`reportDigest`), `autopilot-status.sh`.
Config: `schemas/autopilot-config.schema.json` (`awake`, `maxParallelAgents`,
`gc`, `digest`). Tests: `test/autopilot-awake.test.mjs`,
`test/autopilot-operations.test.mjs`, the `operations:` block in
`test/autopilot-runner.test.mjs`; smokes `smoke-gc-abandoned.sh` (section 7),
`smoke-gui-json-contract.sh` (section 2b).
