# Running multi-agent on a machine nobody is sitting at

This describes what a Mac has to have before an unattended run works, and what
breaks if it does not. **Nothing here is set up by reading it** - `install`
still writes nothing about permissions and no scheduler is loaded unless you
ask for one. That is deliberate: every item below changes what runs on a
machine without a human, and none of it should happen because someone ran an
installer.

Scope is macOS. Linux and Windows were removed on purpose (ADR-0012) and
nothing here reintroduces them.

The mechanical half of this page is `doctor --profile=server`, which checks
four of the conditions below and says which are missing. Read that first; this
page explains why each one matters and what to do about it.

## The failure this page exists for

An unattended run does not fail loudly. It **stops without saying so**, and
every way it does that looks identical from outside: a process that is running,
printing nothing, finishing never.

Four causes, in the order they bite.

## 1. The permission posture

autopilot spawns its child with `--permission-prompts none`. That stops Claude
Code from ASKING, and it does not grant anything - the tools the child then
calls still have to be allowed. On the maintainer's machine they are, because
`Bash`, `Edit`, `Write` and `Agent` sit in the global allow list from years of
interactive use. On a fresh machine they do not, and the first tool call ends
the run.

There is no DEFAULT that fixes this and should not be: writing a permission
posture into someone's settings during an install is the one action an
installer must never take silently. There is now a flag that asks for it:

```bash
npx @mmerterden/multi-agent-pipeline install --unattended --dry-run   # see it first
npx @mmerterden/multi-agent-pipeline install --unattended             # write it
```

It prints the whole profile with a reason per line before writing anything,
and writes it to `~/.claude/multi-agent-unattended.settings.json`, which the
autopilot runner passes to each run with `claude --settings`.
`~/.claude/settings.json`, the posture of every attended session, is not changed.
The profile (`pipeline/schemas/unattended-profile.json`) is
`permissions.defaultMode: "dontAsk"`, a narrow allow list (`Bash(git *)`,
`Bash(xcodebuild *)`, `Bash(xcrun simctl *)`, `Edit`, `Write` and the rest,
each with its reason), deny
rules for merging, pushing and reading `~/.ssh`, `~/.aws`, `~/.gnupg` and
`~/Library/Keychains`, and the toolkit's strict URL policy. `--unattended-sandbox`
adds an OS sandbox block (`sandbox.enabled`, `failIfUnavailable`,
`allowUnsandboxedCommands: false`, a network allowlist, `denyRead` for `~/.ssh`
and `~/.aws`).

It is additive over the profile file: an entry already there survives, a
narrower `Bash(...)` rule is kept alongside without standing in for the
profile's own entry, and a second run changes nothing. Permission lists merge
across settings files, so a whole-`Bash` allow in `settings.json` still reaches
an unattended run; the installer reports it and doctor warns, and neither
removes it. If an earlier install wrote the profile into `settings.json`, the
installer prints the move and removes exactly those entries from it. A profile
file that does not parse is refused rather than overwritten.

The permission list is the second line. The guard hook under
`MULTI_AGENT_UNATTENDED=1` and a separate macOS user are the first and the
boundary: `pipeline/multi-agent-refs/features/unattended-security.md`.

`doctor --profile=server` reports this as `unattended-permissions`, against the
same rule the writer applies.

## 2. The scheduler, and why it is a LaunchAgent

Something has to start the work. The template ships as a **LaunchAgent**
(`install/templates/multi-agent-autopilot.plist.template`), loaded in the
user's session at login - not a LaunchDaemon at boot.

That is a keychain decision, not a preference. Every credential the pipeline
reads lives in the login keychain, and the login keychain is **locked until the
user logs in**. A daemon started at boot gets a locked one: each credential
read fails, and those failures surface much later as 401s that blame the token.
The lock is the cause and nothing in the error says so.

So the machine needs to reach a logged-in session by itself:

- **System Settings → Users & Groups → Automatic login**, for the account that
  owns the agent.
- A **keychain with no lock timeout** for that account, or the agent survives
  a login and dies at the first idle period:
  ```
  security set-keychain-settings -l ~/Library/Keychains/login.keychain-db
  ```
  (no `-t`, so there is no timeout; `-l` still locks on sleep.)

Automatic login means physical access to the machine is access to the
credentials. On a server in a locked room that is the trade being made; on a
laptop it is not.

`doctor --profile=server` reports these as `scheduler` and `keychain-unlock`.

## 3. The unattended contract

`MULTI_AGENT_UNATTENDED=1` is how a run states that nobody is watching.
`multi-agent-refs/unattended-contract.md` lists exactly which entry points read
it and what each resolves to - four of them, named, because a claim of coverage
that is not true is worse than a short list.

The variable does **not** grant permissions and does not suppress errors. It
only stops a process from waiting for an answer that is not coming.

`doctor --profile=server` reports this as `unattended-contract`.

## 4. Which credential each phase needs

A server fails differently from a laptop here: an interactive run asks for a
missing token, an unattended one does not.

| Phase | Needs | Absent |
|---|---|---|
| 0 Init | `github` (issue intake), `jira` (ticket intake) | the run cannot resolve its input and stops at the picker |
| 1-2 Analysis, Plan | `figma` or `figma_mcp` when the task names a design | halts by contract rather than guessing at layout |
| 3 Dev | none beyond git access | - |
| 4 Review | `github` for PR comments | findings are computed and never posted |
| 5 Test | none | - |
| 6 Commit | git push credential | the branch stays local, which reads as "no PR yet" |
| 7 Report | `jira`, `confluence` as configured | the report is composed and dropped |

`credential-store.sh doctor` lists what is mapped. `doctor --probe` additionally
checks each one is alive, which costs network calls and is why it is opt-in.

## 4b. What the runner does when nobody is looking

Three behaviours that are invisible on a laptop and decide whether a server
survives a month.

**The log.** launchd appends the runner's stdout to `runner.log` forever. Each
tick truncates it past `MA_AP_LOG_MAX_BYTES` (5MB by default), keeping the last
256KB in `runner.log.1`. It truncates the SAME file rather than renaming it,
because launchd holds it open with `O_APPEND` - rename it and that descriptor
keeps writing into the renamed file while the new one stays empty, which is the
rotation that looks correct and silently stops logging.

**The breaker.** An expired token, a `claude` that no longer launches, a full
disk: every queued item fails the same way, minutes apart, until the queue is
empty and the ledger is a list of identical failures with no indication which
came first. After `MA_AP_BREAKER_LIMIT` consecutive attempts that produced
nothing (3 by default, 0 disables), the runner stops taking new work, writes the
reason into `queue.json` so `autopilot-status` shows it, and says what to check.
One successful attempt clears it. A run waiting for a person (`needs-input`,
`awaiting-answer`), an item arming refused (`blocked-*`), a run the runner lost
across a restart (`died`) and a rate limit (which has its own backoff) are
deliberately not failures.

**Telemetry.** `ticks.jsonl` gets one JSON line per tick. The human log answers
"what happened just now"; a server is only ever asked "how has this been
behaving for a week".

## 5. Self-hosted CI runner

ADR-0011 rejected a self-hosted runner because "a self-hosted runner on a
public repository executes code from any fork's pull request". Half of that has
changed and half has not, and the difference matters before wiring a runner to
a machine that holds credentials.

Measured, not assumed: the repo is **private**, and it has **one fork**, owned
by another account. "Zero forks" is a claim worth not making - GitHub requires
approval before a fork's pull request runs a workflow, but that protection is a
SETTING, and a self-hosted runner turns "someone approved a workflow run by
reflex" into code on this machine with this keychain.

So: a runner here is defensible on a private single-owner repo, and it is
defensible only while fork-PR approval stays on and every approval is treated
as a decision rather than a formality.

The same machine can host the runner:

```
mkdir ~/actions-runner && cd ~/actions-runner
# download and configure per GitHub's instructions for the repo
./svc.sh install && ./svc.sh start
```

Two things to keep in mind. The runner service runs in the same session as the
autopilot agent, so a long CI job and a pipeline run compete for the same CPU
and the same `~/.claude` state root. And the runner inherits that session's
keychain, which means a workflow can read the credentials - acceptable on a
single-owner machine, not on a shared one.

## What this page does NOT cover

- Turning any of it on. Every step here is the operator's to take.
- Any host other than macOS, and containers. Out of scope and staying there (ADR-0012).
- Multi-user servers. Everything above assumes one account owns the install,
  the keychain and the queue.
