# doctor - the check registry

<!-- toc -->
- [What the exit code means](#what-the-exit-code-means)
- [What BLOCK means, exactly](#what-block-means-exactly)
- [Who calls it, and what they do with the code](#who-calls-it-and-what-they-do-with-the-code)
- [The four severities](#the-four-severities)
- [The line shape](#the-line-shape)
- [It recommends, it never fixes](#it-recommends-it-never-fixes)
- [Checks](#checks)
- [The server profile (`--profile=server`)](#the-server-profile---profileserver)
<!-- /toc -->

Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
`doctor.mjs --list-checks` prints exactly the same set. The equality is checked
in both directions by `smoke-doctor.sh`: a check that ships without an entry
cannot be released, and an entry whose check was deleted cannot linger as orphan
prose. That is the mechanism that keeps this file from rotting into a list of
things the tool used to do.

## What the exit code means

| Code | Meaning |
|---|---|
| 0 | healthy - nothing above INFO |
| 1 | degraded - at least one WARN, no BLOCK |
| 2 | blocked - at least one BLOCK |
| 3 | usage error |
| 4 | indeterminate - the installed layout could not be resolved, so nothing was checked |

4 exists so that "I could not look" never borrows the exit code of "I looked and
it is fine". A consumer that treats 0 as healthy would otherwise read a broken
resolver as a clean bill.

## What BLOCK means, exactly

**A state in which a run will fail or leak. Not a state in which it will merely
be worse.** Without a closed definition "blocked" grows until nobody respects
exit 2, so only six checks can produce it: `host-platform`, `install-present`,
`script-surface`, `state-writable`, `prefs-valid` and `embedded-credentials`.
Every other check tops out at WARN however bad it looks.

## Who calls it, and what they do with the code

| Caller | When | On exit 2 |
|---|---|---|
| `/multi-agent:setup` | before the first question, and again at the end | the closing run is the evidence for or against "setup complete" |
| `/multi-agent:update` | after the install | report that the update left a state a run will fail from, rather than "updated" |
| `/multi-agent:sync` | step 0 | **stop.** Syncing a blocked install copies one fault onto five surfaces, and the copies are what people then debug |

`/multi-agent:sync` stops on exit 4 as well: the layout did not resolve, so
nothing was checked, and syncing from an unknown state is worse than not syncing
at all. Exit 1 never stops anything - a degraded install still syncs correctly,
and a gate that blocked on every warning would be routed around within a week.

## The four severities

| Severity | Meaning |
|---|---|
| BLOCK | a run will fail or leak |
| WARN | a run will work and be worse: a degraded capability, a stale copy, a rejected credential |
| INFO | a capability the user chose not to enable; nothing to fix |
| SKIP | not checked, and why |

`SKIP` is printed in the same list, in the same position, and the summary always
carries its count (`... 2 not checked, 15 checks`). Absence is a finding: a run
without `--probe` exits 0 **and** prints `SKIP credential-liveness`, so "healthy"
and "not looked at" are never the same output.

The INFO / WARN boundary is already decided in `lib/credential-inventory.sh` and
this tool only renders it: an unmapped key is INFO because it is a capability the
user chose not to enable, while `auth-rejected`, `malformed` and
`tier-1-no-grant` are WARN because the configuration made a claim the service
refused. **A missing optional capability is information; a mapped credential the
service rejected is a warning. The difference is whether the configuration
asserted anything.**

## The line shape

```
<SEV> <id>  -  <what is wrong>  -  <single imperative step>
```

The third field is one step, and its first word comes from a closed list:
`run`, `set`, `map`, `revoke`, `install`, `remove`, `free`, `export`, `merge`,
`record`, `re-run`. `consider`, `may want` and `should probably` are refused by
the gate. One problem, one step: a reader with five suggestions does nothing.

## It recommends, it never fixes

Nothing here rewrites a remote URL, edits `settings.json` or touches a token.
A remote whose URL carries a token may be the only credential that repo has, the
remote may be a mirror a script depends on verbatim, and doctor can run inside a
checkout the user does not own. Silent repair breaks all three.

For an embedded credential the honest single step is **not** "hide it". A token
that reached `.git/config` is already burned: it is in the shell history and
readable by anything that can read the working tree. The step is to revoke it at
its host; everything after that is behind `--explain`.

---

## Checks

### host-platform

The host is macOS. BLOCK otherwise: every credential read shells `security`,
every iOS build shells `xcodebuild`, every piece of visual evidence shells
`simctl`, so a run on another platform does not degrade, it fails partway
through with a worktree and a branch already created (ADR-0012).

It runs **first**, before `install-present`, so an unsupported host reads one
line naming the cause instead of a cascade of consequences. The step names the
`MULTI_AGENT_ALLOW_NON_DARWIN=1` escape, which the entry points honour and
nothing tests.

### install-present

The installed tree exists and carries the subtrees a run reads: `commands/`,
`multi-agent-refs/`, `scripts/`, `lib/`, `schemas/`. BLOCK when any is missing -
a run cannot start without them.

### install-version

The installed `.pipeline-version` matches the package version resolved from the
checkout or the published package. WARN on a mismatch: the run works, it just
is not the version the user thinks they have.

### update-policy

Whether a run updates itself before it starts, read the way the run reads it:
`updateCheck.enabled`, `.ttlHours` and `.autoUpdate`. The schema defaults all
three, so a preferences file that omits them behaves exactly like one that sets
them - which is why they are worth reporting. INFO when a value comes from the
schema rather than from the file, so the policy is legible in the place somebody
would change it. Reported, never written: a preferences file is the user's.

### script-surface

Every script a shipped command names by path exists in the installed tree. BLOCK:
the failure lands mid-run, at the call, with the phase already half done.

### skill-siblings

Every command has its copies in the host trees this machine installs. WARN, never
BLOCK: an unsynced host is a fact about the machine, and `skill-siblings.mjs`
itself reports the installed layout as "authored side not checked" rather than
claiming a repo verdict it cannot reach.

### state-writable

`$HOME/.claude/logs/multi-agent` exists or can be created, and a file can be
written there. BLOCK: a run that cannot record its state cannot be resumed,
reported or priced, and it discovers this at Phase 0 after the pickers.

### prefs-valid

`multi-agent-preferences.json` parses and satisfies `schemas/prefs.schema.json`.
BLOCK: Phase 0 reads it before anything else, and a malformed file fails the run
after the user has already answered the pickers.

### identity

A git identity resolves for the account the run would commit as. WARN: the run
reaches Phase 4 and stops there.

### hook-coverage

The blocking `PreToolUse` gates from `templates/claude-hooks.json` are present in
`settings.json`, and none of them is registered twice on the same matcher
(Claude Code runs a duplicate on every matching call). WARN. SKIP when the template is not installed - which was the
permanent state until the installer began copying `templates/`, and is exactly
the case this severity exists to make visible.

### credential-mapping

Which logical keys are mapped, and whether each mapped value is well formed.
Unmapped is INFO. Malformed is WARN.

### credential-liveness

One cheap authenticated request per mapped credential. WARN on `auth-rejected`,
`tier-1-no-grant` or `unreachable` - but with **different steps**, because a
refused credential and a host that never answered need different actions.
`credential-inventory.sh` already draws that line ("unreachable - no response at
all - on a corporate host, almost always the VPN"), and telling someone to
re-onboard a token that was never the problem is the exact failure the keychain
rules warn about one layer up. **SKIP unless `--probe` is passed**, and the
skip is printed, because a network check is the user's decision to spend.

### store-reference

Whether the toolkit's store reference file (`~/.config/multi-agent-toolkit/store.json`,
or `MCP_TOOLKIT_STORE_CONFIG`) names the Keychain items that `keychainMapping`
holds for App Store Connect (`appstore_connect_key_id`, `appstore_connect_issuer_id`,
`appstore_connect_private_key`) and Google Play (`google_play`). The `store_*`
tools and `ios_testflight_validate` read the keys only through that file, so a
key that is in the Keychain but not in the file is invisible to them.

Asked through `write-store-config.mjs --check`, which checks each item's presence
with `credential-store.sh probe` and never reads a value. OK when none of the
four keys is mapped: nothing was claimed. WARN when the file is missing or stale,
with the single step `run node "$HOME/.claude/scripts/write-store-config.mjs"`.
INFO when a store is only partly onboarded (for example the key id and issuer id
mapped without the `.p8`), with setup as the step. SKIP when the writer is not
installed or the credential store does not answer.

### embedded-credentials

A token in a git remote URL or a tracked config file. BLOCK: this one leaks
rather than fails.

The report carries the repo path, the config key, the **host**, a shape label and
a length bucket. Never the value, and not the prefix either: `ghp_` plus a length
is already a fingerprint, and this line is printed to a terminal that is often
shared, which is the whole reason the check exists. The URL is parsed inside a
`python3` heredoc reading stdin, so no shell variable ever holds it.

### task-tools

Whether this session's model carries `TaskCreate` / `TaskUpdate`. Claude Code
provides them by default only on Claude 3.x, Opus 4 through 4.7, Sonnet 4 through
4.6 and Haiku 4.5; on any newer model they are absent unless the user opts in.
INFO, with the opt-in as the step.

A script cannot answer this - only the agent knows its own tool list - so the
caller passes `--task-tools=yes|no` and the default is SKIP. Reporting "absent"
from a script that never looked would be the same defect this check is about.

### mcp-registration

Whether `multi-agent-toolkit` is registered as an MCP server. INFO when it is
not: the tools it provides are optional and a run without them is smaller, not
wrong.

**Registered is not the same as working.** With `--probe` the server is started
and its tools counted; serving none is WARN. Without `--probe` this reports what
is configured, and says so. The distinction is not theoretical: a half-extracted
package in the npx cache left the server dying on `Cannot find module` at
startup, which the client surfaces only as `CONNECTION_CLOSED`, and a check that
read the registration and stopped would have called that healthy.

### mcp-surface

How many MCP servers are charged in the project the caller is standing in,
counted across three places that all cost the same: the global blocks in
`~/.claude.json` and `~/.claude/settings.json`, the per-project block inside
`~/.claude.json`, and a `.mcp.json` committed to the repo. INFO once the count
exceeds `prefs.global.mcpSurface.infoAbove` (default 8); `--explain` lists the
names with the scope each came from.

Servers registered by a marketplace plugin are NOT counted: nothing in the
config names them, and inventing a number is worse than reporting the one that
is countable.

The reason it exists: every registered server sends its tool list on every turn,
the user adds them one at a time, and nobody ever sees the running total - our
own toolkit contributes 118 tools by itself. This is the same argument that made
the pre-run context budget a measured number rather than an intention.

It only reports. It never disables a server, never blocks and never warns: how
many servers are worth their context is the user's call, not a health failure.
The threshold is judgement, which is why it is a pref instead of a constant in
the script - a number nobody can see is a number nobody can argue with.

### disk-space

Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
build is the largest thing a run writes, and ENOSPC mid-run corrupts the state
file it was writing at the time.

### worktree-residue

Worktrees left under `<repo>/.worktrees/` in the repository the caller is
standing in. WARN at five or more, or at 2 GB. A finished task removes its own
worktree at PR time, but a run that stops before Phase 4 never reaches that
step and nothing else collects it: the finalizer only runs on success, and
`gc-worktrees` only sweeps entries git has already forgotten. Each survivor is
a full second checkout, so the total is measured in gigabytes rather than
megabytes. SKIP outside a git repository - a project name in prefs is a name,
not a path, and guessing checkout locations to produce a number is how a
diagnostic starts lying.

## The server profile (`--profile=server`)

Four checks that only run under `--profile=server`. They are ADDITIVE: a
default run is byte-for-byte what it was, and `--list-checks` reports whichever
set the invocation would actually perform, because a list that disagrees with
the run is worse than no list.

None of them may BLOCK. The blocking set is closed at six and these are
readiness, not correctness: a machine that fails all four still runs the
pipeline perfectly well with someone watching it. What they catch is the
failure mode of an unwatched run - stopping without saying so.


### usage-report

The outcome `usage-report.mjs` recorded for its last send attempt, read from
`logs/multi-agent/usage-last-send.json`. SKIP when `usageLog.enabled` is off or
nothing has been sent yet; OK when the last send was accepted; WARN for every
other outcome (`redirected`, `rejected`, `not-accepted`, `failed`, `skipped`),
with the status code or reason and the step that fixes it. The emitter runs
detached from phase-tracker with its output discarded, so this record is the
only surface on which a reporting path that stopped working shows up.
### xcode-skills

Is Apple's `xcode-integration` plugin, which Xcode 27.1 and later materialize
(`xcrun agent plugin path`), linked into Claude Code at the installed Xcode
build. SKIP when this Mac has no such plugin, because the pipeline's own iOS
skills then apply; INFO when it is available but not linked; WARN when the link
was copied from an older Xcode build than the one installed; OK otherwise. The
step is `/multi-agent:update`, which re-runs the installer's link. Apple's skills
are the reference for iPhone Duo and window resizing, SwiftUI, App Intents and
the accessibility audits, so a stale copy is a worse answer, never a failed run.
### unattended-contract

Is `multi-agent-refs/unattended-contract.md` installed. It is the only place
that says which entry points honour `MULTI_AGENT_UNATTENDED=1` and what each
one resolves to. Without it an operator setting up a server has to read the
scripts to find out, which is how the wrong assumption gets made.

### unattended-permissions

Does `multi-agent-unattended.settings.json`, the file the runner passes with
`claude --settings`, carry the unattended profile (`schemas/unattended-profile.json`):
`defaultMode: "dontAsk"`, `Edit` and `Write` allowed, at least one narrow
`Bash(...)` rule, and every deny rule; and is there no whole-`Bash` allow in it
or in `settings.json`, whose permission lists merge into the run's. autopilot spawns its child with `--permission-prompts none`, but the
tools that child then calls still have to be allowed, and on a fresh machine
they are not - the run stops at the first prompt with nobody there to answer
it, printing nothing. This is the single most likely reason a server sits idle
and looks healthy.

### scheduler

Is a multi-agent launchd agent loaded. Something has to start the work; a
server with no agent and an empty queue looks exactly like a server that has
finished everything.

### keychain-unlock

Is a default keychain reachable from this session. Every credential the
pipeline reads lives there, and a LaunchDaemon started before login sees a
LOCKED keychain: each fetch fails, and the failures surface much later as 401s
that blame the token rather than the lock. The check is a proxy - reading a
real credential would be a network call and a side effect - so it asks whether
a login keychain is reachable at all, which is the condition a boot-time
daemon fails.
