# doctor - the check registry

Every check `/multi-agent:doctor` can report has a `### <id>` heading here, and
`doctor.mjs --list-checks` prints exactly the same set. The equality is checked
in both directions by `smoke-doctor.sh`: a check that ships without an entry
cannot be released, and an entry whose check was deleted cannot linger as orphan
prose. That is the mechanism that keeps this file from rotting into a list of
things the tool used to do.

## What the exit code means

| Code | Meaning |
|---|---|
| 0 | healthy - nothing above INFO |
| 1 | degraded - at least one WARN, no BLOCK |
| 2 | blocked - at least one BLOCK |
| 3 | usage error |
| 4 | indeterminate - the installed layout could not be resolved, so nothing was checked |

4 exists so that "I could not look" never borrows the exit code of "I looked and
it is fine". A consumer that treats 0 as healthy would otherwise read a broken
resolver as a clean bill.

## What BLOCK means, exactly

**A state in which a run will fail or leak. Not a state in which it will merely
be worse.** Without a closed definition "blocked" grows until nobody respects
exit 2, so only five checks can produce it: `install-present`, `script-surface`,
`state-writable`, `prefs-valid` and `embedded-credentials`. Every other check
tops out at WARN however bad it looks.

## Who calls it, and what they do with the code

| Caller | When | On exit 2 |
|---|---|---|
| `/multi-agent:setup` | before the first question, and again at the end | the closing run is the evidence for or against "setup complete" |
| `/multi-agent:update` | after the install | report that the update left a state a run will fail from, rather than "updated" |
| `/multi-agent:sync` | step 0 | **stop.** Syncing a blocked install copies one fault onto five surfaces, and the copies are what people then debug |

`/multi-agent:sync` stops on exit 4 as well: the layout did not resolve, so
nothing was checked, and syncing from an unknown state is worse than not syncing
at all. Exit 1 never stops anything - a degraded install still syncs correctly,
and a gate that blocked on every warning would be routed around within a week.

## The four severities

| Severity | Meaning |
|---|---|
| BLOCK | a run will fail or leak |
| WARN | a run will work and be worse: a degraded capability, a stale copy, a rejected credential |
| INFO | a capability the user chose not to enable; nothing to fix |
| SKIP | not checked, and why |

`SKIP` is printed in the same list, in the same position, and the summary always
carries its count (`... 2 not checked, 15 checks`). Absence is a finding: a run
without `--probe` exits 0 **and** prints `SKIP credential-liveness`, so "healthy"
and "not looked at" are never the same output.

The INFO / WARN boundary is already decided in `lib/credential-inventory.sh` and
this tool only renders it: an unmapped key is INFO because it is a capability the
user chose not to enable, while `auth-rejected`, `malformed` and
`tier-1-no-grant` are WARN because the configuration made a claim the service
refused. **A missing optional capability is information; a mapped credential the
service rejected is a warning. The difference is whether the configuration
asserted anything.**

## The line shape

```
<SEV> <id>  -  <what is wrong>  -  <single imperative step>
```

The third field is one step, and its first word comes from a closed list:
`run`, `set`, `map`, `revoke`, `install`, `remove`, `free`, `export`, `merge`,
`record`, `re-run`. `consider`, `may want` and `should probably` are refused by
the gate. One problem, one step: a reader with five suggestions does nothing.

## It recommends, it never fixes

Nothing here rewrites a remote URL, edits `settings.json` or touches a token.
A remote whose URL carries a token may be the only credential that repo has, the
remote may be a mirror a script depends on verbatim, and doctor can run inside a
checkout the user does not own. Silent repair breaks all three.

For an embedded credential the honest single step is **not** "hide it". A token
that reached `.git/config` is already burned: it is in the shell history and
readable by anything that can read the working tree. The step is to revoke it at
its host; everything after that is behind `--explain`.

---

## Checks

### install-present

The installed tree exists and carries the subtrees a run reads: `commands/`,
`multi-agent-refs/`, `scripts/`, `lib/`, `schemas/`. BLOCK when any is missing -
a run cannot start without them.

### install-version

The installed `.pipeline-version` matches the package version resolved from the
checkout or the published package. WARN on a mismatch: the run works, it just
is not the version the user thinks they have.

### script-surface

Every script a shipped command names by path exists in the installed tree. BLOCK:
the failure lands mid-run, at the call, with the phase already half done.

### skill-siblings

Every command has its copies in the host trees this machine installs. WARN, never
BLOCK: an unsynced host is a fact about the machine, and `skill-siblings.mjs`
itself reports the installed layout as "authored side not checked" rather than
claiming a repo verdict it cannot reach.

### state-writable

`$HOME/.claude/logs/multi-agent` exists or can be created, and a file can be
written there. BLOCK: a run that cannot record its state cannot be resumed,
reported or priced, and it discovers this at Phase 0 after the pickers.

### prefs-valid

`multi-agent-preferences.json` parses and satisfies `schemas/prefs.schema.json`.
BLOCK: Phase 0 reads it before anything else, and a malformed file fails the run
after the user has already answered the pickers.

### identity

A git identity resolves for the account the run would commit as. WARN: the run
reaches Phase 6 and stops there.

### hook-coverage

The blocking `PreToolUse` gates from `templates/claude-hooks.json` are present in
`settings.json`. WARN. SKIP when the template is not installed - which was the
permanent state until the installer began copying `templates/`, and is exactly
the case this severity exists to make visible.

### credential-mapping

Which logical keys are mapped, and whether each mapped value is well formed.
Unmapped is INFO. Malformed is WARN.

### credential-liveness

One cheap authenticated request per mapped credential. WARN on `auth-rejected`,
`tier-1-no-grant` or `unreachable` - but with **different steps**, because a
refused credential and a host that never answered need different actions.
`credential-inventory.sh` already draws that line ("unreachable - no response at
all - on a corporate host, almost always the VPN"), and telling someone to
re-onboard a token that was never the problem is the exact failure the keychain
rules warn about one layer up. **SKIP unless `--probe` is passed**, and the
skip is printed, because a network check is the user's decision to spend.

### embedded-credentials

A token in a git remote URL or a tracked config file. BLOCK: this one leaks
rather than fails.

The report carries the repo path, the config key, the **host**, a shape label and
a length bucket. Never the value, and not the prefix either: `ghp_` plus a length
is already a fingerprint, and this line is printed to a terminal that is often
shared, which is the whole reason the check exists. The URL is parsed inside a
`python3` heredoc reading stdin, so no shell variable ever holds it.

### task-tools

Whether this session's model carries `TaskCreate` / `TaskUpdate`. Claude Code
provides them by default only on Claude 3.x, Opus 4 through 4.7, Sonnet 4 through
4.6 and Haiku 4.5; on any newer model they are absent unless the user opts in.
INFO, with the opt-in as the step.

A script cannot answer this - only the agent knows its own tool list - so the
caller passes `--task-tools=yes|no` and the default is SKIP. Reporting "absent"
from a script that never looked would be the same defect this check is about.

### mcp-registration

Whether `multi-agent-toolkit` is registered as an MCP server. INFO when it is
not: the tools it provides are optional and a run without them is smaller, not
wrong.

**Registered is not the same as working.** With `--probe` the server is started
and its tools counted; serving none is WARN. Without `--probe` this reports what
is configured, and says so. The distinction is not theoretical: a half-extracted
package in the npx cache left the server dying on `Cannot find module` at
startup, which the client surfaces only as `CONNECTION_CLOSED`, and a check that
read the registration and stopped would have called that healthy.

### disk-space

Free space on the volume holding `$HOME`. WARN under 2 GB: a worktree plus a
build is the largest thing a run writes, and ENOSPC mid-run corrupts the state
file it was writing at the time.

### worktree-residue

Worktrees left under `<repo>/.worktrees/` in the repository the caller is
standing in. WARN at five or more, or at 2 GB. A finished task removes its own
worktree at PR time, but a run that stops before Phase 6 never reaches that
step and nothing else collects it: the finalizer only runs on success, and
`gc-worktrees` only sweeps entries git has already forgotten. Each survivor is
a full second checkout, so the total is measured in gigabytes rather than
megabytes. SKIP outside a git repository - a project name in prefs is a name,
not a path, and guessing checkout locations to produce a number is how a
diagnostic starts lying.
