# @mmerterden/multi-agent-pipeline

[![GitHub Release](https://img.shields.io/github/v/release/mmerterden/multi-agent-pipeline?color=blue)](https://github.com/mmerterden/multi-agent-pipeline/releases)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Node.js](https://img.shields.io/badge/Node.js-20%20%7C%2022-green)](https://nodejs.org)
[![Zero Dependencies](https://img.shields.io/badge/dependencies-0-brightgreen)](https://github.com/mmerterden/multi-agent-pipeline/blob/main/package.json)
[![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/mmerterden/multi-agent-pipeline/badge)](https://scorecard.dev/viewer/?uri=github.com/mmerterden/multi-agent-pipeline)

🇹🇷 Türkçe: [README.tr.md](./README.tr.md)

A 6-phase AI development pipeline for **Claude Code**, **Copilot CLI** and **Codex CLI**. Drives a Jira issue or GitHub URL to a merged PR in one command - analysis → plan → TDD → review → test → commit → PR - with multi-repo orchestration, a plan-approval gate, CLI-aware parallel review, and store-compliance checks. Component and Figma-to-code work is dispatched to the per-stack marketplace plugins (iOS/SwiftUI, Android/Compose) rather than bundled, so component skills live in one place.

Runs natively on Claude Code, Copilot CLI and Codex CLI. macOS only. Zero runtime dependencies.

📐 **[Architecture diagrams](./docs/architecture.md)** - the 6-phase flow, operating modes, review/triage, Figma subphases, component layout. **[Ecosystem diagram](./docs/ecosystem.md)** - how this repo, the `multi-agent-plugins` marketplace and `multi-agent-toolkit-mcp` compose.

### Prerequisites

- **Node.js >= 20.11** - required; the pipeline's own tooling runs on it.
- **`jq`** - required for nine paths, optional for the rest. 91 shell files call it. The nine that publish or decide - the autopilot queue, Jira comments, PR reviews, issue updates, the plan file, both Figma fetchers, log search and Jira auth - refuse with exit 3 rather than run, because a missing `jq` renders as empty DATA and the work carries on with it. Everywhere else it still degrades. The install prints a note when it is missing.
- **`gh`** - for GitHub issue and PR work. Its built-in `--jq` is independent of the `jq` binary.

## Quick Start

```bash
# from the public registry (no auth) - pick the CLIs you actually use
npx @mmerterden/multi-agent-pipeline install --claude    # Claude Code only (default)
npx @mmerterden/multi-agent-pipeline install --copilot   # Copilot CLI only
npx @mmerterden/multi-agent-pipeline install --codex     # Codex CLI only
npx @mmerterden/multi-agent-pipeline install --all       # all three

# then, once:
/multi-agent:setup          # keychain token scan + git identity + default stack
```

### No `npx` on that machine?

`npx` ships with npm, and npm ships with Node - so "npx: command not found" almost
always means Node is missing from that shell, not that anything is wrong with the
package. Check first, then pick the row that matches:

```bash
node -v; npm -v; command -v node npm npx
```

| What you see                                   | What to do                                                                               |
| ---------------------------------------------- | ---------------------------------------------------------------------------------------- |
| nothing at all                                 | Install Node >= 20.11: `brew install node`, or the LTS installer from nodejs.org         |
| `node` works, `npx` does not                   | `npm i -g @mmerterden/multi-agent-pipeline` then `multi-agent-pipeline install --claude` |
| nvm is installed but the shell does not see it | `source ~/.nvm/nvm.sh && nvm use --lts`, or just open a new terminal                     |
| npm is ancient (< 5.2, which predates npx)     | `npm i -g npm@latest`, or use the global-install row above                               |

And the path that needs neither `npx` nor a global install - clone and run the
installer directly:

```bash
git clone https://github.com/mmerterden/multi-agent-pipeline.git
cd multi-agent-pipeline
node index.js install --claude        # add --dry-run first to see what it would write
```

`npm exec @mmerterden/multi-agent-pipeline install --claude` also works on any npm
7+ without `npx` on `PATH`.

**One error that is not an npx problem.** This package declares `os: ["darwin"]`,
so npm refuses to install it anywhere else and says `npm ERR! notsup Unsupported
platform`. That is deliberate ([ADR-0012](./docs/adr/0012-macos-only.md)), not a
missing tool: every credential read shells `security`, every iOS build
`xcodebuild`, every piece of visual evidence `simctl`.

Tool flags combine (`--claude --codex`). With no tool flag at all, the installer targets Claude Code only. Other flags: `--dry-run` (show what would be written, write nothing), `--platform=ios|android|all` (skip the stack skills you do not need), `--link` (symlink instead of copy, for local development).

Run a task - the input type is auto-detected:

```bash
/multi-agent "PROJ-1234"                              # Jira id → fetch, plan, build
/multi-agent "https://github.com/org/repo/issues/42"  # GitHub issue URL
/multi-agent "my-app#42"                              # repo + issue number
/multi-agent "fix dark-mode contrast on LoginView"    # free-text bug/feature
/multi-agent:jira                                     # browse your open Jira issues → pick
/multi-agent:issue                                    # browse unassigned GitHub issues → pick
```

Every input runs the same short intake - **account → (repo) → maturity check → dev-context** - then enters Phase 0. A Jira id or GitHub URL is fetched and maturity-checked _before_ any code is written; free-text skips the fetch and goes straight to planning. Multi-repo tasks add extra repos at the dev-context step.

Add `autopilot` to skip confirmations (e.g. `/multi-agent:autopilot "PROJ-1234"`). The workspace is not a flag either - it is the one question the run asks about its own shape at Phase 0: where to run, a worktree or your current checkout.

Update later with `/multi-agent:update`. Uninstall (tokens preserved) with `npx @mmerterden/multi-agent-pipeline uninstall`.

**Stack skills are marketplace plugins.** On Claude Code the `ai-<stack>-toolkit` plugins (`multi-agent-plugins` marketplace) are the only stack-skill source - nothing is copied into `~/.claude/skills`. Pick the active stack(s) per repo with `/multi-agent:stack` (multi-select: `ios backend`, or a native picker with no args); each plugin ships a language-aware catalog at `ai-<stack>-toolkit:help`. Copilot CLI and Codex CLI have no plugin loader, so they receive a local copy filtered to the same enabled stacks.

## How it works

One command runs 6 phases, with a gate between the risky ones. Every run
carries the whole set; the one question Phase 0 asks about shape is where the
branch lives, a worktree or your current checkout:

- **0 · Init** - parse the input (Jira id / GitHub URL / free text), pick account + repo(s), fetch the issue, run a maturity check.
- **1 · Plan** - detect the stack, scan the codebase and write the analysis document, then break it into tasks with file-level targets and **stop for your approval** before touching code. Analysis and planning are one decision, so they are one phase ([ADR-0014](./docs/adr/0014-six-phase-consolidation.md)). Codebase scanning runs on the explorer persona (Sonnet).
- **2 · Dev** - TDD: failing test → code → green, following the repo's style + the active stack skills. The phase ends at its own gate: build, lint, tests and a secret scan, run **once**; Review reads that log rather than building again.
- **3 · Review** - a **CLI-aware parallel review** against the logs Dev produced - Claude Code runs 3 models (Fable + Opus + Sonnet; Opus + Sonnet while the fable rung is off), Copilot CLI runs 3 (GPT-5.4 + Opus + Sonnet) - then a **Fable triage** keeps only actionable findings; blockers loop back to Phase 2. The optional user test lives here, keeping its waiting state.
- **4 · Commit/PR** - conventional commit, push (must succeed), open a PR (`Ref: #N`, never auto-close). An unattended run stops at the commit and hands a PR request to the autopilot runner, which verifies it and opens a draft PR.
- **5 · Report** - technical summary + a Jira comment with test scenarios, posted through the channels layer. This is the one step a terminal autopilot run still pauses at; an unattended run ends at the Phase 4 hand-off, and the runner posts the reports you configured once the PR is open.

The phase list lives in `pipeline/schemas/phases.json`, and `smoke-phase-contract.sh` holds every other copy of it - the generator's output, the token budget, the state-schema bounds, the progress fractions in sample output - to that one file.

`/multi-agent:analysis` runs its own shorter chain and reviews what it wrote before publishing it: the draft goes through the same three-reviewer set and triage as a code diff, a blocking finding returns it to synthesis with dispatch closed, and the gaps that survive are either searched, asked about, or recorded with an owner.

### Install verification, cost, and one state directory

- **`multi-agent-pipeline verify`.** The install is a copy, and from the moment it is written the two halves drift independently: an edit in the installed tree is behaviour with no source, and a file the installer skipped is a script the docs describe and nobody has. `verify` compares both against a manifest built at pack time. What a green result proves is stated plainly - the bytes match what the publisher recorded, not who published them.
- **Cost, past the single run.** The per-task ceiling cannot see the two ways a budget actually empties: a drift that trips nothing, and one session that burns a week in an hour while every run stays under its cap. `cost-analyze` projects, finds days out of family by median absolute deviation, and reports acceleration - as LIST-price estimates, which it says on every run rather than in a footnote.
- **A run's state lives in one place.** `{project}/{taskId}/` is canonical, the flat layout is still read, and a run that exists in both is counted once.
- **Server readiness, entirely opt-in.** `doctor --profile=server` adds four checks that only matter when nobody is at the keyboard, and `install --unattended` writes a permission profile after printing it. The default install writes no permissions and the default doctor run leaves the server checks out, because a laptop told it fails a server check learns to ignore doctor.

### Your package manager, your hooks, your MCP surface

Three places where the pipeline looks instead of assuming:

- **Phase 2 uses the repo's package manager.** It is resolved from the repo - an env override, then `package.json#packageManager`, then the lock file, then npm reported as a default rather than as evidence - so a pnpm, yarn or bun repo does not fail after its worktree and branch exist. iOS and Android are untouched.
- **A compaction keeps what a phase learned.** The hooks template flushes captures at `PreCompact` as well as at session end, because an auto-compaction summarizes a long review or development phase while it is still running.
- **`doctor` counts your MCP servers.** Every registered server sends its tool list on every turn and they are added one at a time, so nobody ever sees the total. It reports the count and nothing else: no warning, no blocking, no disabling.

### Maturity: the check has an effect

Phase 0 scores how ready the item is before anything is built. A score that only halts leaves the item exactly as immature as it was found, and the next scan halts on it for the same reason, so the result drives a next step.

An interactive run **asks at that step** - open the item and fix it, continue without it, or abort - and records which gap you waved through rather than just that you continued. An autopilot run parks the item and can be told to **ask on the item itself** (`prefs.global.maturityFollowup.autopilotCommentsOnIssue`, off by default): one comment naming what is missing. Never a status change, never an assignee, never a close, `Ref:` and never `Closes:`.

`resume` re-enters the maturity step with the item re-fetched, and the check decides. An edit is only a reason to look again - a reply reading "will do later" moves the timestamp and fixes nothing. The same question is never asked twice.

### Base branch: evidence, then a question

Phase 0 Step 3 collects the candidates **with the evidence behind each one** before it asks. An issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version - so that becomes a ranked row whose description says why, next to rows that say "the repository's default branch" or "you used this last time". You still choose; the evidence only reorders.

Nothing is hardcoded: any Jira field whose schema resolves to `version` is read whatever the board calls it, and the release-branch naming convention is inferred from the refs that exist rather than read off a table - one repo yields `<prefix>/develop_<version>`, another `release-<version>`, out of the same code. A version whose branch has not been cut yet is reported as a note, never offered as an option you cannot check out.

A failed `git fetch` degrades loudly rather than silently: the list falls back to local refs, the question says so, it gains a retry row, and the exit gate refuses to close Phase 0 if a degraded run recorded its list as the remote's answer. Turn the whole thing off with `prefs.global.baseBranchEvidence.enabled: false` and Step 3 asks with the plain branch list.

### Workspace: worktree or local

Phase 0 Step 5b asks where the branch lives. **Worktree** (`.worktrees/{id}/`) leaves your current checkout untouched; **Local** works in the project root on a new branch, and needs the project root clean. It is a question, not a flag: there is no `:local` command and no `--local` switch, so the answer is always visible in the run rather than buried in how the run was typed. `autopilot` resolves it to a worktree without asking, because an unattended run commits and pushes from wherever it stands and doing that in your own checkout is what worktrees exist to prevent. The manual-test offer at the end of Review follows the same answer: a worktree run is asked whether to check the branch out and test it, a local run already is that checkout, so there is nothing to offer.

Under the hood: each task runs in its own **git worktree** (or the current branch when you choose local), commits use the **git identity routed from the repo's origin URL**, and **multi-repo** tasks get per-repo worktrees plus an integration build. Tokens stay in the OS keychain; nothing is committed or logged. `/multi-agent:review` can also review an existing GitHub/Bitbucket PR - per-finding inline comments anchored to `file:line` + an explicit Approve / Needs-Work state.

The discipline behind all of this - bounded loops, evidence gates, token-budgeted phase docs, immutable tests, fresh-context handoffs - is catalogued in [docs/engineering.md](./docs/engineering.md). The full feature list lives in [docs/features.md](./docs/features.md). How this repo, the `multi-agent-plugins` marketplace, and the `multi-agent-toolkit-mcp` server compose at install time and at run time is diagrammed in [docs/ecosystem.md](./docs/ecosystem.md).

## Modes

| Mode      | Command                              | Flow                                                                                                                                                             |
| --------- | ------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Full      | `/multi-agent "task"`                | All 6 phases, interactive                                                                                                                                        |
| Autopilot | `/multi-agent:autopilot "task"`      | The same 6 phases, no confirmations; the workspace resolves to a worktree and the user-test gate inside Review is skipped                                                                                                       |
| Audit     | `/multi-agent:design-check`          | Mock-mode vs Figma conformance, local-only                                                                                                                       |
| Audit     | `/multi-agent:testflight-validation` | Pre-submission gates for a TestFlight build: static archive audit → Apple's `altool --validate-app` → Review-Guidelines check. Validates only, never uploads     |

`autopilot` is the only knob on the run itself; everything else is its own command. The full catalog is below.

## Commands

`/multi-agent` plus 64 sub-commands. `/multi-agent:help` renders the same catalog in your terminal, in your `outputLanguage`.

### Pipeline entries

| Command                               | What it does                                                                                         |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| `/multi-agent "task"`                 | The pipeline; Phase 0 asks where the branch lives                                                    |
| `/multi-agent:autopilot "task"`       | The same pipeline, unattended: worktree resolved, no confirmations                                   |
| `/multi-agent:resume`                 | Unfinished work, either source: a stopped run picks up where it left off, a branch with no run behind it gets the tail (Review → Build+Test → Commit/PR → Report) |

### Task control

| Command                                 | What it does                                                                                                                                             |
| --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/multi-agent:status`                   | Every task's ID, phase, branch and state                                                                                                                 |
| `/multi-agent:log [#N]`                 | Show a task's `agent-log.md` (most recent by default)                                                                                                    |
| `/multi-agent:resume [#N]`              | Carry a stopped or failed task on from its last phase, or run the tail over a branch with no run behind it                                                |
| `/multi-agent:kill [#N]`                | Stop a task, remove its worktree and branch                                                                                                              |
| `/multi-agent:steer #N "<instruction>"` | Correct a running task without stopping it; applied at the next phase boundary                                                                           |
| `/multi-agent:search`                   | Ranked search across every task log; `--semantic` queries the triage corpus                                                                              |
| `/multi-agent:garbage-collect`          | Sweep leftover scratch, orphan worktrees and offloaded payloads. `--abandoned` also reaps runs that stopped and were never picked back up. Dry-run first |
| `/multi-agent:prune-logs`               | Delete per-task logs by age / project / task. Audit trail and metrics kept                                                                               |
| `/multi-agent:purge`                    | Wipe every worktree, branch, log and state file. Double confirmation                                                                                     |

### Review

| Command                            | What it does                                                                                |
| ---------------------------------- | ------------------------------------------------------------------------------------------- |
| `/multi-agent:review`              | Parallel review of a branch diff or a PR; inline comments + approve/needs-work on PR input  |
| `/multi-agent:review-jira`         | Grade a Jira issue's readiness for development, comment the gaps                            |
| `/multi-agent:review-issue`        | Same grading for a GitHub issue                                                             |
| `/multi-agent:review-analysis`     | Review a written analysis document; findings cite the Locked rule they break                |
| `/multi-agent:diff-explain`        | Map a Phase 3 triage finding back to the diff lines that caused it                          |
| `/multi-agent:estimate`            | Effort range from similar closed work: P25-P75 per role, with the tickets behind it         |
| `/multi-agent:scenario-audit`      | Check test scenarios against the code; each PRESENT on a resolving file:line                |
| `/multi-agent:refactor`            | Best-practice extraction + bug hunt + derived-skill drift + toolkit MCP research → one plan |
| `/multi-agent:scan`                | Skill security scan of local skill directories against a tiered pattern catalog             |
| `/multi-agent:prune-prompts`       | Zero-base review of the always-on instruction footprint; keep / trial / delete per rule     |
| `/multi-agent:ios-coding-standard` | Audit an iOS module against the 99-rule registry, produce a remediation plan                |
| `/multi-agent:security-review`    | Standalone static security review of a diff, branch or repo: threat model, OWASP + CWE findings with CVSS, offline dependency inventory |

### Analysis

| Command                           | What it does                                                                              |
| --------------------------------- | ----------------------------------------------------------------------------------------- |
| `/multi-agent:analysis`           | Standalone feature spec: global (23-section handoff) or corporate (IG/UC/FG) profile      |
| `/multi-agent:analysis-resolve`   | Answer an analysis doc's Section 20 open questions one row at a time                      |
| `/multi-agent:analysis-jira`      | Turn a final analysis document into a Jira story tree: stories from its rule ids, two-way coverage, full preview, creates only what is missing |
| `/multi-agent:graph`              | Build and query the repo's code graph: symbols, imports, references; deterministic, no API tokens |
| `/multi-agent:research <id>`      | Answer an item's gaps from its tracker thread, linked docs and repo before anyone is asked |
| `/multi-agent:complaint-analysis` | Customer-complaint triage with Graylog evidence: client / bff root cause, or core routing |

### Testing on a device

| Command                                  | What it does                                                              |
| ---------------------------------------- | ------------------------------------------------------------------------- |
| `/multi-agent:test`                      | UI Bug Hunter on a booted simulator or emulator: screenshot, tap, analyze |
| `/multi-agent:test-dark-mode`            | Walk every screen light then dark, report contrast and colour bugs        |
| `/multi-agent:test-accessibility`        | VoiceOver labels, sub-44pt tap targets, contrast, traits                  |
| `/multi-agent:test-dynamic-type`         | Re-walk every screen at XL through accessibility-XL, report truncation    |
| `/multi-agent:test-screenshots [locale]` | App Store screenshot set in a locale (defaults to `tr`)                   |
| `/multi-agent:bug-bash`                  | Parallel exploratory sessions by charter; only findings a failing repro test proves are reported |
| `/multi-agent:manual-test`               | Phase 3 standalone: check out the task branch and prepare it for Xcode    |

### Design, build and store

| Command                              | What it does                                                                             |
| ------------------------------------ | ---------------------------------------------------------------------------------------- |
| `/multi-agent:design-check`          | Mock-mode vs Figma audit with a coverage gate; annotated HTML + PDF report               |
| `/multi-agent:store-ready`           | Pre-submission gates for iOS and Android: package audit, store validation, policy review |
| `/multi-agent:testflight-validation` | iOS-pinned alias of `store-ready`. Validates only, never uploads                         |
| `/multi-agent:build-optimize`        | Benchmark an Xcode build, run the analyzers, produce a recommend-first plan              |

### Tickets and reporting

| Command                    | What it does                                                                        |
| -------------------------- | ----------------------------------------------------------------------------------- |
| `/multi-agent:jira`        | Browse your open Jira issues → pick → branch → mode → launch                        |
| `/multi-agent:issue`       | Browse unassigned GitHub issues → pick → auto-assign → launch                       |
| `/multi-agent:create-jira` | Draft a Task / Bug / Story to the project's own conventions, preview before create  |
| `/multi-agent:channels`    | Post the multi-channel report: Jira, Confluence, Wiki, PR description, board status |
| `/multi-agent:feedback`    | Send one message to the maintainer. Only your text is sent, no logs or paths        |

### Your own routines

| Command                      | What it does                                                                       |
| ---------------------------- | ---------------------------------------------------------------------------------- |
| `/multi-agent:save [name]`   | Save a recurring job as a reusable `/multi-agent:<name>`. Local-only, never synced |
| `/multi-agent:routines`      | List your saved routines and what each does                                        |
| `/multi-agent:forget [name]` | Remove a saved routine and its registry entry                                      |

### Setup and maintenance

| Command                          | What it does                                                                   |
| -------------------------------- | ------------------------------------------------------------------------------ |
| `/multi-agent:setup`             | First-run wizard: keychain token discovery, git identity, pipeline preparation |
| `/multi-agent:stack [ids]`       | Enable the marketplace plugin(s) for this repo. Multi-select                   |
| `/multi-agent:scaffold <stack> <name>` | New project skeleton from the toolkit's scaffold skill, green before its first commit |
| `/multi-agent:language [en\|tr]` | Show or set `outputLanguage`; `promptLanguage` stays English                   |
| `/multi-agent:sync`              | One-shot sync: Claude Code, Copilot CLI, pipeline repo, website, toolkit MCP   |
| `/multi-agent:update`            | Update to the latest published npm release and run migrations                  |
| `/multi-agent:doctor`            | Health check of the install: layout, preferences, credentials, hooks, host capabilities; the exit code is the verdict |
| `/multi-agent:uninstall`         | Remove the pipeline from every CLI. Keychain tokens always left intact         |
| `/multi-agent:serve`             | Local contract server for a client: 127.0.0.1 only, token, stops when idle     |
| `/multi-agent:help`              | This catalog, in the terminal, in your `outputLanguage`                        |

### Subagents

Ten are installed alongside the commands: `explorer` and `task-clarifier` (Phase 0-1), `ios-architect` / `android-architect` / `backend-architect` and `plan-critic` (Phase 1, the critic only with the quality gates active), `dev-critic` (end of Phase 2), `code-reviewer` and `security-auditor` (Phase 3), and `bulk-reader`, which `bulk-read.sh` dispatches to summarise one large file when a whole-file read is blocked.

### Skills that are not commands

Two compliance skills install on every host and back the store gates: `apple-archive-compliance` (18-rule Apple review scan with ITMS code mapping) and `google-play-compliance` (21-rule Play policy catalog with Console error codes). Everything else stack-shaped - SwiftUI, Compose, backend, frontend - comes from the marketplace plugins described below.

## Continuous mode

Everything above starts when you start it. Continuous mode is the same pipeline
picking work up on its own, on ONE machine you choose, from repos you choose.

| Command                         | What it does                                                                                                                                                            |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/multi-agent:autopilot-on`     | Pick the repos this machine watches. Labelled GitHub issues and assigned + labelled Jira items then run in a worktree and stop at a draft PR. Re-run to change the list |
| `/multi-agent:autopilot-status` | What is running and at which phase, what is queued, what is waiting for an answer, the PRs of the last day, and the rolling spend                                       |
| `/multi-agent:autopilot-off`    | Remove the schedule. Work already running finishes; the repo selection is kept. `--now` also stops the item in flight, after naming it and asking                      |

### Autopilot (unattended)

The runner is a launchd tick. Each tick takes the head of the queue, runs the
whole pipeline on it in a worktree with `MULTI_AGENT_UNATTENDED=1` set, and
ends at a **draft** pull request. A person reviews it and merges it; nothing in
the pipeline merges.

- **Research before asking.** A run that parks on a maturity blocker or an open analysis question gets a research pass before it is left for a person. `/multi-agent:research <id> --autonomous` reads the tracker thread, linked items and documents, similar closed items and the repository; `research-gate.mjs` closes only what a quoted source or a resolving `file:line` backs, and the maturity check or the open-questions gate re-runs on the result and decides. Whatever is still open parks the run with the candidates found. [Research](./pipeline/multi-agent-refs/features/research.md)
- **One outward writer.** The session never pushes. Phase 4 writes a PR request, and the runner, outside the session, re-runs the stack's build and tests itself, re-checks the gate ledger, the secret scan over the commit range and the outbound gate, pushes the branch the run recorded from a clean staging repository, and opens the PR with `gh pr create --draft`. A check that fails leaves the item `verification-failed` with the verdict on disk. On Bitbucket the branch is pushed and a summary written, since there is no draft PR to open. [Unattended security](./pipeline/multi-agent-refs/features/unattended-security.md)
- **Operations around a run.** A second launchd agent keeps the Mac awake while the mode is on (`caffeinate -s -i`, plus `-d` if you answer yes to the display question at `autopilot-on`), and each run holds a sleep lock until its tick ends. Every credential a configured repo needs is probed before the schedule is offered and again on every tick, so a locked keychain refuses with the key and the fix. A circuit breaker stops launching after consecutive attempts that produced nothing, `costCeilingUsd` caps spend over a rolling 24 hours, `maxParallelAgents` caps sessions working at once, and once a day the runner writes a cleanup report and, when enabled, a digest. [Autopilot operations](./pipeline/multi-agent-refs/features/autopilot-operations.md)
- **A fail-closed guard.** Under the variable, the Bash guard judges a command only when it reduces to simple commands joined by `;`, `&&`, `||` and plain pipes, and refuses anything else unparsed. On what it can parse: no push, PR, issue or tracker write, no Keychain read, no fetch outside the network allowlist, no package install or manifest edit, no write into a protected path. Ticket, page and PR text reaches the model inside untrusted-data delimiters. The permission profile (`install --unattended`) is a second line; the boundary is a separate non-admin macOS user and GitHub rulesets on the default and release branches, which the operator applies. [Unattended security](./pipeline/multi-agent-refs/features/unattended-security.md)
- **Claims are checked, not trusted.** With the quality gates active, a run's claims about the repository are checked before it commits: a localization key or endpoint the change references exists, open analysis questions are answered or park the run, a triage citation quotes the line it cites, a "pre-existing" finding points at a line the diff did not touch, the plan's steps are covered, and a test pass shows a positive executed count read with the stack's own markers (8 stacks: `ios`, `android`, `web`, `backend-node`, `python`, `go`, `rust`, `jvm`) along with a new test that fails without the change. Every verdict goes to a gate ledger, and the commit hook refuses a commit whose mandatory gates did not pass for HEAD. [Unattended gates](./pipeline/multi-agent-refs/features/unattended-gates.md), [stack adapters](./pipeline/multi-agent-refs/features/stack-adapters.md)
- **Spec, plan and review discipline.** `/multi-agent:analysis` writes a project constitution (security, accessibility, architecture and licensing rules, each tied to a source), and a spec-consistency gate checks by id that every requirement has a task and a test and that nothing waives a binding rule. Phase 1 ends with one plan-critic round under four lenses (scope, feasibility, security, a cheaper alternative), answered once and judged by a script. In review, a `blocking` finding stops the run only when two independent reviewers carry it or a failing test shows it; otherwise it is lowered to `important`, never dropped. [Constitution](./pipeline/multi-agent-refs/features/constitution.md), [plan critic](./pipeline/multi-agent-refs/features/plan-critic.md), [review decision](./pipeline/multi-agent-refs/features/review-decision.md)
- **A client contract.** Commands and run questions are declared as data. `/multi-agent:serve` starts a localhost HTTP server over the same JSON the CLI prints, and `pipeline/contract/` ships the manifest, TypeScript types and fixtures a client builds against. Four phone routes read redacted runs, answer a parked question with one of its offered options, and, only when switched on, queue a launch; every request carries an Ed25519 signature from an enrolled device. The phone app is a separate project. [Client kit](./pipeline/contract/README.md), [phone API](./pipeline/multi-agent-refs/features/phone-api.md)

**Attended use is unchanged.** Every behaviour above keys on
`MULTI_AGENT_UNATTENDED=1` or on autopilot mode. With neither, no gate writes a
ledger entry or parks a run, and the guard answers as it did before; a smoke
suite holds each direction. The full contract, entry point by entry point:
[unattended-contract.md](./pipeline/multi-agent-refs/unattended-contract.md).

**Nothing is on by default and nothing is added implicitly.** Installing the
package writes no state and schedules nothing; `smoke-autopilot-default-off.sh`
fails the build if that ever changes. A label is a filter, not a gate - anyone
who can open an issue in a repo you have push on could add one - so the gate is
the picker, and it is per machine.

The biggest win is not parallelism. Measured here, the median run is 44 minutes,
but an item that finishes at 14:00 waits until you sit down again: overnight that
is 16 hours against 40 minutes. Continuous mode removes the waiting, not the work.

**What it will not do.** It does not merge - the runner stops at a draft PR and
the decision stays yours. It does not touch an attended run: per-repo concurrency
is always 1, so the queue steps around a repo you are working in rather than
competing for `.git/index.lock`. There is no cap on PRs; the bounds are
`costCeilingUsd` over a rolling 24 hours, `maxParallelAgents` when you set it,
and what the machine can hold.

**A menu bar indicator**, when `swiftc` is present, draws the same `status.json`
in the top right and refreshes on its own: one row per item with its id, the
phase as a fraction, the elapsed time and the stack. A row disappears the moment
the item finishes and reappears under Reports with its PR. It only draws - it
cannot start, stop or change a run, and the one command it names is spelled the
way your CLI spells it. Its labels render in `prefs.global.outputLanguage`, from
the same table the terminal renders from. ActivityKit is unavailable on macOS,
so this is an `NSStatusItem`, built from source on demand rather than shipped as
a binary that would need signing.

**It survives a restart with no command to run.** launchd loads the job at
**login**, not at boot, and that is correct rather than a limitation: the login
keychain is what unlocks the tokens, so a tick that fired before login could not
reach Jira or GitHub anyway. There is no `autopilot-resume` - a command you have
to remember is a queue that silently stops when you forget it. `doctor` reports
the real failure instead: configured, but launchd holds no job.

**Sleep.** launchd's interval does not wake a sleeping Mac, so while the mode is
on the awake agent holds system sleep off on AC and idle sleep off on battery
too. A laptop left unplugged with the mode on stays awake until it is plugged in
or the mode is turned off; `awake.enabled: false` removes the agent and leaves
only the per-run lock, which macOS honours on AC only.

**The three commands ship to all three CLIs and behave identically there**, because
what they control is not a CLI feature: it is a launchd user agent whose tick
spawns a `claude --bg` child regardless of which CLI you typed the command in.
Continuous mode therefore needs the `claude` binary on `PATH` everywhere. The
queue and the repo selection live in `~/.claude/autopilot/` on every host, on
purpose - two CLIs on one machine must read one queue, not two.

## Which model answers

The ladder is `fable -> opus -> sonnet -> haiku`, and a run walks it on failure.
Routing lets a policy pick the rung instead, per call site, and is off until you
turn it on.

| Command                     | What it does                                                                                                                                        |
| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/multi-agent:route-on`     | Pick a strategy, a scope and the rules that say which rung a call lands on. Validated against `route-config.schema.json` before anything is written |
| `/multi-agent:route-off`    | Disarm it. The rules are KEPT, so turning it back on does not re-ask for the same configuration                                                     |
| `/multi-agent:route-status` | Whether it is armed, which rule applies where, which rung the last dispatches took, and what this run has cost                                      |
| `/multi-agent:model`        | Turn the top rung on or off, and move the cost ledger's pricing with it in the same step                                                            |

**The scope cannot be the whole session.** `scope` accepts `subagent`,
`bulk-read` and `research`; there is no `host-session` value and that is a schema
gate, not a convention. Owning the host's base URL would send every call you make
through a third layer, including work that has nothing to do with this pipeline.

**The honest limit, which `route-status` prints rather than hides.** A rung on a
non-Anthropic provider is refused at dispatch. No call site sends a model
request itself: the one that routes, `bulk-read.sh`, delegates to the `claude`
CLI, which speaks that ladder and nothing else. Sending a read to a third-party
provider would also put the file's full text outside the account that owns it,
so it is a decision a user makes rather than a default. Routing chooses inside
the ladder.

**The top rung is a switch.** `/multi-agent:model` turns `fable` on or off
(`prefs.global.modelFallback.fableEnabled`, shipped off). The review commands,
triage and the plan critic resolve their slots through it: on Claude Code,
`/multi-agent:review` runs Fable + Opus + Sonnet while it is on and Opus + Sonnet
while it is off.

## Stacks

Stack skills ship as versioned plugins in the [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins) marketplace. Select a stack per-repo:

```bash
/multi-agent:stack ios        # or android / frontend / backend / mobile / all
```

This enables the matching plugin (+ the shared `ai-common` plugin) in the repo's `.claude/settings.json`. Phase 1 auto-detects the stack for routing. New repos default to iOS.

## Tool support

The pipeline runs natively on **Claude Code**, **Copilot CLI** and **Codex CLI** - all three install from the same `pipeline/` source and get the same 64 commands.

| Tool        | Flag                 | What it installs                                                                                       |
| ----------- | -------------------- | ------------------------------------------------------------------------------------------------------ |
| Claude Code | `--claude` (default) | slash commands + skills + agents + three `PreToolUse` hooks (secret scan, agent-guard, read-size gate) |
| Copilot CLI | `--copilot`          | instructions + 64 sub-command skills + scripts                                                         |
| Codex CLI   | `--codex`            | one router skill + 64 specs as refs + 10 agent TOML + `AGENTS.md` block + `codex mcp add`               |

Filter skills by stack with `--platform=ios\|android\|all`.

**Why Codex gets one skill and not 64.** Codex assembles every discovered skill's name
and description into a single prompt block and drops entries when it overflows, with no
error. Measured on 0.145: installing one plugin that declares 142 skills surfaced only
75 of them and evicted an unrelated user skill. So on Codex the pipeline ships a single
`multi-agent` router and keeps the sub-command specs as reference files that cost
nothing until read - same commands, same behaviour, a layout the host can actually hold.

Reviewer sets differ because the available models do: Claude Code runs 3 reviewers
(Fable + Opus + Sonnet; 2 while the fable rung is off), Copilot CLI 3 (Opus + GPT-5.4 + Sonnet), Codex CLI 3 (gpt-5.6 at
xhigh, gpt-5.4, gpt-5.6 at medium). Claude Code and Codex are single-vendor panels, so
consensus among their three is weaker evidence than the same consensus on Copilot CLI,
the one host whose panel spans two vendors, and the triage note says so.

## Tokens & integrations

`setup` scans your OS keychain and maps each token by a **logical name** (e.g. `jira`) to its real keychain entry - the pipeline resolves tokens through that mapping (`credential-store.sh`), so literal keychain names never appear in synced files. Tokens stay in the macOS Keychain, are **never committed or logged**, and are all **optional** - the pipeline asks for any it needs at Phase 0.

| Token                 | Used for                                                         | Phase                   |
| --------------------- | ---------------------------------------------------------------- | ----------------------- |
| `jira`                | fetch the issue · post the report comment                        | 0, 5                    |
| `github`              | issues · PRs · `gh` auth                                         | 0, 4                    |
| `bitbucket`           | PR create/update (reviewer-preserving) · diff                    | 4                       |
| `confluence`          | publish analysis / wiki pages                                    | 5                       |
| `figma` + `figma_mcp` | fetch design context                                             | analysis only           |
| `fortify`             | security-scan findings gate                                      | 2                       |
| `firebase`            | Firebase service-account JSON for Firebase projects              | as needed               |
| `jenkins`             | CI trigger / status                                              | build / deploy          |
| `npm`                 | package publish (mostly CI)                                      | release                 |
| `appstore_connect_*`  | TestFlight / App Store pre-submission validation (optional, iOS) | `testflight-validation` |

The **secret scan** runs as a `PreToolUse` hook on Claude Code (hard-blocks a commit on a hit) and as a pre-push check elsewhere.

## Outside a pipeline run

Installing the pipeline is not only useful when you run it. Open an ordinary session and the same three things are available, announced by `rules/outside-the-pipeline.md` which loads with every conversation:

- **Services you already onboarded.** The token `setup` mapped is readable now - resolve the logical name through `credential-store.sh` and fetch the issue, the page, the log. **Reads are ordinary work; writes are not.** Posting a Jira comment, editing an issue or opening a PR goes through the pipeline commands, because the rules that make those safe (never auto-close, `Ref:` not `Closes:`, humanizer on outward prose) live there.
- **The stack skills `/multi-agent:stack` enabled for the repo.** Each toolkit's own `index` skill routes; the pipeline keeps no copy of that table.
- **The `multi-agent-toolkit` MCP.** 80+ tools for a running app - screen state, crash logs, design comparison, store pre-submission.

Uninstall preserves this layer: the tokens, the reader that opens them, the mapping that names them, and the MCP registration. Removing the pipeline should not cost you the credentials you onboarded through it.

## Platform support

Runs on **macOS** only. The package declares `os: ["darwin"]`, so `npm` refuses to install it elsewhere rather than letting a run fail halfway through: every credential read shells `security`, every iOS build `xcodebuild`, every piece of visual evidence `simctl`. Node.js 20.11+ (tested on 20 and 22). Reasoning: [ADR-0012](docs/adr/0012-macos-only.md).

## Companion repos

| Repo                                                                                          | What it is                                                                                                                                                                                                                                                                                                                                             |
| --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| [`mmerterden/multi-agent-plugins`](https://github.com/mmerterden/multi-agent-plugins)         | Marketplace of per-stack skill toolkits (iOS / Android / Frontend / Backend + common). `/multi-agent:stack` enables the matching plugin.                                                                                                                                                                                                               |
| [`mmerterden/multi-agent-toolkit-mcp`](https://github.com/mmerterden/multi-agent-toolkit-mcp) | MCP server for UI testing / simulator capture / xcodebuild - powers the Phase 5 UI Bug Hunter. Published on the public npm registry as [`@mmerterden/multi-agent-toolkit-mcp`](https://www.npmjs.com/package/@mmerterden/multi-agent-toolkit-mcp); the installer registers it with each CLI for you, so `npx` resolves it with no extra configuration. |

## License

MIT - see [LICENSE](./LICENSE). Security issues: see [SECURITY.md](./SECURITY.md) (do not open public issues for vulnerabilities).
