# Features - Full Reference

Comprehensive list of every feature the pipeline ships. The top-level `README.md` only highlights the 5-7 most important; this file is the complete catalog.

## Contents

- [Core Pipeline](#core-pipeline)
- [PR & Review Flow](#pr--review-flow)
- [Review Quality](#review-quality)
- [Autopilot (unattended)](#autopilot-unattended)
- [Safety & Hygiene](#safety--hygiene)
- [Testing & Quality](#testing--quality)
- [Telemetry & Observability](#telemetry--observability)
- [Learning](#learning)
- [User-Defined Routines](#user-defined-routines)
- [Integrations](#integrations)
- [Schemas & Validation](#schemas--validation)
- [File Layout](#file-layout)

## Core Pipeline

### 6-Phase Orchestration (0-5)

```
Phase 0: Init      Project selection, branch setup, identity, worktree
Phase 1: Plan      Stack detection, codebase exploration (parallel Explore agents),
                   task decomposition, architecture review, user approval
Phase 2: Dev       TDD cycle: test → code → build (Sonnet), then the Verify exit
                   gate: build · lint · tests · secrets, run once
Phase 3: Review    Parallel AI review + Fable triage against Dev's logs, then the
                   optional user test + on-demand device audits
                   (Claude Code: Fable + Opus + Sonnet, or Opus + Sonnet with the
                   fable rung off · Copilot CLI: GPT-5.4 + Opus + Sonnet)
Phase 4: Commit    Git commit, push, PR with default reviewers + draft/ready prompt
Phase 5: Report    External: Jira comment · Wiki + Figma screenshots · Confluence
                   Internal: agent-log.md + Quality & Metrics + knowledge + memory
```

Each phase reads its own spec file under `pipeline/multi-agent-refs/phases/phase-N-*.md` - lazy-loaded so the orchestrator only pays the token cost for the phase it's currently in.

### Modifier Flag

| Flag        | Effect                                                                              |
| ----------- | ----------------------------------------------------------------------------------- |
| `autopilot` | Skip all confirmation prompts; still fails safe on review blockers + build retries. Also resolves the workspace to a worktree without asking, and turns on the verification gates described under [Autopilot (unattended)](#autopilot-unattended). |

The workspace is not a flag. **Where the branch lives** is asked at Phase 0 Step 5b - worktree (`.worktrees/{id}/`, your checkout untouched) or local (the project root, which drops the user-test offer because that gate checks the change out of a worktree and there is none). There is no `:local` command and no `--local` switch: an answer given in the run is visible in the run, where a flag is visible only in how the run was typed. Every autopilot entry resolves it to a worktree and never asks, because an unattended run commits and pushes from wherever it stands and doing that in the user's own checkout is what worktrees exist to prevent. `state.workspaceSource` records who decided - `localMode: false` alone is both "the user chose a worktree" and "nothing asked".

There is one pipeline: every mode runs its whole phase set, and no answer during the run adds or removes a phase.

### Outside a Pipeline Run

The install is not only useful while `/multi-agent` is running. `rules/outside-the-pipeline.md` loads with every session and announces three things a plain conversation would otherwise not know it had:

- **Onboarded service credentials.** Resolve the logical name through `credential-store.sh` and read the issue, page or log. Writes route through the pipeline commands, which carry the rules that make them safe - issues are never auto-closed, PR bodies use `Ref:`, outward prose goes through the humanizer.
- **The stack skills enabled for this repo.** Each toolkit's own `index` skill routes. The pipeline reads the effective `enabledPlugins` rather than keeping a stack table, so a seventh toolkit needs no code change.
- **The `multi-agent-toolkit` MCP.** 118 tools for a running app.

Uninstall preserves the whole layer - tokens, the reader that opens them, the mapping that names them, the MCP registration. It is 1.5 kB of always-loaded text; the detail lives in a ref that loads on demand, and a gate keeps both under a ceiling because every byte there is paid by every session.

### Code Graph (`/multi-agent:graph`, opt-in)

A deterministic, LLM-free map of what a repo declares and what refers to what, extracted by regex over comment-stripped source into `~/.claude/knowledge/<project>/code-graph.json`. Four stacks build today (Swift, Kotlin/Java, TypeScript/JavaScript, Python); each is one rules file, and the engine is the same for all of them. Zero runtime dependencies, zero API cost, read-only on the repo.

Phase 1 queries it to hand Explore a ranked starting file set instead of a full scan, and Phase 5 rebuilds it after the branch changed code - a rebuild is seconds, so staleness is a `baseCommit` comparison rather than a date heuristic. Off by default behind `prefs.global.codeGraph.enabled`; with it off the pipeline behaves exactly as before.

`GRAPH_REPORT.md` also ends with **Symbols nothing else references**: symbols no other file in the repo names, split from the ones referenced only by their own tests. Candidates, never verdicts - the extractor is regex, not a parser, so the four classes that could not carry a reference edge either way (a kind outside the stack's `referenceKinds`, a name declared twice, a nested declaration, a test file) are counted and excluded rather than listed, and nothing gates on the result.

Measured on a 4,300-file Swift app against a grep-and-read baseline at the same 30,000-token retrieval budget: 80.4% key-fact coverage at 18,465 tokens against 66.0% at 24,555. The gain is entirely in searches phrased in domain words (63.3% against 32.0%, at under half the cost). When the task already names an exact type, `grep -lw` is still slightly better and slightly cheaper, and the command says so rather than overselling. Reasoning, trade and limits: `docs/adr/0010-own-code-graph.md`.

### Stack Auto-Detection

| Platform       | Detection                                    | Guide Loaded             |
| -------------- | -------------------------------------------- | ------------------------ |
| iOS/Swift      | `.xcodeproj`, `Package.swift`                | `refs/swiftui-guide.md`  |
| Android/Kotlin | `build.gradle[.kts]`                         | `refs/android-guide.md`  |
| Backend        | `requirements.txt`, `package.json`, `go.mod` | `refs/backend-guide.md`  |
| Frontend       | `package.json` + framework detection         | `refs/web-guide.md`      |
| Docker         | `Dockerfile`, `docker-compose.yml`           | `refs/backend-guide.md`  |

Build commands, test runners, lint tools, and review focus areas all adapt to the detected stack.

### Stack Selection (marketplace plugins)

Stack skill sets ship as versioned plugins in the `multi-agent-plugins` marketplace. Selecting a stack enables the matching plugin(s) in the current repo's `.claude/settings.json` `enabledPlugins` - no skill copying, no session restart tricks, no directory shuffling. Two toolkits are always enabled alongside the stack plugin because neither is stack-specific: `ai-common-toolkit` (accessibility audit, humanizer, Firebase) and `ai-analyst-toolkit` (GitHub and package-registry evidence, community signal).

```bash
/multi-agent:stack ios         # ai-ios-toolkit (SwiftUI, Xcode, HIG)
/multi-agent:stack android     # ai-android-toolkit (Compose, Gradle, Hilt)
/multi-agent:stack mobile      # iOS + Android combined
/multi-agent:stack backend     # ai-backend-toolkit (spec-driven APIs)
/multi-agent:stack frontend    # ai-frontend-toolkit (React/TSX)
/multi-agent:stack fullstack   # backend + frontend
/multi-agent:stack all         # every stack plugin
```

### Project Scaffold (`/multi-agent:scaffold <stack> <name>`)

A new project starts green by the same commands every later change is judged by. The command resolves the enabled toolkit's `ai-<stack>-toolkit:scaffold` skill (ios, android, web, backend), dispatches it through the Skill tool, and then only verifies: `scaffold-gate.mjs --phase skeleton` requires a valid `.scaffold.json`, no commit and no remote yet, and the stack adapter's build, test and lint commands to pass, with build and test judged by `evidence-gate.mjs --stack` and a test count. The first commit follows a pass; creating the hosted repository is a separate question asked afterwards, and refused unattended. For a scaffolded repo, Phase 2 runs `--phase story` at entry and exit: build, test, lint and the manifest's demo command must pass, or the next story does not start. Detail: [`scaffold.md`](../pipeline/multi-agent-refs/features/scaffold.md).

### Package Manager Resolution (Phase 2, node-shaped stacks)

Phase 2's web test arm and its build step used to type `npm`. A repo on pnpm, yarn or bun then failed in Phase 2 - with a worktree and a branch already created - or, worse, npm resolved against a lock file it does not own and the run continued on a tree the repo's own tooling would never have produced.

`scripts/package-manager.mjs` resolves it from the repo instead: `$MA_PACKAGE_MANAGER`, then `package.json#packageManager`, then a lock file, then npm - reported AS a default, never as evidence, because "npm because nothing said otherwise" and "npm because the repo committed a package-lock" are different answers. Node core only (ADR-0004): a resolver that shelled out would need a working install of the tool it is identifying. The walk goes up to the directory holding `.git` and stops there, so a monorepo's root lock file is found and a stray one in a home directory is not. Two lock files means a migration left one behind: the newest wins and both are named.

Every manager gets the explicit `run` form (a script named `test` or `add` would otherwise lose to the builtin), only npm gets the `--` separator, and `bun run test` never `bun test`. Exit 3 means the repo declares no such script - the `--if-present` case, answered by an exit code rather than a flag whose support differs per manager. The resolved name is pasted into the phase's `eval`, so it is held to the shape a binary actually has, and the gate proves that by eval'ing the produced line with every manager stubbed out. iOS and Android are untouched. `refs/features/package-manager.md`.

### Maturity Follow-Up (`prefs.global.maturityFollowup`)

The maturity check has always produced a machine-readable gap list - stable codes in `blockers[]` and `warnings[]` - and then thrown most of it away. A blocker halted the run, an autopilot queue moved to the next item, and the item stayed exactly as immature as it was found. Nobody was told, so nothing changed, so the next scan halted on the same item for the same reason.

**Interactive runs ask at the step rather than ending at it**: open the item and fix it, continue without it, or abort. Continuing records which gap was waved through in `state.maturity.accepted[]` - that is what separates an informed continue from a skipped check. An answer typed into a picker improves this run and leaves the item as immature for the next person, so the step offers to write the supplied content back, as a separately approved write.

**Autopilot can ask on the item itself**, behind `autopilotCommentsOnIssue` (off by default, because it is an outward-facing write). One comment naming what is missing, then a halt on the circuit breaker with `state.waitingFor = "maturity"` so `resume` re-enters that step. A question, never a state change: no transition, no resolution, no assignee, no label, no close; `Ref:` never `Closes:`; copy in `outputLanguage`, with the gap wording taken verbatim from the fetcher's own summary rather than re-derived.

**An edit is a reason to look again, never proof the gap closed.** A reply reading "will do later" moves the timestamp and fixes nothing, so a changed item is re-fetched and re-scored and the check decides. Only a different gap set earns a second comment; "cannot tell whether it moved" re-checks rather than waiting, because folding unknown into "nothing changed" parks a run forever on a tracker that omits the field.

Warnings still auto-continue under autopilot - converting them to halts would stall queues on items that ran fine yesterday. `commentOnWarnings` raises them opt-in.

### Base-Branch Evidence (Phase 0 Step 3, `prefs.global.baseBranchEvidence.enabled`)

Step 3 used to ask one question with a list it could not vouch for. `git fetch origin` ran, its exit code was discarded, and `git branch -r` printed the remote-tracking cache either way - so on a restricted network a weeks-old local list was presented as the remote's answer, with nothing saying so. And the answer was usually derivable: an issue carrying a target version, or linking a separate issue that represents the release, already names the branch on a repo whose release branches encode the version.

`base-branch-candidates.mjs` collects candidates **with the evidence behind each one**, ranks them, and the picker row's description IS the evidence - "matches version 1.51.0 from the field Target Version" and "the repository's default branch" are different answers to the same question. A human still chooses; the derivation only reorders the rows.

Nothing is tabled, and that is the design rather than a detail:

- **No Jira field id is hardcoded.** A board's "target version" is a custom field whose id differs per instance. What is stable is the schema: any field resolving to type `version` is read, whatever it is called. A linked issue or parent whose own fix-version or summary names a version is the second source, and `baseBranchEvidence.preferLinkedRelease` raises it above the version field on boards where selecting the release issue is what opens the branch.
- **No branch prefix is tabled.** The release-branch template is inferred from the refs that exist, so one repo yields `<prefix>/develop_<version>` and another `release-<version>` out of the same code. The rule-5 filter carries a version alternative for the same reason: a word list that ranks is fine, one that discards is the prefix table this replaces.
- **A predicted branch is a note, never an option.** "That version has no branch on the remote yet" is a real answer; an option the user picks has to be checkoutable.
- **A failed fetch degrades loudly.** `refProvenance` rides on every candidate, the picker says the refs may be stale and gains a retry row, and `phase0-exit-gate.mjs` refuses to close Phase 0 if a degraded fetch recorded its list as `remote`.

Autopilot resolves `remembered` → `derived` → `default` and records which fired; `derived` requires issue evidence for the branch it chose. With `baseBranchEvidence.autopilotAsksOnIssue` (**off by default**) an ambiguous derivation posts one comment on the Jira or GitHub issue asking which branch, then halts on circuit-breaker trigger 6 and waits for `resume`. A question, never a state change: no transition, no close, `Ref:` never `Closes:`, copy in `outputLanguage`.

### Task Type Detection

Phase 0 Step 9 classifies every task before Phase 1 starts. Deterministic priority order: Figma URL → instruction file path → git diff heuristic → Jira issue type → branch name → description keywords → user prompt (autopilot defaults to `feature`).

Result persisted to `agent-state.taskType`:

| Type        | Downstream effects                                                            |
| ----------- | ----------------------------------------------------------------------------- |
| `component` | Phase 2 dispatches to the marketplace component plugin (create-component) with SubPhase reporting |
| `bugfix`    | Phase 3 emphasizes test coverage + regression; Phase 4 uses `fix(...)` prefix |
| `feature`   | Standard TDD flow; Phase 4 uses `feat(...)` prefix                            |
| `refactor`  | Phase 3 emphasizes behavior preservation; Phase 4 uses `refactor(...)` prefix |
| `chore`     | Lightweight flow; Phase 4 uses `chore(...)` prefix                            |

### SubPhase Convention

When a specialized skill takes over a main pipeline phase, progress is reported as SubPhases (e.g. `SubPhase 3.0: Init`, `SubPhase 3.1: Gather`). The top-level pipeline stays fixed at 6 phases (0-5) - specialized work slots into its parent phase without inflating the count.

## PR & Review Flow

### Default Reviewers

- **Bitbucket**: Fetches via `/rest/default-reviewers/1.0/`. Every PUT must re-send `reviewers`, `fromRef`, `toRef`, `draft` (regression guarded by smoke test).
- **GitHub**: Honors `CODEOWNERS` + falls back to `prefs.projects[].githubDefaultReviewers`.
- PR author always filtered out (Bitbucket: 409, GitHub: GraphQL error).

### Draft vs Ready Prompt

Phase 4 asks `READY or DRAFT?` before creating the PR, recommends READY because every gate has passed by then, and persists the choice in `prefs.projects[].defaultPrMode`. An unattended run does not ask: the autopilot runner always opens a draft, and a person marks it ready.

- Bitbucket: `draft: true` flag (DC 8.x+) with `[DRAFT]` title fallback for older servers.
- GitHub: `gh pr create --draft` + `gh pr ready` for promotion.

### `channels` Command

Multi-channel reporter - Phase 5 delegates to it, and it's also invocable post-hoc for fixes closed outside the pipeline:

```bash
/multi-agent:channels                              # current branch, current PR
/multi-agent:channels https://jira.company/browse/PROJ-12345
/multi-agent:channels #42 --channels pr            # PR only
/multi-agent:channels --message "manual fix description"
/multi-agent:channels ABC-1234 --channels jira,confluence --content test
```

Multi-select **channels** (Jira / Confluence / Wiki / PR description) × multi-select **content** (normal analysis / test scenarios / auto-diff summary / manual note). Each body runs through the humanizer skill per-channel. Bitbucket PR updates use the reviewer-preserving PUT pattern (title + description + reviewers + fromRef + toRef + version mandatory). Replaces the earlier `enrich` command - all its capabilities (diff auto-summarize, manual mode, reviewer-preserving) are preserved; Confluence + Wiki are new.

### Body Preservation Contract (smoke-verified)

Every external-system body (PR description, Jira comment, GitHub issue) uses `jq -n --rawfile body body.md '{description: $body}'` → `curl --data-binary @payload.json`. No literal `\n` strings, no HTML entities (`&amp;`, `&lt;`, `&quot;`). UTF-8 preserved end-to-end. `scripts/smoke-add-detail.sh` runs 14 contract assertions.

### Issue Safety

Never auto-closes issues - uses `Ref: #N` / `Related: #N` / `See: PROJ-12345`, never `Closes` / `Fixes` / `Resolves`. Closure requires team review (configurable, typically 4 approvals).

## Review Quality

### Deterministic Gates (Phase 3 Step 1)

Cheap, objective checks run BEFORE any AI token is spent:

1. Build (acquires xcodebuild lock, isolated DerivedData per worktree)
2. Lint (SwiftLint / detekt / ruff / eslint - stack-dependent)
3. Tests pass
4. Secret scan

If any gate fails, fix first. Don't waste AI tokens reviewing broken code.

### Analysis Document Review (Phase 2.2 + 2.3)

`/multi-agent:analysis` published behind a structural validator alone until v16.12.0: nothing read the
document before it reached Confluence. Phase 2.2 now runs the same reviewer set and triage a code diff
gets, on the draft, before the destination is even chosen. Its first question is what the run skipped -
an input declared missing that nothing searched for, an open question about evidence nobody read, a gap
with no owner, a scope call made without asking. A blocking finding returns to synthesis with dispatch
closed; it never becomes an open question, because "the document is wrong" is not something to ask the
reader.

Phase 2.3 then sorts what is left: reachable evidence is searched (never asked about), decisions the
user owns are asked with `AskUserQuestion`, and only genuinely external gaps enter the document as
`AS-NN` rows with an owner. A gap carrying neither a `searched, not found` nor an `asked, external`
stamp fails the dispatch gate. Autopilot runs both phases; only the asking degrades, into rows stamped
`autopilot: could not ask`.

### CLI-Aware Parallel Review + Fable Triage (Phase 3 Steps 2-3)

| Reviewer   | Model               | Focus                             | Where it runs        |
| ---------- | ------------------- | --------------------------------- | -------------------- |
| Reviewer 1 | `claude-fable-5` (Claude Code) / `claude-opus-4-8` (Copilot CLI) | Deep security + architecture | Both CLIs |
| Reviewer 2 | `gpt-5.4`           | Edge cases, different perspective | **Copilot CLI only** |
| Reviewer 3 | `claude-sonnet-4-6` | Quality + correctness + naming    | Both CLIs            |

The reviewer set is **CLI-aware**: Claude Code dispatches 3 reviewers in parallel (Fable + Opus + Sonnet - Opus fills the slot GPT-5.4 takes elsewhere); Copilot CLI dispatches all 3. Each returns structured JSON for deterministic aggregation. Cross-model diversity catches blind spots that any single model family would miss.

**Fable Triage** (Phase 3 Step 3, Opus on Copilot CLI): Evaluates merged raw findings against task scope. Classifies each as `accepted` (fix now), `deferred` (out of scope, log for later), or `rejected` (false positive / noise). Only triage-accepted blocking items loop back to Phase 3.

### Runtime Triage Validator

After triage returns, output is validated by `validate-triage.mjs`:

| Exit  | Meaning                                                      |
| ----- | ------------------------------------------------------------ |
| **0** | Valid and clean - act on triage output                       |
| **1** | Invalid structure - retry once, then fallback                |
| **2** | Over-rejection guard tripped - pause for human               |
| **3** | Contradiction auto-corrected - proceed with corrected output |

### Bidirectional Approved↔Blocking Auto-Correction

If triage returns `approved: false` but has no blocking items, the validator forces `approved: true`. Conversely, if `approved: true` but blocking items exist, it forces `approved: false`. Hardened with an `if`/`then` constraint in the schema itself.

### Verify-by-Test Triage (Phase 3 Step 3.7, opt-in)

A triage verdict is a judgment call; a failing repro test is proof. When `prefs.global.verifyByTest.enabled` is on, one verifier agent (default Sonnet) writes a minimal repro test per accepted blocking finding (cap: `maxFindings`=3) and runs only that test. Fails as predicted -> finding confirmed, the repro test becomes the Phase 2 rework RED test. Passes under `evidence-gate.mjs` -> finding downgraded to `deferred`. Compile error / timeout -> `inconclusive`, judgment stands. Timeout-bounded, never blocks. Full spec: `refs/features/verify-by-test.md`.

### Immutable-Test Rule + `test_lines_removed` Signal

Existing tests are immutable during a task: deleting, renaming, or weakening an assertion to reach green is a violation (`refs/rules.md`, Phase 2 GREEN step). A test changes only when the task changes the spec it encodes, named in the commit body. Deterministic backstop: `diff-risk-score.mjs` emits `test_lines_removed` (w=3.0) for any test-classified file whose diff removes more lines than it adds.

### Update Check at Run Start

Phase 0 Step 0.6. Once per `ttlHours` window (cached, 3s-bounded curl to the npm registry), the installed version is compared against two dist-tags.

**`latest` - automatic**, since v16.5.0. `prefs.global.updateCheck.autoUpdate` defaults to `true`: a newer version is installed before the run starts, in interactive modes and autopilot alike, and the run continues. Set `autoUpdate: false` to be asked once per `ttlHours` instead, or `updateCheck.enabled: false` to silence the check (neither disables the required-version floor).

**`required` - blocking** (v15.14.0+). Most releases do not publish this tag and nothing changes for them. A release that changed a contract a run depends on is promoted with `npm dist-tag add <pkg>@<version> required`, and an install below that floor is not behind, it is wrong: `require-supported-version.sh` exits 3, the run halts, `/multi-agent:update` runs, and the user re-issues the command on the new version rather than continuing on docs already loaded from the old one. Interactive and autopilot behave identically. The gate fails open on every undeterminable answer (offline, blocked registry, no tag), `updateCheck.enabled: false` does not disable it, and the single override is the env var `MULTI_AGENT_ALLOW_OUTDATED=1`, which is logged in the run record. Exemptions: `update`, `setup`, `uninstall`, `help`, `status`, `log`, `search`, `routines`, `forget`, `language`.

### Structured Handoff Blocks

Every phase transition appends a `## Handoff` block (Done / Remaining / Decisions / Open findings / Next) to `agent-log.md` - orchestrator-written from existing state, no LLM call. `/multi-agent:resume` and post-`/compact` re-grounding read the latest handoff first, so long runs re-enter from durable artifacts instead of conversation memory (fresh-context discipline from Anthropic's long-running-agent harness guidance).

### Accessibility Code Review (Phase 3 Step 1.5)

If changes include UI files, reviewers check for:

- Missing `.accessibilityLabel` / `contentDescription` on interactive elements (→ blocking)
- Small tap targets (<44×44pt iOS / <48×48dp Android) (→ important)
- Missing identifiers + Dynamic Type support (→ suggestion)

Pure code analysis - no simulator needed. Device-level audits run in Phase 5 when requested.

### Status Enforcement

Phase 2 treats the issue-tracker status update as a required step with a post-mutation verify step that re-reads the field and retries once on silent `VALIDATION` failures (e.g. stale Projects V2 option IDs after a board rebuild).

## Autopilot (unattended)

Continuous mode (`/multi-agent:autopilot-on`) runs the pipeline with nobody watching. The runner is a launchd tick: it takes the head of the queue, runs the whole pipeline on it in a worktree with `MULTI_AGENT_UNATTENDED=1` set on the child, and ends at a draft pull request that a person reviews and merges. Every behaviour in this section keys on that variable or on autopilot mode; an attended run is unchanged. Contract, entry point by entry point: [`unattended-contract.md`](../pipeline/multi-agent-refs/unattended-contract.md).

### Research Before Asking

A run that parks on a maturity blocker or on open analysis questions gets a research pass before it waits for a person. `/multi-agent:research <id> --autonomous` reads the tracker thread, issue links, similar closed items, linked documents and the repository; `research-gate.mjs` decides what closed. `inference` alone never closes a gap, `evidence` must be quoted from a fetched source, `repo` must resolve at `file:line` or in a commit, and acceptance criteria and business rules only ever yield candidates. The maturity check or the open-questions gate then re-runs on the result: `proceed` resumes the run, anything else parks it with the gaps left. `maxAskRounds` bounds the rounds per item. Detail: [`research.md`](../pipeline/multi-agent-refs/features/research.md).

### The Runner Publishes

The session commits and stops. Phase 4 writes a PR request through `pr-request.mjs`, and `autopilot-publish.mjs`, in the runner process, decides whether it leaves the machine: the repository comes from the runner's config, the branch must be the one the run recorded and must be unprotected, the stack adapter's build and test commands are re-run in the worktree with token-shaped variables removed, the gate ledger, the secret detectors over the commit range and the outbound gate are re-checked, and the push goes from a fresh bare staging repository that carries none of the worktree's git config. GitHub gets `gh pr create --draft` with the body from a file; Bitbucket gets the branch and a summary (`pushed-awaiting-pr`). Any failed check is `verification-failed`, with the verdict in `pr-requests/<session>.verdict.json`. Post-PR reports (`reportChannels`, `reportIssueUpdates`) run once, after the PR opened, and are off by default.

### Operations Around a Run

| Operation | Default | What it does |
|---|---|---|
| Awake agent | on | `com.multi-agent.autopilot-awake` runs `caffeinate -s -i` from `autopilot-on` to `autopilot-off`; `awake.display` adds `-d` |
| Sleep lock | on | `caffeinate -s -w <runner pid>` from the launch to the end of the tick |
| Credential probe | on | `autopilot-arming.mjs probe` before the schedule is offered, `check --probe` on every tick; `credential-store.sh probe` answers `readable`, `missing` or `locked` without reading a value |
| Circuit breaker | on | stops launching after consecutive attempts that produced nothing, and lets one attempt through per cooldown |
| Cost ceiling | on | `costCeilingUsd` over a rolling 24 hours, each run counted once at its largest recorded spend |
| Parallel cap | unlimited | `maxParallelAgents` caps runner-launched sessions working at once |
| Cleanup report | dry run | `gc-report-<date>.json` / `.md` over the worktrees the runner created; `gc.autoDelete` removes |
| Daily digest | off | items, outcomes, PRs, parked items and cost, through `reportChannels` after the outbound gate |

Detail: [`autopilot-operations.md`](../pipeline/multi-agent-refs/features/autopilot-operations.md), [`autopilot-circuit-breaker.md`](../pipeline/multi-agent-refs/features/autopilot-circuit-breaker.md).

### Security Base

- **Fail-closed guard.** `agent-guard.sh` is the PreToolUse hook on `Bash`, `Edit|Write|NotebookEdit` and `WebFetch|mcp__multi-agent-toolkit__.*`. Under the variable a Bash command is judged only when it reduces to simple commands joined by `;`, `&&`, `||` and plain pipes between non-interpreter commands; a subshell, group, nested substitution, unquoted heredoc, background job, shell keyword, interpreter reading stdin, shell `-c`, inline interpreter program (`node -e`, `perl -e`, `php -r`, ...), program file the run wrote or modified, or runtime-built command name is refused unparsed. On what it parses: no push, PR, issue or tracker write, no Keychain read, no package install or manifest edit, no write into a protected path. The guard arms its own 5-second deadline and answers "block" on expiry when unattended.
- **Network allowlist.** `curl`, `wget`, `WebFetch` and the toolkit's `web_*` tools reach only hosts in the allowlist (`network.staticHosts` plus the configured service hosts); a URL built at run time, a proxy flag or a request body is refused, and `web_eval` is refused outright. The runner starts the toolkit with `MCP_TOOLKIT_URL_POLICY=strict` and `MCP_TOOLKIT_INDEX_DENY`.
- **No package installs.** A restore from the committed lockfile is allowed; adding a package is not.
- **Untrusted data.** Fetched bodies, ticket text, comments and PR text enter prompts inside `<untrusted-data>` delimiters that defuse a forged closing tag.
- **Permission profile.** `install --unattended` writes a `dontAsk` profile with a narrow allow list and deny rules into `~/.claude/multi-agent-unattended.settings.json`, printed in full before anything is written; the runner passes it with `claude --settings`, so `~/.claude/settings.json` and attended sessions are unchanged. A default install writes no permissions.
- **Machine setup, applied by the operator.** A separate non-admin macOS user for the runner, per-repo GitHub credentials limited to contents and pull requests, a Jira token that can only read and comment, and GitHub rulesets on every default and release branch (no force-push, no deletion, PR and code-owner review required, the bot in no bypass list).

Detail: [`unattended-security.md`](../pipeline/multi-agent-refs/features/unattended-security.md).

### Verification Gates

While the quality gates are active (`MULTI_AGENT_UNATTENDED=1`, or `state.autopilot` for a terminal autopilot run), each gate appends its verdict to `state.gates[]` through `gate-ledger.mjs`, and under the variable the commit hook refuses a commit unless every mandatory gate passed or was not applicable for HEAD.

| Gate | What it checks |
|---|---|
| `symbol-existence` | every localization key and endpoint the added lines reference exists in the repository's truth sources |
| `open-questions` | open analysis questions are answered, assumed or park the run |
| `verify-citations` | a triage citation quotes the line it cites (at least 8 non-space characters, or the whole line) |
| pre-existing claims | a "pre-existing" finding cites a line the diff from the base does not add or change |
| `plan-coverage` | every plan step is accounted for; a planned run with no todos is a gap |
| `evidence-gate`, `test-summary` | build and test claims read with the stack's own markers and a positive executed test count |
| `test-strength` | a new test goes red without the change it claims to cover |
| `review-decision` | a `blocking` finding carries two independent reviewers or a failing test |
| `spec-consistency`, `plan-critique` | recorded, not mandatory at commit (see below) |

Stack adapters cover 8 stacks: `ios`, `android`, `web`, `backend-node`, `python`, `go`, `rust`, `jvm`, plus `unknown`. Detail: [`unattended-gates.md`](../pipeline/multi-agent-refs/features/unattended-gates.md), [`stack-adapters.md`](../pipeline/multi-agent-refs/features/stack-adapters.md).

### Constitution and Spec Consistency

`/multi-agent:analysis` writes the project's invariant rules (security, accessibility, architecture, licensing) to `constitution.md` and `constitution.json` under the project's knowledge directory. Each rule cites one labelled source: `repo` (a file and line), `evidence` (a standards document or tracker item) or `inference`, which is always `proposed` and never binding. `spec-consistency-gate.mjs` checks by id, with no model call, that every requirement has a task and a Test Plan row, every task cites a defined requirement, and no row or task waives a binding rule. Attended it prints an advisory report and writes nothing. Detail: [`constitution.md`](../pipeline/multi-agent-refs/features/constitution.md).

### Plan Critic

With the gates active, Phase 1 ends with one adversarial critic (`agents/plan-critic.md`) under four lenses: scope, feasibility, security and a cheaper alternative. Each objection carries evidence anchors; the planner answers each once, accepting and revising or rebutting with evidence, and `plan-critique-gate.mjs` judges with no model call. An objection citing a binding constitution rule blocks unless its reply resolves it with an anchor that holds; the rest is advisory and appears in the plan render and the PR summary. One round, because further rounds move agents toward each other rather than toward the right answer. The critic follows the fable switch. Detail: [`plan-critic.md`](../pipeline/multi-agent-refs/features/plan-critic.md).

### Review Decision Rule

Every reviewer on Claude Code is a Claude model, so with the gates active a `blocking` finding keeps its severity only when two independent reviewer outputs carry the same finding, a confirmed verify-by-test log shows a failing test, or the test-integrity gate produced it. Anything else is lowered to `important` in place, never dropped, with the reason in the report and the ledger. The rebuttal round does not run under the gates. Attended review is unchanged. Detail: [`review-decision.md`](../pipeline/multi-agent-refs/features/review-decision.md).

### Client Contract and Phone API

Commands (`commands.mjs --json`) and run questions (`launch-request.mjs questions`) are declared as data, and every surface is reachable two ways that print the same bytes: the CLI scripts, and `contract-server.mjs`, started with `/multi-agent:serve` on 127.0.0.1 behind a bearer token. `pipeline/contract/` ships the manifest, generated TypeScript types and fixtures ([kit README](../pipeline/contract/README.md)). A launch request's input is a reference, not a prompt: a Jira key, an issue URL, `repo#N` or `#N`, or desktop free text passed as one quoted argument; op and mode words are refused. Four phone routes under `/v1/phone/` (runs, one run, answer, launch) accept only Ed25519-signed requests from devices enrolled with `phone-devices.mjs`, always return the redacted view, and keep launch off until `phone-devices.mjs launch on`. The phone app is a separate project. Detail: [`phone-api.md`](../pipeline/multi-agent-refs/features/phone-api.md).

## Safety & Hygiene

- **Pre-Commit Secret Detection** (12 patterns): `PreToolUse` hook scans staged files for API keys/tokens, AWS access keys, private keys, `.env` files, service account JSON. Commit **blocked** if found.
- **Read-Size Gate** (opt-in, `prefs.global.bulkRead.mode`): a `PreToolUse` hook inspects `Read` and the shell commands that read a file whole. In `observe` it only logs what it would have caught - the baseline you measure before routing anything. In `enforce` a file over `minLines` (default 350) is blocked and delegated to a haiku-rung worker (`bulk-read.sh`), which returns a line-numbered summary so the follow-up is a bounded `Read(offset:limit:)` instead of the whole file; the full text is parked under `.multi-agent/refs/`. The development phase and any file the run has already touched are exempt, because Claude Code's `Edit` requires its own `Read` first.
- **Capture Hooks** (`SessionEnd`, `PreCompact`, `SessionStart`): every durable write used to live in Phase 5, the phase a run is least likely to reach. `SessionEnd` flushes a run that never got there; `PreCompact` flushes before an auto-compaction summarizes a long phase mid-flight, which is the same loss one level down; `SessionStart` prints at most two lines about an unfinished run. None calls a model, none reads a payload, and all exit 0 on every path - a hook that fails a session over bookkeeping is worse than the bookkeeping.
- **Operational Reporting** (`prefs.global.usageLog`): coarse run metadata - task id and type, current phase, run status (in progress, completed, halted, parked), phase durations, whether a PR opened, token counts - sent at every phase transition and at completion, halt and park, and never prompts, code, diffs or absolute paths (task title and PR URL only when `includeTitle` / `includePrUrl` are set). Runs are reported under the GitHub account on the machine that can read the pipeline repository, found with no question asked through gh, the git credential helper or an SSH key (`usage-identity.mjs`, cached in `usageLog.login` / `loginMethod`, `--refresh` to re-check); with no such account nothing is sent. The per-machine token is REQUESTED from the endpoint by `usage-register.mjs` (setup, update, and the Phase 0 exit gate as a backstop), is write-only, and lives in the OS credential store; prefs hold only the entry name and the switch. `usageLog.optOut: true` blocks registration permanently and is checked before the network call. Registration proves the login with that account's gh or credential-helper token, or with an SSH signature over a server challenge. The server issues a token only to the repository owner or a collaborator, pins the login to it, and shows only those accounts' records on the panel. An unreachable endpoint leaves reporting off with one line and exit 0 - a run is never failed over bookkeeping.
- **Build Queue**: All `xcodebuild` calls acquire a lock. Each worktree uses own `-derivedDataPath`. Stale locks auto-clean after 15 min. Non-Xcode builds don't need the lock.
- **Context Management**: `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=65` - compaction at 65% usage (prevents degradation in 6-phase sessions).
- **3-Iteration Hard Kill**: Any retry loop stops after 3 attempts, then pauses for user. No infinite loops.

## Testing & Quality

### Schema-Validated State

All critical state files are schema-validated at read and write time:

- `agent-state.schema.json` - validates `$HOME/.claude/logs/multi-agent/.../agent-state.json`
- `prefs.schema.json` - validates `$HOME/.claude/multi-agent-preferences.json`
- `triage-output.schema.json` - validates triage output (contradiction `if`/`then` constraint built in)

### Smoke Test Suites

100+ suites, auto-discovered from `smoke-*.sh` (no hardcoded list). Representative: add-detail (body-preservation assertions), review-triage (validator exit codes), prefs (schema round-trip), state (agent-state lifecycle), metrics (telemetry emission), sync (instruction parity), secret-scan (hook patterns), phase-banner (terminal UI), token-budget (per-phase limits), phase-tracker (progress tracking).

### Adversarial Eval Fixtures

Adversarial fixtures that test triage resilience against adversarial reviewer output: over-rejection, hallucinated findings, contradictions, invalid JSON, schema violations, duplicate findings, scope creep, empty results, timeout simulation, and combined edge cases.

### Sync Parity Check

Detects drift between Claude Code instructions (`~/.claude/commands/multi-agent/SKILL.md`), Copilot CLI instructions (`~/.copilot/copilot-instructions.md`), and the repo's pipeline spec files. Reports discrepancies during Phase 0 Init.

### Exploratory Testing and Bug Bash

`/multi-agent:test` runs every exploratory session under one contract (`features/exploratory-testing.md`, `explore-findings.mjs`, `explore-run.schema.json`): each finding has expected, observed, repro steps, severity 1-5, kind `issue` or `warning`, and evidence on disk; each run has a step, minute and token budget and stops after three consecutive blocked steps that reported nothing; steps blocked by environment, credentials, seed data or the automation itself are reported apart from product failures. Credentials are typed by `secret_ref`, and locators plus expect steps replace coordinates when the toolkit offers them.

A visual or semantic assertion is decided by a fresh judge subagent that receives only the assertion and the current screenshot or UI tree (`features/assertion-judge.md`); `explore-findings.mjs judge-packet` refuses any field describing the run, and an inconclusive or uncited answer fails as `ASSERTION_INCONCLUSIVE`.

`/multi-agent:bug-bash` fans out 5-10 charters, each with a persona and its own budget, merges the findings, triages them against the source, and proves every candidate with a reproducing test. `bug-bash.mjs report` confirms a finding only when its repro failed with `ASSERTION_FAILED` on the assertion that encodes it; everything else is listed as rejected with the reason (artifact, environment, design, fixture, did-not-reproduce, unproven, unverified). The `e2e-testing` knowledge skill in ai-common-toolkit carries the toolkit side: locators, expect verdicts, traces, replay, parallel web sessions and synthetic MRZ document fixtures.

### Token Budget Enforcement

Per-phase token budgets prevent runaway sessions. If a phase exceeds its budget, the pipeline pauses and offers: continue (extend budget), skip phase, or abort. Budgets are configurable in `prefs.global.tokenBudgets`.

## Telemetry & Observability

- **Pipeline Metrics**: Structured metrics to `metrics.jsonl` via `log-metric.sh`. Aggregated by `aggregate-metrics.mjs`.
- **Cost Telemetry**: Per-phase token cost tracking (`tokens_in`, `tokens_out`, `model`, `duration_ms`). Omitted fields handled gracefully.
- **Phase Tracker**: Cross-CLI visual progress (current phase, elapsed time, iteration count).
- **Phase Banner**: Terminal UI for phase transitions with Unicode box-drawing characters.
- **Per-task Cost Breakdown in agent-log.md**: Phase 5 appends a 4-column block (Phase · Model · Tokens in/out · Est. USD) to every run's `agent-log.md`. Sourced from `phase-tracker.sh tokens` accumulators × `cost-table.json` prices. Independent of the channels-side `reportContent.costSummary` toggle. The `LOG_METRIC_FORWARD_TO_TRACKER=1` env flag mirrors `tokens_in`/`tokens_out`/`model` from `log-metric.sh` into the tracker so JSONL metrics and the cost block stay in sync from one call site.

### Diff Risk Scoring

`pipeline/scripts/diff-risk-score.mjs` runs at Phase 3 Step 1.75 - before reviewer dispatch. Heuristic, deterministic, sub-second, no LLM. Top-N risk-ranked files inject into each reviewer's prompt as a `${PRIORITY_FILES}` block; reviewers read those files first but still review the entire diff.

Signals + weights: `security_path` ×3, `migration` ×4, `public_api` ×2, `no_test_change` ×2.5, `test_lines_removed` ×3 (test file shrinks - immutable-test backstop), `complexity_delta` ×1.5, `ui_critical` ×1.5, `loc_changed` ×1. Toggle via `prefs.global.diffRiskAdvisory` (default ON).

### Test Gap Detection

`pipeline/scripts/test-gap-scan.mjs` runs at Phase 3 Step 0. Walks the diff for newly added public symbols and reports those with no paired test. Stack-specific rules ship for iOS, Android, Python, Node.js. iOS Views and Android `@Composable` symbols default to `important`; other public API additions to `suggestion`. Optional gating via `prefs.testGap.blockingThreshold` - when set, the report becomes a Phase 4 rework finding once `important + blocking` count exceeds the threshold.

### Visual Evidence (UI changes)

A UI change carries its own picture. `state.visualEvidence.required` is decided mechanically from `taskType` plus the changed-file list, never from a reading of the task.

**Stills.** The "before" is the reporter's own ticket attachment, harvested in Phase 0; the pipeline never rebuilds the old state to photograph it. The "after" is captured in Phase 2 right after the build goes green, not in the user test, which autopilot and both local modes drop. `capture-evidence.sh` cleans the status bar and downscales to 1242px so two captures of one screen differ by the change and not by the clock.

**The flow video rides on a test run.** `probe-evidence-capability.sh` measures the UI test target, the tests matching this change, the device, the recorder and the MCP registration; Phase 0 Step 7.7 then asks the depth with the options built from that measurement, and a closed option keeps its row and states why. Tier 1 runs the repo's own UI test and records around it, tier 2 drives the flow through `agent_run_steps`, tier 3 records nothing and says so. The tier is re-checked before the recording starts, because a simulator booted at intake can be gone by Phase 2.

UI test detection keys on `XCUIApplication` rather than on a folder named `*UITests`: in a real app the overwhelming majority of files under such a path are snapshot tests, which never launch the app and would produce a still frame filed as a flow.

**Where it lands.** Jira takes both stills and video as attachments. With no Jira the stills go to an orphan `evidence/<task-id>` branch and the PR body embeds them, or links them with a blob permalink when the repo is private (GitHub's image proxy has no credentials for a private repo, and a broken image reads as missing evidence). Phase 4 blocks when a required artefact is neither published nor explained; the gate is against silence, not against an honest "the ticket carries no image".

Toggle via `prefs.global.visualEvidence.enabled` (default ON), `visualEvidence.githubHost`, `visualEvidence.maxAttachmentMb`, `visualEvidence.maxVideoSeconds`, `prefs.global.testDepth.default`.

### Triage Memory

Per-repo append-only JSONL corpus at `~/.claude/memory/multi-agent/<repo-slug>/triage-corpus.jsonl`. Phase 5 ingests every triage output (idempotent), Phase 1 enriches the analysis with similar past tasks, Phase 3 triage attaches prior-art hits to each raw finding with an explicit bias hedge. Token-overlap recall, zero deps, Node-18-compatible. `/multi-agent:search "<text>" --semantic` routes the query to the corpus instead of agent-log grep. Toggle via `prefs.global.priorArtEnrichment.enabled` (default ON).

## Learning

### Knowledge Base (per project)

Incremental learning. Phase 5 captures architecture, patterns, gotchas, and decisions into `$HOME/.claude/knowledge/{project}/`. Phase 1 reads it on the next run. Token cost decreases over time as the base grows.

### Memory Capture (cross-session)

Pipeline learns behavioral signals (feedback corrections, project constraints, external references). Phase 5 saves, Phase 1 injects. Max 3 new memories per run. Merge-over-duplicate. Stale memories verified before use.

**What does NOT go in memory**: architecture, code patterns, build gotchas, design decisions - those belong in the knowledge base.

### Lesson Diagnosis (Reflexion)

Phase 3's lesson-memory loop records the causal root cause of each fix (`--diagnosis`), not just the outcome: the verbal "why" that prevents recurrence (Reflexion). `learnings-ledger.mjs brief` renders it as `(why: ...)` back into Phase 1 + triage on the next run, so the reason re-enters the loop, not only the symptom.

### Corpus Freshness Gate

Each triage-corpus row is stamped with the file's git `file_sha`; a query annotates `stale=true` when the file changed since the lesson was recorded, so lessons about code that has since moved on stop resurfacing.

### Learning Curve

`learning-curve.mjs` renders a time-bucketed trend over `metrics.jsonl` (first-pass clean rate, review cycles, rework per task, tokens per task, cache ratio) so a repo's runs can be shown getting better and cheaper over time. Flags: `--bucket=<days>`, `--since`, `--json`, `--markdown`.

## User-Defined Routines

Turn a recurring, project-specific job into a first-class `/multi-agent:<name>` command.

- **`/multi-agent:save`** distills candidate routines from the work just done this session (and named procedures in `~/.claude/CLAUDE.md`), offers them in a multi-select picker (pick one, combine several into one, or free-text a new one), and registers the chosen routine.
- **`/multi-agent:routines`** lists saved routines (in `outputLanguage`); **`/multi-agent:forget`** removes one (guarded: never touches a shipped command).
- Backed by `routine-registry.mjs`. Saved routines are `local-only: true` command dirs + a `prefs.global.routines` entry: preserved across `/multi-agent:update` by install snapshot/restore, never synced to the public repo, and never counted in the command inventory.

## Integrations

### Figma / Component Generation (dispatched to marketplace plugins)

Component + Figma-to-code work is no longer bundled in this repo. When Phase 0 classifies a task as `component`, Phase 2 dispatches it to the per-stack marketplace plugins (`ai-ios-toolkit` / `ai-android-toolkit` in the `multi-agent-plugins` marketplace) via the Skill tool. The plugin's component skill generates `{Name}Configuration.swift`, `{Name}View.swift`, `{Name}+Modifiers.swift`, `{Name}.figma.swift`, and `FIGMA.md` with a variant matrix, then runs a 14-item pre-commit checklist covering design tokens, accessibility, tests, and Code Connect.

The plugin's cross-cutting integration skills feed component detection + implementation when the design triggers them (content: form / price / ui-patterns; interaction: navigation / overlays / bottom-sheets). Each is native-SwiftUI-first and reads project specifics (token namespaces, component paths, UI systems) from `figma-config`, including the optional `ui.navigationSystem` / `ui.overlaySystem` / `ui.sheetSystem` hooks (absent -> stock SwiftUI), so the same capabilities work on any SwiftUI codebase. The plugin's evolve-component skill reconciles an existing component against current Figma (drift-heal) and additively extends it, behind a human gate.

### UI Bug Hunter + Audits

Automated visual testing and compliance audits via direct Bash (no MCP server dependency):

| Audit                 | When                | Command                      |
| --------------------- | ------------------- | ---------------------------- |
| iOS Accessibility     | Phase 3, on request | `swift ui-tree-dumper.swift` |
| Android Accessibility | Phase 3, on request | `adb shell uiautomator dump` |
| iOS Biometric         | Phase 3, auth flow  | `xcrun simctl keychain`      |
| Android Launch Time   | Phase 3, perf       | `adb shell am start -W`      |
| iOS Archive           | Phase 4, release    | `codesign`, `plutil`, `nm`   |
| Android APK           | Phase 4, release    | `aapt2`, `apksigner`         |

Audits are **on-demand** - triggered by user, never automatic.

### Jira + Confluence

- Phase 2: transition issue to `In Progress` (verified post-mutation).
- Phase 5: post analysis + test scenarios as Jira comment (Turkish by default, configurable).
- Phase 5 (optional): create Confluence page under chosen parent, cached per project.

### Keychain

Token registry maps logical names (`jira`, `bitbucket`, `github`, `confluence`) to Keychain item names. Tokens never land in a config file, never synced to the repo.

## Schemas & Validation

- `pipeline/schemas/agent-state.schema.json` - validates agent state lifecycle
- `pipeline/schemas/prefs.schema.json` - validates preferences
- `pipeline/schemas/triage-output.schema.json` - validates triage output (contradiction `if`/`then` constraint built in)

## File Layout

Each subcommand is its own directory: `commands/multi-agent/<name>/SKILL.md` → slash command; the dispatcher is `commands/multi-agent/SKILL.md`. The user-facing invocation `/multi-agent:<name>` is unchanged. Guides, phase specs, and rules live under `pipeline/multi-agent-refs/**` (kept out of `commands/` so they are not invocable as slash commands). Modifier flags and ops stay inline in the dispatcher SKILL.
