### Phase 3: Review (parallel + triage + user test)

> **TLDR**  -  Two review stages over evidence Phase 2 already produced: the deterministic gates (build + lint + test + secret scan) are the Phase 2 exit gate now, and this phase inherits `.build.log` and `.test.log` rather than building again. Stage 1: 3 reviewers in parallel per host, second slot **CLI-aware** (Claude Code Fable + Opus + Sonnet; Copilot CLI GPT-5.4 + Opus + Sonnet). Stage 2: Fable triage (Opus on Copilot CLI) filters raw findings for false positives and out-of-scope items. The user test step closes the phase. Only triage-accepted blocking items loop back to Phase 2.

<!-- progress-contract: applied -->
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md`  -  lines for each gate, each reviewer dispatch + finish, triage start, triage verdict, fix dispatch.

#### Step 0  -  Analysis mode branch

`state.mode === "analysis"` has no diff: the artefact is the document. Replace Steps 1-3 with (1) `validate-analysis-doc.mjs <file> --strict` per platform, (2) the same CLI-aware reviewer set reading the document against one question - **could an implementer build the right thing from this alone?** a finding is anything that would force a guess - and (3) `$HOME/.claude/multi-agent-refs/analysis/resolve.md` over the Section 20 rows still `Acik / Open`; deferred rows report `review_blocking`. Then go to Phase 4.

Log: `Phase 3: Document review  -  validator:{pass/fail} | {N} findings | {M} resolved, {K} deferred`
#### Step 1.45  -  Test plan coverage cross-check (Phase 1 modes with a test plan)

Every `state.dev.testPlan[]` row must map to a real test in the diff, matched on the planned name. Missing -> `important` ("planned test <name> (BR-<id>) has no implementation"); present but asserting something else -> `blocking`, plan and code disagree about the rule. This is what makes "analysis quality is output quality" measurable. Skipped when `docStatus === "not-applicable"` or the mode has no Phase 1.

Log: `→ test plan coverage: <N>/<M> planned rows implemented`

#### Step 1.5  -  Platform Compliance Review (skill-based, no device needed)

If changes include UI files (iOS: `*View.swift`, `*Screen.swift`, `*Cell.swift`; Android: `*Screen.kt`, `*Content.kt`, `*Composable.kt`), run platform-specific checks:

**Accessibility** (both platforms):
- Missing `.accessibilityLabel` / `contentDescription` → **blocking**
- Tap target < 44×44pt (iOS) / 48×48dp (Android) → **blocking**
- Missing `.accessibilityIdentifier` / `testTag` → **important**
- Dynamic Type / font scaling not supported → **important**

**iOS  -  Apple HIG compliance** (skills: `ai-ios-toolkit:hig-patterns`, `ai-ios-toolkit:hig-components-layout`, `ai-ios-toolkit:hig-foundations`):
- Navigation pattern mismatch (e.g. custom back button instead of system) → **important**
- Non-standard gesture without discoverability hint → **suggestion**
- Missing safe area / keyboard avoidance → **important**
- Hardcoded colors instead of system/semantic colors → **suggestion**

**iOS  -  SwiftUI interaction & accessibility conventions.** Gated to changed SwiftUI files. These are rule-registry territory as of v14.0.0, not a list transcribed here: Step 1.78 resolves them from whichever registry declares SwiftUI scope, so the criteria and their severities live in one place instead of drifting between this doc and the skill. Reviewers receive the resolved rule IDs. Native-SwiftUI-first unless the project's `figma-config` `ui.*` declares a custom system, in which case check against that system. Reference skills, when no registry covers the change: the enabled stack plugin's `navigation`, `overlays` and `bottom-sheets` skills (`ai-ios-toolkit:*` / `ai-android-toolkit:*`), plus `figma-to-swiftui`.

**Android  -  Material Design compliance** (skills: `ai-android-toolkit:compose-components`, `ai-android-toolkit:android-architecture`):
- Non-Material3 component when M3 equivalent exists → **suggestion**
- Missing `contentDescription` on icons/images → **blocking**
- Hardcoded dp values instead of Material spacing tokens → **suggestion**

**App Store / Play Store readiness** (skills: `ai-ios-toolkit:app-store-review`, `ai-android-toolkit:play-store-review`):
- Privacy: API usage without purpose string / permission rationale → **blocking**
- Deprecated API usage flagged by latest SDK → **important**

Device-level audit runs in Phase 3 (`audit-guide.md`).

#### Step 1.6  -  Repo Map Injection (advisory, opt-in)

**Gated by `prefs.global.repoMap.enabled`** (default: `false`). Same pattern as Phase 1 Step 2.5  -  runs `$HOME/.claude/scripts/repo-map.mjs` and injects the result as `${REPO_MAP}` into each reviewer's prompt context. Reviewers treat it the same way the priority files block is treated: advisory hint, not gospel.

Phase 3 reuses the cached map from Phase 1 when both phases run in the same task (avoid recomputing the same scan twice). The orchestrator caches under `state.repoMap.<sha>` keyed by `git rev-parse HEAD`; cache miss → regenerate. When `enabled=false`, both phases skip the script entirely and no `${REPO_MAP}` placeholder appears in prompts (no empty-string artefact).

Cost ledger: `phase-4.repo_map_emitted bytes=N budget=B cache_hit=true|false`  -  Phase 5 cost rollup distinguishes a cache hit (free) from a regeneration (~150-300ms wall, 0 LLM cost either way).

#### Step 1.8  -  Platform parity cross-check (advisory, read-only)

Runs when a counterpart app repo resolves AND the diff touches a screen,
service, request model or localization file. Resolution is automatic and
mirrored (ios ↔ android); the order and the learned pref live in the contract.
No counterpart, or an unattended run facing several, skips silently - no
section, no placeholder, run unaffected.

Load `$HOME/.claude/multi-agent-refs/platform-parity.md` and follow it. The
counterpart repo is **read only**, and parity findings are **never blocking**.

#### Step 1.75  -  Diff Risk Scoring (advisory)

Before dispatching reviewers, run the deterministic diff risk scorer and inject the top-N files into each reviewer's prompt as a priority hint. **Advisory only  -  never gates the pipeline.** Heuristic, zero LLM, runs in well under a second.

Score the diff **once, in full** (no `--top`): Steps 1.76/1.77 need every scored file (a shrinking test file ranked 20th; whether ANY file is high-stakes). The top-N hint is derived from that report, not a second git walk.

```bash
RISK_FULL=$(node $HOME/.claude/scripts/diff-risk-score.mjs \
  --base "$BASE_BRANCH" --head HEAD \
  --task-id "$TASK_ID" 2>/dev/null)
echo "$RISK_FULL" | node $HOME/.claude/scripts/validate-diff-risk.mjs - >/dev/null 2>&1 || RISK_FULL=""

# Priority hint for the reviewer prompts: top 5 of the full report.
RISK_JSON=$([ -n "$RISK_FULL" ] && jq -c '.files |= (sort_by(-.score) | .[:5])' <<< "$RISK_FULL" || echo "")
```

Persist the totals as `state.diffRisk` (Phase 4 `risk` section, Phase 5, `run-metrics.mjs` read them):

```bash
[ -n "$RISK_FULL" ] && jq -c '{diffRisk: (.totals + {signals: ([.files[].signals[]?.name] | unique)})}' <<< "$RISK_FULL" | node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
```

**Signals & weights** (see `$HOME/.claude/schemas/diff-risk.schema.json`):

| Signal | Weight | Triggers when |
|---|---|---|
| `security_path` | 3.0 | path matches Auth/Keychain/Credential/Token/Crypto/Networking glob |
| `migration` | 4.0 | path matches DB schema / migration glob (.sql, /Migrations/, alembic/, prisma/migrations) |
| `public_api` | 2.0 | added line declares `public func/class/struct/enum`, `@objc`, `open fun`, `@Composable`, or `export function/class/...` |
| `no_test_change` | 2.5 | source file changed, no paired test file (`{Base}Tests.{ext}`, `{Base}.test.{ext}`, etc.) appears in the diff |
| `test_lines_removed` | 3.0 | test-classified file whose diff removes more lines than it adds (immutable-test backstop, see `$HOME/.claude/multi-agent-refs/rules.md`) |
| `complexity_delta` | 1.5 | added control-flow tokens (`if`/`guard`/`switch`/`while`/`for`/ternary/`&&`/`\|\|`) |
| `ui_critical` | 1.5 | path matches `*View.swift`, `*Screen.kt`, `*Configuration.swift`, etc. |
| `loc_changed` | 1.0 | base sensitivity to total `+/-` lines |

**Reviewer prompt injection**: when `$RISK_JSON` is non-empty, the orchestrator builds a `${PRIORITY_FILES}` block (numbered list of top-N files with their score + signals) and injects it once per reviewer. Reviewer prompt template (`code-reviewer.md`) treats it as advisory and does not echo it back. Triage does not see the priority list  -  its job is to filter the merged findings, not the diff.

**Gate behavior**: **never blocking**. If risk scoring fails (git, parse, validator), continue with no priority hint: reviewers get the full diff in default order. Failures are logged:

```bash
[ -z "$RISK_JSON" ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.diff_risk_skipped reason=$REASON
```

On success, emit a single summary metric:

```bash
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.diff_risk \
  top_files=$(jq '.files | length' <<< "$RISK_JSON") \
  max_score=$(jq '.totals.max_score' <<< "$RISK_JSON") \
  loc_added=$(jq '.totals.loc_added' <<< "$RISK_JSON") \
  loc_removed=$(jq '.totals.loc_removed' <<< "$RISK_JSON") \
  files=$(jq '.totals.files' <<< "$RISK_JSON")
```

**Opt-out**: `prefs.global.diffRiskAdvisory = false` skips this step (no script, no priority block). Default `true`: the cost is bounded and the signal-to-noise is measured against the golden-task fixtures.

#### Step 1.76  -  Test-integrity gate (produces BLOCKING findings)

Step 1.75's `test_lines_removed` hint is too weak for what it detects: a suite made green by deleting tests. This makes it blocking findings triage must adjudicate. Pure function of the full report. `--out` always writes the file, the empty result included.

```bash
TI_FILE="$WORKTREE/.pipeline/test-integrity.json"
printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/test-integrity-gate.mjs --out "$TI_FILE" 2>/dev/null
TI_COUNT=$(jq -r '.count // 0' "$TI_FILE" 2>/dev/null || echo 0)
[ "$TI_COUNT" -gt 0 ] && $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.test_integrity findings="$TI_COUNT"
```

`findings[]` are reviewer-shaped (`test_integrity`, `blocking`), so they merge into the reviewer findings at Step 3.0 and need no triage-prompt or `validate-triage.mjs` change. Triage keeps each blocking unless the removal is justified per the immutable-test rule (spec changed AND commit body names the test) → `deferred[]`.

The gate never blocks the phase; it *emits* blocking findings (unreadable input: zero). Feed it the FULL report: `--top` hides a shrinking test file below the cut. **No opt-out**: a run that can switch off its own anti-reward-hacking control cannot be trusted to report a pass.

#### Step 1.761  -  Owned-path gate (BLOCKING findings)

Phase 2 Gate 0 again, same base:

```bash
OP_FILE="$WORKTREE/.pipeline/owned-path.json"
node $HOME/.claude/scripts/owned-path-gate.mjs --repo "$WORKTREE" --base "origin/$BASE_BRANCH" \
  --state "$STATE_FILE" --out "$OP_FILE"
```

`--out` drops the old file; exit 2 writes none. A missing `$OP_FILE` is a gate error, never "no findings" (`review-decision-gate.mjs` exit 3 under the gates). Findings merge at Step 3.0. Contract: `features/repo-profile.md`. Step 1.762: `features/valven.md`.

#### Step 1.77  -  Reviewer scope (cost gate)

On a trivial diff every reviewer agrees and the extra models plus triage are paid for nothing. Reviewer count comes from the same report, no LLM.

```bash
SCOPE_JSON=$(printf '%s' "$RISK_FULL" | node $HOME/.claude/scripts/review-scope.mjs 2>/dev/null \
  || echo '{"scope":"full","reason":"no risk report - failing safe"}')
REVIEW_SCOPE=$(jq -r '.scope // "full"' <<< "$SCOPE_JSON")
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.scope scope="$REVIEW_SCOPE"
```

`single` (Reviewer 1 only) requires **all** of: churn <= 20 lines, `totals.max_score` < 3.0, and no `security_path` / `migration` / `public_api` / `no_test_change` / `test_lines_removed` on any file. Anything else → `full`.

Fails safe in one direction only: empty report, parse failure or validator rejection all resolve to `full`, because skipping a reviewer trades coverage for cost. `consensus.reviewerCount` records what actually ran, so a single-reviewer run never reads as cross-model agreement. Opt-out: `prefs.global.reviewScopeGate = false` forces `full`.

#### Step 1.78  -  Criteria resolution (skill conformance, required)

Resolves WHAT the changed code was supposed to honour, before any reviewer sees it. Zero LLM. Full contract: [`$HOME/.claude/multi-agent-refs/features/skill-conformance.md`]($HOME/.claude/multi-agent-refs/features/skill-conformance.md).

```bash
node $HOME/.claude/scripts/skill-conformance.mjs \
  --diff "$WORKTREE/.review-diff.txt" --state "$WORKTREE/agent-state.json" \
  --repo "$WORKTREE" --out "$WORKTREE/.pipeline/criteria-manifest.json"
CRIT_RC=$?
CRITERIA=$(cat "$WORKTREE/.pipeline/criteria-manifest.json" 2>/dev/null || echo '{}')
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 review.criteria \
  rules="$(jq -r '.selectedRuleCount // 0' <<< "$CRITERIA")" \
  ledger="$(jq -r '.ledger.source // "derived"' <<< "$CRITERIA")"
[ "$CRIT_RC" != "0" ] && HALT "criteria could not be resolved (rc=$CRIT_RC)  -  see resolutionFailure / unparseableRegistries"
```

Output conforms to `$HOME/.claude/schemas/criteria-manifest.schema.json`. The four parts that matter downstream:

- **`selectedRules[]` is the denominator.** Rule IDs from every registry whose declared `scope` matches the diff, persisted BEFORE the reviewers run so the set cannot be renegotiated after one has seen the diff. That is what makes "applied completely" answerable rather than "looks fine": a reviewer finding nothing must still return a verdict per ID. Discovery is declared via `standards-registry:` frontmatter, so no stack-specific skill is named here.
- **`coverage.declaredGaps[]`, `droppedReasons`, `ledger`.** Every gap, dropped rule and unbindable declared skill is reported with a reason; coverage is never computed from self-report. Rationale: the contract above.
- **`findings[]`.** Reviewer-shaped, merging at Step 3.0 alongside test-integrity: expired / unexplained / unknown-ID exception markers, plus any registry whose delegated linter is not wired here (those rules are unverified, so reporting no violations reports that nothing was measured).

**Any non-zero exit halts** (1 = setup error incl. a bad `--skills-root`; 2 = no root resolved, unparseable registry, or a declared path escaping its skill dir): continuing would drop a rule set from the denominator, and since the reviewer validator skips the checklist when zero rules were selected, the run would report clean over criteria never loaded. A coverage gap halts only under `prefs.global.skillConformance.blockOnCoverageGap`. No opt-out for the stage or the exception-expiry check, on the same grounds as Step 1.76.

#### Step 1.8  -  Figma visual-fidelity context (when task carries a Figma reference)

**Missing inputs.** A section the evidence did not support is absent from the analysis document, so three inputs this phase was written around can be missing. Substitute them and RECORD the substitution  -  a step that could not run and one that passed must not read the same, or the completeness claim cannot be checked:

| Absent input | Substitute |
|---|---|
| `detectedStack` (Phase 1 Step 2) | the language census in `criteria-manifest.json` → `languages`, from the diff's file extensions |
| Phase 1 analysis summary + Phase 1 plan (triage scope, cache prefix) | the task description plus the inline task list Phase 2 generated for itself |
| `state.evidence.figma[]` (this step), Step 2.8 visual conformance | nothing. Record each as `not-applicable (no Phase 1 evidence)` in `consensus.visualConformance`; never silently omit |

When `state.evidence.figma[]` is non-empty, the reviewer subagents MUST receive the captured screenshot URLs / paths and the canonical-component name (from `state.evidence.figma[i].screenshotUrl` and `state.evidence.figma[i].codeConnectSnippets[0].componentName` when present) so they can compare visual fidelity. Pass them inline in each reviewer prompt under a `## Figma evidence` block, one row per frame.

Visual-fidelity mismatches against the captured screenshot are BLOCKING findings, not nits:

- Canonical component usage: the Code Connect-mapped component is used verbatim  -  a sound-alike substitute, a forked copy, or ad-hoc inline UI where a mapping exists is blocking (Phase 2 "Design fidelity contract")
- Inter-component spacing: gaps, paddings, and alignment BETWEEN components match the design's measured values mapped to spacing tokens  -  invented numeric values are blocking
- Everything else: the 30-group catalog in `$HOME/.claude/multi-agent-refs/features/design-conformance.md`. Two of its rules carry into review: a height or inset is **measured**, never read from the token under test; an unrun check is a fail-to-verify, not a pass.

When `state.figmaAccess.tier === 3` (user-attached screenshot, no Code Connect snippet), the reviewer additionally sets `findings[i].severity = "blocking"` and `findings[i].tag = "review_blocking_tier3"` on every UI atom that lacks a confirmed canonical-component mapping. The triage step preserves these findings unless the user has explicitly cleared the open question.

#### Step 1.85  -  Review file set (the denominator, required)

Decides what the reviewers read before the cap decides what fits: the cap truncates the largest files first, so a lockfile survives while real code is cut. Zero LLM. Full contract: `$HOME/.claude/multi-agent-refs/features/review-file-set.md`.

```bash
git -C "$WORKTREE" diff --name-only "$BASE_BRANCH"...HEAD \
  | node $HOME/.claude/scripts/review-file-filter.mjs \
  > "{worktreePath}/.pipeline/review-files.json"
```

`reviewed[]` is the denominator, fixed here for the same reason `selectedRules[]` is. `excluded[]` reaches the run report with its reason and the glob that matched. Exit 2 means the pattern list is unreadable and everything is reviewed: continue, logging `review.file_filter_failed`.

#### Step 1.9  -  Context economy (cache prefix + diff cap)

Phase 3 sends the same diff to every reviewer and then to triage, so the diff is the dominant token cost. Two measures keep it bounded:

**Shared cache prefix.** Build the reviewer and triage prompts so the large invariant context  -  the full diff, the `${CRITERIA}` block from Step 1.78, the `${REVIEW_FILES}` list from Step 1.85, the Phase 1 analysis summary, the Phase 1 plan  -  is a byte-identical leading block across all dispatches in this iteration. Only the per-reviewer focus + skill line varies, and it goes AFTER the shared block. `${CRITERIA}` goes in the prefix, identical for every reviewer: subsetting it per reviewer would invalidate the prefix for the whole panel and re-bill the largest block in the phase. Per-reviewer emphasis stays a one-line pointer in the suffix. When the host supports prompt caching, the 2nd/3rd reviewer and the triage call then read that prefix at the discounted cache-read rate instead of re-billing it as fresh input. Forward the host-reported cache-read count as `tokens_cached` per the Token telemetry contract so the saving lands in the cost ledger. The `<scope-self-check>` block and, from iteration 2, the `<previous-round-findings>` block (Step 2.1) close the shared block, after the plan and before the per-reviewer suffix.

**Single-repo diff cap.** Applied to the Step 1.85 `reviewed[]` set only. If it exceeds the Phase 3 token allowance (`token-budget.json`), truncate the largest files and append a footer `[truncated  -  full diff in file://$WORKTREE/.review-diff.txt]`, writing the full diff to that path. Reviewers and triage receive the same capped view + the marker so they can flag "review the full diff manually." Log `review.diff_truncated bytes_dropped=<N>`. (Multi-repo already caps the combined diff at 80% of budget; this is the single-repo equivalent.)

#### Step 2  -  Parallel AI Review (CLI-aware reviewer set)

Launch Agent instances **in parallel** using the shared `code-reviewer` subagent definition (`~/.claude/agents/code-reviewer.md`). Every host runs three reviewers; only the second slot differs, because GPT-5.4 exists on Copilot and Codex but not on Claude Code, where Opus fills it. Three reviewers on Claude Code cost more than two, and the cost buys the thing a second opinion cannot: a finding two independent readers both miss is what triage has no chance to catch.

**Scope from Step 1.77.** `$REVIEW_SCOPE == "single"` → dispatch **Reviewer 1 only**, and skip Step 2.5 + 3.6 (both no-ops with one reviewer). `"full"` (default + fail-safe) → the whole set below. Either way record the count in `consensus.reviewerCount`.

| Reviewer   | subagent_type   | Claude Code | Copilot CLI | Codex CLI | Focus | Skills Referenced |
| ---------- | --------------- | --- | --- | --- | --- | --- |
| Reviewer 1 | `code-reviewer` | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `xhigh` | Deep security + architecture | `api-security-best-practices`, `architecture` |
| Reviewer 2 | `code-reviewer` | `claude-opus-5` | `gpt-5.4` | `gpt-5.4` @ `high` | Edge cases, different perspective | cross-model diversity |
| Reviewer 3 | `code-reviewer` | `claude-sonnet-5` | `claude-sonnet-5` | `gpt-5.6` @ `medium` | Quality + correctness + naming | `ai-backend-toolkit:clean-code`, stack-specific skill |
| Triage     | triage persona  | `claude-fable-5` | `claude-opus-5` | `gpt-5.6` @ `max` | Filter false positives + out-of-scope | - |

Reviewer count per host: **Claude Code 3, Copilot CLI 3, Codex CLI 3**  -  **2** on Claude Code with the fable rung off, per `$HOME/.claude/multi-agent-refs/features/model-fallback.md`.

#### Codex CLI  -  two constraints that fail silently

Both measured against Codex 0.145, both produce a review that looks like it ran. Full
statement: the managed block in `~/.codex/AGENTS.md`, always loaded on that host.

1. **`fork_turns: "none"` on every `spawn_agent` that sets `model` or
   `reasoning_effort`** - a full-history fork discards the override and collapses the
   panel onto one model, silently.
2. **Three concurrent children is the ceiling** - 4 slots including the orchestrator,
   so the reviewer count on Codex is capability-derived, not a preference.

Sub-agent delegation itself is authorized by that same managed block; without it Phase 3
degrades to a single in-thread review.

**Single-vendor caveat.** Claude Code and Codex both run a one-vendor panel, so their
consensus is weaker evidence than Copilot CLI's; say so in the triage note on a
borderline finding. Where the diversity budget goes instead:
`cross-cli-contract.md`, "Panel diversity per host".

Each reviewer inherits the `code-reviewer` agent's focus areas (Security, Architecture, Quality, Performance) and output contract. The orchestrator overrides only the model and the stack-specific skill per-reviewer  -  no prompt duplication.

**Model override wiring:** `code-reviewer.md` declares `preferredModel: fable`, so Reviewer 1 uses the persona default (Fable 5). Reviewer 2 (`claude-opus-5` on Claude Code, `gpt-5.4` elsewhere) and Reviewer 3 (`claude-sonnet-5`) set `PHASE_MODEL_OVERRIDE=<model>` before dispatch  -  the orchestrator exports `CLAUDE_CODE_SUBAGENT_MODEL` on Claude Code, or passes `--model` on Copilot CLI. Full precedence rule: `skills/shared/core/multi-agent/SKILL.md#agent-dispatch--per-persona-model-routing`. Fable dispatches are subject to the fallback contract (`$HOME/.claude/multi-agent-refs/features/model-fallback.md`): dispatch-error retry walks `fable -> opus -> sonnet` and budget-ceiling downgrade.

**Stack-specific skills loaded per reviewer** (from Phase 1 `detectedStack`). All three columns are used on every host; Reviewer 2 reads them as Opus on Claude Code and as GPT-5.4 elsewhere.

| Stack | Reviewer 1 (Fable / Opus on Copilot) | Reviewer 2 (Opus on Claude Code, GPT-5.4 elsewhere) | Reviewer 3 (Sonnet) |
|-------|-------------------|-----------------------------------------|---------------------|
| iOS/Swift | `ai-ios-toolkit:ios-security`, `ai-ios-toolkit:swiftui-performance`, `ai-ios-toolkit:hig-patterns` | `ai-ios-toolkit:swift-concurrency`, `ai-ios-toolkit:ios-accessibility` | `ai-ios-toolkit:swiftui-pro`, `ai-ios-toolkit:swift-testing` |
| Android/Kotlin | `ai-android-toolkit:android-security`, `ai-android-toolkit:android-performance` | `ai-android-toolkit:compose-testing`, `ai-android-toolkit:android-architecture` | `ai-android-toolkit:compose-components`, `ai-android-toolkit:kotlin-coroutines-expert` |
| Python | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:fastapi-pro` | `ai-backend-toolkit:python-patterns` |
| Node.js | `ai-backend-toolkit:api-security-best-practices` | `ai-backend-toolkit:nodejs-backend-patterns` | `ai-frontend-toolkit:typescript-patterns` |
| Docker | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:docker-expert` | `ai-backend-toolkit:ci-cd-pipelines` |
| Generic | `ai-common-toolkit:security-review` | `ai-backend-toolkit:clean-code` | `ai-backend-toolkit:clean-code` |

##### 2.1 Previous-round findings (iteration >= 2) and 2.2 scope self-check (every iteration)

From iteration 2, render the previous round's accepted blocking/important findings (`.pipeline/triage-round-$((ITERATION-1)).json`, max 40) into a `<previous-round-findings>` block at the end of the shared prefix. Every iteration also renders `.pipeline/scope-check.json` (Phase 2 Step 3.7) plus `scope-check-gate.mjs --advisory` output as `<scope-self-check>`: file reasons, unjustified files, and `notDone[]` (never re-raised as findings); a missing record logs `review.scope_check=missing`. Block text and recipes: `$HOME/.claude/multi-agent-refs/features/review-delta.md`.

#### Step 2.8  -  Visual conformance gate (component / screen work only)

Runs when `state.taskType == "component"` **or** the diff touches SwiftUI UI files
AND the task carried a Figma reference. Two checks, in order:

1. **`ai-ios-toolkit:figma-review`** over the implemented frames  -  the
   plugin's own component review, including the 14-item checklist that covers design
   tokens, accessibility identifiers, previews and Code Connect.
2. **`/multi-agent:design-check`** for pixel + spacing + typography + colour
   conformance against the Figma variants, with its coverage gate: a variant that is
   neither audited nor skipped-with-a-reason fails the run.

**Code Connect must be published, not merely written.** A `*.figma.swift` file on
disk with `Code Connect: Not published` in Figma means the binding does not exist for
anyone but the author. Assert the publish step ran; an unpublished binding is a
blocking finding.

Why this is a gate and not advice: `design-conformance.md`, "Why this runs as a gate".

Skip only when the diff has no UI change. Record the outcome in
`consensus.visualConformance` so Phase 5 reports whether it ran.

**iOS/Swift  -  interaction & convention checks (conditional).** Step 1.78 resolves these. Where no registry covers the change, reviewers fall back to the analysis doc (Section 14 Code Connect mapping) and, when `ai-ios-toolkit` is enabled, that plugin's navigation / overlay / bottom-sheet + accessibility conventions.

**Module review guides (conditional, all stacks).** Step 1.78 resolves them into `criteria-manifest.json` → `moduleGuides`. Inject with the directive: read each guide, apply its rules to the changed files under its directory  -  a guide governs only its own subtree, and its violations are findings triaged like any other. Same contract as `/multi-agent:review` Step 2b.

**Dispatch timeout (required, mirrors triage 3.3).** Reviewers run in parallel and triage waits on all of them, so one stalled reviewer hangs the phase. Bound each reviewer dispatch by `REVIEWER_TIMEOUT_SECONDS` (default 180). If a reviewer has not returned by the budget: log `review.reviewer_timeout reviewer=<name>`, treat that reviewer as absent, and proceed to triage with the reviewers that did return. The merged-findings count and `consensus.reviewerCount` reflect only the reviewers that returned. If **zero** reviewers return, retry Reviewer 1 once; on a second total failure HALT with `ERR: no reviewer returned within ${REVIEWER_TIMEOUT_SECONDS}s; resume with /multi-agent:resume #N.`. The Step 2.5 rebuttal round uses the same per-dispatch timeout. Never block indefinitely on a slow or dead reviewer dispatch.

#### Output contract  -  reviewer step

Step 2 produces N reviewer-output objects (one per dispatched reviewer), each conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`. They are persisted to `state.reviewIterations[<iteration>].reviewers[]` and consumed by Step 3 (Fable triage)  -  never by Phase 4 directly. The triage step (below) is the producer of the only review artifact Phase 4 reads, conforming to `$HOME/.claude/schemas/triage-output.schema.json`.

#### Step 2.7  -  Security audit (conditional, produces reviewer-shaped findings)

Runs when Step 1.75 scored `security_path`, on a release branch, or when `/multi-agent:security-review` dispatches here. The `security-auditor` returns reviewer `findings[]`, each with a `security` envelope (`security-finding.schema.json`) and severity from the CVSS band, held in `$SECURITY_AUDIT_JSON` for the merge so a `blocking` one reaches triage and blocks Phase 4. Mechanics: `~/.claude/multi-agent-refs/features/security-audit.md`.

**Subagent return format**  -  each reviewer returns JSON conforming to `$HOME/.claude/schemas/reviewer-output.schema.json`:

```json
{"findings":[{"severity":"blocking|important|suggestion","file":"...","line":N,"issue":"...","fix":"...","ruleId":"SEC-01","criteriaSource":"ios-coding-standard"}],
 "conformance":[{"ruleId":"SEC-01","verdict":"conformant|violated|not-applicable","file":"...","line":N,"reason":"..."}],
 "fileCoverage":[{"path":"src/App.swift","verdict":"reviewed|skipped","reason":"..."}],
 "approved":true|false}
```

`ruleId` + `criteriaSource` appear on a finding that cites a rule from `${CRITERIA}`. `conformance` is required whenever Step 1.78 selected at least one rule, and `fileCoverage` whenever Step 1.85 left at least one file in `reviewed[]`: one row per ID, one row per path, none outside either set.

**Required: validator gate (deterministic)  -  run immediately after each reviewer returns, before merging findings.** Persist each reviewer's output and validate the file  -  the validator's exit code decides, not the LLM turn:

Write it with the Write tool to `$REVIEWER_FILE` = `$WORKTREE/.pipeline/reviewer-$N.json`, then:

```bash
node $HOME/.claude/scripts/validate-reviewer.mjs "$REVIEWER_FILE" \
  --criteria "$WORKTREE/.pipeline/criteria-manifest.json" \
  --coverage "$WORKTREE/.pipeline/review-files.json" \
  && node $HOME/.claude/scripts/verify-citations.mjs "$REVIEWER_FILE" --repo "$WORKTREE" --worktree \
  && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$REVIEWER_FILE"
```

`verify-citations.mjs` resolves each finding's `file:line` against the checkout,
not a commit: a round's fix is uncommitted, and HEAD would call it invented.
Exit 1 takes the single rework below; exit 2 means not a repository.

Progress line: `    → checking validator validate-reviewer ({reviewer})`

`finding-fingerprint.mjs` stamps each finding with its cross-round id once the validator passes; an echoed one is kept, and anonymization leaves it intact.

Exit 0 = valid. Exit 2 = contradiction (approved=true with blocking findings)  -  flip `approved` to `false`, continue. With `--criteria`, exit 1 also covers the conformance checklist: a selected rule ID with no verdict, a verdict for an ID that was never selected, a `conformant` row with no file evidence, or a `violated` row with no matching finding. Those are the four ways a review can look complete without being complete, and the validator is what makes the checklist more than decoration  -  it is hand-written and does not apply `additionalProperties`, so an unchecked array would otherwise pass. Exit 1 = malformed; gate protocol (fails CLOSED, same handling as the evidence gate): emit the validator stderr + `errors[]` verbatim, attempt ONE self-correction rework (re-invoke that reviewer with the errors quoted, overwrite the file), re-run the validator. If it fails again -> HALT the phase (no merge, no triage). Recovery hint: `ERR: reviewer output failed validate-reviewer.mjs twice. Inspect $REVIEWER_FILE against $HOME/.claude/schemas/reviewer-output.schema.json, then resume with /multi-agent:resume #N.`

#### Step 2.5  -  Disagreement-round loop (opt-in)

**Gated by `prefs.global.reviewDisagreementRound`** (default `false`). **Never in autopilot:** `node $HOME/.claude/scripts/review-decision-gate.mjs --rebuttal-allowed --state "$STATE_FILE"` exits 1 with the quality gates active; go to Step 3, reviewers stay blind (`$HOME/.claude/multi-agent-refs/features/review-decision.md`). Otherwise:

1. Reviewers agree iff all return `approved=true` with no `blocking` findings, OR all return `approved=false` with overlapping `blocking` findings. Agreement → Step 3.
2. Disagreement → one rebuttal round: re-prompt each reviewer with (a) its original output, (b) the OTHER reviewers' blocker findings **anonymized** through `node $HOME/.claude/scripts/anonymize-findings.mjs` (labels `Source A/B/C`, no model name, order deterministic per `taskId:iteration`), (c) *"Given the opposing arguments, keep / withdraw / modify each of your findings. You may also newly agree with a finding you previously missed. Return the SAME JSON schema  -  this is a revision, not a new review."* Launch all in parallel (the Step 2 set), max one round; results replace the originals with `roundCount: 2`.

**Parity:** every host runs it identically; telemetry `review_round_count={1|2}` per reviewer. **Cost:** one more Step 2 on a disagreeing run (~8%).

#### Step 3  -  Fable Triage (filter before acting)

**CRITICAL**: Reviewer findings are **raw signals**, not commands. Never auto-loop on every "blocking" tag  -  reviewers hallucinate, misread scope, or repeat each other. Run Fable triage (Opus on Copilot CLI) to evaluate merged findings against task scope.

Optional: when `ai-analyst-toolkit` is enabled and a finding blames a third-party library rather than this diff, ask `evidence-github` whether it is already open upstream. A confirmed one is `deferred` with its `GitHub:<owner>/<repo>#<n>` citation, not `accepted` and handed to Phase 2 to fix code that is not ours.

Opt-in empirical layer: when `prefs.global.verifyByTest.enabled` is `true`, accepted blocking findings additionally go through Step 3.7 (verify-by-test), which tries to reproduce each one with a minimal failing test before the Phase 2 rework loop fires. Full wiring: `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`.

##### 3.0 Anonymize the reviewer findings, then merge the deterministic ones

**Anonymize first (required).** On both CLIs the triage model is also a reviewer (Fable on Claude Code, Opus on Copilot), and a judge that can see which findings are its own is marking its own homework:

```bash
ANON=$(jq -n --argjson r "$REVIEWERS_JSON" --arg t "$TASK_ID" --argjson i "$ITERATION" \
        '{taskId: $t, iteration: $i, reviewers: $r}' \
      | node "$HOME/.claude/scripts/anonymize-findings.mjs" --map "/tmp/review-$TASK_ID-$ITERATION-map.json")
```

`$REVIEWERS_JSON` is `state.reviewIterations[i].reviewers`. Findings come back with `foundBy: "Source A|B|C"` and every identity key removed. Persist the map to `state.reviewIterations[i].anonymizationMap` for Phase 5 per-reviewer telemetry, and **never put the map in a prompt**.

Then append the deterministic findings (Steps 1.76, 1.761, 1.762, 2.7):

```bash
SA_FILE="$WORKTREE/.pipeline/security-audit-$ITERATION.json"
MERGED=$(printf '%s' "$ANON" | cat - "$TI_FILE" "$OP_FILE" "$VG_FILE" "$SA_FILE" 2>/dev/null \
  | jq -s '.[0] + ([.[1:][] | .findings // []] | add // [])')
```

Deterministic findings keep their `tag` (`test_integrity`, `owned_path`, `valven`) and no `foundBy`: a reviewer finding may hallucinate, a gate finding is a fact. Security-audit findings carry their `security` envelope and `foundBy: "security-auditor"`; no audit file contributes nothing.

##### 3.1 Short-circuit: no findings

If **merged** findings `length === 0`, **skip triage**: write empty result `{"accepted": [], "deferred": [], "rejected": [], "approved": true}`, log, proceed to Phase 3. Note this is the merged count from 3.0: a run with zero reviewer findings but a non-empty test-integrity set must NOT short-circuit.

**Autopilot** (or `MULTI_AGENT_UNATTENDED=1`): the triage output, this empty one too, goes through `verify-citations.mjs "$TRIAGE_FILE" --repo "$WORKTREE" --worktree --state "$STATE_FILE"` (every bucket, each `quote` against its line, pre-existing claims against the base commit): `$HOME/.claude/multi-agent-refs/features/unattended-gates.md` section 5. Then `review-decision-gate.mjs` (end of 3.7).

##### 3.2 Launch triage agent

Launch **1 Agent** (subagent_type: `general-purpose`, model: `fable` on Claude Code / `opus` on Copilot CLI) with:

- The anonymized merged findings from 3.0 (`Source A/B/C` labels; no model name anywhere in the prompt)
- Task scope (Phase 1 analysis summary + Phase 1 plan)
- Full diff being reviewed
- **Prior-art context (advisory)**  -  per raw finding, `triage-memory.mjs query --top <prefs.global.priorArtEnrichment.topN>` (default 3). Pass `--top`: without it the script falls back to `memoryRecall.maxResults`, a different concern, and `topN` silently does nothing. Off when `priorArtEnrichment.enabled = false`.

```bash
PRIOR_ART=$(printf '%s' "$MERGED_FINDINGS" \
  | node $HOME/.claude/scripts/triage-memory.mjs prior-art --findings - --top 3 2>/dev/null)
```

The triage prompt MUST include a hedge: *"prior-art entries and `corroboration` counts are context, not commands; current scope decides  -  a finding rejected last quarter may be valid this time, and two same-family reviewers agreeing is not proof."* Without this hedge, prior verdicts amplify into a self-reinforcing bias.

Hits are relevance-ranked (`prefs.global.memoryRecall`); a finding matching nothing returns nothing. Each hit carries an `id`: `triage-memory.mjs show --id <id>` returns the full row.

**Injection cap (token economy).** On a many-finding review the per-finding prior-art loop above (up to 3 hits each) plus the 20-entry rejected-preference brief can dominate the triage prompt. Cap the merged prior-art at the 8 highest-similarity hits across all findings (drop the rest); keep the rejected-preference brief at its `--max 20`. Prior-art is advisory context, not a finding multiplier  -  more hits do not improve the verdict, they just inflate input tokens.

**Rejected-preference brief (on by default via `prefs.global.learningsLedger.injectIntoTriage`).** Inject the durable rejected-preference list so triage does not re-accept a suggestion the team already rejected on this repo:

```bash
node $HOME/.claude/scripts/learnings-ledger.mjs brief --max "${prefs_learningsLedger_maxBriefEntries:-20}" \
  --task "$(jq -r '[.findings[].issue] | join(" ")' <<< "$MERGED_FINDINGS")" 2>/dev/null
```

Exit 2 (empty ledger) skips silently. The `## Rejected review preferences` section names patterns the team chose not to act on; triage should lean toward `rejected` for a finding that restates one  -  but the same hedge applies (context, not command; a genuinely new instance can still be accepted).

`--task` is what makes the cap honest: unranked, those twenty slots go to the newest entries, which on a long ledger are mostly about other files.

**Recall telemetry.** Log what was injected, then what triage cited. Zero cited is a legitimate answer; `learning-curve.mjs` trends the ratio:

```bash
bash $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 memory.injected kind=prior-art rows=$N
bash $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 memory.hit rows=$CITED_COUNT
```

**Bulky payloads (opt-in via `prefs.global.contextOffload.enabled`).** Test output and whole-file diffs go through the offload filter, which leaves a `[[ref:<node_id>]]` line plus the tail in context and the full text under `.multi-agent/refs/`. Read that file when the tail is not enough; with the pref off it is a pass-through.

```bash
<test-command> 2>&1 | bash $HOME/.claude/scripts/offload-ref.sh --phase 4 --label tests
```

**Triage prompt skeleton:**

```
You are the Review Triage agent. Three reviewers (fewer on a single-scope run) returned findings on this diff.
Your job: separate signal from noise. Do NOT add new findings. Do NOT re-review code.
Preserve `fingerprint` verbatim on every finding you carry through; it is the finding's identity across rounds.

For each finding, decide:
- ACCEPTED: real issue, in scope, must be fixed now
- DEFERRED: real issue but out of current task scope → log for later, do not block
- REJECTED: false positive, duplicate, style-only nitpick, or already correct

A finding carrying a `ruleId` cites a written rule from a registry the project
adopted, so it is not a matter of taste: it can be DEFERRED or REJECTED as out of
scope for THIS task, but "I would have written it differently" is not available as
a reason. Preserve `ruleId` and `criteriaSource` on every finding you keep.

#### Output contract  -  triage step

Step 3 produces a single triage-output object conforming to `$HOME/.claude/schemas/triage-output.schema.json` and persists it to `state.reviewIterations[<iteration>].triage`. This is the **only** Phase 3 artifact Phase 4 reads. Phase 4 commits MUST cite only `accepted` findings that were resolved; `deferred` items get linked in the PR description as follow-up work; `rejected` items never appear in any user-facing output.

Return ONLY valid JSON conforming to $HOME/.claude/schemas/triage-output.schema.json:
{
  "accepted":  [{ "severity": "blocking|important|suggestion", "file": "...", "line": N, "issue": "...", "fix": "...", "reviewer": "fable|opus|sonnet|gpt" }],
  "deferred":  [{ "finding": {...}, "reason": "..." }],
  "rejected":  [{ "finding": {...}, "reason": "..." }],
  "approved":  true|false,  // true if no accepted blocking items remain
  "consensus": { "reviewerCount": N, "verdict": "unanimous-pass|unanimous-block|split|unverified", "disagreements": [{ "file": "...", "line": N, "issue": "...", "note": "Fable blocking, Sonnet approved" }] }  // optional, see 3.6
}
```

##### 3.2.1 Required: validator gate (deterministic)

Run on the persisted file immediately after the triage agent returns, before acting on the verdict; the validator's exit code decides, not the LLM turn:

Write it with the Write tool to `$TRIAGE_FILE` = `$WORKTREE/.pipeline/triage-round-<N>.json` (N = `jq '.reviewIterations | length' "$STATE_FILE"`), then:

```bash
node $HOME/.claude/scripts/validate-triage.mjs "$TRIAGE_FILE" \
  && node $HOME/.claude/scripts/finding-fingerprint.mjs annotate --in-place "$TRIAGE_FILE" \
  && cp "$TRIAGE_FILE" "{worktreePath}/triage-output.json"
```

Progress line: `    → checking validator validate-triage`

One file per round; `triage-output.json` is the latest copy every downstream reader expects (Phase 5, finalize, work-summary, diff-explain). Step 3.7 rewrites `$TRIAGE_FILE`; repeat the `cp` after it.

| Exit  | Meaning                      | Action                                            |
| ----- | ---------------------------- | ------------------------------------------------- |
| **0** | Valid and clean              | Act on triage output as-is                        |
| **1** | Invalid structure            | Gate protocol (strict, fails CLOSED): (a) emit the validator's stderr and `errors[]` JSON into the log verbatim; (b) attempt ONE self-correction rework  -  re-prompt triage with the errors quoted, overwrite `$TRIAGE_FILE`; (c) re-run the validator. Second failure -> HALT the phase with the recovery hint: `ERR: triage output failed validate-triage.mjs twice. Inspect $TRIAGE_FILE against $HOME/.claude/schemas/triage-output.schema.json, fix or regenerate it, then resume with /multi-agent:resume #N.` |
| **2** | Over-rejection guard tripped | Pause for human confirm (autopilot: log + accept) |
| **3** | Contradiction auto-corrected | Proceed with `result.corrected`                   |

Capture stdout into `state.reviewIterations[-1].validatorResult` for Phase 5 audit.

##### 3.3 Edge case handling

Failure fallback (timeout >120s, or agent crash before any JSON is produced): retry triage ONCE → on second failure treat ALL raw findings as `accepted`, log cause. Structural validator failures (exit 1) do NOT take this fallback  -  they follow the section 3.2.1 gate protocol (one self-correction rework, then halt with the recovery hint).

| Scenario                                       | Action                                                                      |
| ---------------------------------------------- | --------------------------------------------------------------------------- |
| Over-rejection (>80% rejected, ≥5 findings)    | Pause for user; autopilot: log `triage=high-rejection-rate`, accept verdict |
| Contradiction (`approved: false`, no blockers) | Force `approved: true`, log `triage=contradiction-corrected`                |
| Hallucinated findings (not in raw input)       | Strip; log `triage=hallucinated-finding-stripped`                           |

##### 3.4 Telemetry emission (required)

Emit metrics per review pass for Phase 5 cost rollup:

One `review.reviewer_call` per dispatched reviewer, one `review.triage_call`, one `review.completed` to close the pass. Reviewer 1 is `fable` on Claude Code and `opus` on Copilot CLI; Reviewer 2 is `opus` on Claude Code and `gpt-5.4` elsewhere:

```bash
LOG_METRIC_FORWARD_TO_TRACKER=1 bash "$HOME/.claude/scripts/log-metric.sh" "$TASK_ID" 3 review.reviewer_call \
  model=<fable|opus> duration_ms="$R1_DURATION" tokens_in="$R1_IN" tokens_out="$R1_OUT"
LOG_METRIC_FORWARD_TO_TRACKER=1 bash "$HOME/.claude/scripts/log-metric.sh" "$TASK_ID" 3 review.reviewer_call \
  model=<opus|gpt-5.4> duration_ms="$R2_DURATION" tokens_in="$R2_IN" tokens_out="$R2_OUT"
LOG_METRIC_FORWARD_TO_TRACKER=1 bash "$HOME/.claude/scripts/log-metric.sh" "$TASK_ID" 3 review.reviewer_call \
  model=sonnet duration_ms="$SONNET_DURATION" tokens_in="$SONNET_IN" tokens_out="$SONNET_OUT"
LOG_METRIC_FORWARD_TO_TRACKER=1 bash "$HOME/.claude/scripts/log-metric.sh" "$TASK_ID" 3 review.triage_call \
  model=fable duration_ms="$TRIAGE_DURATION" tokens_in="$TRIAGE_IN" tokens_out="$TRIAGE_OUT"
bash "$HOME/.claude/scripts/log-metric.sh" "$TASK_ID" 3 review.completed raw_count=$RAW accepted=$ACC \
  deferred=$DEF rejected=$REJ approved=$APPROVED duration_ms=$DURATION
```

`LOG_METRIC_FORWARD_TO_TRACKER=1` mirrors `tokens_in`/`tokens_out`/`model` into `phase-tracker.sh` (see `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`). On non-zero validator exit (1/2/3), also emit `triage.edge_case` with cause. Omit `tokens_in`/`tokens_out` if unavailable. Best-effort  -  never fails the pipeline.

##### 3.5 Optional cross-check (single-point-of-failure mitigation)

Opt-in via `prefs.global.triageCrossCheck.enabled` (default `false`). Sampled runs dispatch a **Sonnet** triage agent as second opinion, validated via `validate-triage.mjs` (same fallback rules). Disagreements logged as `triage.cross_check_diff`; `blockOnDisagreement` pauses for user (autopilot: proceed with the Fable verdict). Doubles triage cost on sampled runs.

##### 3.6 Consensus surfacing (anti-correlation)

**Rationale:** Claude Code and Codex CLI run one-vendor panels, Copilot CLI's is two-thirds Anthropic, so unanimity on a *judgment call* is not independent confirmation: same-family models drift alike. Autopilot compensates with executable evidence (`features/review-decision.md`). Triage records a `consensus` block (schema v3.1.0) and surfaces disagreement and unverified agreement instead of burying it.

After the triage verdict is computed, populate `triage.consensus`:

1. `reviewerCount` = reviewers that actually dispatched, not the configured max: `3` normally, `2` with the fable rung off, lower if one timed out or was skipped.
2. Classify the iteration `verdict`:
   - `unanimous-block` -> all reviewers returned at least one overlapping `blocking` finding.
   - `split` -> reviewers disagreed on existence or severity of one or more findings (the Step 2.5 disagreement definition). List each split in `disagreements[]` with a `note` naming who held which position (e.g. "Fable blocking, Sonnet approved").
   - `unanimous-pass` -> all reviewers approved AND the diff is low-risk (no security/auth/concurrency surface per Phase 1 `touchedAreas`). Clear-cut; trust it.
   - `unverified` -> all reviewers approved BUT the diff touches a judgment-heavy surface (security, auth, concurrency, money, data migration). Agreement here may be correlated; do NOT treat it as a confirmed pass. Surface it.
3. `disagreements[]` is populated for `split` and is also used to carry `unverified` notes (e.g. "both approved a keychain change  -  agreement unverified, confirm manually").

**Surfacing (Step 4 + Phase 5):** When `verdict` is `split` or `unverified`, the disagreements are shown to the user at the Step 4 checkpoint (interactive modes) and always written to the Phase 5 report. Autopilot does not block on `unverified` (it logs `review.consensus=unverified` and proceeds), matching the maturity-check model  -  but the report records it so a human review can catch it. This is additive: omitting `consensus` is valid and means "not computed."

##### 3.7 Verify-by-test (opt-in, empirical validation of blocking findings)

A triage verdict is judgment; a failing repro test is proof. Runs only when `prefs.global.verifyByTest.enabled` is `true` AND `accepted` contains a `blocking` finding; otherwise skip silently. **Full contract (verdict table, cleanup invariant, prompts): `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`  -  read it before executing this step.**

Compressed flow: dispatch ONE verifier agent (model `verifyByTest.model`, default `sonnet`) for up to `maxFindings` (default 3) accepted blocking findings. Per finding it writes ONE minimal repro test and runs ONLY that test (Phase 2 single-test invocation, build lock, log tee'd to `$WORKTREE/.pipeline/verify-<i>.test.log`). Outcomes: test FAILS as predicted -> `confirmed`, finding stays blocking and the test is KEPT in `redTests[]` as the Phase 2 rework RED test; test PASSES on every one of `verifyByTest.repeatCount` runs (default 3) -> `not-reproduced` ONLY if `evidence-gate.mjs --claim test --status passed` exits 0 on each log; a run that disagrees with the others -> `inconclusive` with `flaky: passed k/N`, finding moves to `deferred[]`, test deleted; compile error / timeout / not unit-testable -> `inconclusive`, judgment verdict stands. Stamp findings with `verification` (schema v3.2.0), persist `state.reviewIterations[-1].verifyByTest = {attempted, confirmed, downgraded, inconclusive, redTests[]}`, recompute `approved`, re-run `validate-triage.mjs` under the 3.2.1 gate. Whole step bounded by `stepTimeoutSec` (default 600); on breach or crash remaining findings keep judgment verdicts  -  never blocks. Telemetry per 3.4: `review.verify_by_test attempted= confirmed= downgraded= inconclusive= duration_ms=`.

**Autopilot decision rule (every round, after 3.7):** run the call in `$HOME/.claude/multi-agent-refs/features/review-decision.md` (Wiring: one `--integrity` per gate file, `--source "$SA_FILE"`), merge its JSON into `state.reviewIterations[-1].reviewDecision`, repeat the `cp`. A blocker backed by neither two reviewers nor a failing test becomes important; on exit 1 or 3 run `gate-ledger.mjs park --outcome verification-failed --gate review-decision`.

##### 3.8 Cross-round delta + circuit-breaker trigger 2 (iteration >= 2)

**Full contract (state merge, telemetry line, picker wording): `$HOME/.claude/multi-agent-refs/features/review-delta.md`.**

```bash
TRIP=$(jq -r '.global.autopilotCircuitBreaker.identicalFindingCycles // 2' "$PREFS_FILE")
DELTA_JSON=$(node $HOME/.claude/scripts/review-delta.mjs --rounds-dir "$WORKTREE/.pipeline" --iteration "$ITERATION" --trip-cycles "$TRIP"); DELTA_RC=$?
```

Merge the JSON into `state.reviewIterations[-1].delta` via `write-state.mjs`; emit `review.delta iteration= new= still_present= resolved= plateau= tripped=`. Progress line: `    → comparing round {N} vs {N-1} (new={n} still={n} resolved={n})`.

| Exit | Action |
|---|---|
| **0** | Continue to Step 4 (also iteration 1 and a missing previous round). |
| **3** | Autopilot with `prefs.global.autopilotCircuitBreaker.enabled` (default true): write `state.circuitBreaker = {tripped: true, trigger: 2, detail, checkpoint: {phase: 4, step: "3.8", iteration: N}, trippedAt, counters}`, then the `operations.md` halt protocol with `haltReason="4:circuit-breaker:identical-finding"`. Interactive: show `delta.stillPresent`, ask `Continue rework` / `Escalate to me` / `Accept as deferred`. |
| **1** | Log `review.delta_skipped reason=invalid`, continue; the delta never blocks on its own failure. |

`plateau` is logged, not acted on; trigger 3 (the rework cap) is recorded by the Phase 2 re-entry.

#### Step 4  -  Consensus + Action (triage-driven)

If `triage.consensus.verdict` is `split` or `unverified`, surface `consensus.disagreements[]` to the user before acting: interactive modes show the split and ask whether to treat the unverified agreement as a pass (picker-contract: Trust / Re-review / Treat-as-blocking); autopilot logs `review.consensus=<verdict>` and proceeds on the triage verdict. Never silently average a split into a pass.

Act **only on triage.accepted**:

- **accepted.blocking** → back to Phase 2 (max 3 iterations, with reflection prompt citing only accepted items). The reflection prompt names findings by `fingerprint` and quotes `delta.stillPresent` first, marked `STILL PRESENT after round N-1's fix`. When Step 3.7 ran and `state.reviewIterations[-1].verifyByTest.redTests[]` is non-empty, the reflection prompt cites each red test: "a failing repro test already exists at <testRef>; make it green; do not delete or weaken it."
- **accepted.important** → fix and re-review
- **accepted.suggestion** → apply if reasonable
- **deferred** → append to Phase 5 report as "follow-up items" (do not block)
- **rejected** → log reasons for audit; do not touch

##### Lesson memory loop (required, end of each fix/rework round)

At the end of every fix/rework round (each Phase 2 re-entry that resolved accepted findings, including the final one), append ONE one-line root-cause lesson per resolved blocking/important finding to the existing learnings ledger (`learnings-ledger.mjs`  -  the store Phase 1 and triage already replay; never invent a parallel store):

```bash
node $HOME/.claude/scripts/learnings-ledger.mjs add --kind fact \
  --statement "<one-line rule/what-to-do so this class of finding does not recur (<=140 chars)>" \
  --diagnosis "<the causal WHY the failure happened, one line>" \
  --scope "<file-or-area glob from the finding>" --task "$TASK_ID"
```

Statement shape: the durable rule/root cause, not the symptom  -  "force-unwrapped optional in async callback path", not "fixed crash in LoginView". Always pass `--diagnosis` with the causal why (Reflexion: the verbal reason prevents recurrence; a bare outcome does not). Dedup built in (exit 2 = already stored). Lessons re-enter Phase 1 + triage via `learnings-ledger.mjs brief` (renders diagnosis as "(why: ...)").

Progress line: `    → writing lesson to learnings ledger ({N} entries)`

Log: "Phase 3: Review  -  raw={N1+N2+N3} accepted={Na} deferred={Nd} rejected={Nr} approved={bool} consensus={verdict}"

---

#### Multi-Repo Mode

**Only when `state.projects[].length > 1`.** Load `$HOME/.claude/multi-agent-refs/features/review-multi-repo.md` and follow it: per-repo diff scoping, how one merged finding set is attributed back to the repo that owns each file, and the per-repo `buildStatus` contract. A single-repo run skips this section entirely.

## Token telemetry  -  invoke after every LLM call

```bash
bash $HOME/.claude/scripts/phase-tracker.sh tokens 3 <input_count> <output_count> [cached_count]
```

The optional 4th `cached_count` is the prompt-cache-read token count when the host reports it (Anthropic `cache_read_input_tokens`); it defaults to 0 and is priced at the cheaper `cacheReadPerMtok` rate in the Phase 5 cost ledger. The tracker accumulates the totals additively, so multiple calls in the same phase compound. The render output then shows live cost on the active phase tile (e.g. `Phase 2  Dev   2m 14s · 12.4k tok`). This satisfies the contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` and the `smoke-tracker-tokens-invocation.sh` enforcement gate. Skipping this call is the #1 cause of "I can't see how much it cost" complaints.

Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.

---

## User test

The tail of Review rather than a phase of its own: it judges work that already
exists, which is what Review does.

> **TLDR**  -  Optional test gate. Offers to boot the simulator/emulator (UI Bug Hunter) or hand off to the user for manual QA. Needs an interactive prompt AND a worktree checkout, so it runs on an attended `/multi-agent` whose workspace is a worktree, and is skipped on `autopilot` or when the user chose to work locally. If issues found, loops back to Phase 2.

<!-- progress-contract: applied -->
Progress emission per `$HOME/.claude/multi-agent-refs/progress-contract.md`  -  lines for local-test prompt render, user-answer capture, repo checkout (if selected).

#### Step 0  -  Test Gap Report (advisory)

`state.testPolicy: none` → skip the gap scan (the gap IS the recorded policy) and run only pre-existing test targets; none → recorded no-op. Otherwise, before the local-checkout prompt, run the static test-gap detector. Heuristic, deterministic, no LLM, sub-second. The report ends up in `agent-log.md` under "Test Scenarios" and surfaces public symbols added in this branch that have no paired test.

```bash
GAP_JSON=$(node $HOME/.claude/scripts/test-gap-scan.mjs \
  --base "$BASE_BRANCH" --head HEAD --stack-from "$STATE_FILE" 2>/dev/null)
echo "$GAP_JSON" | node $HOME/.claude/scripts/validate-test-gap.mjs - >/dev/null 2>&1 || GAP_JSON=""
```

Skipped when `prefs.global.testGap.enabled` is `false`. Add `--scan-tree` when `testGap.scanTree` is true and `--severity-promote` when `testGap.promoteSeverity` is. Exit 3: the analysed stack (`ios|swift`, `android|kotlin`, `python`, `node|typescript|js`) has no gap rules, so there is no report.

**What the report contains** (per `$HOME/.claude/schemas/test-gap.schema.json`):

| Field | Meaning |
|---|---|
| `gaps[].sourcePath` | source file with the unprotected symbol |
| `gaps[].symbol` | symbol name |
| `gaps[].kind` | rule id (e.g. `public_func`, `composable_fun`, `named_export`) |
| `gaps[].severity` | `blocking` / `important` / `suggestion` (see severity table below) |
| `gaps[].expectedTestPaths` | likely test paths the user should land at, in priority order |
| `gaps[].hint` | stack-specific testing reminder (e.g. swiftui-qa.md 3-layer) |

**Severity defaults**:

| Symbol kind | Severity |
|---|---|
| `composable_fun`, `view_struct`, `config_struct`, `interface`, `objc_export`, `public_proto` | important |
| Other public API additions (`public_func`, `open_fun`, `named_export`, `default_export`, ...) | suggestion |

**Gating** (opt-in): if `prefs.testGap.blockingThreshold` is set and `gapBySeverity.important + gapBySeverity.blocking` exceeds it, Phase 3 surfaces the report as a Phase 3 rework finding and loops back. **Default off**  -  gaps render as advisory under "Test Gap Report" only.

**Telemetry**:

```bash
LOG_METRIC_FORWARD_TO_TRACKER=0 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 test_gap.scanned \
  stack=$SCAN_STACK \
  sources=$(jq '.totals.sourcesScanned' <<< "$GAP_JSON") \
  gaps=$(jq '.totals.gapCount' <<< "$GAP_JSON")
```

(No tracker forwarding  -  the scanner has no token cost.)

**Figma reference panel (when `state.evidence.figma[]` is non-empty).** Before the local-checkout prompt, print a single block listing each captured frame so the user has a side-by-side reference during manual test:

```
Figma evidence (tier=<n>):
  <fileKey>:<nodeId>  <canonicalComponentName>
    screenshot: <screenshotUrl or local path>
  <fileKey>:<nodeId>  <canonicalComponentName>
    screenshot: <screenshotUrl or local path>
```

Tier 1 / Tier 2 records print `screenshotUrl` from the captured evidence (Tier 2 URLs expire after 30 days, re-fetch on the spot if needed). Tier 3 records print the local path to the user-attached screenshot. The block is informational; it never blocks the prompt.

1. Ask with a native `AskUserQuestion` picker (never a typed y/N prompt), per `$HOME/.claude/multi-agent-refs/picker-contract.md`. The options MUST make the local-checkout side effect explicit  -  testing removes the worktree and checks the branch out into the main repo:
   - `question`: "Check out locally to test now?" (rendered in `outputLanguage`)
   - `header`: "Test" (English, <=12 chars)
   - `options`:
     - `{ label: "Test now", description: "Removes the worktree and checks the branch out into the main repo for Xcode / manual test" }`
     - `{ label: "Skip", description: "Stay in the worktree and go to Phase 4" }`
   - **Skip** → set `state.phases["3"].status = "skipped"` (so Phase 4 can offer the local-checkout prompt) → Phase 4
   - **Test now** → set `state.phases["3"].status = "in_progress"` → continue:
2. **Commit changes in worktree BEFORE removing** (WIP commit to preserve work):
   ```
   git -C {worktree-path} add -A
   git -C {worktree-path} commit -m "WIP: {jiraId}  -  changes for user test"
   ```
3. Remove worktree, checkout branch in main repo:
   ```
   git worktree remove .worktrees/{jiraId} --force
   git checkout {branch-name}
   ```
   Branch now has the WIP commit  -  all changes are preserved.
4. Show test instructions:
   ```
   Switched to branch: {branch-name}
   To test: Xcode -> build -> run -> manual test
   "ok" -> proceeds to Phase 4 (WIP commit will be replaced via git reset HEAD~1 + proper commit)
   "fix: ..." -> worktree is recreated, returns to Phase 2
   ```
5. Mark the tracker as waiting, then wait for the user response:
   ```bash
   bash $HOME/.claude/scripts/phase-tracker.sh now 3 "awaiting local test (user)"
   bash $HOME/.claude/scripts/phase-tracker.sh render
   ```
   The waiting state persists in `tracker-state.json` across the handoff; `/multi-agent:resume` and `/multi-agent:manual-test` CONTINUE this state file and never re-init it (`$HOME/.claude/multi-agent-refs/tracker-contract.md` "Continuation runs").

   **"ok" is a structured result, not a word.** Before "ok" is accepted, the run writes `$WORKTREE/.pipeline/manual-test.json`: one entry per acceptance criterion, the criteria taken from the analysis doc test plan (Section 15 / 20), the plan tasks, and the user's own words in the reply. Every criterion records what was seen; a criterion that was not tried says so with a reason.
   ```json
   {"criteria":[{"spec":"<quote>","source":"analysis 15.2 | plan task 3 | user","observed":"<what was seen>","verdict":"pass|fail|not-tested","reason":"<required when not-tested>","screenshot":"<path or null>"}],"verdict":"passed|failed"}
   ```
   Then gate it:
   ```bash
   node $HOME/.claude/scripts/evidence-gate.mjs --claim manual --status passed --evidence "$WORKTREE/.pipeline/manual-test.json"
   ```
   Add `--require-screenshot` when `state.visualEvidence.required` is true: a passing criterion then has to name a file that exists, because a `screenshot` key pointing nowhere is not evidence.

   Exit 1 means the "ok" is not accepted: tell the user which criterion is missing evidence (a `fail` verdict, or `not-tested` without a reason) and wait for the next reply. Exit 0 marks Phase 3 completed with `Result "local test passed (user)"`. The "fix: ..." path below is unchanged.
6. If fix needed:
   - Branch already has WIP commit (from step 2)  -  changes are safe
   - **Heal stale admin state first** (same contract as Phase 0  -  step 3's
     `worktree remove` or an interrupted run can leave a stale entry, so a bare
     re-add fails with `already exists`/`already registered`):
     ```bash
     git -C "$PROJECT_ROOT" worktree unlock "{worktree-path}" 2>/dev/null || true
     git -C "$PROJECT_ROOT" worktree prune 2>/dev/null || true
     ```
     Unlock first: prune skips a locked entry. Phase 0's repo residue guard is already in `.git/info/exclude`  -  no re-add.
   - Recreate worktree from branch: `git -C $PROJECT_ROOT worktree add {worktree-path} {branch}`
   - Re-set git identity: `git -C {worktree-path} config user.name/email` (from state)
   - Go back to Phase 2
7. Log: "Phase 3: Review (user test)  -  {result}"

**CRITICAL**: Never remove a worktree with uncommitted changes. Always WIP commit first.

#### Automated Device Checks (on-demand)

Before or during user testing, run device-level audits via Bash if user requests. See `audit-guide.md` for commands.

| Check               | When              | Command                                   |
| ------------------- | ----------------- | ----------------------------------------- |
| UI flow video       | `state.visualEvidence.required` AND Phase 2 recorded none | `capture-evidence.sh video start` -> drive the flow -> `video stop` -> `fit`. See below |
| Accessibility audit | UI changes        | `mcp__multi-agent-toolkit__{ios,android}_accessibility_audit` |
| Biometric test      | Auth flow changes | ios: `mcp__multi-agent-toolkit__ios_biometric` (android: manual) |
| Launch time         | Perf-sensitive changes | ios: app-launch instrument · android: `mcp__multi-agent-toolkit__android_launch_time` |
| Visual test         | Any UI changes    | `/multi-agent test` (sim-test, both platforms) |
| Snapshot regression | Component / pixel-stable UI changes | ios: `mcp__multi-agent-toolkit__ios_visual_diff` · android: `mcp__multi-agent-toolkit__android_screenshot` + compare |
| Store screenshots   | `taskType === screenshot` | ios: `ios_status_bar({preset: "clean"})` · android: `android_screenshot` |

##### UI flow video, when Phase 2 produced none

Phase 2 Step 3.55 is the primary host and runs in every mode. Phase 3 is the richer one where it exists: the device is up and a person is watching, so the flow is one somebody confirmed. It adds, never replaces.

Run only when `visualEvidence.required` and `visualEvidence.video.file` is empty: re-check the device with `$HOME/.claude/scripts/probe-evidence-capability.sh --only device`, then `$HOME/.claude/scripts/capture-evidence.sh video start` -> drive the flow (`run-ui-tests.sh run`, or the Phase 3 scenarios by hand) -> `video stop` -> `fit`. The cap is `visualEvidence.maxVideoSeconds`, read through `capture-evidence.sh limits` so one value serves both. Update `videoTier` and `videoTierReason` with what ran.

When the intake answer was `unit` there is no recording here either: overriding it in a phase the user may not be watching makes the question decorative.

Results included in Phase 5 report. MCP tools preferred when available  -  concise structured output, lower token cost.

**Snapshot regression flow (optional):** when the task changes a stable component, capture a screenshot before the change (baseline) and after (current), then call `ios_visual_diff({baseline, current, max_diff_pct: 1.0})`. Threshold can be relaxed for animated / non-deterministic regions  -  keep `max_diff_pct ≤ 1.0` for static layouts.

#### Telemetry  -  token forwarding

When the security-auditor or any other Phase 3 sub-agent runs, forward its token totals so Phase 5's Cost Breakdown captures Phase 3:

```bash
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 3 audit.completed \
  model=opus tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
```

If Phase 3 is purely user-driven (no sub-agent ran), no token forwarding is required and the cost block stays empty for this phase. Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.
