### Phase 2: Dev (Sonnet)

> **TLDR**  -  Sonnet executes the plan task-by-task with TDD (red→green→refactor). Required: issue-tracker status moved to "In Progress" before any code, with a post-mutation verify step (re-reads the field, retries once on silent VALIDATION failures). Build verification after each task (up to 3 retries). Build-queue lock serializes concurrent xcodebuild. `taskType === component` **short-circuits the TDD path** and delegates the whole phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`)  -  see next subsection.

## Phase 2 Pre-flight (BLOCKING, v9.0.0)

Per Locked decision 30, Phase 2 Dev consumes the analysis document as the sole design source. **Figma** MCP / REST / URL fetches are forbidden; the toolkit MCP is not, and Steps 3.4 and 3.55 call it.

Pre-flight steps (run in order, abort on failure).

**Steps 1, 2, 3, 5 and 6 read the analysis document.** Phase 1 always runs, so the document always exists; what varies is how much evidence it carries. A section the evidence did not support is absent by the Locked 2 omission rule, and a step whose section is absent is recorded `not-applicable (no <section> in this document)` rather than aborting. Steps 4, 7, 8 and 9 read nothing from it and apply always.

1. **Analysis document presence** (Phase 1 modes only): read `state.analysis.docStatus` and `state.analysis.docPath[]`, both set by Phase 1 Step 4.
   - `produced` | `reused` -> read the active platform's file; multi-repo runs need one per selected repo.
   - `not-applicable` -> the analysis mode exported a document this run does not consume; record it for steps 1, 2, 3, 5, 6 and skip them.
   - **Abort**: `produced` but unreadable -> `ERR: analysis doc at <path> unreadable. Resume with /multi-agent:resume #N.` Producing it is Phase 1's job.

2. **Parse YAML front-matter** into `state.analysis.frontMatter`: `feature`, `platform`, `language`, `mode`, `evidence_digest`, `template_version`. **Abort** when `template_version` < `v3`; there is no degraded mode.

3. **Spec freshness**: two checks, both non-blocking, both flagged for Phase 3.
   - **Digest**: `state.run.lastAnalysisDigest` against `frontMatter.evidence_digest`. Mismatch = written from different evidence. Key absent -> `not-verifiable`, never `fresh`.
   - **Repo drift**: `git diff --name-only <state.run.analysisBaseCommit>..HEAD` intersected with Section 14 paths. Non-empty = the repo moved under the spec. This is the one that fires on a `reused` doc, where the digest still matches because the evidence inputs did not change.

4. **Code Connect mapping lookup**: for each component named in analysis Section 6 (Bileşen Envanteri), search the repo for matching `*.figma.swift` / `*.figma.kt` files. Persist hits to `state.dev.codeConnect[<componentName>]`. Missing mappings are not blockers but the Phase 3 reviewer will flag them.

5. **Standards binding citations**: read `analysis Section 21 References` table. For each row with `Rol: bağlayıcı / binding`, persist the source path into `state.dev.standardsBindings[]`. Phase 2 Dev tasks MUST cite at least one binding source per architectural decision (pre-existing Locked decision 8).

5b. **Test plan handoff (the RED input)**: read `analysis Section 15` into `state.dev.testPlan[]` (15.1 unit rows with name/arrange/act/expected/`BR-` id, 15.2 snapshot variants, 15.6 UI flows, 15.7 manual scenarios). **RED writes these tests, not invented ones** - development is TDD, so the analysis matrix is literally the first thing written. A row too vague to write from is an analysis defect: open a Section 20 question, do not improvise.

6. **Conventions handoff**: read `analysis Section 13.1 Concept Table` (Pass B output with footnotes). Persist concept-to-realization mapping into `state.dev.conventions[<concept>]`. Phase 2 implementation uses these names verbatim, and a non-empty `sharedUtilities` bucket is binding: bind those formatters/rule-facades/tokens, never hand-roll a duplicate (e.g., if Section 13.1 says "State holder: OrderSummaryViewModel", Phase 2 names the class exactly `OrderSummaryViewModel`).

7. **MCP forbidden**: calling `mcp__claude_ai_Figma__*` in Phase 2 is a violation. Record any call in `state.telemetry.mcpCalls[]` with its phase; the maintainer regression check `smoke-no-mcp-in-dev-phases.sh` flags a Figma entry at phase 2 or later.

8. **Criteria ledger (required, every mode)**: the moment this phase consults a skill, a marketplace plugin skill, a stack guide or a module `CLAUDE.md` in order to write code, append an entry to `state.telemetry.skillCalls[]`:

   ```json
   {"skill": "ios-coding-standard", "phase": 3, "targetFiles": ["Sources/Login/LoginViewModel.swift"], "timestamp": "<ISO-8601>"}
   ```

   `targetFiles` is required  -  without it a skill applied to the wrong files still reads as "applied". Append at the moment of consultation, not at the end of the phase. Phase 3 Step 1.78 treats this as self-report only and resolves criteria independently; it is the one signal separating "applied to the wrong files" from "never opened".

9. **Stack skill routing (every `taskType`, always)**: write the task title + intent to `$TASK_FILE` with the Write tool (never a shell string), run `node $HOME/.claude/scripts/skill-candidates.mjs resolve --dir "$PROJECT_ROOT" --state "$STATE_FILE" --task-file "$TASK_FILE"`, and store its `mode`, `toolkits`, `excluded`, `unscoped`, `hints` and `fallbacks` in `state.telemetry.skillRouting`. Ask each kept toolkit's own `index` which skills govern this task, load them BEFORE writing code, and record each into `state.telemetry.skillCalls[]` with `phase: 2` and `routedBy: "<toolkit>:index@<version>"`. Precedence: repo-local skill (`source: "repo"`) > detected-stack toolkit skill > common toolkit; `overlaps[]` names the owner of a skill two toolkits share. Unattended or autopilot runs load `fallbacks[]` read-only (`source: "marketplace-fallback"`) and never route to `unscoped[]`; a missing stack toolkit is a hint, never another stack's toolkit. The routing table stays in the plugin. A screen-creation task loads the routed toolkit's `workflow/create-screen` when one exists. No toolkit is a recorded no-op, not a halt. Contract: [`features/stack-skill-routing.md`]($HOME/.claude/multi-agent-refs/features/stack-skill-routing.md).

10. **Autopilot** (or `MULTI_AGENT_UNATTENDED=1`): a run that Phase 1's `open-questions-gate.mjs` parked (`waitingFor: "question"`), or that `spec-consistency-gate.mjs` or `plan-critique-gate.mjs` failed (`verificationFailed.gate: spec-consistency` / `plan-critique`), does not enter this phase. When `state.research.md` exists, read it before writing code: what research closed, each answer with its source ([`features/research.md`]($HOME/.claude/multi-agent-refs/features/research.md)).

The analysis document is the SOLE design source in Phase 2. Variant choices, padding values, color tokens, copy strings, accessibility identifiers, and test method names all come from the rendered Pass B cells. If something is missing in the analysis doc, the fix is to re-run `/multi-agent:analysis`, not to fetch from Figma.

<!-- progress-contract: applied -->

**Progress (per `$HOME/.claude/multi-agent-refs/progress-contract.md`):** every non-trivial step emits one `→ <verb> <object>` line at immediate flush. Canonical set for standard TDD: `→ reading <file>`, `→ writing test <name>`, `→ running xcodebuild RED`, `→ writing code <file>`, `→ running xcodebuild GREEN`, `→ committing WIP <summary>`. Component-task dispatch emits the line `→ dispatching create-component <componentName>` immediately before the Skill-tool call and then relays the plugin skill's progress lines verbatim under indent `    →`.

#### Input contract

Phase 2 consumes the Phase 1 output object conforming to `$HOME/.claude/schemas/planning-output.schema.json`  -  the task graph (`tasks[]` with `id`, `title`, `type`, `files`, and optional `dependsOn` / `acceptanceCriteria`) plus the architecture review notes. Tasks execute in dependency order; the schema's `dependsOn` field drives the ready-task picker.

**Plan Todo iteration (opt-in)**: gated by `prefs.global.planTodos.enabled` (default: `false`). When enabled and Phase 1 Step 11 emitted a `plan.todos[]`, Phase 2 iterates via `$HOME/.claude/lib/plan-todos.sh next/start/complete/fail` instead of walking `tasks[]` directly. When disabled, the loop walks `tasks[]` from `planning-output`  -  TDD contract is unchanged. Full helper loop + state semantics: `$HOME/.claude/multi-agent-refs/features/plan-todos.md`. A todo with `sourceTag: Reuse` binds the file analysis already found; `Modify` edits in place. Writing a new file over a `Reuse` step is a Locked 11 violation and Phase 3 flags it.

**Shadow-Git checkpoints (opt-in)**: gated by `prefs.global.shadowGit.enabled` (default: `false`). When enabled, the orchestrator snapshots the worktree via `$HOME/.claude/lib/shadow-git.sh` so sub-phase rollback is possible without touching the project's real `.git` history. Lifecycle: `shadow-git.sh init` (Phase 0 baseline), `shadow-git.sh snapshot` (per step after `plan-todos complete`), `shadow-git.sh restore <sha> --files` (rollback). Modes: `per-todo-step` (default) or `per-tool-call`. Full wiring + storage cap: `$HOME/.claude/multi-agent-refs/features/shadow-git.md`.

#### Component tasks  -  delegated dispatch (taskType === "component")

When Phase 0 Step 7 classified the task as `component`, Phase 2 delegates the entire phase to the enabled `ai-<platform>-toolkit` marketplace plugin's component skill (`create-component`, fallback `create-ui-component`) via the Skill tool and does NOT run the TDD loop below. The dispatch layer passes the plugin skill the analysis Section 6 (Bileşen Envanteri) entry + Section 13.1 conventions for the named component as context. Because plugin skills do not write pipeline state, the **dispatch layer** (not the skill) owns `state.phases["2"].subphases[]`, recording a coarse component-build row - multi-agent's `phase-tracker` reads that array with no special case. Plugin resolution (dual-name), dispatch call, failure/resume, multi-repo, and the intentional cross-CLI divergence live in `$HOME/.claude/multi-agent-refs/component-dispatch.md` - read it before editing component-task behaviour here. Phase 2 still owns: progress line `-> dispatching create-component <name>`, `retryCount` cap at 3, and fallthrough to the TDD path when dispatch prerequisites are missing (`taskType` absent OR the plugin is not enabled in this repo -> log anomaly, halt or run TDD per component-dispatch.md).

For non-component taskTypes (`bugfix`, `feature`, `refactor`, `chore`), continue with the standard TDD section below.

#### Design fidelity contract (BLOCKING when the task carries a Figma reference)

Applies to EVERY task whose analysis doc carries design ground truth (Section 6 rows, node IDs, screenshots)  -  component-dispatch AND TDD-path UI work alike. MCP stays forbidden here (analysis-only rule); the analysis doc IS the design:

1. **1:1, not interpretation.** The captured frames are the reference for every UI line. Layout, variant states, copy, and colours come from the analysis doc's measured values  -  never eyeballed, never "close enough".
2. **Code Connect decides the component.** A mapped component (`CodeConnectSnippet` / repo `*.figma.swift` / `*.figma.kt` row) MUST be used verbatim  -  sound-alike substitutes are forbidden; a missing modifier is added to the component's `+Modifiers` extension, never forked. No mapping exists -> write a NEW component per the Configuration/View/Modifiers architecture; never inline ad-hoc UI into the consumer screen.
3. **Inter-component spacing is part of the design.** Gaps, paddings, and alignment BETWEEN components must match the frame's measured values, mapped to spacing tokens (`Spacing*`)  -  never invented numbers. A spacing/layout deviation is a `review_blocking` finding in Phase 3 Step 1.8, not a nitpick.

Missing design data (variant, padding, copy) → HALT and instruct the user to re-run `/multi-agent:analysis`; never guess and never call Figma MCP from this phase.

#### Re-entry from Phase 3 triage

Phase 2 runs twice in the pipeline lifetime: first for initial development, then optionally for rework after Phase 3 review. **Phase 2 never acts on raw reviewer output.** It only consumes `triage.accepted` findings  -  Fable triage in Phase 3 already filtered false-positives, deferred out-of-scope items, and rejected noise.

When re-entering from Phase 3:

1. Read `state.reviewIterations[-1]`  -  the latest iteration's **accepted** list.
2. For each `accepted` finding where `severity === "blocking"` or `severity === "important"`:
   - Treat it exactly like a Phase 1 plan task (TDD cycle, build verify).
   - Quote the original `issue` + `fix` text in the task description so the dev agent has full context.
3. `accepted.suggestion` items are applied opportunistically (no TDD loop required) unless the user asked for suggestions to be treated strictly.
4. `deferred` items are NOT actioned in this re-entry  -  they surface in Phase 5's "Follow-up items" section.
5. `rejected` items are never touched. Log their IDs + triage reasons for audit only.
6. After rework, increment `state.phases["2"].retryCount`; hard-kill at `retryCount === 3` and escalate to the user. Do not loop indefinitely. When `prefs.global.autopilotCircuitBreaker.enabled` is true (default), record the hard-kill before escalating as `state.circuitBreaker = {tripped: true, trigger: 3, detail: "rework cycles exhausted", checkpoint: {phase: 3, step: "re-entry", iteration: <n>}, trippedAt, counters: {reworkCycles: <n>}}`: the rework-storm trigger; `autopilotCircuitBreaker.maxReworkCycles` (default 3) is bounded above by this cap.
7. When `state.reviewIterations[-1].delta` exists, `stillPresent[]` findings come first in the task list, quoted `STILL PRESENT after round N-1's fix`; repeating the previous attempt is the loop the circuit-breaker stops.

If the latest iteration has `triage.approved === true` AND `accepted === []`, Phase 2 was entered by mistake  -  log the anomaly and return to Phase 3.

**Telemetry**: at the start of every re-entry, emit:

```bash
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 rework.started \
  iteration=$ITERATION accepted_blocking=$BLOCKING accepted_important=$IMPORTANT
```

---

For each task (respecting dependency order), on start and on end: `bash $HOME/.claude/scripts/phase-tracker.sh sub 2 <task_id> "<title>" in_progress|completed|failed`.

1. **Update task status: `in_progress` (required  -  do not skip or defer)**
   - **Why mandatory**: the board/tracker visually stays in "Todo" until this runs. Skipping it means reviewers and the rest of the team cannot see that work has started, and automations downstream (status reports, kanban swimlanes) will be wrong.
   - **Where it applies**:
     - Jira: transition issue to "In Progress" via REST API (`/rest/api/2/issue/{id}/transitions`)
     - GitHub Projects V2: `updateProjectV2ItemFieldValue` on the Status field (singleSelectOptionId)
     - Internal state: `agent-state.json` → `phases["2"].status = "in_progress"`
   - **Post-update verify** (catches silent VALIDATION errors from stale option IDs after board rebuilds):
     - Re-read the Status field after mutation.
     - If the value is NOT `In Progress` (Jira) / `In Develop` (Projects V2), re-fetch current option IDs and retry ONCE.
     - If verify still fails after retry → log error, do NOT proceed to TDD. The mutation silently failed and continuing would leave the task mis-tracked.
2. **Multi-submodule awareness**: If the project has submodules (detected in Phase 0 Step 7), development may span multiple submodules. For example, a UI component might need:
   - New types/tokens in the **common** submodule
   - The component itself in the **uicomponents** submodule
   - Work in BOTH submodules in the same task  -  this is normal and expected
3. **TDD cycle** (Launch Agent with `model: "sonnet"`)  -  gated by `state.testPolicy`: `tdd` = the loop below; `tests-after` = skip RED, author the same tests once green; `none` = author no tests (existing ones never weakened), report "tests not written - project policy" rather than a gap.

   **RED  -  Write ONE failing test first:**
   - **Rework re-entry**: if `state.reviewIterations[-1].verifyByTest.redTests[]` exists (Phase 3 Step 3.7 ran), those failing repro tests ARE the RED step for their findings  -  make them green, do not write a duplicate failing test and do not delete or weaken them. See `$HOME/.claude/multi-agent-refs/features/verify-by-test.md`.
   - Test framework: use whatever the project already uses. Detect from existing test files:
     - `@Test` / `#expect` → Swift Testing
     - `XCTestCase` / `XCTAssert` → XCTest
     - `describe` / `it` → Quick/Nimble
   - Test naming: `test{Scenario}_{Expected}` (e.g. `testKeychainReturnsNil_doesNotCrash`)
   - One test per behavior change  -  not one test per file
   - Run test to confirm RED with the stack's command. ios, under the build lock:
     ```bash
     bash $HOME/.claude/scripts/build-lock.sh acquire "$TASK_ID"
     xcodebuild test -scheme "{scheme}" -destination "platform=iOS Simulator,name={simulator}" \
       -derivedDataPath "{worktreePath}/.DerivedData" -only-testing:"{testTarget}/{testClass}/{testMethod}" 2>&1 | tail -5
     bash $HOME/.claude/scripts/build-lock.sh release "$TASK_ID"
     ```
     android: `./gradlew test --tests "{testClass}.{testMethod}" 2>&1 | tail -5` (same lock); backend: `pytest "{test_file}::{test_name}" 2>&1 | tail -5`; web: `node $HOME/.claude/scripts/package-manager.mjs test --dir "{worktreePath}" --pattern "--testPathPattern={file}"` prints the manager's command, then run that command `2>&1 | tail -5`.
   - The web step resolves the manager instead of typing `npm`; exit 3 means the repo
     declares no such script - say so, never substitute one, and never let an empty
     command pass for a pass (`features/package-manager.md`).

   - Must fail for the RIGHT reason (expected assertion, not compilation error)

   **GREEN  -  Minimal code to pass:**
   - Smallest change that makes the test green  -  no extras
   - **Existing tests are immutable**: deleting, renaming, skipping, or weakening an existing assertion to reach green is a violation. A test may change only when the task itself changes the spec that test encodes, and the commit body must name the changed test and the spec change. (Deterministic backstop: the `test_lines_removed` signal in Phase 3 Step 1.75 flags test files that shrink.)
   - Run the same test again → must PASS
   - Run full test suite → no regressions:
     ```bash
     bash $HOME/.claude/scripts/build-lock.sh acquire "$TASK_ID"
     xcodebuild test \
       -scheme "{scheme}" \
       -destination "platform=iOS Simulator,name={simulator}" \
       -derivedDataPath "{worktreePath}/.DerivedData" \
       2>&1 | tail -20
     bash $HOME/.claude/scripts/build-lock.sh release "$TASK_ID"
     ```

   **REFACTOR (if needed):**
   - Only if duplication or naming is poor
   - Re-run tests after refactor → still GREEN
   - **Stability (required):** run every test added or changed in this diff `prefs.global.testStability.repeatCount` times (default 3; 1 disables) with the single-test invocation above. Disagreeing outcomes are not GREEN: log `test.flake_signal file=<f> passed=<k> of=<N>` and fix the test or the code first. A pass only on retry is a flake signal, not a pass. Same rule on the Phase 3 rework re-entry.

   **Target resolution** (auto-detect once per project, cache in `agent-state.json`; ios resolves scheme + simulator, android resolves module + variant, backend/web need none):
   ios (prefer the scheme matching the project name, then the first non-test scheme):
   ```bash
   xcodebuild -list -json -project "{projectPath}" 2>/dev/null || xcodebuild -list -json -workspace "{workspacePath}" 2>/dev/null
   xcrun simctl list devices available -j | jq '.devices | to_entries[] | select(.key | contains("iOS")) | .value[0].name'
   ```
   android: `./gradlew projects 2>/dev/null | grep -E "^\+--- Project"` and `./gradlew tasks --all 2>/dev/null | grep -m5 "assemble.*Debug"`.

4. **Build verification** (per stack; ios/android under the build queue lock, see below):
   - **ios, preferred (MCP, multi-agent-toolkit >= 3.0.0)**: `build-lock.sh acquire "$TASK_ID"` → `mcp__multi-agent-toolkit__ios_xcodebuild({project|workspace, scheme, configuration: "Release", destination: "generic/platform=iOS", derived_data_path: "{worktreePath}/.DerivedData"})` → `build-lock.sh release "$TASK_ID"`. Returns one line `Build: SUCCESS|FAILURE (E errors, W warnings) [xcresult-<id>]`; on failure drill in via `mcp__multi-agent-toolkit__ios_xcresult({id, mode: "errors"})`, never dump the full log.
   - **ios, fallback (raw)**: same lock pair around `xcodebuild build -scheme "{scheme}" -destination "generic/platform=iOS" -derivedDataPath "{worktreePath}/.DerivedData" 2>&1 | tail -5`.
   - **android**: lock pair around `./gradlew assembleDebug 2>&1 | tail -5` (the Gradle daemon and `build/` outputs contend across parallel worktrees exactly as DerivedData does  -  the lock applies).
   - **backend / web**: `python -m compileall .` / run the command `node $HOME/.claude/scripts/package-manager.mjs run --dir "{worktreePath}" --script build` prints, `2>&1 | tail -5` (exit 3: no build script); no lock.
5. If build fails → fix → rebuild (max 3 attempts, track `retryCount` in state).
6. **Intermediate commit** (after each completed task in the plan):
   ```bash
   git -C "{worktreePath}" add -A
   git -C "{worktreePath}" commit -F "/tmp/wip-message-$TASK_ID.txt"
   ```
   Write `wip({scope}): {task summary} [{jiraId}]` to `/tmp/wip-message-$TASK_ID.txt` with the file tool first. The summary comes from the issue, and inside a double-quoted `-m` the shell would run any `$(...)` in it.
   WIP commits are squashed in Phase 4 before PR. This prevents work loss if a later task fails or the session crashes.
7. Update task status: `completed` (apply the same verify rule  -  re-read the Status field and retry once on mismatch)
8. Log: "Phase 2: {task}  -  Test + Build passed"

---

#### Build Queue (parallel task serialization)

**Problem**: Multiple parallel worktrees may reach build/test phase simultaneously. `xcodebuild` cannot run in parallel  -  DerivedData, simulators, and project locks cause conflicts.

**Solution**: Lock-file based queue using `mkdir` atomicity (`build-lock.sh`). Before ANY `xcodebuild` call, acquire the lock; release after completion:

```bash
bash $HOME/.claude/scripts/build-lock.sh acquire "{jiraId}"
xcodebuild -derivedDataPath "{worktreePath}/.DerivedData" ...
bash $HOME/.claude/scripts/build-lock.sh release "{jiraId}"
```

**Key details:**

- Lock path: `/tmp/claude-xcodebuild.lock` (shared across all Claude instances; `MA_BUILD_LOCK` overrides)
- Stale lock timeout: 15 minutes (auto-cleanup if a session crashes)
- Each worktree uses its OWN `-derivedDataPath` to avoid cache poisoning
- `sleep 5` between retries  -  not aggressive polling
- Lock owner tracked for visibility: which task is currently building
- `release` takes the same task id and leaves a lock held by another task in place
- A lock whose owner file is missing is taken over only after a 5-second grace window; with `MA_BUILD_LOCK_PID` set, a lock whose recorded process is gone is taken over at once

**This applies to ALL xcodebuild calls in the pipeline:**

- Phase 2 Step 4 (build after development)
- Exit gate Step 1 Gate 1 (build gate before review)
- Exit gate Step 1 Gate 3 (test gate before review)

**Android**: same lock discipline  -  the Gradle daemon and `build/` outputs contend across parallel worktrees. **Backend/web** (Python, Node.js): no lock needed  -  these build/test in parallel without conflicts.

---

#### Step 3.5  -  Dev Critic (Evaluator-Optimizer, opt-in)

Gated by `prefs.global.devCritic.enabled` (default: `false`). When enabled, after the generator's last edit and BEFORE Phase 3: dispatch `dev-critic` sub-agent (Sonnet), run 4 deterministic gates (build / lint / test / secrets), then platform checklist. STRICT loop cap  -  **max 2 iterations** (round 2 re-checks round-1 failures only); round 3+ returns `escalate: true`. Severity routing: `blocking` → generator must fix; `important` → SHOULD fix or pass through to Phase 3; `suggestion` → generator's judgement. Full agent contract, schema, telemetry, when-to-enable: `$HOME/.claude/multi-agent-refs/features/dev-critic.md` + `$HOME/.claude/agents/dev-critic.md`.

---

#### Step 3.55  -  Visual evidence capture (UI changes only)

Re-decide `visualEvidence.required` from the real diff first - Phase 0's `bugfix` verdict is provisional (section 1a). When required, capture with the build that just went green - here, not Phase 3 (section 3). Exit 4 anywhere below is a gap, not a failure. Contract: `features/visual-evidence.md`.

`EVIDENCE_PLATFORM=$(jq -r '.visualEvidence.platform // ""' "$STATE_FILE")`. Empty closes every tier below with that reason; passing it on is exit 2.

1. **Still**: `$HOME/.claude/scripts/capture-evidence.sh after --task "$TASK_ID" --platform "$EVIDENCE_PLATFORM" --label <slug>`.
2. **Re-check the device.** `evidenceCapability` was measured at intake; a simulator booted then can be gone now. Re-measure the one volatile row with `$HOME/.claude/scripts/probe-evidence-capability.sh --platform "$EVIDENCE_PLATFORM" --repo "$WORKTREE" --only device`, which skips the repo scan the full probe does; gone means fall to the next open tier and write `visualEvidence.videoTierReason` as the transition (`tier 1 -> 3: simulator no longer booted`).
3. **Recording**, by `state.testDepth`:
   - `unit+ui`: `video start` -> `$HOME/.claude/scripts/run-ui-tests.sh run --platform "$EVIDENCE_PLATFORM" --repo "$WORKTREE" --changed "$CHANGED_CSV"` -> `video stop`. Runner exit 3 (no matching test) or 4 (no target) -> stop, discard, fall to tier 2. Exit 1 is a red UI test: keep the recording, it shows the failure.
   - `unit+mcp`: `video start` -> drive with `mcp__multi-agent-toolkit__agent_run_steps` -> `video stop`. The only MCP-dependent path; MCP absent -> tier 3 gap.
   - `unit`: no recording; the gap reason is the user's own answer.
4. **Fit**: `$HOME/.claude/scripts/capture-evidence.sh fit --file <path>` per artefact. Quality degrades, the artefact is never dropped.

`video stop` warns when the recording is under a second: both recorders encode on change, so a screen that never moved yields a valid two-frame file. Keep it, record the warning as a gap reason, never present a still as a flow.

Persist `state.uiTest` and the artefact entries. A UI test that ran is subject to the default-FAIL rule like the build: run `evidence-gate.mjs --claim test --status passed --evidence "$WORKTREE/.pipeline/ui-test.log"` first, because a runner that died before reaching the tests also exits non-zero.

#### Step 3.6  -  Code-simplifier pass (required diff shrink, before Phase 3 handoff)

After the build/test green step and BEFORE Phase 3 handoff, run one diff-shrink round (bloated diffs waste Phase 3 reviewer tokens and unfocus the PR).

1. **Dispatch ONE subagent** (`subagent_type: "general-purpose"`, model: `sonnet`); input is ONLY the working diff (`git -C "$WORKTREE" diff "origin/$BASE_BRANCH"...HEAD`) + task summary. No correctness re-review (Phase 3's job), no new features. Exactly four smells:
   - **Comment bloat**  -  comments restating code, scaffolding notes, TODO chatter
   - **Unrelated rewrites**  -  hunks outside task scope (re-formatting, unrequested rename sweeps)
   - **Dead code**  -  unused symbols, unreachable branches, commented-out blocks from this diff
   - **Over-abstraction**  -  single-call-site protocols/wrappers/helpers from this diff

   Progress line: `    → dispatching code-simplifier diff-shrink`
2. **Return contract**  -  a list of shrink edits `{ "file", "lines", "edit", "rationale", "risk": "safe|unsafe" }`; the subagent never writes files. Keep the `rationale` strings: Step 3.7 writes them into `scope-check.json`.
3. **Apply safe edits only**  -  skip `unsafe` (logged). Zero edits is normal  -  log and continue. Progress line: `    → applying shrink edits ({N} applied, {M} skipped)`
4. **Re-run build + tests** (same build-queue lock + evidence-gate rule as Step 4). Any breakage -> revert shrink edits wholesale and proceed pre-shrink; the simplifier must never cost a green state.
5. **Record tokens in the cost ledger** so Phase 5's Cost Breakdown captures the pass:

```bash
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.simplifier_pass \
  model=sonnet tokens_in=$IN tokens_out=$OUT duration_ms=$DUR \
  edits_returned=$RET edits_applied=$APPLIED edits_skipped=$SKIPPED
```

Scope guard: a single pass, never looped. Component tasks (`taskType === "component"`) skip it (the figma skill owns its own checklist).

---

#### Step 3.7  -  Scope self-check (required handoff artifact)

Phase 3 cannot reconstruct why each file was touched or what was left out on purpose, so Dev states both before the handoff: write `$WORKTREE/.pipeline/scope-check.json` (`schemas/scope-check.schema.json`: one `{path, reason}` per file in the diff, `notDone[]` with `why` in `out-of-scope | follow-up | rejected-abstraction`, and the Step 3.6 `simplifier` rationales), then run the gate. **Record rules, consumers and the full contract: `$HOME/.claude/multi-agent-refs/features/scope-check.md`.**

```bash
git -C "$WORKTREE" diff --name-only "origin/$BASE_BRANCH"...HEAD \
  | node $HOME/.claude/scripts/scope-check-gate.mjs --check "$WORKTREE/.pipeline/scope-check.json" --diff-files -
```

Exit 1 lists `unjustified[]`: complete the record once and re-run. A second exit 1 does not block; log `dev.scope_check=incomplete unjustified=<n>` and hand off, and Phase 3 shows the gate output to reviewers inside `<scope-self-check>`. Progress line: `    → checking scope self-check ({n} files, {m} not done)`.

---

#### v2.1.0+ Multi-Repo Mode

Active when `state.projects[].length > 1` (set by Phase 0 multi-select). Single-repo flow above is preserved verbatim  -  this section adds the deltas.

**Todo tagging**: Phase 1 plan items in multi-repo mode carry a `repo` field naming the target project root. The dev agent uses this tag to dispatch each todo:

```json
{
  "id": "T-3",
  "title": "Add tokenResolver public API",
  "repo": "common",                         // ← matches state.projects[i].name
  "files": ["Sources/Common/TokenResolver.swift"],
  "tdd": { "test": "...", "code": "..." }
}
```

If a todo has no `repo` tag in multi-repo mode → log warning + ask user, do not auto-pick.

**Per-todo worktree switch**: Before each todo's TDD cycle, resolve `WT_PATH = state.projects[ findIndex(name === todo.repo) ].worktreePath` and run **all** subsequent commands with `git -C "$WT_PATH"` and `cd "$WT_PATH"` for build/test. Identity (`user.name`/`user.email`) was already pinned per worktree by Phase 0 Step 8  -  no re-config needed here.

**Status update (single ticket, multi-repo)**: The Jira/GitHub issue is shared across all repos in the group  -  call the status-update API exactly **once per todo** (not once per repo per todo). The post-mutation verify rule still applies.

**Build serialization**:
- Xcode worktrees still share `BUILD_LOCK = /tmp/claude-xcodebuild.lock`  -  multi-repo just means more callers contending for the same lock; the existing `mkdir`-based queue handles it correctly.
- Non-Xcode repos may build in parallel across worktrees  -  but ONLY when builds are independent. If repo B's build depends on repo A's freshly-published artifact (e.g. SPM local path dependency), serialize them in dependency order. Plan items declare dependencies via Phase 1's `dependsOn` field; multi-repo dispatch respects that order.

**Failure isolation**: A build/test failure in one repo's worktree must NOT silently leave the other repos in a half-built state. On failure:
1. Capture the failure into `state.projects[i].buildStatus = { ok: false, attempts: N, lastError: "..." }`
2. The ENTIRE Phase 2 retries within that repo's worktree (max 3, same as single-repo)
3. After 3 attempts, escalate; do NOT proceed to Phase 3 with any unbuilt repo

**Recording a pass (default-FAIL evidence gate):** before setting `buildStatus.ok = true`, the build output must be tee'd to a log and that log must substantiate the success  -  a zero exit code alone is not trusted. Run the evidence gate; on exit 1, do NOT record a pass:
```bash
<build-command> 2>&1 | tee "{worktreePath}/.build.log" \
  | bash $HOME/.claude/scripts/offload-ref.sh --phase 3 --label build --root "$WORKTREE"
node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log"
```
Non-zero from the gate: the build pass is unverified, so treat it as a failure and keep `buildStatus.ok=false`.
This closes the gap where an agent records "built" without ever producing build output.

**Why the pipe (opt-in via `prefs.global.contextOffload.enabled`).** `tee` decides where the log is written, not how much of it the model reads. The filter parks the full text at `.multi-agent/refs/<node_id>.md` and prints a stub plus the tail, where a failing build's error already is; read that file before re-running a failed build. The evidence gate still reads the whole `.build.log`, so what counts as a verified pass is unchanged. Pref off = pass-through.

**Telemetry**: Per-repo metrics in addition to per-task metrics:
```bash
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=common duration_ms=$D status=ok
$HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 build.completed repo=uicomponents duration_ms=$D status=ok
```

**Token forwarding:** every TDD round (red, green, refactor) that hits the dev model MUST forward token totals into the tracker so Phase 5's Cost Breakdown captures Phase 2:

```bash
LOG_METRIC_FORWARD_TO_TRACKER=1 $HOME/.claude/scripts/log-metric.sh "$TASK_ID" 2 dev.tdd_round \
  model=<sonnet|opus> step=<red|green|refactor> \
  tokens_in=$IN tokens_out=$OUT duration_ms=$DUR
```

Model is Sonnet unless routing names another rung (`multi-agent-refs/features/model-fallback.md`). Best-effort. See `$HOME/.claude/multi-agent-refs/progress-contract.md#token-telemetry-forwarding`.

---

## Token telemetry  -  invoke after every LLM call

```bash
bash $HOME/.claude/scripts/phase-tracker.sh tokens 2 <input_count> <output_count>
```

Contract and rationale: `progress-contract.md` -> Token telemetry forwarding.

#### Generated trees are not yours to edit

Many repos generate part of their source: a service client from an OpenAPI spec, mock
scenario indexes, localization keys, testing identifiers, design tokens. A generated
file is regenerated on the next build, so an edit there is lost silently, and the
matching hand-authored tree is the one that takes the change.

Before writing into any path, check whether it is generated or owned; the repo
profile answers both:

```bash
node $HOME/.claude/scripts/repo-profile.mjs resolve generators --repo "$PROJECT_ROOT" --state "$STATE_FILE"
node $HOME/.claude/scripts/repo-profile.mjs resolve ownedPaths --repo "$PROJECT_ROOT" --state "$STATE_FILE"
```

A target under a `generators[].output` goes to that entry's `input` and is
regenerated with its `command` (in the same commit when `sameCommit`). A target
under an `ownedPaths[].glob` is not edited at all: report the `owner`, and the
`bypass` label when there is one; exit Gate 0 enforces this. Without a profile,
look for the signals directly, case-insensitively:

```bash
# a Generated/ segment, or a header saying so, is the signal
find . -type d -iname Generated -not -path './.*' | head
grep -rliE "do not edit|auto-?generated|generated by" --include="*.swift" --include="*.kt" . | head
```

The pairing is usually `Generated/<x>` for output and `Custom<X>/` or
`CustomSources/` for input. Two concrete shapes seen in the wild:

| Want to | Wrong place | Right place |
|---|---|---|
| add a mock fixture / named scenario | a `Fixtures/` file under a generated tree | the repo's custom fixture tree, plus registering the scenario in the generated index the build reads |
| add or change a service endpoint | the generated client method | the OpenAPI source the generator consumes, then regenerate |

When the analysis doc has not recorded which trees are generated, that is a Phase 1
gap  -  say so rather than guessing, since guessing wrong is invisible until the next
regeneration.

---

## Exit gate: Verify (BLOCKING)

These deterministic gates run once, at this phase's exit, so the build runs once:
Review inherits the logs instead of regenerating them. A failure here does not
reach Review at all.

#### Step 1  -  Deterministic Gates (run BEFORE AI review)

```bash
# Gate 0: Owned paths (repo profile), before the build
node $HOME/.claude/scripts/owned-path-gate.mjs --repo "$WORKTREE" --base "origin/$BASE_BRANCH" --state "$STATE_FILE"
# Gate 1: Build (xcodebuild/gradle assemble/tsc/py compile  -  stack-dependent; Xcode uses the build queue lock, see Phase 2)  -  tee output to a log
<build-command> 2>&1 | tee "{worktreePath}/.build.log"
# Gate 2: Lint (swiftlint/ktlint/ruff/eslint  -  stack-dependent)
# Gate 3: Tests pass (xcodebuild/gradle/pytest/the resolved node command)  -  tee output to a log
<test-command> 2>&1 | tee "{worktreePath}/.test.log"
# Gate 4: Secrets  -  run the scanner against the staged diff
bash $HOME/.claude/scripts/pre-commit-check.sh
# Gate 6: features/valven.md
```

**Gate 0** exit 1 blocks the phase: move each change where its `fix` points, revert the owned path, re-run. Exit 2 is a call error, never a pass: an unknown `--base` ref means fetch the base branch. Exit 0 with `skipped: true` gives its reason in `note`; log it. Contract: `features/repo-profile.md`, "Owned-path gate".

**Default-FAIL evidence gate (required before recording any pass):** a green exit code is not enough  -  the captured log must actually show success. Before marking build/test passed, run the evidence gate against the tee'd log; it fails CLOSED when the log is missing, empty, or shows failure markers:

```bash
node $HOME/.claude/scripts/evidence-gate.mjs --claim build --status passed --evidence "$WORKTREE/.build.log" || BUILD_PASS=false
node $HOME/.claude/scripts/evidence-gate.mjs --claim test  --status passed --evidence "$WORKTREE/.test.log"  || TEST_PASS=false
```

This prevents a false "it built" claim with no log behind it. On exit 1, treat the gate as failed (do NOT proceed to AI review) and surface the gate's `reason`.

**Autopilot runs** (`state.autopilot` or `MULTI_AGENT_UNATTENDED=1`; interactive runs skip this paragraph) also run the stack-aware checks, each recording its verdict in `state.gates[]`: `evidence-gate.mjs --stack`, `test-summary.mjs` (zero tests executed parks the run as `verification-failed` instead of reworking), `test-strength.mjs --head worktree` and `symbol-existence-gate.mjs`. With the variable set, the Phase 4 commit hook refuses a commit whose ledger lacks them. Commands and verdicts: `$HOME/.claude/multi-agent-refs/features/unattended-gates.md` section 4.

**Scaffolded repo** (`.scaffold.json` at the worktree root): at this phase's entry and exit, `scaffold-gate.mjs --dir "$WORKTREE" --phase story --story <id>`; exit 1 or 3 halts and the next story does not start; 2 is a call error, never a pass (`$HOME/.claude/multi-agent-refs/features/scaffold.md`).

**Inherited failures (when `state.baseline.tests` exists).** Phase 0 Step 7.6 recorded whether the suite was already red, so Gate 3 blocks on what this work broke, not what it walked into:

| baseline status | Gate 3 |
|---|---|
| `green` | unchanged; every failure is this run's |
| `red` + `failing[]` | subtract those ids. Nothing left -> pass, logged `test:pass (inherited {N})`. A NEW failure still blocks. |
| `red`, empty `failing[]` | do NOT pass and do NOT silently block: report `test:inherited-red (not attributable, <logPath>)` and ask. Inventing a set here masks regressions. |
| `unknown` / absent | unchanged from today |

The subtraction never widens: match on identifier only, and when identifiers cannot be compared fall to the `not attributable` row.

**Gate results:**

- All pass (including the evidence gate) -> proceed to AI review
- Any fail -> fix, re-run gates (no AI review until clean)
  Log: "Phase 3: Gates  -  owned:{ok/fail} build:{pass/fail} lint:{pass/fail} test:{pass/fail/inherited-red} secrets:{clean/found} evidence:{ok/unverified}"

##### Gate 5  -  Fortify SSC findings (runs when `state.contextLinks[]` contains a `fortify` entry, or when `prefs.global.fortify.alwaysCheck === true`)

If the task description referenced a Fortify version, or named a bare issue instance id, `~/.claude/lib/fetch-fortify.sh` already populated `state.fortifyFinding` in Phase 0 (`alwaysCheck` needs `prefs.global.fortify.versionIds` to know what to scan). Phase 3 reuses that payload and applies the deterministic gate:

```bash
jq -r '.fortifyFinding as $f | if ($f.gateOutcome // null) == null then "n/a (no Fortify URL referenced)" elif $f.gateOutcome.blocking == true then "BLOCKED (\($f.gateOutcome.reason), critical=\($f.severityCounts.Critical))" else "pass (high=\($f.severityCounts.High // 0) warnings carry into the channel summary)" end' "$STATE_FILE"
```

Print it as `→ fortify gate: <line>`. `BLOCKED` is a deterministic gate failure: fix the critical findings before AI review.

Gate semantics:

| `gateOutcome.reason` | Phase 3 action |
|---|---|
| `critical-findings` | **Blocking**  -  Phase 2 re-dispatches with the finding list; pipeline does not advance until Critical count is 0. |
| `high-findings-warning` | Pass-with-warning  -  High findings appear in the channels summary (`## Security Scan (Fortify)` section in PR + Jira); they don't block the merge. |
| `clean` | Pass silently  -  section omitted from the channels summary. |

When `state.fortifyFinding.status === "skipped"` (token missing, VPN unreachable, or host mismatch) the gate logs `→ fortify gate: skipped (<reason>)` and never blocks  -  the user already saw the structured Save Flow signal at Phase 0 and chose to proceed.
