# Multi-Agent Pipeline  -  Phase Reference

<!-- toc -->
- [Phase Files](#phase-files)
- [Pipeline Flow](#pipeline-flow)
- [Phase entry  -  pending steer (every phase, every mode)](#phase-entry---pending-steer-every-phase-every-mode)
- [Visual Phase Tracker](#visual-phase-tracker)
- [Preferences File](#preferences-file)
- [Host Configuration](#host-configuration)
- [Token Budget](#token-budget)
- [SubPhase Convention](#subphase-convention)
<!-- /toc -->

## Phase Files

| Phase                             | File                                                                 |
| --------------------------------- | -------------------------------------------------------------------- |
| Modes (autopilot, analysis) | `$HOME/.claude/multi-agent-refs/phases/modes.md`            |
| Operations (kill, purge, resume)  | `$HOME/.claude/multi-agent-refs/phases/operations.md`       |
| Phase 0: Init                     | `$HOME/.claude/multi-agent-refs/phases/phase-0-init.md`     |
| Phase 1: Plan                     | `$HOME/.claude/multi-agent-refs/phases/phase-1-plan.md`     |
| Phase 2: Dev                      | `$HOME/.claude/multi-agent-refs/phases/phase-2-dev.md`      |
| Phase 3: Review                   | `$HOME/.claude/multi-agent-refs/phases/phase-3-review.md`   |
| Phase 4: Commit & PR              | `$HOME/.claude/multi-agent-refs/phases/phase-4-commit.md`   |
| Phase 5: Report                   | `$HOME/.claude/multi-agent-refs/phases/phase-5-report.md`   |
| Log format                        | `$HOME/.claude/multi-agent-refs/phases/log-format.md`       |

## Pipeline Flow

```
           0-Init -> 1-Plan -> 2-Dev -> 3-Review -> 4-Commit -> 5-Report
Local:     The same set with no worktree  -  Phase 0 Step 5b answered local, so
           work happens directly on a branch in the project root

One pipeline: every mode runs its whole phase set, and no answer during the run adds or removes a phase.
```

## Phase entry  -  pending steer (every phase, every mode)

Before a phase does any work, read `pendingSteer` from `agent-state.json`. It
is how a correction reaches a run that is already going: `/multi-agent:steer #N
"<instruction>"` queues one, and the next phase to start consumes it.

```bash
STEER=$(node -e 'try{const s=require(process.argv[1]);const p=s.pendingSteer;if(p&&p.text&&!p.appliedAt)process.stdout.write(p.text)}catch{}' "$STATE_FILE")
```

The `catch` is not decoration: a state file that does not exist yet (Phase 0
before its first write) makes `require` throw, and an unguarded read prints a
stack trace at every phase entry that reads like a failed phase.

Empty → carry on, nothing to do. Non-empty:

1. Apply it to this phase's context before starting. It is the user's own
   words; treat it as an instruction from them, arriving now.
2. Say what changed, in one line, so the correction is visible in the run and
   not only in the state file.
3. Mark it consumed. The record stays as a trace; `appliedAt` is what stops it
   being read again, and the next steer overwrites all four keys:
   ```bash
   printf '{"pendingSteer":{"appliedAt":"%s","appliedPhase":<N>}}' "$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
     | node $HOME/.claude/scripts/write-state.mjs "$STATE_FILE"
   ```

   `<N>` is this phase's number, substituted literally and unquoted - it is a
   JSON number. Leaving it as an unset shell variable produces
   `"appliedPhase":}`, which is invalid JSON and exits 1.
4. Record it on the tracker and in the log: `bash $HOME/.claude/scripts/phase-tracker.sh meta <N> steer "applied"`, plus a `Steer applied: <text>` line in `agent-log.md`.

Two limits, both deliberate. It is read at phase **entry** only, never mid-phase:
a phase that changes target halfway discards the work it already did, which is
the outcome steer exists to avoid. And an instruction that contradicts an
approved plan halts for the user rather than silently rewriting the plan -
steering corrects a run, it does not re-authorize one.

## Visual Phase Tracker

Two channels run in parallel at every phase boundary. Both are required in their target CLIs  -  skipping either is the #1 source of "I see no progress" complaints.

| Mechanism                        | Available in     | Status        | Style                                                |
| -------------------------------- | ---------------- | ------------- | ---------------------------------------------------- |
| **`phase-tracker.sh`** (state)   | Every CLI        | required     | Drives `tracker-state.json`; powers `:resume`/`:log`/`:status` |
| **`TaskCreate` / `TaskUpdate`**  | Claude Code only | required here | Native sticky TaskList widget  -  the only progress signal Claude Code surfaces |
| **`phase-tracker.sh render`**    | Every other CLI  | required here | Bordered ANSI card printed as last tool result so the user sees the phase table |
| **`phase-banner.sh`**            | Every CLI        | Optional      | One-shot ANSI banner per phase boundary (extra emphasis) |

**Cross-CLI contract**: every phase boundary MUST update both the state channel (`phase-tracker.sh`) AND the visual channel (TaskList in Claude Code, render in every other CLI). The tracker is the cross-CLI source of truth; the visual is what the user actually sees.

### Tracker bootstrap (Phase 0, mandatory)

Phase 0 MUST initialize the tracker and register the active mode's phase set:

```bash
$HOME/.claude/scripts/phase-tracker.sh init "$TASK_ID"
for p in 0:Init 1:Plan 2:Dev 3:Review 4:Commit 5:Report; do
  $HOME/.claude/scripts/phase-tracker.sh add "${p%%:*}" "${p#*:}"
done
```

This produces an initial card stack printed by both CLIs.

No mode is an exception: the phase set is a property of the command, so every
tile is created in this one batch. Full contract: `tracker-contract.md`.

### Tracker updates (every phase boundary)

As each phase enters/exits:

```bash
$HOME/.claude/scripts/phase-tracker.sh update <N> in_progress    # phase starts
# ...do the work...
$HOME/.claude/scripts/phase-tracker.sh update <N> completed      # phase ends OK
# or:
$HOME/.claude/scripts/phase-tracker.sh update <N> failed         # phase failed
$HOME/.claude/scripts/phase-tracker.sh update <N> skipped        # e.g. 2 in analysis mode
```

After every LLM call (counts are additive; skipping this is why runs end with durations but no cost  -  nothing reconstructs spend afterwards):

```bash
$HOME/.claude/scripts/phase-tracker.sh tokens <N> <in> <out> [cached]
```

For sub-phase progress (e.g. Phase 1's parallel Explore agents, Phase 3's reviewer dispatch + triage + validator gate, Phase 4's commit + push + PR sub-steps):

```bash
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" pending      # register
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" in_progress  # advance
$HOME/.claude/scripts/phase-tracker.sh sub <N> 1 "<sub name>" completed    # done
```

State persists at `$HOME/.claude/logs/multi-agent/<task_id>/tracker-state.json`  -  atomic writes mean concurrent worktrees don't corrupt each other.

### Optional banner (single-event flair)

Call alongside tracker update for extra emphasis:

```bash
$HOME/.claude/scripts/phase-banner.sh start <N> "<name>" "<one-line detail>"
$HOME/.claude/scripts/phase-banner.sh end   <N> done   "<name>" "<short result>"
$HOME/.claude/scripts/phase-banner.sh sub   <N> 2     "<sub>" "<detail>"
```

Banner status enum on `end`: `done` | `failed` | `skipped`. Anything else exits 64.

### TaskCreate registration (Claude Code  -  required)

In Claude Code the agent MUST register one TaskCreate tile per phase at Phase 0 startup (one per phase the current COMMAND runs  -  `/multi-agent` and the autopilot entries = 0..5, `:analysis` = 0/1/3/4/5). Capture the returned `taskId`, persist it via `phase-tracker.sh meta <N> tasklist_id "<taskId>"` so `:resume` can rebuild the widget.

Per phase boundary:

```text
# Phase entry  -  flip the tile to in_progress alongside the state update
TaskUpdate({ taskId: <saved>, status: "in_progress" })
bash phase-tracker.sh update <N> in_progress

# Active sub-step inside a phase  -  keeps the spinner header live
TaskUpdate({ taskId: <saved>, activeForm: "Editing TopBarView.swift" })
TaskUpdate({ taskId: <saved>, activeForm: "Running xcodebuild test" })

# Phase exit  -  flip both channels
TaskUpdate({ taskId: <saved>, status: "completed" })
bash phase-tracker.sh update <N> completed
```

A phase outside the command's set gets no TaskCreate at all, and the set is known before the tracker boots.

**(strict) TaskCreate ordering**: All TaskCreate calls MUST fire in strict phase-number order BEFORE any TaskUpdate is applied. The native widget renders by creation order, not by phase number  -  out-of-order calls produce visually scrambled tile stacks (e.g. `1 ✓ · 3 ✓ · 4 ✓ · 0 ▶ · 2 ☐`) even when the underlying state is correct. Pre-marking phases as completed/skipped before Phase 0 starts is FORBIDDEN  -  register the tile in order with default `pending` status, then flip status via TaskUpdate when the phase actually short-circuits. Full contract in `$HOME/.claude/multi-agent-refs/tracker-contract.md` section "TaskCreate ordering (strict)".

**Copilot CLI / plain shell**: do NOT call TaskCreate  -  the tool does not exist on these CLIs. Instead, after every state update, call `bash phase-tracker.sh render` so the bordered ANSI card prints as the last tool result. That's the equivalent visual signal there.

## Preferences File

Path: `$HOME/.claude/multi-agent-preferences.json`  -  persistent state across sessions.

## Host Configuration

Pipeline uses placeholder hosts for corporate services. Set these before first use:

| Placeholder         | Purpose                    | Example                  |
| ------------------- | -------------------------- | ------------------------ |
| `{JIRA_HOST}`       | Jira server hostname       | `jira.company.com`       |
| `{BITBUCKET_HOST}`  | Bitbucket server hostname  | `bitbucket.company.com`  |
| `{CONFLUENCE_HOST}` | Confluence server hostname | `confluence.company.com` |
| `{CORP_DOMAIN}`     | Corporate domain           | `company.com`            |

Configure via preferences or environment variables. If not set, VPN check (Phase 0) skips gracefully.

Read modes.md first (always), then the phase file for the current stage.
For operations (status, kill, resume, purge): read operations.md.

## Token Budget

Each phase doc is lazy-loaded  -  only the current phase's spec is in context. Budget enforced by `smoke-token-budget.sh`.

The live ceilings are in the repo's `schemas/token-budget.json`, derived from
its `schemas/phases.json`. **They are deliberately not repeated here.**

A table of the same numbers stood in this spot for most of the project's life and
said the total was 17,000 tokens. The file the gate actually reads said 63,150.
Nothing compared the two, so the copy drifted by a factor of four and stayed
wrong through every release. A second copy of an enforced number is not
documentation; it is a claim with no gate behind it.

Token estimate: `ceil(chars / 4)`. The budget covers `phase-*.md` files only.
Guides, rules, and agents are loaded separately on demand. The repo's
token-budget gate prints each doc's measurement next to its ceiling, which a
table here could never do  -  that is where to look for a current number.

**Prompt caching (token learning-curve):** assemble each phase prompt as a stable, cacheable prefix (phase instructions + repo learnings brief + conventions + repo evidence) followed by the volatile task suffix, so repeated runs on a known repo pay cache-read price on the prefix. Full contract: `$HOME/.claude/multi-agent-refs/prompt-assembly.md`. The effect (rising cache ratio, falling tokens/task) is what `learning-curve.mjs` trends.

## SubPhase Convention

When a specialized skill takes over a main pipeline phase, progress is reported as **SubPhases** nested under the parent. Top-level stays 6 phases (0-5); specialized flows get sub-phase detail.

**Visual example:**

```
■ Phase 0: Init              completed
■ Phase 1: Plan              completed
* Phase 2: Dev (figma)       in progress
  ■ SubPhase 2.0: Init            completed
  ■ SubPhase 2.1: Gather          completed
  ■ SubPhase 2.4A: Configuration  completed
  ■ SubPhase 2.4B: View           completed
  * SubPhase 2.4F: Wiki           writing wiki pages...
  □ SubPhase 2.5A: ViewInspector  pending
  □ SubPhase 2.6: Code Connect    pending
  □ SubPhase 2.7: Issue Update    pending
□ Phase 3: Review            pending
□ Phase 4: Commit            pending
□ Phase 5: Report            pending
  □ SubPhase 5.1: Jira comment      pending
  □ SubPhase 5.2: Wiki + screenshots pending
  □ SubPhase 5.3: Confluence        pending
  □ SubPhase 5.4: Report + log      pending
  □ SubPhase 5.5: Knowledge capture pending
```

**TaskCreate pattern:** Parent phase + SubPhases linked via `addBlockedBy`. SubPhases block parent from completing.

**Why SubPhases, not separate phases:**

- Main pipeline stays a fixed 6-phase contract (0-5) regardless of task type. Wiki, Confluence, and Figma screenshots all live as SubPhases under Phase 5 Report  -  external delivery grouped logically, internal report + knowledge after.
- Conditional phases are a code smell  -  they force every reader to learn which phase runs when.
