# Cross-CLI Contract (Claude Code · Copilot CLI · Codex CLI)

<!-- toc -->
- [1. Command Inventory (64 commands)](#1-command-inventory-64-commands)
- [2. Canonical Placeholder Vocabulary](#2-canonical-placeholder-vocabulary)
- [2.6 Intentional structural divergence  -  thin dispatcher vs inlined orchestrator](#26-intentional-structural-divergence---thin-dispatcher-vs-inlined-orchestrator)
- [2.7 One command, three different meanings: `/multi-agent:model`](#27-one-command-three-different-meanings-multi-agentmodel)
- [3. Frontmatter Transform Rules (Claude ↔ Copilot)](#3-frontmatter-transform-rules-claude-copilot)
- [4. Progress Signalling Parity](#4-progress-signalling-parity)
- [5. Argument Parsing Invariants](#5-argument-parsing-invariants)
- [6. Output Format Expectations](#6-output-format-expectations)
- [7. Platform Guards (macOS)](#7-platform-guards-macos)
- [8. Enforcement](#8-enforcement)
- [9. Change Control](#9-change-control)
<!-- /toc -->

> **Non-negotiable**. Any change that breaks this contract blocks merge. Validated by `smoke-cross-cli-behavior.sh`.

**Purpose**: every pipeline command must produce identical artifacts (state, logs, outputs) and respect identical placeholder vocabulary regardless of which of the three host CLIs invokes it. This file is the source of truth for "what must stay the same."

---

## 1. Command Inventory (64 commands)

```
analysis, analysis-jira, analysis-resolve, autopilot, autopilot-off,
autopilot-on, autopilot-status, bug-bash, build-optimize, channels, complaint-analysis,
create-jira, design-check, diff-explain, doctor, estimate, feedback, forget,
garbage-collect, graph, help, ios-coding-standard, issue, jira, kill,
language, log, manual-test, model, prune-logs,
prune-prompts, purge, refactor, research, resume, review,
review-analysis, review-issue, review-jira, route-off, route-on,
route-status, routines, save, scaffold, scan, scenario-audit, search,
security-review, serve, setup, stack, status, steer, store-ready, sync, test, test-accessibility,
test-dark-mode, test-dynamic-type, test-screenshots, testflight-validation,
uninstall, update
```

Categories:

- **Interactive pickers** (single-purpose, not modes): `jira`, `issue`
- **Issue generator** (one-shot, no worktree, asks type Task/Bug/Story, hard approval gate before create): `create-jira`
- **Pipeline entries**: `autopilot` (plus the bare `/multi-agent` in the dispatcher). The phase set is a property of the command, and every mode runs its whole set.
- **Tail path** (the pipeline tail over work already on a branch, no run behind it): `resume`
- **Ops commands** (one-shot, no worktree): `status`, `log`, `kill`, `steer`, `purge`, `uninstall`, `resume`, `review`, `review-jira`, `review-issue`, `analysis`, `analysis-resolve`, `research`, `scaffold`, `complaint-analysis`, `bug-bash`, `build-optimize`, `channels`, `scan`, `search`, `diff-explain`, `garbage-collect`, `graph`, `prune-logs`, `prune-prompts`, `serve`
- **Local audits** (worktree only to build; no commit, push, PR or channels): `design-check`, `testflight-validation`, `ios-coding-standard`. `testflight-validation` additionally never invokes `altool --upload-app`  -  a validation run must not be able to ship a build by accident.
- **Meta-ops**: `setup`, `sync`, `update`, `help`, `refactor`, `test`, `stack`, `manual-test`, `language`
- **Routines** (user-defined routine registry; the routines they create are local-only and never synced): `save`, `routines`, `forget`

> **Inventory drift is a contract violation.** Adding a slash command under `pipeline/commands/multi-agent/` without updating this list + its counterpart Copilot dir (`pipeline/skills/shared/core/multi-agent-<cmd>/`) is a merge blocker. `smoke-commands-skills-parity.sh` enforces command ↔ skill directory parity; `smoke-cross-cli-behavior.sh` enforces behavior parity. This doc is the authoritative command list  -  bump the count + table together.

### 1.1 Figma / component work (plugin-based on Claude Code; NOT parity-enforced)

Component work is **no longer a bundled pipeline skill set**. Claude Code dispatches `taskType === "component"` to the enabled `ai-<platform>-toolkit` **marketplace plugin** (`create-component`, fallback `create-ui-component`) via the Skill tool  -  full contract in `$HOME/.claude/multi-agent-refs/component-dispatch.md`. The pipeline no longer ships `pipeline/skills/figma-ios|figma-android|figma-common`; the pipeline-unique component skills (iterate loops, performance harness, validate/review, commit, adapters, wiki) were absorbed into the `ai-ios-toolkit` plugin so component skills live in one place.

**How each host receives the plugin's skills.** Only Claude Code loads the marketplace
plugin natively. The other two are served by the installer, so all three end up with the
same active skill set:

| Host | Delivery | Discovery |
|---|---|---|
| Claude Code | the marketplace plugin, loaded natively | matched on description |
| Copilot CLI | `install/copilot.mjs` copies the **enabled** stack plugin's authored skills (`index`, `reference/`, `workflow/`, `tools/`) into `~/.copilot/skills/` | matched on description |
| Codex CLI | `install/codex.mjs` copies the same set **plus** `skills/shared/external` into `~/.codex/multi-agent-refs/skills/` as reference files | located by path; the AGENTS.md block carries the `ls` / `grep` recipe |

Three rules make that work:

1. **`knowledge/` is never copied to Copilot.** A stack plugin's `knowledge/` tree is
   generated from this pipeline's `skills/shared/external`, which the installer already
   lays down there. Copying it again would duplicate ~5 MB the host has.
2. **Only the ENABLED plugins are delivered.** `fix-bug` and `branch-and-pr` exist in
   four stack plugins, `component`, `state` and `create-component` in three, and the
   destination is flat  -  delivering every stack would be last-write-wins, so an iOS
   repo could end up running the Android `create-component`. The enabled list in
   `~/.claude/settings.json` is the source of truth, exactly as on Claude Code;
   `--platform` is the fallback when no such list exists.
3. **Codex takes refs, not list entries.** Measured on 0.145: one plugin declaring 142
   skills surfaced 75 and evicted an unrelated user skill, and every plugin-provided
   skill renders with an empty description, so a block entry buys nothing even when it
   fits. Parity on that host is **reachability**, not an identical listing.

What still MUST match across CLIs for component tasks: the `taskType === "component"` classification and the `state.phases["2"].subphases[]` shape the dispatch layer writes. Resolution/routing is Claude-plugin vs Copilot-local-copy by design.

### 1.2 Store-compliance skills

Two parallel shared/core skills under `pipeline/skills/shared/core/` wrap external tooling for pre-submission store compliance:

| Skill | Path | Tool wrapper | Rules |
|---|---|---|---|
| `apple-archive-compliance` | `pipeline/skills/shared/core/apple-archive-compliance/` | `ios_app_store_audit` MCP tool (in `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.0.0) | 18 (Apple ITMS + App Store Review Guidelines) |
| `google-play-compliance` | `pipeline/skills/shared/core/google-play-compliance/` | bundletool + aapt2 + apksigner | 21 (4 categories: Technical / Security / Privacy / Hygiene) |

Both skills are wired to 4 consumers: `/multi-agent:test "store-ready"` (primary), Phase 3 Security Auditor (`pipeline/agents/security-auditor.md`), `/multi-agent:review` + SKILL.md counterpart, `/multi-agent:channels` PR-body auto-augmentation. Contract enforced by `smoke-compliance-skills.sh`.

### 1.3 Figma-skill routing from multi-agent (superseded)

Superseded by section 1.1: Claude Code resolves component dispatch to the marketplace plugin (dual-name `create-component`/`create-ui-component`; full contract in `$HOME/.claude/multi-agent-refs/component-dispatch.md`); Copilot CLI receives the enabled plugin's authored skills through `install/copilot.mjs` (section 1.1), not a standalone `figma-*` tree  -  the installer prunes those. No shared filesystem-path routing table or figma-skill frontmatter/inventory parity remains under enforcement.

### 1.4 End-to-end testing tools

`/multi-agent:test` (`commands/sim-test.md`), `/multi-agent:bug-bash` and the
`e2e-testing` knowledge skill (ai-common-toolkit) call these toolkit tools when
the server has them, and fall back to UI-tree reads and frame-centre taps when
it does not. Contracts: `features/exploratory-testing.md`,
`features/assertion-judge.md`.

| Capability | Tools / fields | Minimum |
|---|---|---|
| Element locators | `ui_locate`, `ios_tap_element`, `android_tap_element`; codes `LOCATOR_NOT_FOUND`, `LOCATOR_AMBIGUOUS` | `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.19.0 |
| Assertion steps and verdicts | `agent_run_steps` steps `expect_text`, `expect_element`, `expect_url`; `verdict: passed / failed / blocked` with `ASSERTION_FAILED`, `CREDENTIAL_MISSING`, `ENVIRONMENT_UNAVAILABLE`, `SEED_DATA_MISSING`, `BUDGET_EXCEEDED`, `REPLAY_DIVERGED`; failure `trace.md` | ≥ v3.19.0 |
| Secrets by reference | `secret_ref` on the type tools | ≥ v3.19.0 |
| Parallel web sessions, replay | `session` on the web tools; `save_as` / `replay` | ≥ v3.19.0 |
| Travel document fixtures | `document_generate`, `document_mrz_build` | ≥ v3.19.0 |

---

## 2. Canonical Placeholder Vocabulary

**Problem we prevent**: same concept referenced with different placeholder names in Claude vs Copilot files (e.g., `{github-username}` vs `{owner}`). Every new command must pull from this table.

### 2.1 Identity / account placeholders

| Placeholder | Meaning | Example real value |
|---|---|---|
| `{owner}` | Public GitHub org/user owning the pipeline repo | `mmerterden` |
| `{npm-scope}` | npm scope for publishing | `@mmerterden` |
| `${USER}` | Local OS username, used for per-user keychain keys | `alice` |
| `{author-name}` | Git author name (comes from preferences.identities[].name) | `Full Name` |
| `{author-email}` | Git author email | `foo@example.com` |

### 2.2 Host / URL placeholders

| Placeholder | Meaning | Example |
|---|---|---|
| `{jira-host}` | Jira server hostname | `jira.company.com` |
| `{bitbucket-host}` | Bitbucket server hostname | `bitbucket.company.com` |
| `{confluence-host}` | Confluence server hostname | `confluence.company.com` |
| `{corp-domain}` | Corporate email domain | `company.com` |
| `{website-host}` | Public website hostname | `company.dev` |

### 2.3 Project name placeholders

Use these specific names (from the Copilot set)  -  **never invent new ones**. They correspond to realistic example project types.

| Placeholder | Represents |
|---|---|
| `my-ios-app` | An iOS native project |
| `my-figma-app` | A figma-driven SwiftUI component project |
| `my-ui-components` | A UI component library (cross-platform possible) |

### 2.4 Task / tracker placeholders

| Placeholder | Meaning |
|---|---|
| `{JIRA_KEY}` | Jira project key (e.g., `PROJ`, `MOBILE`) |
| `PROJ-12345` | Example Jira issue (placeholder, always use `PROJ-12345` for examples) |
| `#N` | GitHub issue number placeholder |

### 2.5 Deprecated placeholders (DO NOT USE)

Legacy names found during the v3.7 audit  -  replaced per this contract.

| Deprecated | Replace with | Reason |
|---|---|---|
| `{github-username}` | `{owner}` | Shorter, consistent with `{owner}` pattern |
| `{website-repo}` | `website` (literal, since the repo name IS "website" under `{owner}`) | Hardcoding the repo name inside `{owner}` scope is fine |
| `{your-website}` | `{website-host}` | Match `{*-host}` convention |
| `mmerterden` (hardcoded in examples) | `{owner}` | Personal handle, public but must be a placeholder in generic docs |
| `my-project`, `my-figma-project`, `my-mobile-app` | `my-ui-components`, `my-figma-app`, `my-ios-app` | Canonical project names per 2.3 |

---

## 2.6 Intentional structural divergence  -  thin dispatcher vs inlined orchestrator

The top-level orchestrator has two deliberately different shapes. This is **not
drift**  -  it reflects a real capability gap between the host CLIs, and auditors
must not flag it as a parity violation.

| Surface | File | Shape | Why |
|---|---|---|---|
| Claude Code (colon-form) | `pipeline/commands/multi-agent/SKILL.md` | ~300 lines  -  thin dispatcher that routes to `$HOME/.claude/multi-agent-refs/phases/phase-N-*.md` on demand | Claude Code lazy-loads reference files, so only the active phase's docs enter the context window. Saves tokens. |
| Copilot CLI (dash-form) | `pipeline/skills/shared/core/multi-agent/SKILL.md` | ~830 lines  -  full inline orchestrator with all 8 phase specs embedded | Copilot CLI loads the whole SKILL.md once the skill is dispatched; it has no equivalent of Claude's ref-file loading. Embedding keeps behavior identical without relying on a feature Copilot lacks. |
| Codex CLI | installed as `~/.codex/skills/multi-agent/SKILL.md`, generated from the Claude dispatcher | thin dispatcher, refs under `~/.codex/multi-agent-refs/` | Codex reads files on demand, so it takes the Claude shape  -  **and it has to.** See below. |

### Why Codex takes the thin-dispatcher shape, and must keep it

Codex assembles every discovered skill's name + description into a single prompt
block and **silently drops entries when that block overflows**. Measured against
Codex 0.145 during the install design: installing one plugin that declares 142
skills took the block from 11 skills / 4,710 bytes to 83 skills / 22,111 bytes  -
only **75 of the 142** surfaced, **and an unrelated user-scope skill was evicted**.
Removing the plugin brought it back.

So shipping every sub-command as a peer skill on Codex would silently lose
pipeline commands next to any stack toolkit, with no error anywhere. The pipeline
therefore contributes **exactly one** skill on Codex (`multi-agent`) and keeps every
sub-command spec as a reference file that costs nothing until read.

**Do not "fix" this by adding per-command skills on Codex.** The layout is
capability-derived, and `smoke-install-layout.sh` fails if the Codex skills tree
gains a second pipeline entry.

### Parity axis differs per host

Claude Code and Copilot CLI are compared on their **skill directory sets**. Codex is
compared on its **ref set**: every command spec must exist under
`~/.codex/multi-agent-refs/commands/<cmd>/SKILL.md`, and
`smoke-codex-install.sh` asserts the count against the source tree. Comparing Codex
on skill directories would demand exactly the layout that breaks it.

**What must stay identical** (byte-level) across the two files:

- Phase canonical labels (Phase 0-5)
- Input parsing table (Section 5 of this doc)
- Routing decisions for each input type
- Placeholder vocabulary (Section 2 of this doc)
- Mode flag semantics (Section 5 of this doc)
- Output artifact schemas (Section 6 of this doc)

**What's allowed to differ**:

- File length and section grouping
- Whether phase docs are inlined vs referenced
- Language of operational commentary (TR helpers in dash form are OK; key identifiers / commit messages / PR bodies stay English everywhere)

Future changes that break an item in the "stay identical" list must update **both** files in the same commit. `smoke-cross-cli-behavior.sh` enforces the identity-preserving axis (input parsing, routing, output shape); structural differences are left to manual review because enforcing them would require forcing the files to the same shape, which we intentionally don't want.

### Panel diversity per host

Phase 3 runs three reviewers everywhere, but the diversity those three buy is not the
same on every host. Copilot CLI gets cross-VENDOR disagreement for free: GPT-5.4 sits
beside two Claude models. Claude Code and Codex each run a one-vendor panel - three
Anthropic models on one, three OpenAI models on the other - so the same three-way
agreement is weaker evidence there, and Phase 3 says so in the triage note on a
borderline finding.

Where the budget goes instead, when vendor diversity is unavailable:

| Host | Reviewer 1 | Reviewer 2 | Reviewer 3 |
|---|---|---|---|
| Copilot CLI | Fable/Opus, security + architecture | GPT-5.4, edge cases (cross-vendor) | Sonnet, quality |
| Claude Code | Fable, security + architecture | Opus, edge cases | Sonnet, quality |
| Codex | `xhigh`, security + architecture | a different family member, edge cases | `medium`, quality |

On Codex the axis is reasoning effort as much as model identity, because the family
members available there are closer to each other than Fable and Sonnet are. That is a
weaker axis, not an equivalent one, and treating it as equivalent is the error this
section exists to prevent.

## 2.7 One command, three different meanings: `/multi-agent:model`

Most commands do the same thing on all three hosts. This one does not, and the
divergence is stated here rather than discovered at runtime, because a command
that quietly no-ops is worse than one that says it cannot act.

| Host | What `model on|off` does | Why |
|---|---|---|
| Claude Code | live switch: flips `modelFallback.fableEnabled` and realigns `costBudget.pricingModel` | the only host where the `fable` rung is Fable 5 |
| Copilot CLI | writes the preference, changes no dispatch | Fable 5 is not offered there; its personas never sat on this rung |
| Codex CLI | writes the preference, changes no dispatch, **deliberately** | the `fable` rung on Codex means `gpt-5.6 @ xhigh` - a different model on a different account. A knob named after an Anthropic model must not silently retune a Codex run |

`model-rung.sh` prints which of the three it is on every invocation. The rule it
follows: report the host's actual behaviour, never claim a change that did not
happen.

The routing trio (`route-on`, `route-off`, `route-status`) has no such split -
it writes `prefs.global.modelRouting` identically everywhere. Its honest limit is
different and applies on every host: a subagent cannot be dispatched to a
non-Anthropic model, because subagent dispatch belongs to the host. External
providers are reachable only where the pipeline makes the HTTP call itself
(`bulk-read.sh`, `research_ask`). `route-status` prints that on every run.

## 3. Frontmatter Transform Rules (Claude ↔ Copilot)

Each file has a different frontmatter schema. The sync flow transforms between them:

### 3.1 Claude Code  -  `commands/multi-agent/{cmd}/SKILL.md`

```yaml
---
description: "<one-line purpose>"
allowed-tools: <comma-separated tool list>
---
```

### 3.2 Copilot CLI  -  `skills/multi-agent-{cmd}/SKILL.md`

```yaml
---
name: multi-agent-{cmd}
description: "<one-line purpose>"
user-invocable: true
argument-hint: "<input hint>"
---
```

### 3.3 Transform Contract

| Direction | Action |
|---|---|
| Claude → Copilot | Add `name`, `user-invocable: true`, `argument-hint`; remove `allowed-tools`; `description` stays English (localization happens only in the installed Claude tree via `description-tr` + `localize-commands.mjs`, never in repo or Copilot files) |
| Copilot → Claude | Remove `name`, `user-invocable`, `argument-hint`; add `allowed-tools` (inferred from command's actual tool calls) |
| Both directions | `not-for: <sibling>` carries across with the same value: top-level in the command, under `metadata:` in the skill, where every key no host reads lives (`language` too) |

**`not-for`** is optional and lists sibling commands this one must not be chosen for, as bare names (`not-for: jira, review-jira`). It is written for the model reading the routing surface, so it belongs beside the description rather than in prose further down. `lint-skills.mjs` also consumes it: a trigger-vocabulary collision either side has named drops off the warning list and is reported as settled instead, and a name that resolves to no sibling on the same surface is an error - a typo would otherwise leave the pair unanswered while the author believes it is handled. Both trees spell the value the same way and the lint reads it from either place; the `multi-agent-` prefix on the Copilot side is stripped before matching.

**Command-only keys stay top-level.** `parameters`, `gui`, `destructive`, `confirm` and `description-tr` sit at the top level of a command's frontmatter, not under `metadata:`. Claude Code ignores a key it does not recognize, command files are never packaged for claude.ai or the Skills API (where unknown keys are rejected), and these keys are read by the command contract, the GUI catalog and the installer's localizer, which all parse the top level. A shared skill keeps every key no host reads under `metadata:`, because shared skills are the files that do get packaged.

**`disable-model-invocation: true`** marks every command whose `destructive` is true: such a command runs only when the user types it. It is set on the Claude command only; the Copilot mirror does not carry it. Commands that only confirm (`confirm: required` without `destructive`), such as `sync`, stay model-invocable because they are routinely run on request and confirm inside their own flow. `test/command-contract.test.mjs` holds the rule.

**Invariant**: the English `description` must have the same meaning on both sides, and a `description-tr` line (when present) must be a faithful translation of it. Semantic drift is not allowed. `description-en` sidecars are an installed-tree artifact and must never appear in repo or Copilot files.

---

## 4. Progress Signalling Parity

### 4.1 TaskCreate ↔ phase-tracker.sh

| Concept | Claude Code | Copilot CLI | Codex CLI |
|---|---|---|---|
| Print the registration calls | `phase-tracker.sh tiles` (emits the TaskCreate list) | `phase-tracker.sh tiles` (emits the reprint instruction) | `phase-tracker.sh tiles` (emits the update_plan payload) |
| Register a phase | `TaskCreate` tool call with subject/description | `phase-tracker.sh add <N> <name>` | `update_plan` step, `status: pending` |
| Mark a phase in-progress | `TaskUpdate` → `in_progress` | `phase-tracker.sh update <N> in_progress` | `update_plan` step → `in_progress` |
| Mark a phase complete | `TaskUpdate` → `completed` | `phase-tracker.sh update <N> completed` | `update_plan` step → `completed` |
| Mark a phase failed | `TaskUpdate` → `completed` + log failure in agent-log | `phase-tracker.sh update <N> failed` | `update_plan` step → `completed` + failure in agent-log |
| Sub-phase progress | TaskCreate with `addBlockedBy` | `phase-tracker.sh sub <N> <sub> <name> <status>` | `phase-tracker.sh sub ...` (the plan tool has no nesting) |

`update_plan` takes the FULL step list, not a delta, so each boundary rewrites the
whole plan. It must never be called in parallel with another tool, and it is
unavailable in Codex plan mode  -  fall back to `phase-tracker.sh render` there.

### 4.2 Contract

Every phase boundary MUST call EITHER TaskCreate/Update (Claude) OR `phase-tracker.sh` (Copilot). `phase-tracker.sh` is the cross-CLI source of truth (writes to `~/.claude/logs/multi-agent/<task_id>/tracker-state.json`). Claude Code users get BOTH: TaskCreate tiles AND tracker state.

### 4.3 Banner Parity

`phase-banner.sh` runs on both CLIs. Same output format, no Claude-only or Copilot-only flair.

### 4.4 Continuous mode

`autopilot-on`, `autopilot-off` and `autopilot-status` ship to all three hosts and
behave identically there, because the thing they control is not a CLI feature:
it is a launchd user agent whose tick spawns a `claude --bg` child regardless of
which CLI you typed the command in. Continuous mode therefore requires the
`claude` binary on `PATH` on every host, Copilot and Codex included, and there is
no Copilot or Codex equivalent of the child.

| Concept | Claude Code | Copilot CLI | Codex CLI |
|---|---|---|---|
| Turn the mode on | `/multi-agent:autopilot-on` | `/multi-agent-autopilot-on` | `multi-agent autopilot-on` via the router skill |
| Scripts, libs, templates | `~/.claude/{scripts,lib,templates}` | `~/.copilot/...` | `~/.codex/...` |
| Which one is read | `ma_ap_asset <rel>` resolves the caller's own tree first, then the other two | same | same |
| Queue and config state | `~/.claude/autopilot/` | `~/.claude/autopilot/` | `~/.claude/autopilot/` |
| In-flight rows on the widget | `autopilot-status --subjects` into `TaskUpdate` | into the reprinted card | into `update_plan` |

The state row is the one that looks wrong and is not. `~/.claude/autopilot/` is
shared across hosts on purpose, exactly like `logs/`, `knowledge/` and
`multi-agent-preferences.json` (2.6, "shared state is deliberately NOT
retargeted"): two CLIs on one machine must read ONE queue. A per-host state root
would give a Copilot session a second, invisible queue and the same ticket would
be taken twice.

Everything that is NOT state resolves per host, and that half had to be fixed:
`templates/` was laid down only by `install/claude.mjs`, so `autopilot-on` on a
Codex-only machine rendered a launchd job from a file the host did not have.
`smoke-autopilot-hosts.sh` holds the line.

---

## 5. Argument Parsing Invariants

Given the same user input, both CLIs must produce identical parsed state:

| Input pattern | Parsed as |
|---|---|
| `PROJ-12345` (bare Jira ID) | `input.type = "jira-id"`, `input.jiraId = "PROJ-12345"` |
| `https://{jira-host}/browse/PROJ-12345` | `input.type = "jira-url"`, `input.jiraId = "PROJ-12345"` |
| `https://github.com/{owner}/{repo}/issues/42` | `input.type = "github-issue-url"`, `input.issueNo = 42`, `input.owner = "{owner}"`, `input.repo = "{repo}"` |
| `#42` or bare `42` | `input.type = "github-issue-number"`, `input.issueNo = 42` |
| Any other string | `input.type = "free-text"`, `input.summary = <string>` |

One modifier exists:
- `autopilot`  -  skip confirmations, and resolve the workspace to a worktree without asking.

The workspace is not a flag. Where the branch lives is a Phase 0 Step 5b question on attended runs.

---

## 6. Output Format Expectations

| Artifact | Both CLIs must produce |
|---|---|
| `agent-state.json` | Same JSON schema (see `~/.claude/schemas/agent-state.schema.json`); same field names; same values for identical input |
| `agent-log.md` | Same section order, same phase headers (`📊 Phase 1`, `🧠 Phase 2`, etc.), same timeline columns |
| PR body | Same template; no CLI-specific headers |
| Jira comment | Same template; real newlines (no literal `\n`); no HTML entities (no `&amp;` / `&lt;` / `&quot;`) |
| Commit messages | Format `{type}(scope): description [{JIRA_KEY}-{id}]`; identical on both CLIs (see `~/.claude/rules/git-conventions.md`) |

---

## 7. Platform Guards (macOS)

All three CLIs are installed from this package, and `package.json` declares
`os: ["darwin"]` (ADR-0012), so every command executes on macOS. That is not a
licence to stop being careful about shell: the runtime is **BSD userland and
bash 3.2**, which is a narrower target than "portable", and three constructs
below exist because the BSD form silently differs rather than failing.

For Keychain I/O the canonical path is **`~/.claude/lib/credential-store.sh`**
(or `~/.copilot/lib/credential-store.sh` / `~/.codex/lib/credential-store.sh` on
those installs). Call sites never invoke `security` themselves - the wrapper is
what keeps a secret off argv and what writes the audit entry.

| Purpose | Canonical | Underlying backend (for debugging only) |
|---|---|---|
| Keychain read | `~/.claude/lib/credential-store.sh get <key>` | `security find-generic-password -a "$USER" -s <key> -w` |
| Keychain write | `~/.claude/lib/credential-store.sh set <key> -` (the secret streams over stdin and stays off argv; `set <key> <value>` is for non-secret values) | `security add-generic-password -a "$USER" -s <key> -w` |
| Clipboard read | `pbpaste` | -  |
| Clipboard clear | `pbcopy < /dev/null` | -  |

**The three that bite on BSD**, and why each is a rule rather than a preference:

- PCRE mode is absent from BSD grep, and is banned here. It exits 2, which under
  `2>/dev/null` reads as "found nothing" - two gates shipped green for exactly
  that reason. Use `-E`.
- `stat -c %Y` FIRST, then `stat -f %m`. GNU-first ordering is required because
  `stat -f` is a valid GNU flag (`--file-system`) that succeeds and prints
  something that is not a timestamp.
- `sed -E`, never `-r`; `sed -i ''` when editing in place.

Clipboard callers may use `pbpaste` / `pbcopy` directly; the three-way
`command -v` guard they used to carry was for hosts this package no longer
installs on.

---

## 8. Enforcement

This contract is validated by:

- `smoke-autopilot-hosts.sh`  -  asserts continuous mode resolves its scripts, libs and plist template into whichever host tree is installed, and that the state root is NOT retargeted (4.4)
- `smoke-cross-cli-behavior.sh`  -  asserts every command behaves identically, pulls from Section 2 (placeholder vocab), Section 5 (argument parsing), Section 6 (output formats); also regression-locks the 8-persona agent deployment
- `smoke-commands-skills-parity.sh` (two assertions per command)  -  enforces colon-form command ↔ dash-form skill directory parity
- `smoke-compliance-skills.sh`  -  enforces store-compliance skill catalog + 4 consumer wiring
- `smoke-personal-data.sh`  -  extended in 0.5.5 to treat deprecated placeholders (`{github-username}`, `{your-website}`, `{website-repo}`) as leaks; adds `mmerterden` to public-handle blocklist for generic docs
- `pre-push-check.sh`  -  runs the cross-CLI + personal-data smoke tests before any push that touches `pipeline/skills/shared/core/` commands
- `sync.md` (Section 2 REPO step)  -  genericization lookup table MUST match Section 2 of this contract

---

## 9. Change Control

To change anything in this contract:

1. Amend this file with rationale
2. Update `sync.md` genericization table (if Section 2 changes)
3. Update `smoke-cross-cli-behavior.sh` fixtures (if Section 5/6 changes)
4. Run `smoke-cross-cli-behavior.sh` and `smoke-personal-data.sh` locally  -  both must pass
5. Announce in pipeline repo `CHANGELOG.md` under `## [Unreleased]` → `### Contract`

No silent changes. No exceptions.
