# Design decisions

Why combo is shaped the way it is, and what was reversed along the way.
This page records **decisions**, not the state of the code: read the relevant
section before undoing a design choice, and add to it when you take one.

`AGENTS.md`, at the root of the repository, holds the short version - the
invariants an agent must not violate. The rest of [docs/](index.md) explains how
to use what these decisions produced.

## Structural decisions (do not undo without discussion)

1. **In-process execution through the pi SDK** (`createAgentSession()` from
   `@earendil-works/pi-coding-agent`), not `spawn("pi", ["--mode", "json"])`.
   Each subagent is an `AgentSession` with its own context, model and tools. No
   NDJSON parsing, no process startup cost, native events, testable without a
   network. It is also what makes **persistence** possible: a session we keep
   alive.
2. **Two surfaces, one core**: the logic lives in a pure TS library (`src/`),
   and a **pi extension** (`extension/`) exposes it as a tool in the TUI. Every
   feature must be usable from a script *before* it is exposed in the extension.
3. **Agents and flows are data; our code decides what runs next.** An agent is
   declared in Markdown + frontmatter (pi's convention); a flow is YAML +
   Markdown built from a closed set of nodes; a workflow is written in
   TypeScript with combinators, for what a file cannot say. A model produces
   values, never the next node. **Reversed:** this used to read "Agents are
   data, workflows are code. No YAML DSL: we want composable code, not a
   configuration engine." See [Flows: a closed language](#flows-a-closed-language)
   for why the line moved.
4. **Display is an observer, never a participant.** No workflow may depend on a
   UI being present. *Reporters* (herdr, pi TUI, silent) subscribe to an event
   stream; unplug them all and the result is identical.
5. **The pi API lives in one file.** `src/session.ts` is the only place that
   imports from `@earendil-works/pi-coding-agent`. Everything else talks to
   `SessionPort`, which runs one turn and reads it, and `sessionPort()` adapts
   `AgentSession` to it. When pi moves, one file moves - and tests inject a
   fake session with no network, no disk, no `~/.pi`.
6. **A subagent inherits nothing from the user's environment.** Its system
   prompt goes through our own `StaticResourceLoader`: no extensions, no
   skills, no context files. `DefaultResourceLoader` would re-read the disk on
   every spawn and pull in non-deterministic context nobody asked for.
7. **English everywhere** - documentation, comments, public API, agent system
   prompts, and **git commit messages**. Everything that lands in the repository
   or in its history is English.

## Mental model

```
definition (.md)   ──►  Agent      "who"    : prompt, model, tools
spawn(agent)       ──►  Subagent   "alive"  : a session, a memory, a state
subagent.ask(task) ──►  Result     "one turn of work"
combinators        ──►  Workflow   "how"    : chain, fanOut, orchestrate, loop
```

Two levels of API, the second built on the first:

```typescript
// Low level: a live subagent whose lifetime you control
const coder = await spawn(agents.coder, { lifetime: "workflow" });
await coder.ask("Implement the parser");
await coder.ask("Apply these remarks: …");   // it remembers the previous turn
await coder.close();

// High level: disposable, everything is handled (spawn → ask → close)
const result = await run(agents.scout, "Find the authentication code");
```

`Result` is the one contract shared by everything else:

```typescript
type Result = {
  agent: string;
  output: string;          // last assistant text
  messages: AgentMessage[];
  usage: Usage;            // time, tokens, cost, turns, context
  ok: boolean;
  error?: string;
};
```

A workflow is a function `(input) => Promise<Result | Result[]>`. Workflows
compose because they share that signature - that is all.

## Subagent lifetime

**The central point of the project.** A subagent is a session: keeping it open
means keeping a context, a memory, and a token cost that accumulates. Closing it
means starting clean but amnesic. Both are legitimate; the choice must be
**explicit and local**.

| `lifetime` | The subagent… | Cost / context | When |
|-----------|----------------|-----------------|-------|
| `"task"` *(default)* | is born and dies with each task | minimal context, reproducible | exploration, fan-out, independent tasks |
| `"workflow"` | lives for the workflow | remembers iterations, growing context | coding↔review loop, iterative refinement |
| `"session"` | lives as long as the pi session | long memory, watch it | "companion" agent consulted several times |

### Coding ↔ review loop: the two regimes

Same workflow, two behaviours, one parameter:

```typescript
// "Team" regime: coder and reviewer remember the previous turns.
// The reviewer does not repeat its remarks, the coder knows what it was told.
await loop({
  steps: [agents.coder, agents.reviewer],
  lifetime: "workflow",
  until: (r) => r.output.includes("LGTM"),
  maxIterations: 5,
});

// "Freshness" regime: brand new subagents at every iteration.
// No accumulated bias, every review starts from the code alone. More expensive
// in re-reading, more honest about the result.
await loop({
  steps: [agents.coder, agents.reviewer],
  lifetime: "task",
  until: (r) => r.output.includes("LGTM"),
});
```

Rules:

- **`"task"` is the default.** Persistence is asked for, it is never obtained by
  accident.
- **A persistent `Subagent` is an explicit object** with `ask()`, `usage`,
  `close()`. There is no global session cache hidden behind `run()`.
- **Whoever opens, closes.** The owner of a `Subagent` is whoever `spawn()`ed
  it. A workflow that creates its subagents closes them in a `finally`,
  cancellation included. A workflow that *receives* live subagents **never**
  closes them.
- **Persistent subagents do not share their history.** "Working together" means
  passing `Result`s as input, not merging contexts. If an agent must see another
  one's work, you **tell** it in the task.
- **Context growth is visible**: `subagent.usage.contextTokens` is reported to
  the TUI. A `"workflow"` agent approaching its limit must either compact
  (`session.compact()`) or fail cleanly - never truncate silently.
- **No shared mutable state** between fan-out branches, whatever the lifetime.

### A turn is the pool's interface

`SubagentPool` used to hand out subagents: `acquire`, `release`, `closeAll`.
Every combinator then wrote the same six lines around them - check the signal,
acquire under a key, `try`, `ask(task, { signal, timeoutMs })`, `finally`
release - ten times over, and the pool's constructor, which absorbed eight of
the shared options, dropped the two that a turn needs. Twenty-four call sites
threaded `signal` and `timeoutMs` by hand, and the file holding the pool had no
test of its own: the only proof that a deadline reached a turn was one
assertion per combinator, each testing the same thing.

The pool now plays the turn. `turn(agent, task, { key })` is the whole of the
sequence: it refuses without spawning on a signal already aborted, runs the turn
with the workflow's signal and deadline, and gives the subagent back in a
`finally`. `hold(agent, { key })` is the other shape a combinator needs - a
conversation of several turns, for the interviewer and for a swarm's members -
and it returns an `id` and an `ask`, not the `Subagent`: closing stays the
pool's job, so nothing a combinator receives can be forgotten. `release` is
private. What was `common.ts` is three files, one concept each: `options.ts`,
`pool.ts`, `concurrent.ts`.

Two abort checks survive outside the pool, and both are about what happens
*before* a turn: `pair` would otherwise make a working copy for a pair that will
never run in it, and `interview` would spawn a held interviewer it never asks.

The tests moved with the behaviour: the pool's contract - lifetime by key, the
options on every turn, refusal without a spawn, release on a throw, `closeAll`
surviving a close that fails - is asserted once in `test/pool.test.ts`, and the
per-combinator assertions that each proved a forwarded `timeoutMs` are gone. A
combinator's tests now say what the combinator decides, not what it relays.

### The trail is the pool's too

Playing the turns left each combinator keeping what the turns added up to. Six
of them held a `steps` array they pushed every turn into, seven kept a
`performance.now()` of their own, five wrote the same `sumUsage(steps, now -
startedAt)`, three derived `ok` and `error` from the first failed step, and two
wrapped their progress hook in the same seven-line `try`/`catch`. None of that
was the combinator's to decide, and one of the copies had a hole: `turn`
refused a signal already aborted before spawning, `hold` did not, so `swarm`
held its whole roster and only then looked at the signal - a swarm called off
before it started spawned every member. No swarm test said so; the other eight
combinators each had a cancellation test.

`Trail` (`src/workflows/trail.ts`) is the envelope, written once: the results
recorded in order, a clock opened when the trail is, `usage()` over that clock,
`broken()` for the first failure. The pool owns one and records every turn on
it - a refused turn included, because a chain that was called off should say so
where the answer would have been. A combinator reads `pool.trail.steps` and
`pool.trail.usage()` and keeps nothing beside them. `hold` refuses the way
`turn` does: a held subagent is spawned before it is asked anything, so the
check on its first `ask` would have come after the session it was meant to
spare. A refused hold is addressed by its key, since there is no subagent to be
named after, and every `ask` of it answers the refusal.

The trail is the pool's second constructor argument rather than a member the
pool creates, for two callers: `pair` opens its trail before its working copy,
because the time the copy takes is part of what the pair took, and its pool
cannot exist until the copy does; `orchestrate` has no pool of its own - the
planner's, the branches' and the reducer's are three - and records the three
workflows' results on one trail. Making the trail members of the pool would have
left both keeping their own again.

Of the two abort checks that survived outside the pool, `interview`'s went with
`hold` refusing. `pair`'s stays, for the working copy. `audit` checks before the
turn too, for a different reason: a refused turn would be recorded as an audit
round that never ran. `swarm` still reads the signal between rounds, because
that is where `stoppedBy: "signal"` is decided. `deliver` keeps its own sum: its
usage is the build's, kept tasks and recorded audits of a resumed run included,
and that is a property of the build's progress rather than of the turns this run
played.

The progress hooks go through `notify` in `events.ts`, the same swallow the bus
gives every reporter: a hook is an observer, and a throwing observer has never
been the workflow's problem.

The tests moved with the behaviour. The trail's arithmetic is asserted once in
`test/trail.test.ts` and its place in the pool - every turn recorded, held turns
too, refusals too, a refused hold without a spawn - in `test/pool.test.ts`. The
usage sums `pair` and `interview` each proved through their own outcome are
gone; `orchestrate`'s stays, because what it records on its trail is its own
decision. `swarm` has the cancellation test it was owed.

### Every workflow is a Result

`WorkflowResult` was the contract five combinators honoured - `chain`, `loop`,
`reduce`, `route`, `pair` - and six did not: `fanOut`, `orchestrate`,
`deliver`, `interview`, `swarm` and `audit` each returned a shape of their own,
`usage` and `ok` included but no `output`, no `agent`, no `steps`. The two
places that read every combinator recoupled the gap by hand and did not agree.
The pipeline runner made a `Result` out of a fan-out by taking the first failed
branch or the last one and replacing its output with every branch's joined,
made one out of an orchestration by naming the planner and copying its
messages, made a third out of a delivery the same way; the `subagent` tool
read an orchestration's `planning` when nothing else was there. A fix to what
a fan-out amounts to reached one of the two.

The reading now lives with the combinator, and both callers do the same thing
for every kind: keep the result as it came. A fan-out speaks through the branch
that failed, or the last one, with every branch's output labelled where its own
would be; an orchestration through its synthesis when it has one and its
planner otherwise; a delivery through its planner, over every subtask's report;
an interview through its last turn, whose output is the brief; a swarm through
the member that failed or the last on the roster, each member's last word
labelled by the name it posted under; an audit through its last review. `ok`
keeps its one meaning, every turn ran, and what a workflow says beyond that
stays in a field of its own - `approved`, `converged`, `plan` - which is why a
`deliver` step keeps the whole delivery beside its `Result` rather than folding
`approved` into `ok`.

`joinOutputs` moved from the runner into `result.ts` and learnt to mark a
failure: the runner's copy printed a failed branch as a heading over nothing,
and a step reading six sections when eight ran would take the silence for
completeness. The runner also records each step's result on a `Trail`, so its
usage is the trail's rather than a sum kept by hand. One pipeline rule stays in
the runner because it is the pipeline's: a `loop` that never converged fails
its step, since handing unconverged work to the next step is the silent failure
`converged` exists to expose.

`FanOutResult.results` survives beside `steps`, the same list under two names,
because `results` is what every reader of a fan-out already calls its branches
in task order and `steps` is what the contract calls them. `InterviewResult.brief`
survives beside `output` for the same reason. Two aliases were judged cheaper
than a rename through every caller and every page.

### One shape for where a build stands

**Removed:** `deliver`, its progress and `BuildState` went with the linear
pipeline; a flow run's journal is the one record of where it stands. See
[The linear pipeline is removed](#the-linear-pipeline-is-removed). The reasons below are why the shape was one while it lasted.

"What a delivery has done so far" was spelled four ways: `AuditProgress` from
the audit cycle, `BuildProgress` for the hook and the resume, `BuildState` on
disk, and `DeliverResult` at the end. `deliver` held a fifth in four `let`s,
filled by a `take()` that copied the audit's progress field by field so that
`report()` could rebuild a `BuildProgress` and `outcome()` a `DeliverResult`.
The disk had drifted from the rest: a saved audit kept the review's text, its
`ok`, `approved` and the fixes, and dropped the verdict, the check as it stood
and what the fixes produced, so a resumed cycle read a thinner history than the
one it had lived; and the review came back as a `Result` with `agent:
"auditor"` written in, whatever the auditor was called.

`BuildProgress` is now `AuditProgress & { plan }`, defined by `deliver`, which
is the workflow that reports it - `resume.ts` imports it rather than the other
way round. `DeliverResult` is that progress as it ended, plus what only the end
can say: the brief, the planning turn, the landings, `approved`, and the
`Result` reading. `deliver` keeps one `progress` and moves it forward by spread:
the plan once made, the tasks and the check once settled, and after every
audit round whatever the cycle reports - which is why `AuditResult` now carries
its cycle under `progress`, the same shape every round reported, so a caller
absorbs the end of the cycle in the same spread as its rounds. `take()` and the
four variables are gone.

`done` left the progress. A build's progress does not know whether the build is
over; the moment it is reported does. So `onProgress(progress, done)` says both,
`toBuildState` takes `done` with the rest of what the state says about the
build, and a `resume` no longer carries a `done` nobody read. `notify` grew
variadic for it, which costs the bus nothing.

The saved round keeps what the live round has: the auditor's name, the review's
usage and error, its verdict, the check that stood, and the fixes' results as
saved tasks. `BUILD_STATE_VERSION` is 2, because that is what the version is
for: an older file is refused whole rather than read into a shape it does not
fill. The two conversions are written once each, `saveTask` and `loadTask`,
shared by the subtasks and the fixes. What a state still drops is the trail -
messages, turns, the review a pair kept, the working copy - and the test that
proves it is now an identity: a full progress through `toBuildState` and back
is itself, less that trail.

### The record states its own terms

The record joined the verdict to the ledger, and both callers still did the
same three things around it. They built it alike, reading `declaresVerdict` off
the agent and wrapping `saysWord` in a one-line predicate each, with its own
constant. They wrote the same ternary to offer the tool - `record.tool ? … :
…`, cast included - one to the reviewer alone, one to the auditor. And they
handed `byTool` and `open` back out to their prompt, where `reviewPrompt` and
`auditPrompt` each rendered "Still open, from your earlier rounds:" over
`openList` and the same fork between calling the tool and answering the word
alone. `auditPrompt` had grown to seven positional parameters carrying it.
The record was deep on reading a round and shallow on asking for one, which is
why its interface exposed `byTool`, `open` and `tool` raw.

`reviewRecord(reviewer, { word, approved?, restored? })` now takes the agent
and the word. Whether it decides by tool is read off the agent here, and a
caller's own `approved` stands in for the tool and the word alike, because the
caller's rule is the nearer one. Two things it renders itself. `terms()` is
what every round ends on: the open lines by id, then the instruction - a call
to the tool, with the `resolved` sentence when anything is owed, or the word
alone. `offer(others)` is what goes through `customTools`: the tool to this
reviewer and nobody else, `others` to everyone else, `others` untouched when
the reviewer decides in prose. The prompts take `terms` as a string, so they
stay functions of data: `reviewPrompt(goal, work, round, terms)` and
`auditPrompt({ brief, tasks, round, maxAuditRounds, verification, workers,
terms, byTool })`, the latter keeping `byTool` because the fix lines go where
the decision goes.

Two changes of behaviour rode along, both towards the rule that an agent gets
what its file names. A pair whose reviewer held the tool offered its worker
nothing, dropping whatever the caller had passed in `customTools`; and the
audit's pool dropped the caller's `customTools` whatever the auditor held. Both
now pass `others` through. The prose instruction reads "Answer WORD alone when
you have nothing left to ask for" for the auditor as for the reviewer; what
differed between them was wording, not meaning.

The rules of the ledger are asserted once, in `test/review.test.ts`, where the
terms and the offer are too. The pair's "an obligation a round does not name
stays open" and the audit's "a yes over an open obligation does not approve"
each proved a record rule through a workflow, and are gone; what the two
workflows still assert is that the terms reach their prompt.

## Workflows to cover

| Workflow | Shape | Semantics | Status |
|----------|-------|------------|--------|
| `chain` | 1→1→1 | output of step *n* is the input of *n+1* | done |
| `fanOut` | 1→N | N subtasks in parallel, bounded concurrency | done |
| `loop` | 1→1 | iterates until a criterion (judge, test, regex) is met | done |
| `reduce` | N→1 | one agent synthesises a fan-out's results | done |
| `interview` | user→1 | an agent questions the *user*, one question at a time, and writes a brief | done |
| `pair` | 1→1 | a worker and a reviewer discuss until the work is accepted | removed, a loop of the `build` flow |
| `deliver` | brief→? | plan, a pair per subtask, a check, an audit, fixes | removed, the `build` flow |
| `orchestrate` | 1→? | an agent *decides* the split, then delegates (dynamic fan-out) | done |
| `route` | 1→1 | a classifier agent picks the destination agent | done |

Rules:

- **Every workflow is an exported function**, not a class. No inheritance, no
  global registry.
- They all accept `{ lifetime, signal, timeoutMs, openInHerdr, onEvent, bus, cwd,
  sessionDir, exportDir, spawn }` - same names, same defaults (`"task"`, then none) - plus whatever is
  specific to them (`concurrency` and `failFast` for `fanOut`, `until` and
  `maxIterations` for `loop`).
- **`spawn` is an injectable parameter**, never a hard import inside a
  combinator. That is what makes workflows testable without a network.
- **Cancellation propagates**: the `AbortSignal` reaches every turn and closes
  the sessions that were opened.
- **`timeoutMs` is a per-turn deadline, with no default.** pi's agent loop is a
  `while (true)` (`pi-agent-core/dist/agent-loop.js:84`) with **no step cap**: it
  runs as long as the model keeps requesting tools. A weak model that
  hallucinates a tool name, gets "unknown tool" back and asks again will loop
  until something stops it - observed in the wild, 79 calls to a non-existent
  `run` tool, ~500k input tokens in a single turn. A signal alone is not enough:
  something has to fire it. No default value, though - the library does not get
  to decide that a legitimate task took too long.
- **`loop`'s `maxIterations` does have a default (5).** That is not
  inconsistent with the above: an iteration is a discrete, expensive unit with a
  meaningful small default, whereas any default wall-clock deadline would be
  arbitrary. The two guards sit at different levels - `timeoutMs` inside a turn,
  `maxIterations` between turns - and "loop forever" must not be reachable by
  forgetting an argument.
- **Reaching a cap is not success.** `loop` reports `converged` separately from
  `ok`: `ok` says the last turn ran without a model error, `converged` says the
  work reached the bar. A loop that burns through `maxIterations` with every
  turn technically fine is `ok: true, converged: false`, and collapsing those
  two into one boolean would hide the only thing worth knowing.
- **A failure does not crash the workflow**: it becomes a `Result` with
  `ok: false`. It is the caller (or an explicit `failFast` option) that decides
  to stop.
- **A reduction shows its failed branches, it does not drop them.** A synthesis
  built from six branches when two of them crashed, with nothing saying so, is a
  confident lie - and the caller can no longer tell a thin answer from a thin
  body of evidence. `formatBranches` labels them; a caller who really wants only
  the successes filters the array, which needs no option.
- **`reduce` returns the branches in `steps`**, followed by the synthesis. The
  cost of an N→1 is the cost of everything that produced it, and a `Result.usage`
  is always one turn - so the total has to be summable from `steps`.
- **Lifetime cannot change the shape of every combinator.** `reduce` is one
  agent and one turn: `"task"` and `"workflow"` both spawn once. The lifetime
  test then asserts what is actually observable (the option reaches the spawn,
  and the subagent is closed either way) rather than inventing a spawn count
  difference that does not exist.
- **How the decision of a deciding agent is read: a parsed convention.** That
  was `orchestrate`'s one open question, and the answer is `parsePlan`. The
  alternatives were weighed: a **tool call** is possible (`createAgentSession`
  takes `customTools`) and would give validated arguments, but it means teaching
  `SessionPort` about tool definitions and betting the combinator on a model
  that reliably emits tool calls - the weak models this library is run against
  do not. **Structured output** is not uniformly available across providers, and
  `prompt()` returns text either way. A parsed convention costs one function,
  works everywhere, and the check that actually matters - is this a real agent
  name? - is a lookup no schema would have replaced.
- **The parsers are lenient, and only they.** `parsePlan` and `pickDestination`
  read what a *model* wrote, not what a caller passed; everywhere else a
  malformed input is an error. The leniency is not a guess, though: an
  unrecognised agent name is **dropped**, never remapped onto a plausible
  neighbour, and an ambiguous routing answer resolves to nothing rather than to
  the first match. Silently doing the wrong work is worse than failing.
- **Leniency is decided by real runs, not by taste.** Asked for a JSON array,
  the planner answered with bare objects and no brackets - a green suite and a
  reasonable-looking prompt had said nothing. `parseJsonPlan` therefore collects
  every `{…}` block that carries `agent` and `task`, in order, so an array, a
  lone object, several objects on their own lines and a fenced block all reduce
  to the same plan.
- **Routing reads the agents' `description`.** That field is already mandatory,
  so routing needs no second vocabulary to maintain - and a vague description
  produces vague routing that no parser can repair.
- **`orchestrate` caps the plan (`maxTasks`, default 8) and fails before
  spawning.** Every subtask is a session and a bill; a plan of two hundred steps
  must not be reachable by a hallucination, and losing a run costs less than
  paying for a runaway one. Same reasoning as `loop`'s `maxIterations`, one
  level up.
- **Independence cannot be enforced, only asked for.** The planning prompt
  insists that subtasks run in parallel, and a weak planner still produced a
  step beginning "review the code identified by the scout". Nothing in the
  combinator can check that, and adding a dependency graph would be building
  `chain` a second time. When the work is sequential, use `chain`.
- **Reading code is not running it.** `deliver` takes a `verify` port, and when
  one is given **its verdict is final**: no approval makes a failing check a
  success. This is not a precaution, it is a bug that shipped - in a real run a
  pair wrote a helper and its tests, the reviewer approved, the auditor
  approved, and the test file imported `./slugify.js` for a file named
  `slugify.ts`. The suite never even loaded. Both agents had read the code.
- **An agent that decides must see the roster.** The planner was given the list
  of workers and the auditor was not, so it answered `agent: fix the quote` -
  literally the word "agent" - and every fix was dropped as an unknown name.
  Found by a real run, not by a test.
- **A refusal in prose still has to reach someone.** An auditor that explains the
  fix in English and names nobody is refusing all the same. With exactly one
  worker the whole review goes to them - there is no ambiguity to resolve. With
  several, dropping it stays right: guessing who owns a fix is how the wrong
  file gets rewritten.
- **No speculative abstraction**: a combinator is added when a real example
  needs it. `reduce` is deliberately not chunked: folding branches in batches to
  fit a context window is a real need when it appears, and until it does it
  would be a configuration knob nobody asked for.

### Reading what a model wrote lives in one file

`src/text.ts` holds `truncate`, `firstLine`, `scalar`, `saysWord` and
`jsonObjects`: everything that turns free-form assistant text into something a
workflow can act on. Only `truncate` is on the public surface, because the
extension needs it and an extension imports from `src/index.ts`, never from an
internal file.

**This reverses a decision.** `interview.ts` carried a comment saying its brace
scanner was kept apart from `plan.ts`'s on purpose, since "merging them would
mean a generic find-me-some-JSON utility that neither caller could read". The two
had since become character-identical, and what they share is only the **scan** -
each caller still keeps its own filter (`readStep`, the question reader), which
is where the readability actually lives. The same held for the `LGTM` /
`APPROVED` / `READY` matcher, written three times, and for `truncate`, written
five times and already drifting: one copy trimmed, one did not.

The rule that survives is the one that made the original call defensible: **a
shared helper takes the part that is identical, never the part that is
interpretation.**

### Writing for a model lives in one place too

The other direction had the same history. "Several results, laid out for a
reader" was written six times - the tool's answer, the runner's join, the
delegate's report, `reduce`'s branches, the audit's reports, the swarm's
answer - with four heading grammars (numbered or not, `(failed)` or `- failed`,
one blank line after the heading or none) and three spellings of nothing
(`unknown error`, `(nothing)`, `(no output)`). Whoever compared a `/run`
answer with a tool result read two conventions for one fact. Beside it, four
truncations with a marker, two identical but for a word; `plural` written and
then walked past at four sites, one of them in a file that imported it; and
the list of lifetimes kept in three places, each checking it its own way.

`joinOutputs(results, { numbered?, note? })` in `result.ts` is the one layout:
a heading naming the agent, numbered when the caller says so, a note in
parentheses that defaults to `failed` on a failure, a blank line, then the
output or `(no output)` or the error. The two knobs are the two things the six
sites varied for a reason - a reducer refers to branches by number, an audit
says what a review made of each part - and nothing else was. The swarm's
answer labels by member id through `membersOutput`, shared with the swarm's own
`Result`. `head` and `tail` in `text.ts` are the two cuts, one for what goes
into a prompt and one for a check's output, and `git.ts`, `worktree.ts` and
`verify.ts` call them. The four counts go through `plural`. `pipeline.ts`
reads `LIFETIMES` from `agent.ts`, so the tool's schema, the definition parser
and the pipeline parser name the same three words.

Two visible changes rode along. Every heading is followed by a blank line, as
Markdown wants, where the tool and the delegate had none; and an audit's failed
report says `(failed)` beside the name with the error where the output would
be, as every other failed section does, rather than `(failed: error)` over
`(no output)`.

### A result is built in one place, and a double is built whole

`failed()` had six callers and no pendant: the six-field literal of a turn that
ran was spelled out in the core twice, in a workflow, in a resume and three
times in the fake, each free to drift - the close event's `Result` carried
`messages: []` and a cumulative usage, a shape no `ask` returns, and nothing
said so. `succeeded(agent, output, usage?, messages?)` is the pendant, and the
six sites call it. A failed review or a failed fake turn that has text keeps
it, by spreading `output` over `failed(…)`, because that text is the evidence.

The `spawn` event was constructed field by field in the core and in three test
files, each with its own launch counter; when `order` and `parentId` were
added, all four moved. The core keeps its one, since it is where the event is
born; the reporters' test now derives its two events from `test/fixtures/
picture.ts`, which is where the other tests already took them from.

The tests of `/build` and `/run` fed `runPipeline` and `interview` doubles a
corner of what the real thing returns, and cast the rest away: forty-nine
`as never` in two files, and a test author who had to know which corner a
command reads. `test/fixtures/results.ts` builds an `InterviewResult`, a
`DeliverResult` and a `PipelineRunResult` whole, from the few fields a test
cares about, the way `fixtures/picture.ts` builds a snapshot. The casts are
gone, and a double that stops matching the real shape now fails to compile
rather than passing on a lie. **Removed:** the file went with the linear
pipeline; no test had used it since the commands took flows.

### An offer of tools composes, and a combo tool shares its constant parts

`SpawnOptions.customTools` is a list or a function of the id to come, and
`WorkflowOptions.customTools` a function of the agent answering either. Adding
a tool to what a caller offered therefore needed a flattener, and the only one
in the tree lived in a test fixture. `swarm` did it with a cast -
`options.customTools?.(agent) as never[]` - which read a caller's offer as a
list whatever it was: an offer written in the function form, as the `subagent`
tool writes its own, would have thrown the moment it was spread. Latent only
because `/swarm` passed no offer.

`toolsOffered(offer, id)` in `subagent.ts` is the one reader of the two shapes:
`spawn` reads through it, the fixture is a one-line caller of it, and so is the
composition. `ToolOffer` names what a workflow offers each of its agents, and
`offerBoth(first, second)` is two of them as one, asked for the id whenever
either needs it. `swarm` offers the board beside the caller's offer with no
cast. The two shapes stay: a list is what most callers write, and a function is
what a tool that must know its holder needs, and the reader is what makes the
union safe to hold.

Around each combo tool the same three helpers were written: whether an agent
declares it, an answer, a refusal - one body three times, and two of them
inline in the verdict tool's body. They are `src/tool.ts`: `declares(tools,
name)`, `said(text)`, `refuse(text)`, with `declaresVerdict`, `declaresBoard`
and `declaresDelegate` kept as the named one-liners the extension reads. A
tool body is now the decision it records or the act it performs.

### A verdict is a tool call, and prose is the argument for it

A reviewer's answer carries two things: an argument, which is prose and belongs
in the transcript, and a decision, which is a boolean and does not. `saysWord`
recovers the second from the first, and how well it does that depends entirely
on the agreement about how to write the word. `includes` accepted "I cannot say
LGTM yet" as an approval; the whole-line rule that replaced it rejects that one
and still rests on a convention a model is free to miss.

So a reviewer whose `tools:` names `verdict` is handed that tool, and its call is
the decision. `pair` reads the collector behind the tool, never the text beside
it.

Three options were weighed.

**Keep parsing prose, more carefully.** No new machinery, and the ceiling is the
convention: every refinement buys one more phrasing and leaves the next one.

**Ask for structured output**, a JSON object read by `jsonObjects`. Cheaper than
a tool, and it still arrives inside the prose channel, so "the model wrote
something that parses" and "the model decided" stay the same event.

**A tool call**, which is what was taken. It is a discrete event with a schema,
on its own channel, so *did it decide* and *what did it decide* are separate
closed questions. It costs the `customTools` seam in `src/session.ts` and one
file, `src/review/verdict.ts`.

Three consequences worth stating plainly.

A reviewer that holds the tool and calls nothing has **not** approved, and
`PairResult.verdict` is absent rather than `false`: a turn that failed to answer
is not a refusal, and guessing which one it was from the prose is the reading
this tool exists to retire.

**The tool takes the decision and leaves the argument where it was.** What the
worker receives between rounds is the reviewer's prose, not `verdict.remarks`:
an agent's definition disciplines its prose, and a field the model fills a second
time says the same thing worse. `remarks` is the short form attached to the
decision, and it is what an obligation will be opened from.

And the word survives for every reviewer that holds no tool, because an agent
nobody offered one to still has to be able to say yes.

None of this makes a model's judgement deterministic. It makes reading that
judgement deterministic, which is the only part of it that was ever ours.

### A decision is not lost to the bookkeeping beside it

Reversed, and the measurement is what reversed it. The verdict tool used to
refuse the whole call when `resolved` named an id nothing had raised, on the
grounds that a closure the ledger would refuse is better refused where the agent
is told and can call again. Told, the agent called again with the same invented
id, was refused again, and answered `APPROVED` in prose that nothing reads. The
delivery ended unapproved with no fix raised and no reason a user could see,
over an id that closed nothing either way.

So an unknown id is dropped from the verdict and named back in the same result,
beside `Recorded: approved.` The ledger is exactly as honest as before - an id
nobody raised closes nothing, here or in `ledger.close` - and what changes is
that the decision the call carried survives the mistake sitting next to it.

The refusal that stays is the one about the decision itself: `approved: false`
with no remarks and nothing raised is still refused, because there the missing
part *is* the answer.

### Finished is a ledger, not an opinion

Measured on the run that shipped the verdict tool: the reviewer called `verdict`
correctly and approved a function that computes `a - b` while claiming to add. A
clean channel does nothing about a wrong judgement, so `approved` stops being
what the reviewer said.

Everything a reviewer raises becomes an **obligation** in
`src/review/ledger.ts`, with an id combo assigns and that never changes.
Finished means the reviewer has nothing further to ask *and* nothing it raised
is still open.

The ledger belongs to this code rather than to the agent. Asking a reviewer to
re-emit its remarks each round puts us back to matching one round's prose against
another's, where "is this the same remark as last time" is a guess. A round is
handed the open ids and answers a closed question per id.

Three rules, in code rather than in a prompt:

- **Only whoever raised an obligation may close it.** A worker cannot declare its
  own work accepted, and `close` refuses with `ok: false` rather than throwing -
  an agent naming the wrong id is a runtime outcome, not a programming error.
- **An obligation a round does not name stays open.** A model that forgets has
  not approved, and failing closed is the only default that cannot be talked
  round.
- **Nothing is rewritten.** An obligation keeps the text it was raised with, so a
  reworded one is a new one.

Closures are applied before anything new is raised, so a round cannot raise and
close the same obligation in one call.

An id the ledger has nothing open for is refused by the **tool**, not silently
by the ledger afterwards. Measured with a small open-weight model: an auditor
with an empty ledger sent `resolved: [{ id: "1" }]`, inventing both the line and
the id format. The ledger refused it and the outcome was right, but nothing said
so anywhere a reader would look. Refusing in the tool tells the agent which ids
it may close and lets it call again inside the same turn, which is the only
moment it can still repair the mistake.

A boolean could only ever say that the work stopped. A ledger says which lines
are open and since which round, which is the difference between a pair making
progress and a pair that is stuck.

It buys nothing against a reviewer that closes an obligation it should not have.
That is still a judgement, about a single sentence the reviewer wrote itself
rather than about the whole of the work, and attributable to it.

The auditor in `deliver` signs the same way, and its fix lines are what go on the
ledger. Its prose is then not read at all: `deliver` has a concession for an
auditor that refuses in English without naming anyone, and that concession is
for prose-only auditors. Applied to one holding the tool it turned the word
`APPROVED`, written beside a call that said otherwise, into a fix a coder was
sent away to make. `DeliverResult.approved` then needs three things: the auditor signed off,
nothing it raised is still open, and the check passed. The check keeps the last
word it already had.

Obligations are the one thing besides the plan that survives a resume. A subtask
that was still being argued over runs again, because nobody signed off on the
tree it left; an obligation that was open is still open, and the resumed run
keeps its id rather than raising a duplicate of it. `BuildState.obligations` is
optional for a state written before the ledger existed, and an empty ledger is
the honest reading of that rather than a reason to refuse the file.

### The record is the one that joins the verdict to the ledger

The tool and the ledger were built in two places, `pair` and `deliver`, with the
same two callbacks wiring one to the other written out in each; and each round
of each workflow then made the same join by hand - take the last verdict, apply
its closures, raise what it raised, decide `said && settled`. `pair` had it in a
helper, `deliver` inline. The rule that makes the join safe, closures before
raises, was tested through `pair` and nowhere near `deliver`.

`src/review/review.ts` is the record: one per reviewer, both the tool and the
list. `reviewRecord(name, { byTool, inProse, restored })` builds the ledger and,
when the reviewer decides by tool, the tool wired to it. `record.tool` is what
the reviewer is offered, `record.open` what a round is asked about, `record.all`
what a result reports, and `record.close(review, round)` answers the one
question a round has: what was declared, whether the reviewer said yes, whether
that finishes anything, and what was raised. A review that did not run to
completion decided nothing, and the collector is drained regardless so a
verdict left behind by a failed turn cannot be read as the next round's.

`byTool` is the caller's to say, not the record's to infer: `pair` lets a
caller's own `approved` predicate stand in for the tool whatever the agent
declares, and that is `pair`'s rule to keep readable in `pair.ts`.

`lastVerdict` is gone. It was one line, its only callers were the two joins,
and the rule it carried - the last call wins, because an agent that calls again
has changed its mind - is the record's to state. `verdict.ts` and `ledger.ts`
stay as they are, each with its own rules and its own tests: the record joins
them, it does not absorb them.

### The flow grants the verdict, and a definition does not name it

The shipped `reviewer` and `auditor` named `verdict` in their `tools:` and told
their model to call it "when you are given the tool". Only a `verdict:` node
gives it, and the runner adds it to its agent's `tools:` itself, as it adds
`submit` for a typed node, so the name in the file granted nothing anywhere.
What it did was send the model after the tool where nothing offers it.
Measured against `ilaas/gemma-4-31b`: a `/step reviewer` called `verdict` three
times, each answered `Tool verdict not found`, before its prose, and the
reviewer in the tool's `loop` did the same beside the `LGTM` the loop read.
The reviewer was held to two contracts at once, a word for a loop and a call
for a node, and outside a node only the word was ever read.

Three fixes were weighed.

**Refuse, before the spawn, an agent that names a tool nobody provides.**
Strict, and readable in the file. It would have made the shipped reviewer
unusable in a `/step` and in a `loop`, where it runs most, and it puts the
fault on the person running it rather than on the definition.

**Give a reviewer outside a node a `verdict` bound to a throwaway ledger.** The
calls would succeed and nothing would read them: a `loop` stops on a word, so
the tool would answer "Recorded" for a decision no code looks at. The verdict
tool exists so that a decision is read off its own channel, and this one
would have no reader.

**The definitions stop naming it, and the flow grants it.** Taken. The node
already adds the tool, and the closing part of its turn already says the call
is the decision. The definitions keep the word, `LGTM` and `APPROVED`, which is
what a `loop`'s `until` reads. Each contract is now stated once, by whoever
reads it: the word by the definition, for anything that matches words, and
the call by the node that reads the call. What an agent can do stays readable:
its file says what it holds everywhere, and a `verdict:` node says what it
adds, as `output:` does for `submit`.

The auditor's line about putting its fixes in `raised` moved to the `audit`
section of the shipped `build`, the one place it applies. `test/agent.test.ts`
holds every shipped agent to the rule: neither `verdict` nor `submit` in its
`tools:` or in its prompt. The tool's `until` description names the two words,
so a session calling `loop` knows what the shipped agents answer with. A
user's agent that still names `verdict` is not refused; it runs as before, and
[Agents](guide/agents.md) says what the name costs.

### A review that answers in prose is asked again, and its call ends the turn

Once the definitions stopped naming `verdict`, reviews in real `/run build`
runs on `ilaas/gemma-4-31b` ended in prose: in three of four pair reviews
across two tutorial captures the reviewer wrote `LGTM`, wrote a remark, or
typed the call out as a code block in its text, and never called the tool.
Each review failed `schema`. `pair` is `on-fail: continue`, so the work
landed, but the remark (`slugify("hello.world")` returning `"helloworld"`)
was never raised, and no later round went after it.

The failure is rare in fresh runs and frequent in some contexts. Fresh runs of
a loop with a ledger, a coder and a `verdict:` reviewer, and of the shipped
`build`, called the tool in 109 review and audit turns out of 110. The one that
did not ran out of output tokens. Replaying the four failed turns, each in its
own recorded context and six times over, is what tells the candidates apart:

| candidate | turns that called `verdict` |
|---|---|
| the turn asked again, as it was | 13 of 24 (`LGTM` context: 0 of 6) |
| (a) a closing line saying prose, `LGTM` included, is not read | 13 of 24 |
| (b) the reviewer's file: "when you are given a tool to record your decision, that call is your answer" | 18 of 24 |
| (b) the same, naming `verdict` | 14 of 24 |
| (c) the failed answer kept, then the runner's retry turn | 24 of 24 |

**(a) changes nothing.** The closing line is already the last thing before the
language line, and a stronger one is read no better.

**(b) halves the failures and breaks a rule.** A definition that speaks of the
tool, by name or by description, describes something that a `/step` and a
`loop` never hold, which is why the definitions stopped naming it.

**(c) is taken:** `retry: 1` on the `review` and `audit` nodes of the shipped
`build`. A turn that ends without the call fails `schema`, `retry:` covers
`schema`, and the retry turn names the failure to the same subagent. The
definitions stay as they are, and the flow file says that the node is asked
twice. When the tutorials were captured again with it, two of the four
reviews answered `LGTM` in prose first, and both retries came back as the call.

The retry turn exposed a second defect. In 2 of those 24 retries the model
called `verdict`, read `Recorded: not approved.`, and called it again, more
than 250 times, until the turn was cut. Fresh build runs showed the same loop
once, 59 calls until the turn's deadline, and the last call, which the record
reads, approved a change 33 of the others had refused. So a `verdict:` node's
recorded call now ends the turn (`terminate` in pi's tool result). Nothing is
lost: the node's output is the call, and its prose is never read. With that,
the retries called the tool once each, 24 of 24, in a median of 13 seconds,
and eight fresh runs of the shipped `build` ended each of their 24 review and
audit turns on a single call. The TypeScript `pair` keeps a turn going after the call, because it hands the
worker the reviewer's prose, and that prose mostly comes after the call.

A `/step reviewer` still calls a `verdict` tool nobody offered, in four runs
of four with this change and two of two without it, although no file it reads
names the tool. That is the model's own habit, and there a refused call costs
a line of the display, not a decision.

### A turn cut by the output limit fails, and every agent node of `build` is asked twice

Measuring real `/run build` runs on `ilaas/gemma-4-31b` turned up two holes.

- **A turn that ended on the output limit passed.** pi ends such a turn with
  `stopReason: "length"`, and `lastTurn` read only `error` and `aborted` as a
  failure. On a node with neither `verdict:` nor `output:`, a turn that spent
  its whole output budget thinking passed as `ok: true` with an empty answer,
  and the next node read nothing. Reproduced with a scratch agent directory
  (`PI_CODING_AGENT_DIR`, the user's `~/.pi` untouched) giving the model
  `maxTokens: 200`: 200 output tokens, `output: ""`, `ok: true`. Two other
  models on 24 tokens passed a sentence cut mid-word the same way.
- **`length` is now a failure in `lastTurn`, text or not.** That is the one
  place pi's message shape is read, so `Subagent.ask`, `run()`, every
  combinator and the flow runner see it alike. What is there is kept as the
  message's text, but it is not an answer: a cut plan or a cut report handed on
  as whole is the same silent failure with more words. In a flow it is
  `provider`, the kind for a turn that failed on the provider's side, so
  `retry:` covers it with no new kind in `ERROR_KINDS`: a new kind would be a
  thirteenth that every condition reading `x.error.kind` has to learn, for a
  failure no flow has been measured treating differently.
- **`plan` gets `retry: 1`.** 2 of 16 real builds failed there, one on
  `schema` and one on a provider error; both are what `retry:` covers, and a
  plan that fails ends the build before any work. `review` and `audit` got
  theirs for the same reason.
- **So do `locate` and `report`.** Neither failed in those 16 builds, but the
  first hole closed turns a cut answer, which used to pass, into a failure on
  exactly those two untyped nodes. `report` runs after every patch has landed,
  so a failure there fails a finished build for want of one more turn. `locate`
  failing costs less, since nothing has run yet, and one retry costs less
  still. Every agent node of `build` is now asked at most twice.

### Every agent node of the shipped flows is asked twice

`build` alone had `retry: 1` on every agent node, while `explore`, `split`,
`interview` and `build-attended` failed on the first provider error, cut
answer, deadline or missing `submit`. Each node was weighed on its own, and
none had a reason to go without:

- **`explore`'s `find` and `split`'s `act`.** Their map has `on-fail:
  continue`, so a failed scout or worker did not fail the run, but it reached
  the answer as a hole: a third of `explore`'s evidence, or a whole task of
  `split`'s. The failures `retry:` covers are transient, and one more turn
  costs less than an answer written around a gap. A second failure still
  reaches the answer as a failed report, as before.
- **`split`'s `plan`.** It is typed, so a plan that forgets `submit` fails on
  `schema`, the failure `build`'s `plan` was measured failing on. It fails the
  run before any worker starts.
- **`explore`'s and `split`'s `answer`.** It runs last, so a failure there
  throws away every report already paid for, the case `build`'s `report` got
  its retry for.
- **`interview`'s `ask_next`.** The loop's `on-fail: continue` absorbs a
  failure, but it absorbs it by ending the questions: one prose answer where a
  `submit` was due would silently hand `brief` an interview cut short. It has
  `memory: flow`, so the retry resumes the same interviewer, which keeps what
  it asked and what it was told. No question is put twice: the card is its own
  node, after `ask_next`.
- **`interview`'s `brief`.** It is the flow's output, written once the person
  has answered every question. Failing there loses their answers, and in
  `build-attended` fails the run before the confirm.
- **`build-attended`'s `message`.** It runs after the whole build, so a failed
  commit message fails a finished build. The `commit` node refuses `retry:`
  and names the node writing its message as the one to take it.

Each of these nodes is now asked at most twice, and every bound counts the
second turn.

### A working copy belongs to the work, not to the subagent

`deliver` pins `concurrency` to 2 because its workers write to the same tree.
A git worktree each turns that into a question about the tasks rather than about
the filesystem, and `pair` takes `worktree: true` to ask for one.

The copy is the **pair's**, not each agent's. A reviewer given its own would be
reading the code the worker did not touch, which is the one arrangement that
looks right and reviews nothing.

Two consequences that were found by writing it rather than by reasoning about
it. `git worktree remove` refuses a copy holding changes and knows nothing about
the patch a caller is holding, so `patched: true` is both our guard and the only
thing that lifts git's. And `pair` cannot return from inside its round loop any
more: the copy is released in the `finally`, and a result built before that
carries no patch.

A copy that cannot be made stops the pair. Carrying on would write into the tree
the caller asked to spare, which is the failure the option exists to prevent.
And a copy that cannot be **released** fails it too: the first version dropped
that error, so an approved pair whose patch never came back was indistinguishable
from one that wrote nothing, while the work sat in a temporary directory nothing
named. The path goes in the error.

`worktree` is an option of `pair`, not of `WorkflowOptions`. On the shared type
every workflow accepts it and one honours it, which is a silent no-op for the
rest - `fanOut({ worktree: true })` would typecheck and run in the caller's tree.

The work is **committed** on the copy's branch before the copy goes. A first run
left the branches pointing at the base commit, holding nothing, one pair of them
per run: the patch was the only copy of the work and `PairResult.worktree` named
something empty. A branch that turns out to hold nothing is cleared with
`git branch -d`, which refuses any that holds something.

### A shared directory is a channel, so several writers get copies by default

`deliver` used to share one working tree unless somebody asked otherwise, and
the reason given was the race: two workers writing over each other. A probe
measured something worse. Four subagents given **one** directory, told only to
write a file and list what they saw, each read the other three's files inside a
single turn without being asked to look. The same four in four directories
crossed nothing. A shared `cwd` is not a hazard the workers might hit, it is a
channel they use.

So the default flipped: with more than one subtask, every pair gets a copy.
`worktree: true` and `worktree: false` are still obeyed exactly as written; what
changed is what *nothing* means. A delivery of one subtask keeps writing where it
was told - it has nobody to leak to, and `/build` on your own repository is the
case that would be ruined by isolating it.

The decision needs the plan, so it is taken after planning rather than in the
options: the number of writers is the whole question, and it is not known before.

**Choosing it also means checking it is possible, before the work.** Copies come
home through `land`, and `land` refuses a tree that already has changes in it -
which is the normal state of a repository somebody is working in. Discovering
that after two subtasks have run is paying for them twice, so a delivery that
turned isolation on by itself asks `landable()` first and stops with both ways
out in the message. Asked for explicitly, nothing is second-guessed: the failure
stays where it always was.

The flag had to grow a third answer for any of this to survive the trip through a
command. `--worktree` absent used to arrive as `false`, which is an answer; it
now arrives as nothing at all, and `--worktree=false` is how a person says no.
An option that a caller may leave unsaid must not be coerced on its way down, or
the default is decided by the plumbing.

A copy bounds writes, not reads. `..`, `/tmp` and everything else the `read` tool
reaches are outside any worktree, and a subagent that goes looking still finds
them. What closes is the channel that opens by accident.

One thing had to be true elsewhere for any of this to work on a real
repository: **a run's own exports are invisible to git**. They land in `runs/`
inside the tree the delivery is about to land in, so while git counted them the
answer to "can the patches come back" was always no.

### Copies of one repository are made and removed one at a time

git keeps the list of a repository's copies under `.git/worktrees/` and takes
no lock on it. A `worktree add` writes the new entry one file at a time, and
`add`, `remove`, `list` or `branch -d` walking the list meanwhile can read it
half written and die. A `remove` that takes the last copy can also delete the
directory under an `add` about to write into it. Two copies made and released
at once are enough, which is exactly what `copies: true` does. A test caught it
once in 94 suite runs under load, and a stress run of the same pattern failed 8
times in 14,400 copies.

So those four commands wait in a queue per repository (`src/git/registry.ts`).
It is keyed by the common git directory, not by `cwd`, because a copy and the
repository it came from share one list. Three alternatives were set aside:

- **Retrying the command when it fails.** The failure shows up as several
  different messages, and a retry keyed on those messages would hide a real
  error that happens to look the same.
- **A lock file shared across processes.** One run makes all its copies in one
  process, and nothing has yet needed two processes to share a repository's
  copies. The limit is stated in the guide instead.
- **Serialising every git command.** Only the commands that touch the list race.
  Queuing `apply`, `commit` or `diff` would make parallel branches wait on each
  other for nothing.

### Patches go in one at a time, and nothing is undone

Two patches that each apply cleanly on their own can still be wrong together:
one renames what the other calls, both add the same helper under two names, or
the second simply overlaps the first. Three ways of dealing with that were
weighed.

**Hand them back and merge nothing.** Honest, and it leaves `deliver` stopping
one step short of where it stops today. It also makes the common case, subtasks
that really were disjoint, cost a human a manual step every time.

**Apply them all, then check once.** Cheapest, and it answers the wrong
question: a red tree after three patches says only that one of them broke it.

**One at a time, with the check between them**, which is what `land` does. Which
patch broke the tree is then a fact rather than a bisection, and the cost is one
run of the suite per patch, paid only by callers who gave a `verify`.

A patch is checked before it is applied, so one that does not fit touches
nothing. `--3way` is deliberately not used: it writes conflict markers into the
files and calls that success, and a caller left to find markers in a tree it
believed clean is worse off than one told the patch was refused.

**Nothing is rolled back.** A failure stops the rest where it is and what landed
stays landed. Undoing would mean discarding work that was expensive to produce,
and every patch is also on a branch, so nothing is lost by leaving the tree
readable. The tree must be clean to start with, for the same reason: landing
onto somebody else's changes makes "which patch broke this" unanswerable, which
is the one question the whole arrangement exists to answer.

Landing adds no commit and moves no ref. What goes in stays in the working tree,
the way `deliver` already leaves its work for a human to read.

`deliver` takes `worktree` and does both halves: a copy per subtask, then a
landing per batch of them. That is why the two were built in this order -
`deliver`'s check had nowhere to run while the work sat in copies nothing put
back, so the policy had to exist before the delivery could use the mechanism.

`approved` gains a third condition there: the auditor signed off, nothing it
raised is open, the check passed, **and** every patch reached the tree. Work
nobody could land is not delivered, whatever was said about the reports.

A delivery lands more than once: the planned subtasks, then each round of audit
fixes. Only the first of those meets a tree it did not write, so the clean-tree
precondition is the **first** caller's and not every call's. Measured: with it
on every call, a real delivery whose audit asked for one fix had that fix
refused with "refusing to land onto a tree that already has changes in it" -
changes the delivery had put there itself one step earlier.

It still does not combine with `resume`. A resumed delivery finds a tree holding
what a previous process landed, which it has no record of, so it cannot tell
that work from somebody else's. That refusal is right rather than missing.
**Reversed for flows:** a flow's journal names every copy it opens and every
patch it lands, so a resume tells its own work from somebody else's, and takes
back the copies it left open. See
[A resume goes as deep as the journal](#a-resume-goes-as-deep-as-the-journal-and-replays-only-what-did-not-end).

### How the work reaches the tree is one policy, asked once

**Removed:** `settle.ts` went with `deliver`, its one caller; a flow's copies
are the file's (`copies: true`), landed by the runner. See [The linear pipeline is removed](#the-linear-pipeline-is-removed).

`deliver` decided its copies in five places: whether to isolate, from the
number of writers; the pre-flight on the tree, only when that decision was its
own, with a refusal in two spellings; the `worktree` it handed each pair; a
`settle` that landed a batch or only ran the check; and a list of landings whose
first entry alone had to meet a clean tree, and whose every entry weighed on
`approved`. Five rules of one concept, between the plan and the audit.

`settle.ts` is the concept. `settling({ cwd, worktree, writers, verify })`
decides, pre-flights and refuses; what it returns says whether pairs are
isolated, settles a batch and answers with the tree's check, and keeps the
landings and whether every patch reached the tree. `deliver` asks once after
the plan, hands `isolate` to `pair`, settles after the subtasks and after each
round of fixes through the audit's `fix`, and reads `landings` and `landed` for
the result. It names neither `land` nor `landable` any more: `land.ts` stays the
mechanism any caller may use, and this is what a delivery does with it.

The `worktree` knob does not move. It is `pair`'s and `deliver`'s, for the
reason the decision above gives, and `pair`'s side - one copy for one piece of
work, released in the `finally`, the path named when it could not be - was
already the shape a single copy wants. What this settles is the batch.

**Two `git status` on the same tree, kept on purpose.** A review read the
pre-flight in `settling` and the clean-tree check `land` makes on the first
landing as one question asked twice, and proposed that `settle` own it once
and `land` take the answer. They are the same question at two moments. The
pre-flight answers before any work, so that no subtask is paid for whose patch
cannot come home; the landing-time check answers on the tree as it stands then,
minutes later. Nothing of combo's writes into that tree in between - the pairs
write in copies - but the person whose tree it is may, while `/build` runs,
and a patch landed onto their edits would make "which patch broke this"
unanswerable, which is the one question `land` exists to answer. The first
check is a warning the second can only be anticipated by; removing the second
would trade the guard for one `git status`. The comment now sits at the call,
so the next reader does not have to rediscover it.

`settling` keeps returning a `GitResult`. Its refusal is what git said about the
tree, with both ways out appended; a ninth two-armed outcome type for the same
fact would be one more shape to learn, not one less.


`timeoutMs` has no default in the library, deliberately: it cannot know how long
a task should take. A command can, and `/interview` has to. The interviewer
reads the repository between questions, pi's agent loop has no step cap, and the
person waiting for the next question cannot tell a slow turn from a stuck one.
Five minutes is what `NEXT.md` measures these models at; 120s fails roughly half
the turns, so a shorter one would cut work that was going to finish.

The interview writes its transcript too, into the same `runs/<timestamp>/` the
pipeline uses, and that folder is now made **before** it rather than after. A
failed interview used to leave nothing behind, which is the one moment the only
question worth asking is what was actually sent - a report of `[node9] prompt
blocked` from a guard sitting in front of the provider had no record to check it
against. The failure names the folder.

The interview is also watched like everything else. It had a fixed
`interviewing…` and no reporter, so its first turn - 35 seconds and seven file
reads, measured on this repository - looked exactly like a turn that had hung.
Display is an observer and unplugging it changes no result, which is why adding
it here costs nothing and why leaving it out cost the only thing it could: the
user's confidence that anything was happening.

The same place had `--model` checked before the interview and then not used by
it, so the interviewer ran on whatever `~/.pi/agent/settings.json` named. That is
the hole invariant 5 exists to close, one command away from where it was closed
everywhere else.

### The interview answers in the language it was asked in

A question the user reads less precisely is one they answer less precisely, and
the specification is the artefact they are handed to correct before anything is
built on it. Both come back in the language of the request.

It does not contradict the English-everywhere rule: that rule governs what is
written into the repository, and none of this is. The **specification** does
travel on to the planner, the coder, the reviewer and the auditor, whose own
prompts stay English. A model takes a request in any language without trouble,
and the person signing off on the spec reading it exactly is worth more than a
uniform pipeline.

Two things are exempted by name, and both would break silently otherwise.
`parseQuestion` reads the question by its JSON **keys**, so a translated
`options` is a question nobody can display. And the loop ends on `READY`: a model
told to write French writes `PRÊT`, which is an interview that never finishes and
burns every one of its questions first.

This was the first place output left English, and for a while the only one -
[every subagent answers that way now](#every-subagent-answers-in-the-language-it-was-given).
The per-turn reminder stays: it is the one prompt that names its own two
exemptions, and an interview that loses `READY` costs six questions before
anybody notices.

## A subagent may have subagents

An agent whose `tools:` names `subagent` is handed `delegateTool()`, and can
split its task across children of its own. `examples/14-delegation-tree.ts` is
the case it was built for: one explorer that cannot read a large repository in a
turn, three scouts that each read a part, one note back.

**This is the second exception to invariant 5**, and the only one besides
`situate()`. Granting it is not inheritance: the tool comes from combo rather
than from the user's machine, the roster is the one the caller passed, and an
agent that does not name the tool in its own definition cannot have it. What an
agent can do stays readable in its file, which is the part of that invariant
that was ever load-bearing. It is still an exception to the letter, which is why
`AGENTS.md` names it.

The tool is built outside `spawn` and handed over through
`SpawnOptions.customTools`, the seam the verdict tool already uses. `spawn` never
learns what a roster is, and the recursion stays in one file: a child that
declares the tool is handed one built at `depth + 1`.

**How wide, and how deep, are decided by different people.** The roster and the
bound are the caller's: what an agent may reach, and how far a tree may grow,
are facts about the run. `concurrency:` is the agent's own, read from its
frontmatter, because how many pieces a task is worth splitting into follows from
how the agent was told to think about it. An explorer asked for two to four
tasks wants three in flight, and that belongs beside the instruction that asked
for them rather than at every call site.

A child that delegates in turn is read the same way, from its own file.

**The depth guard ships with the feature.** Delegation that can go on forever is
a bill discovered afterwards. Two levels by default - the session, a child, a
grandchild - which is where a split stops paying, because a grandchild rarely
knows enough about the whole to split anything usefully.

At the bound the tool is still handed over and refuses when called, saying how
deep it is and how deep it may go. Withholding it would leave a model calling a
tool that does not exist, getting "unknown tool" back and trying again, which is
the runaway turn `timeoutMs` exists to survive rather than one to cause.

The bound travels in a closure. Nothing reads the environment for it: an ambient
variable is how the model hole in invariant 5 existed, and that mistake is not
worth making twice.

## Delegation reaches pi the way it reached a script

The `subagent` tool wires `delegateTool` for **any agent whose definition names
it**, and for nobody else. No flag turns it on: a call that enabled delegation
would be the caller granting a capability the agent's file does not admit to,
which is the whole of invariant 5 read backwards. `maxDepth` only tightens the
bound the library already has.

The children inherit the **call's** terms - model, deadline, export directory,
signal - and not its lifetime: a delegated child is disposable, whatever the
parent was asked to be. `orchestrate` gets it too, through the same
`customTools`, so a planner's worker that declares the tool can split its share.

The drawing follows `parentId` rather than the spawn order, everywhere at once:
the dots above the prompt, the collapsed tool row, the expanded view. A
`WidgetRow` carries a `depth` and the terminal applies the indent - the
collector lays out and never draws, the same split the rest of the display
already keeps. A run with no delegation renders exactly as it did, which is the
property the tests pin.

## A fan-out reads in the order it was launched

Three scouts launched together drew as `scout#2, scout#1, scout#3`, and stayed
that way for the whole run. The collector kept arrival order and the comment
above it claimed launch order, so the defect was one line of documentation away
from being invisible.

`spawn()` takes the id synchronously and emits the `spawn` event only after
`await createSession()`, because the event carries the model pi resolved. Rows
therefore land in the order sessions came *up*, which is a property of the
provider and not of the run.

The fix is a number on the event: `nextSubagentId` hands out the id and the
launch order together, since they are one fact - the moment the subagent was
asked for - and a second counter kept elsewhere would drift the day one of the
two calls moved. The collector sorts on it and keeps it to itself: where a row
is drawn is the display's business, and `SubagentSnapshot` stays about the
subagent.

Two alternatives were worse. Sorting by id breaks a chain, where
`planner#1, coder#1, reviewer#1` is chronological and alphabetical order is
nonsense. Emitting the event before the session exists means emitting it without
the model, and the model on the spawn event is what makes a run say on its face
what it ran on.

## What a delegated run costs is a tree

A delegation with no tree-shaped measurement is a cost discovered on the
invoice: two children and their six grandchildren read as eight peers, and the
agent that caused the bill reads as the cheapest row in it.

The link is **the parent's id, on the `spawn` event**. Not a name, which two
explorers running at once already share. Not a lookup into the core either -
invariant 4 stands, a reporter observes and never queries, the same reason
`openInHerdr` travels on the event.

**The id is minted by `spawn`**, so a tool that will spawn children cannot be
built before the subagent that holds it exists. Hence `SpawnOptions.customTools`
accepting a function of the id to come, next to the list it already took. The
alternatives were worse: generating the id outside `spawn` scatters the one
place a subagent is named, and a mutable box filled in after the fact is a
lifetime bug waiting for a second caller. The list form stays because most tools
- a verdict, a check - have no use for an id, and paying for delegation
everywhere would be the tax invariant 11 exists to refuse.

**The measurement stays a flat list carrying a link, rather than nesting.**
`usage.json` keeps `subagents` flat with a `parentId` on each row, ordered so a
child follows its parent; `summaryTable` indents. Nesting the document would
force every reader to walk a tree to sum it, for a total that is a sum over
every row either way. The tree is one pass away for whoever wants one -
`treeOrder` is that pass - and no reader written against the flat shape breaks.

Two properties are pinned by tests because losing either would be silent:
**nothing is ever dropped** - a subagent whose parent is not in the list reads
as a root, and even a pair pointing at each other is still reported - and **the
total is the whole tree**, failed children included.

## A subagent may name skills

An agent's `skills:` is an allowlist, read exactly like `tools:`. What it names
is resolved before pi is opened and handed to the resource loader; what it does
not name is absent, whatever the machine holds.

**This is the third exception to invariant 5**, and the widest one: a name
resolves against the repository's `.pi/skills/` and the user's
`~/.pi/agent/skills/`, so two machines can disagree about what a name means.
That was the choice, and it is a trade rather than an oversight. The reason to
take it: a skill is the unit people already have - they write them, share them,
and hold directories full of them - and a library that refused to look there
would be answered by pasting a skill's text into a prompt, which is the same
dependency with none of the versioning. The reason it is survivable: the *name*
is still written in the definition, so what an agent can do is still readable in
its file, and `agents/<name>/skills/` is searched first, so anything that must
be reproducible ships beside its definition and wins the name.

Three ways a declared skill could have been dropped in silence, and none of them
is:

- A name matching nothing **throws at spawn**, naming the three directories.
  Prose that quietly lacks a step is the expensive failure, not a stopped run.
- A skill needs `read`: pi omits the whole section from a system prompt whose
  toolset cannot open a file. Handing the tool over instead was rejected -
  widening an allowlist to satisfy another field is exactly the silent grant
  invariant 7 exists against.
- `disable-model-invocation` is refused rather than honoured. pi keeps such a
  skill out of the prompt entirely, and combo has no `/skill:` for the model to
  reach it with, so accepting one would offer a capability that cannot arrive.

Nothing is loaded eagerly. pi advertises a name, a description and a path; the
model opens `SKILL.md` itself with `read`. A skill costs a line of prompt until
it is used.

## Pipelines: a workflow written down

**Removed:** the linear format, its parser, its runner and the shipped
`pipelines/` are gone, and a file left in an old `pipelines/` directory is
refused with `pipeline-format-removed`. See [The linear pipeline is removed](#the-linear-pipeline-is-removed). What follows is kept for
its reasons, several of which the flow format took over: a broken file is
refused, everything is resolved before the first spawn, what a turn is handed
arrives under headings of its own, and an agent does not write the file.

`src/pipeline/pipeline.ts` parsed one, `src/pipeline/load.ts` found it, and
`src/pipeline/run.ts` walked it. `/build` ran one.

- **Reversed: a branch no longer makes a run a TypeScript workflow.** A pipeline
  was linear on purpose, and "the moment a run needs a branch, it is a
  TypeScript workflow" was the rule. The flow format replaces it with a closed
  language that has branches and loops; see
  [Flows: a closed language](#flows-a-closed-language). The linear format keeps
  its rules for as long as it ships.

- **`/build` has no built-in behaviour any more, it has a default file.**
  `DEFAULT_BUILD_PIPELINE` is a pipeline like any other, parsed by the same
  parser and run by the same runner. Had the command kept a hard-coded path
  "for the simple case", that path and the pipeline path would have drifted
  within two changes, and the file would have become the untested one.
- **Only the middle is a pipeline** (the two stops have since gone, see
  [`/build` asks nothing](#build-asks-nothing)). The interview and the commit stay stops of
  the command: a question card owns the terminal, and "the agent writes the
  message, our code makes the commit" is a boundary a file must not be able to
  move. A pipeline describes *work*, not *acts on the world* - which is also why
  `runPipeline` takes a `Verify` port and never builds one from the `verify`
  field itself. Naming a command and running it are two decisions, and the
  second belongs to whoever owns the working tree.
- **An agent does not write pipelines.** It was considered and dropped: what a
  generated pipeline buys is a reviewable artefact, and `orchestrate` already
  gives that with a parser and a cap. A second, larger place where a model
  decides the shape of a run is more surface for the same benefit. A user writes
  the file; the file is data; the run is ours.
- **A broken `build.md` is refused, never silently replaced by the default.**
  The whole point of `findPipeline` reporting a parse error is that a file
  sitting right there and quietly not being used is the failure nobody detects.
- **Everything is resolved before the interview.** Parsing, shape checks and
  every agent name, so a typo costs a second rather than a conversation and
  three steps of real work. That is the same reasoning as `orchestrate`
  validating a plan before spawning, one level up.
- **Resuming is keyed by step id.** `BuildState.step` is optional: a state
  written before pipelines existed has none, and a single-delivery pipeline has
  nothing to disambiguate. It earns its place the day a pipeline delivers twice,
  where handing the second delivery the first one's approved subtasks would
  resume the wrong work.
- **A `loop` that never converges fails its pipeline.** Passing unconverged work
  to the next step is exactly the silent success `converged` exists to expose;
  in code the caller reads the flag, and in a file there is nobody to read it.
- **`reduce` folds the step before it.** That is what the linear rule buys: no
  templating, no `${{ steps.x.output }}`, and the one N-to-1 case that matters
  still works. A `reduce` with nothing to fold fails without spawning. It is
  handed those branches **once**: `reduce` formats them itself from `results`,
  so passing the previous output as text too printed every report twice - and
  the synthesiser duly reported "duplicate reports, verbatim duplicates".
- **The request reaches every step, not only the first.** A step that sees only
  the previous output cannot tell what the run was for. Found by a real run of
  `explore`: the synthesiser answered "there is no question asked in the
  prompt", because there was not - the user's question had been overwritten by
  the fan-out's output. The same bug silently starved the shipped `build`
  pipeline, whose delivery step saw the scout's report and never the brief. The
  dataflow is therefore two named sections, `## Request` and `## Output of step
  <id>`, and not one anonymous blob: a model asked to answer a question it
  cannot distinguish from the evidence answers about the evidence.
- **A subagent is told where it is.** One line appended to its system prompt by
  `situate()`. It sits oddly beside "a subagent inherits nothing from the user's
  environment", and it is not the same thing: the working directory is not
  inherited context, it is the ground every tool call stands on. A scout that
  was not told called `ls /Users/loic/gouarin/…` - a name with a dot turned into
  a slash - got "no such path", and gave up without trying a relative one. One
  branch of three, spent on a fabricated path.
- **Reversed: `/run` and `/build` are one command.** `/run build` starts the
  shipped build, since the interview and the commit are nodes of a flow now;
  see [`/run` runs flows, and `/build` is `/run
  build`](#run-runs-flows-and-build-is-run-build). The reason below is why
  there were two.
  **`/run` exists because `/build` delivers a change.** An interview settles what
  "done" means and a commit stop protects history; a pipeline that only reads
  needs neither, and putting one through `/build` means being interviewed about a
  request that wants no decision and then told there is nothing to commit. `/run`
  is the pipeline and its answer, nothing around it - and it is **lighter, not
  safer**: what a step writes is still written, because what an agent may do is
  its toolset, never the command that started it.
- **A finished `/run` leaves its answer in the conversation, not in the prompt
  editor.** It still does, for every flow, and says how the run ended under
  it; see [`/run` runs flows](#run-runs-flows-and-build-is-run-build). The editor is right for `/interview` - a brief is read, edited and
  sent - and wrong for an exploration, which is read and then *asked about*:
  putting it where the user types means they have to send their own report back
  before the model knows anything about it. It is a **custom** message and not an
  assistant one because pi has no door for the latter: `sendMessage` (custom, in
  context), `sendUserMessage` (a user message, always triggers a turn) and
  `appendEntry` (drawn, invisible to the model) are all there is, and
  `convertToLlm` turns a custom message into the **user** role. Hence the
  framing line naming the pipeline: unattributed findings in a user slot read as
  an instruction.
- **`/agents` answers the same question `/pipelines` does**, for the roster, and
  it is grouped by source rather than sorted by name because the question behind
  it is nearly always a scope: the agent is *there*, and the call that could not
  find it was loading somewhere else. A source that contributed nothing still
  names the directory it read, since "where do I put mine" is the other half of
  the same question. Its first run against this repository found `explorer`
  shipped with no symlink into `.pi/agents/`, which a test now pins.
- **Replaced: `/flows` lists what `/pipelines` listed, for flows**, and names
  what is left in an old `pipelines/` directory as refused; see [The package
  exports the flow API, and `/flows` replaces
  `/pipelines`](#the-package-exports-the-flow-api-and-flows-replaces-pipelines).
  The reason below is why `/flows` exists too.
  **`/pipelines` was there because the error message was not enough.** A pipeline
  of one repository is invisible from another - by design - and the failure
  reported where pipelines live without saying what had been loaded, with no way
  to ask. Found by running it in a scratch directory, not by reading the code.
  The listing shows the broken files **beside** the good ones: a file that does
  not parse is the most likely reason anyone is looking.
- **The extension does bring its own agents and pipelines, at the lowest
  priority.** This **reverses** the rule written one commit earlier ("an
  extension never brings its own"), and the reversal is the honest one: loading
  an extension already runs its code - pi's own documentation says so - so
  reading Markdown from the same directory adds no risk that installing it did
  not already accept. What the old rule was really protecting is *not silently
  losing a name*, and precedence protects that directly: shipped, then the
  user's, then the repository's, so a `scout.md` of your own replaces ours
  without removing anything. `builtin` is **off by default** in the library: a
  script asking for the user's agents must not be handed ours as well. Found the
  hard way - `/build --pipeline explore` in a scratch directory found nothing at
  all, because the definitions only existed in this repository.
- **One default, and it is a file.** `DEFAULT_BUILD_PIPELINE` lived exactly as
  long as it took to ship `pipelines/build.md`: a default written in TypeScript
  *and* a default written in Markdown would have differed within two changes,
  which is the drift the constant was introduced to prevent in the first place.
  The shipped file names **no `verify`**, deliberately - imposing `npm test` on a
  project that has none is worse than asking, and `/build` asks when the pipeline
  is silent. It no longer asks: `--check` names the command instead, see
  [`/build` asks nothing](#build-asks-nothing).
- **Replaced: a flow run paints its plan**, through the same `liveRun`; see
  [`/run` runs flows](#run-runs-flows-and-build-is-run-build).
  **`/build` and `/run` paint the same run the same way**, through one
  `liveRun` in `extension/ui/run.ts`: two call sites, two timers and two ways of
  clearing a widget is exactly how the one nobody is watching that day drifts.

## Flows: a closed language

A flow is a task graph written in YAML + Markdown: branches, parallel forks,
bounded loops, human nodes, sub-flows. It is built in `src/flow/` beside the
linear pipeline, and the package exports it from the switch on, see [The
package exports the flow API, and `/flows` replaces
`/pipelines`](#the-package-exports-the-flow-api-and-flows-replaces-pipelines).

- **The line was drawn in the wrong place.** What protected a run was never
  "no branch", it was "our runner decides what runs next". The linear format
  kept that guarantee only by hiding the real shape of `build` (a loop, a
  coder/reviewer pair, a check, an audit) inside `deliver`, a combinator the
  file called without showing. A closed language keeps the guarantee and puts
  the shape in the file. What follows from it: validation, a rendering and a
  dry run exist only for what is written as data.
- **Closed, not general.** Every construct comes from a closed set, every loop
  and every `map` has a bound written in the file, and every condition
  terminates and gives the same answer for the same values, so the worst case
  is known before the first spawn. A new key has to pass the same test.
- **An agent still never writes one.** Nothing in the format stops it, which is
  why the rule stays in the invariant rather than in the parser.

### A condition is CEL, cut down

`src/flow/condition/` parses, type-checks and evaluates the expressions a
`choice` case and a loop read.

- **A subset of CEL syntax, parsed by us.** CEL is the one condition language
  whose specification guarantees termination and determinism, and the project
  takes no dependency without discussion. Every expression the module accepts
  is valid CEL, so a real CEL library could replace it without breaking a file.
  The subset only removes productions from CEL's grammar, it never adds one, so
  precedence is CEL's: `!a == b` is `(!a) == b`.
- **What it leaves out, and why.** No string functions: a node that decides
  declares an enum, and nothing is read out of prose. No arithmetic: a count is
  what `max` is for. No ternary: branching is what `choice` is for. No clock and
  no randomness. What is left is literals, addresses, comparisons, `&& || !`,
  `in` on a list, `size()`, `has()`, and the `all`/`exists` macros.
- **Checked whole before the first spawn, and enums are strict.** Every address
  exists and is typed, the result is a boolean, and compared sides agree. A
  string literal compared with an enum must be one of its values:
  `status == "aproved"` would otherwise be `false` on every visit and spin its
  loop to the cap. Problems are collected rather than thrown at the first, and
  a part already refused is read no further, so one typo is one problem.
- **A condition that cannot be evaluated fails, never reads as `false`.** A node
  that failed has no `output`, an optional field can be absent, and `previous`
  is empty on a first iteration; each of these would otherwise read as "not
  yet". `&&`, `||`, `all` and `exists` follow CEL's rule for errors, which goes
  further than short-circuit: the side that decides wins over a side that errs,
  whichever comes first, so `audit.output.approved && audit.ok` is `false` for a
  failed audit just as `audit.ok && audit.output.approved` is. A module that
  only short-circuited would disagree with the CEL library meant to be able to
  replace it.
- **A name read by a condition has no dash.** CEL reads `ask-next.output` as
  `ask - next.output`, so a node id a condition reads is letters, digits and
  `_`, and the dash is refused where it is written, with the name to use.

### A schema is written short, or as a JSON Schema subset

`src/flow/schema.ts` reads what a node declares after `output:` (and a flow
after `input:`) into the one type model conditions are checked against.

- **The short notation is YAML read as a type.** `{ subtasks: [{ text: string }] }`
  is already a mapping holding a list holding a mapping once YAML has parsed
  it, so the notation needs no parser of its own: a string is a type name or an
  enum (`scout | reviewer`), a one-element list is a list, a mapping is an
  object whose `?`-suffixed keys are optional. `Question` is the one named type,
  the shape an `ask` card draws.
- **The long notation is for descriptions.** `{ json-schema: ... }` takes
  `type`, `properties`, `required`, `items`, `enum` and `description`, and
  refuses every other keyword by name. A model filling a value reads the
  descriptions; nothing else the short form lacks has been needed. Both land in
  the same model, so whatever reads a type never asks which notation wrote it.
- **Nothing reads as "any".** An unknown type name, an empty object, a list of
  two schemas and a keyword outside the six are refused with the path inside
  the schema, all at once. A value whose shape a flow cannot name is text, and
  only an agent reads text.
- **A field name is a CEL identifier.** An address splits on dots and CEL reads
  a dash as subtraction, so `next-step` would be a field no condition could
  read. It is refused where the schema declares it.

### A flow is refused whole, and each mistake once

`src/flow/file.ts` reads a flow's text into nodes and `src/flow/check.ts`
resolves its names against a catalogue; `checkFlow` is the one door, and a
`CheckedFlow` is the one thing the runner will take.

- **A fault is data with a stable code.** `{ code, file, at, message }`, never
  an exception, so every fault of a file comes back at once and a caller formats
  them. Tests assert codes, not wording, and `docs/guide/flows.md` explains each
  code in a table that a test holds to `FAULT_CODES`. There are no warnings:
  what deserves flagging deserves refusing.
- **One mistake is one fault.** A node that is refused is left out of its
  sequence and its id remembered, so a read of it, a section for it and a
  condition on it say nothing more. A broken `input:` is treated the same way.
  Faults come back in file order (the flow's keys, then each node with its own
  section, then the body), whatever order the checks ran in.
- **A key valid elsewhere is still unknown here.** `max:` on an `agent` node is
  refused with the keys the kind has; a typo is offered the nearest key within
  two edits.
- **An id is a CEL identifier, everywhere.** Letters, digits and `_`, and
  neither an address word nor a word CEL reserves. A node id with a dash would
  be readable by `reads:` and not by a condition, which is two rules for one
  name; one rule means no flow ever holds a node its conditions cannot name.
  The message offers `ask_next` for `ask-next`.
- **A flow is found by its file name, and its `name:` must agree.** A file that
  does not parse still has a name then, so it is reported as broken rather than
  as unknown; an author asking for `build` and told "unknown flow" while
  `build.md` sits right there is the failure this prevents.
- **An agent's text is its own type.** `text` is what an agent with no
  `output:` wrote: `reads:` hands it on whole, and a condition refuses it. The
  alternative, typing it `string`, would let `plan.output == "done"` read a
  decision out of prose.
- **An address is typed by the condition checker.** `plan.output.first` is a
  CEL field selection, so `reads:`, `agent-from:` and a condition share one
  parser and one checker and cannot disagree on what an address names.
- **A `##` inside a fenced code block is prose.** A section may show the
  Markdown it asks for, and the linear format's reader would have cut it there.

### A block hands on what its branches ended with

`src/flow/blocks.ts` reads `choice`, `parallel` and `map`, and
`src/flow/check-blocks.ts` checks them in lexical scope (`src/flow/scope.ts`).

- **A block's nested nodes are readable inside it and gone after it.** What
  leaves a block is its own output, so a node never reads into a sibling block
  and nothing can be addressed that only a run could resolve. Inside a `map`,
  `item` is the current item, and a nested `map`'s `item` hides the outer one.
- **A `choice` names the case that ran by its position.** `case` is `"1"`,
  `"2"`... or `"default"`, an enum a condition can compare strictly. Cases have
  no names of their own: adding one would be a second id for the same thing.
  `output` is typed only when every case that runs a node ends on the same
  type, since which one ran is only known at run time, and it is optional when
  a case runs nothing.
- **A `parallel` has at least two branches, and a case runs a node.** A
  `parallel` of one is a sequence, and a case with nothing to run is what
  `default: []` is for.
- **A literal `map:` list is strings.** That is what the shipped flows need;
  anything typed comes from an address, where a schema declares it.
- **The copies rule reads the agents' files.** Branches that run together (a
  `parallel`, or a `map` with `concurrency` above 1) need `copies: true` as soon
  as one of them holds an agent with `write`, `edit`, `bash` or `subagent`. It
  can refuse a flow whose agent would never actually write, and that is the
  side to err on: the other side is two branches editing one tree.
- **A refused node refuses itself, not the block around it.** The block is
  still checked, but its output cannot be typed with a node missing, so nothing
  that reads the block is reported again, and neither are the nodes read under
  the refused one.

### A loop reads its body one iteration back

`src/flow/loop.ts` reads a `loop`, and `src/flow/check-loop.ts` checks it.

- **`until` and `max` are both required.** A loop with no condition would
  repeat a body a fixed number of times, which no flow needs; one with no cap
  is an unbounded cycle, which a flow cannot write.
- **The body is checked twice, the first time in silence.** Inside it,
  `<loop>.previous.<node>` is a body node one iteration back, so the body's
  output types are needed before the body is checked. A node's output type
  never depends on what it reads, so the silent pass cannot disagree with the
  one that reports.
- **A carry's type is what both of its sides share.** `first` is read before
  the loop and `next` after each iteration; only the fields both have, with the
  same name and type, are readable, so `build` carries the plan's `text` and
  then the ledger's, and reading the ledger's `id` is refused. Sides with
  nothing in common are refused once, as `carry-mismatch`.
- **A ledger is named after the node that opens it.** `ledger: deliver` on
  `deliver`, and a `verdict: deliver` inside it writes there. A `map` may keep
  one too, one per item. A `verdict:` node's output is the verdict,
  `{ approved, remarks? }`, so an `output:` beside it is refused.

### A flow's catalogue keeps the agent files that are not agents

`src/flow/catalogue.ts` reads flows and agents from the package, the user's
directory and the repository, least specific first, as agents and pipelines
are. `src/flow/agents.ts` resolves the names a flow gives.

- **A broken agent file is kept, with its cause.** `loadAgents` still drops it
  in silence, which is pi's behaviour and what a TypeScript caller has always
  had. A flow is checked before it runs, and "unknown agent" about a file that
  is right there sends its author looking in the wrong place.
- **A broken file wins its name like a valid one.** A repository's `scout.md`
  that stopped parsing shadows the user's `scout`, and the flow is refused with
  the path. Falling back to the user's would run an agent the repository did
  not mean, and nobody would be told.
- **A broken file is named by its `name:` when it has one.** Agents are asked
  for by `name:`, not by file name, so `terse.md` holding `name: scout` and no
  `description:` is the `scout` a flow meant. Only a file whose YAML does not
  parse falls back to its file name.
- **Everything spawn refuses about skills, the flow stage refuses first.**
  That is `skills:` without `read`, a skill found nowhere, and a skill set to
  `disable-model-invocation`, the last one included because a checked flow
  should never throw at spawn. `findSkills` in `src/skills.ts` is the one lookup:
  spawn throws its first problem, and a flow reports all of them. The throws stay
  for TypeScript workflows, which have no validation before they run.
- **The catalogue carries its `cwd`.** A skill resolves from the repository a
  run starts in, so the catalogue is loaded for one directory and says which.

### The runner composes every turn, and reads a typed value from a tool call

`src/flow/run/` walks a checked flow: `runFlow`, and `dryRunFlow` on scripted
sessions. It walks every kind of node. It first threw before the first spawn
on `copies: true`, until the `git` port came (see below).

- **`submit` is added to the agent's `tools:` for a typed node.** `tools:` is
  an allowlist that covers the tools combo offers, so without it the tool built
  from `output:` would never be enabled. It is granted by the flow file, which
  asked for a typed value, and it hands a value back without acting on
  anything; an untyped node's agent gets exactly its own tools.
- **A call off the schema is refused back to the model.** The turn goes on, the
  model can call again, and the last call accepted is the answer; the node
  fails `schema` only when the turn ends with none. The `verdict` tool already
  works this way, and a turn thrown away over a fixable call costs a whole
  retry. A value that is not an object travels as `{ value }`, since a tool's
  parameters are an object.
- **Nodes sharing a subagent declare one `output:`.** A subagent's tools are
  fixed when it is spawned, so two typed nodes resuming one subagent through
  `memory:` with different schemas cannot both have their `submit`. The flow
  stage refuses it as `memory-output-mismatch`. A text node beside a typed one
  is fine.
- **Our code classifies a failure, not a message.** Each attempt gets a
  deadline signal of its own, and a failed attempt is `stopped` when the run's
  signal or the subagent's stop fired, `timeout` when its deadline did, and
  `provider` otherwise. The dry run gives each attempt a deadline its scripted
  session fires itself, so a scripted timeout goes through the same abort with
  no clock to wait on.
- **A retry on the same subagent is asked the failure, not the turn again.**
  The turn is already in its history; it is told what failed, and a typed node
  is reminded to call `submit`. A fresh subagent, after a timeout with no
  `memory:`, is asked the whole turn.
- **A read typed `string` goes as it is.** `input: string` is what a person
  typed, and a JSON string with its quotes and escaped newlines would change how
  a model reads it. Text goes as it is, every other value as JSON. An optional
  field left out keeps its heading, as an empty text does.
- **`agent-from:` that cannot be read fails with `condition`.** A value that
  picks an agent is read like an address in a condition, and a condition that
  meets a failed node or an absent field fails with that kind.
- **The runner does not write the language line.** `Subagent.ask` closes every
  turn with it already, so a flow's turn gets it last like any other.
- **A visit event carries no subagent id.** A visit is a node's, and a `choice`
  has no subagent, so `picture`, herdr and the mirror skip them (`isVisit`).
  `visit_end` carries the agent a visit ran and its model, which the journal
  will record, so `Subagent` now exposes the model pi resolved.
- **The input and the kinds are checked before the first spawn, by a throw.**
  An input off `input:` is the caller's mistake; the run stage will refuse it
  as a fault once it exists.
- **A dry run's script is keyed by node address.** `gate/look`, as a fault's
  `at` names it, never a bare id; the refusal offers the address when a key is
  an id. An array is always a list of answers, since an answer can itself be a
  list, and one value answers every attempt. A script's faults are a closed set
  of their own, `ANSWER_CODES`, apart from the flow's.


### Blocks join what their branches ended with, and a ledger is a scope's

`parallel`, `map` and `loop` run in `src/flow/run/blocks.ts` and `loop.ts`,
and the ledgers and `verdict:` nodes go through the review record.

- **A block names the first branch in order that failed of its own.** Branches
  finish in any order, and the failure a flow reports must not depend on it. A
  `cancelled` branch is named only when every failure was one, which happens
  when the cut came from further out.
- **`on-fail: continue` on a `parallel` or `map` ends it `ok: true`.** Its output
  keeps each failed branch as `{ ok: false, error }`, the way a loop under the
  same key ends `ok: true, converged: false`. Without it, the output of a
  failed block could never be read, and "failed branches are kept" would say
  nothing.
- **A `parallel` branch opens a memory scope of its own, like a `map` item.** A
  subagent takes one turn at a time, so two branches running at once cannot
  share one. Branches running together that resume an outer scope's subagent
  take turns on it instead of failing: its history is still one conversation.
- **A `fail-fast` cut is a signal of the block's, carrying its reason.** Our
  code reads `cancelled` off that signal, like `stopped` off the run's, rather
  than off a message. A branch not started runs no visit: its entry is
  `cancelled` at the path of its first node.
- **`copies: true` was refused at run time until the `git` port existed.** A
  runner that ran the branches in one tree instead would have made the
  `copies-needed` fault a lie. The refusal went with the commit node, below.
- **A `verdict:` node is a review record handed the scope's ledger.**
  `reviewRecord` takes a `ledger` read at every use, so the three rules and
  the terms a reviewer is asked on stay written once. It is read at every use
  because a subagent a `memory:` scope keeps can outlive the ledger it first
  wrote to. The runner adds `verdict` to the agent's `tools:`, as it adds
  `submit`; `approved` is the record's own (it said yes, and nothing is left
  open); and a turn with no call fails `schema`, like a typed node's.
- **`<scope>.ledger` is read when the address is.** A verdict earlier in the
  same iteration changes the open obligations, and a snapshot taken when the
  scope opened would hand the next node a stale list.
- **A condition, a carry or a `map-from` that cannot be read fails with
  `condition`.** It is the kind an address that meets a failed node or an
  absent field already has.
- **A dry-run key is walked against the tree.** An exact visit path numbers
  every loop iteration and map item, names every branch, and stays within each
  bound. A list is refused as `answer-past-max` when it holds more answers
  than the node can be asked for in its enclosing path: `1 + retry` times the
  `max` of every loop around it. Each attempt tells the script its node's
  address, so a path is never parsed back into one.
- **A subagent's `spawn` event carries its `visit`.** That is the one link the
  live view needs between a plan line and the subagents that ran it. A scope's
  subagent names the first visit that asked for it.

### A check runs what was read before the run, and the run stage reads it

A `check` node names a script of the project; `checkRun` is the run stage,
and `runFlow` takes the `CheckedRun` it returns in place of a flow and a
`cwd` option.

- **The `check` port sits beside `Verify` in `src/verify.ts`.** `CheckScript`
  takes the script's content, a directory and a bound; `Verify` takes nothing,
  and the linear pipeline still calls it. Both are the project checking itself,
  so they share a file, and `Verify` goes when the pipeline does.
- **The content runs as `bash -c <content> <path>`.** `checkRun` reads each
  script once, and the `CheckedRun` holds what it read, so an agent editing the
  file mid-run changes nothing and no file is read twice. `$0` is the path, so
  a script that finds its siblings by `dirname "$0"` still does. Its stdin is
  closed: nobody can type into a check, and one that waits for input ends.
- **The script leads a process group, killed whenever it ends.** A timeout or a
  stop that killed only `bash` would leave a test runner's workers holding the
  pipes, and the check would not end until they did; killing the group at a
  normal exit too clears what a script left in the background.
- **The report is a rolling window over both streams, in arrival order.** The
  output of a check is not bounded, and 8000 bytes of it is all that is kept, so
  no more than twice that is ever held. `tail` takes how much was already let go,
  and its mark counts the whole.
- **A check's bound is its own.** The run's `timeoutMs` and the flow's
  `timeout:` bound agent turns: a slow provider, which the operator sees and the
  author cannot. How long a project's suite takes is the project's, and a
  `--timeout` raised for a slow model should not let a hung suite hang the run.
- **`retry:` on a check is refused with its reason, not as an unknown key.**
  `retry-refused` says to raise `timeout:` or make the check stable; a plain
  unknown key would offer the kind's keys to someone who meant something.
  `RETRY_REFUSED` holds one reason per kind, so `commit` and `flow` add theirs.
- **The path is checked twice.** An absolute path, or one leaving the
  repository, is a flow-stage `key-type`, since no project can make it right.
  Whether the file is there is the run stage's, against the `cwd` it is given,
  which the launch sets to the repository root.
- **A missing port or script is one fault.** Two nodes naming one missing
  script, or a flow of checks launched with no `check` port, is one mistake,
  reported at the first node that needs it.
- **The port says `stopped`, the runner reads the signals.** `CheckScript`
  knows nothing of `fail-fast`, so a stop and a cut are read off the run's and
  the block's signals, as for an agent visit.
- **A dry run answers a check by the same keys as an agent turn.** An answer
  is `{ passed, report }` or `{ fail: "unavailable" | "timeout" }`, checked
  before the start, and a check no answer covers is `unscripted`. It takes a
  `CheckedFlow`, so it reads no script.
- **`checkRun` takes `somebodyThere` already.** The launch states it once, and
  the `CheckedRun` holds it, so the `ask` node reads it rather than changing
  every caller of `checkRun`.

### A commit is our act, and copies start from the tree as it stands

The `commit` node, the `diff` address and `copies: true` reach git through one
port, `gitPort()` in `src/git/port.ts`, built on the rest of `src/git/`. The
runner never runs git itself.

- **`checkRun` is async.** Whether `cwd` is in a repository is a question for
  git, so the run stage asks the port it was handed rather than guessing from
  a `.git` on disk, and that is a process.
- **The port holds no state; the run's branch lives in the run.** A port can
  serve several runs, and a branch belongs to one. The first commit opens
  `combo/<slug of input>` (a typed input is slugged from its JSON), suffixed
  while the name is taken, and every later commit first checks `HEAD` is still
  on it: committing onto a branch somebody switched to is the one thing a
  branch of its own is there to prevent. It opens even when the tree is clean,
  so `branch` always names the run's.
- **A commit's message is an earlier node's output.** A text, or a `string`
  field of a typed one. `input`, `item` and `diff` are refused where they are
  written: a message is what an agent node wrote, and an `ask` can sit between
  the two. An empty or absent message fails `empty-message` before git is
  asked anything, so nothing is opened for a commit that cannot happen.
- **`diff` goes through an index of its own.** A copy of the repository's
  index in a temporary file, with every change added, so untracked files show
  and the person's `git status` does not change; `git add -N` on the real index
  would have marked each one as added. It is a text, so no condition reads it,
  and a node that cannot have it fails `unavailable`.
- **A copy starts from a snapshot of the tree, not from `HEAD`.** A commit of
  the tree as it stands, which no ref points to. From `HEAD`, the second
  iteration of a loop around a `copies: true` block would not see what the
  first landed, and its patches would conflict with it on the way back.
- **Each branch's entry says whether its patch landed.** `landed`, and
  `refused` with git's reason on the patch that stopped the landing; the
  branches after it read `landed: false`. A patch that could not be taken stops
  the landing like one that does not apply, and a branch that changed nothing
  counts as landed. It is a value, like a red check: the `check` after the
  block, or a condition, decides what to do about it.
- **A stopped run lands nothing.** Its copies are still released, which
  commits each one's work on the copy's own branch before removing it, so what
  a branch wrote is kept and nothing half-finished reaches the run's tree.
- **A commit writes, and copies are no way out for it.** Beside branches
  running at once, `copies-needed` offers to commit after the block (or
  `concurrency: 1` on a `map`), since a commit inside a copy is refused as
  `commit-in-copies`.
- **A memory scope used inside a copy opens inside it.** A subagent works in
  the tree it was spawned in, so one kept by `flow`, or by a block around the
  copies, would work in a tree that is not its branch's, and outlive the copy
  it was spawned in. `memory-outside-copies` refuses it.
- **A dry run makes no copy and runs no git.** A commit is answered by the
  script, `{ committed, sha?, branch }` or `{ fail: "unavailable" }`; an empty
  message comes from the node that writes it. `diff` reads as an empty text,
  and every branch of a copies block reads as landed.

### An ask says what not answering gives, and a model's labels compare to nothing

The `ask` node puts a question through the `ask` port, the same `AskUser` the
interview uses. Nothing of the extension is wired to it yet.

- **The port grows an optional second argument, not a second port.** `Asking`
  says the card's form, the reads shown above the question, the visit asking,
  the label of "enough" or `false`, and a signal that takes the card down. An
  existing `AskUser` ignores it and keeps working. `undefined` stays "the
  person declined": "that's enough" where it is offered, the stop key where it
  is not.
- **Declining a card with no "enough" stops the whole run.** The card owns
  `esc` while it is up, so that key press is the run's stop, and it ends the
  run `stopped` the way `stopSwitch.all()` does. `on-fail: continue` does not
  catch it. Ending only the node would turn the stop key into a way of failing
  one question.
- **Literal options make a closed card.** `answer` is then a strict enum of
  the labels, so the card offers no typed answer beside them: a typed one
  would be a value off its own type. `custom` stays in the output for one
  shape across choice cards, and is `false` there.
- **`ask-from:` is a choice card.** A `Question` carries its options, so
  `options:` and `confirm:` beside it are refused, and `enough:` is refused on
  a yes or no and on a free text, which are always answered.
- **A model's labels are marked on the type, not on the node.** The answer of
  an `ask-from:` is a `string` with `free: true`, and the condition checker
  refuses a literal on the other side of `==`, `!=`, `in` and the ordered
  comparisons, as `condition-free-string`. One flag on the existing type model
  travels through every address and field the checker already follows. A free
  text's answer is a plain `string`, as the format says: a person typed it.
- **`default:` is a literal of the node's form.** A label of its options on a
  literal card, any text after `ask-from:` (its `custom` is worked out at run
  time against the question), `true` or `false` on a yes or no, and a text on
  a free text.
- **A `timeout:` starts when the card is shown.** A card waiting its turn in
  the queue is not a person failing to answer, as an agent turn waiting for a
  shared subagent is not a slow one.
- **No `ask` port is nobody there, whatever the launch says.** The run stage
  refuses the flow, as `unattended-ask`, when an `ask` with neither
  `default:` nor `enough:` could be reached. At run time, nobody there takes
  the same rule as a timeout.
- **An `ask` may read `diff`.** "Commit these changes?" wants the changes
  above it, so the read goes through the code an agent's does and needs the
  `git` port the same way. What a turn shows and what a card shows are one
  function, `showRead`, and differ only in the fence around JSON.
- **A dry run answers an ask with its output.** `{ fail: "nobody" }` and
  `{ fail: "timeout" }` take the node's own path, its `default:` or `enough:`
  included, and `{ fail: "stopped" }` is the card declined. Each is accepted
  only where the node's keys allow it, and a choice card's answer is held to
  what a card can give: `answered: false` only with `enough:`, an `answer`
  whenever it is answered.

### A call is checked once, and only `input` crosses it

A `flow` node runs another flow of the catalogue as one node. Nothing of the
extension is wired to it yet.

- **A callee is checked on its own, once per check.** Its faults stay in its
  file, and the caller gets one `broken-flow` naming the callee's file and its
  first fault. `/flows` lists the callee with its own faults, so copying them
  into every caller would report each one twice, at addresses of another
  file. A callee called from two nodes is one checked flow, attached to both.
- **A cycle is found from the files, before the callee is checked.** The call
  graph is read from what the files write, behind a `choice` too, and a call
  that leads back to the flow being checked is refused as `call-cycle`, with
  the path from that flow. The callee is not checked, so the check always ends
  and no depth constant is needed. A cycle that does not pass through the
  flow being checked makes the callee broken, which is what the caller is
  told.
- **Only `input` goes in, and only the last root node comes out.** A callee
  is checked in a fresh root scope and run in a fresh `flow` memory scope,
  opened by the visit of the call and closed in a `finally` when it ends. A
  callee taking `string` takes any value: a string as it is, anything else as
  JSON, so an enum's value goes in as the word it is. A typed input takes the
  same type and nothing else. `input: diff` reads the visit's tree, and needs
  the `git` port like any read of `diff`.
- **A callee's run is the run's own, one call further down.** It keeps the
  run's ports, signal, stop and event stream, so there is one queue of cards
  and one run branch. What changes is the stack of flows `model:` and
  `timeout:` are read from, callee first, and the memory scopes it shares.
- **The world is told a node by its address through the calls.** A
  `visit_start`'s `node`, an attempt's deadline and a dry run's keys use
  `spec/round/look`, not the callee's own `round/look`. A flow called from two
  places has two addresses, and whatever keys a node by its address has to
  tell them apart. The rules about copies, commits, checks and unattended
  questions read the same addresses, so their faults land in the caller's
  file at the call path.
- **A dry run's key on a call and a key under it are refused together**, as
  `answer-flow-overlap`, even when they would name different visits. Telling
  visits apart would need the loop iterations resolved against each other,
  and a script that does both for one address is almost always a slip.

### A run keeps what its check read, and writes each fact once

A run given a run directory keeps a snapshot there and appends a journal,
which a resume reads back (below); nothing of the extension uses them yet.

- **The journal is a port, written by the runner.** A real run appends to
  `journal.jsonl` in its run directory, a run given none writes nowhere, and
  a dry run keeps the entries in an array. A reporter never writes it, so
  unplugging every reporter leaves it the same. Each append is synchronous:
  the fact is on disk before the run goes on, and two branches never
  interleave a line.
- **A visit's entries are its `visit_start` and its `visit_end`, each written
  before it is told.** One object goes to the journal, then to the event
  stream, so a reader folding the two never sees an event the journal lacks.
  The start is there for a reader holding the file alone, such as a harness
  following a `/run` from outside the process: without it, that reader can
  name the last visit to end but not the one running. Every reader that
  restores a run folds the ends and skips the starts.
- **Only the last line can be torn.** Whatever follows the last newline is
  ignored. Any other line that is not an entry throws: that file was not
  written by a run, and guessing past it would resume a different run.
- **A ledger writes its own obligations down.** The ledger a loop or a `map`
  item opens is wrapped to append each raise and each close it accepted, keyed
  by the visit whose scope keeps it (`fix`, `work[2]`). The review record and
  the `verdict` tool are unchanged, and a refused close writes nothing.
- **Every fact is keyed by the visit it belongs to.** A `carry` names the
  iteration that reads it (`fix#2`). A `map`'s list is written for a literal
  list too, so every `map` is restored the same way. A copy is written when it
  is made, with its directory and git branch, and its landing once the block
  landed; a dry run writes both with no directory. The run's end is what
  `runFlow` returned, written down.
- **What validation read is on the checked flow.** `sources` holds the flow
  file, every file it reaches, and each agent it names with the skills they
  resolved to. Each checked flow composes its own from its callees', so a
  callee's is right on its own too. `checkRun` already held each check
  script's content. The snapshot copies these two and nothing else.
- **An agent is kept as it was parsed, a skill as its files.** Reading an
  agent's file again for the snapshot could see a file changed since the
  check, so the parsed definition is stored. Validation holds only a skill's
  metadata, so its directory is copied, since the model may open what sits
  beside `SKILL.md`.
- **Read back, an agent lives in the run directory.** Its `filePath` becomes
  `<runDir>/agents/<name>.md`, so its own `skills/` directory, looked up
  first, is the copy. This holds for the check and for the spawn, and it is
  the one difference between a check of the snapshot and the first check. A
  flow file keeps its own path, which a resume needs to say what changed on
  disk.
- **A run directory holds one run.** `runFlow` refuses a directory that
  already holds a snapshot, and checks the input before writing anything, so
  a refused input leaves no directory behind. The settings kept are what the
  run started with: `cwd`, `somebodyThere`, `model`, `timeoutMs`.

### A resume goes as deep as the journal, and replays only what did not end

`resumeFlow(runDir, ...)` carries a run on from its snapshot and journal, and
`resumePoint(checked, journal)` is the reading both it and a dry run's `from:`
take, so a resume is tested offline through the same door.

- **A visit that ended with a value survives; nothing else does.** `ok`, or
  failed under `on-fail: continue`: the flow read either as a value and went
  on. A visit stopped or cut by `fail-fast` had no end of its own, and one a
  failure travelled up through ended by the run's failure, so both run again.
  That one rule is the failure chain, and its fresh `retry:` budget.
- **Structural visits are walked again, their children skipped.** The walk
  itself decides, visit by visit: a surviving visit hands back what it ended
  with and is neither told nor written again. A loop, a `map`, a `choice` or
  a call that did not end is entered again, and what it reads comes back from
  the survivors, so a condition decides the same. A `carry` and a `map`'s
  list are taken from the journal rather than read again, and not written
  twice. `resumePoint`'s `from` is only a sentence for whoever asked.
- **A failure the flow decided is refused.** A cap, `give-up`, a condition that
  could not be read, a list past `max:` and a commit with no message are
  decided from values that survive, so replaying them decides the same. A
  provider, a timeout, a schema, nobody answering and a person's stop are
  worth another attempt.
- **An obligation is written with the visit that decided it.** A verdict
  visit cut after its decision but before its end runs again, and its
  obligations would be raised twice. `obligation_raised` and
  `obligation_closed` name their `visit`, and a ledger is restored from the
  entries whose visit survives.
- **A copy is kept with its base, and taken back only while its work is not
  in the tree.** `copy_opened` records the commit the copy started from, so a
  resume can take its patch against the same one. An open copy still there
  carries on; a copy gone or moved, or one whose patch never landed, is
  forgotten with `copy_lost`, written in the journal so a later resume
  forgets the same; a copy that landed put its work in the tree, so its
  branch's visits survive and what is left runs in a fresh copy. A stop
  releases every copy and lands none, so after Ctrl+C a `copies: true` block
  starts its branches over: their work is on their git branches, not lost.
- **Only `model` and `input` are refused by name.** They are what a new run is
  for. `timeoutMs` replaces the run level. `ports` and `somebodyThere` are
  where the resume runs, like the tree at a start, and the run stage holds the
  snapshot to them: somebody at the keyboard may resume a run started
  unattended.
- **The disk is compared for flow files only.** A flow file differing from the
  snapshot is named in `changed`. An agent is kept parsed and a script by its
  content, and the run uses those; the one line is about the file a person
  edits and would expect to see run.
- **The result's `usage` is what this resume spent.** A visit that survives
  costs nothing now, and adding what the journal says it cost would count a
  cut visit's partial cost nowhere and the rest twice across resumes. The
  journal holds every visit's own, for whoever sums them.
- **The lock is a file made with `wx`.** It holds the pid and the host. A
  process of this host is asked with signal 0, and one we may not signal is
  alive. `runFlow` takes it too, after the snapshot, which is what refuses a
  directory holding a run.
- **A stale lock is replaced under `lock.json.takeover`.** Removing it and
  making it again let two resumes that both read it as stale both take it:
  the second removed the lock the first had just made. The takeover file is
  made with `wx` too, so one taker at a time gets through. That taker reads the
  lock again, since another may have taken it over in between, and replaces it
  with a rename, so no plain `wx` finds it missing. A taker that finds the
  takeover file refuses, with the pid in it when that process lives, or its
  path when it died there. The rename is atomic on a local filesystem, and a
  run directory is expected to be on one. `takeLock` takes the read as a
  parameter, which is how the test puts a second taker between two steps of
  the first.

### Bounds are a worst case, and a rendering draws a checked flow only

`checkFlow` computes a flow's bounds, and `planOf` and `mermaidOf` take what it
returned, so a broken file is never drawn: its faults are what is shown.

- **The bounds belong to the checked flow.** They are computed once, from the
  resolved tree, and the plan, `/flows` and the dry run read the same figures.
  A callee's own `bounds` are its figures as a root; under a call, its nodes
  are counted in the caller's, by their address through the call, since a
  callee with no `timeout:` takes its caller's.
- **The time is the sum of the bounded waits.** Each turn, check and ask runs
  to its bound, branches running together take the longest one, and a `map`
  runs in waves of `concurrency`. A shared subagent or a queue of cards can
  make branches wait for each other beyond that, and the guide says so: a
  figure that added every branch would be true and useless for a `map` of ten
  run ten at a time.
- **An ask with no `timeout:` is a flag, not infinity.** `waits` says a person's
  answer is part of the time, and the rest of the figure still says what the
  machine can take. One unbounded question would otherwise hide every other
  bound of the flow.
- **A `choice` is its worst case, turns and time apart.** Which case runs is
  decided at run time; each figure is the worst any case reaches, even when no
  single case reaches both.
- **The plan is data, and text is one reading of it.** Each line carries its
  kind, id and the visit path it stands for, `#n` and `[i]` where a run will
  write numbers, so a live run fills the plan instead of drawing a second
  picture. A call is one line: its callee's plan is the callee's own.
- **One turn bound, read in one place.** The default of 30 minutes and the
  order node, nearest flow, default moved beside the bounds, and the runner
  reads the same function, so what the plan says a turn may take is what the
  run gives it.
- **Mermaid's ids are numbered, not the file's.** A node named `end` would
  close a subgraph, and a branch name could collide with a node id. Numbering
  in drawing order keeps the text the same for the same flow, and every label
  carries the id. What Mermaid or its HTML would read as markup in a label is
  written as an entity code.
- **The reference always has an index.** `docs/reference/flows/index.md` is
  written even while no flow ships, saying so, so the navigation entry exists
  before the first flow does and the staleness test covers the directory from
  the start. It draws the package's flows against the package's agents only,
  files named from the repository root, so a page reads the same on every
  machine.
- **A plain `mermaid` fence.** `myst_fence_as_directive` makes MyST read it as
  the directive, so the same file is drawn by GitHub and by the site.

### The live view folds the plan by the state of each visit

`livePlan(checked, journal, events)` is one pure fold beside `picture.ts`,
which is unchanged, as is `tree.ts`. The picture is still what a TypeScript
workflow with no plan draws, and what `usage.json` and the tool's result
read.

- **The journal is the earlier lives and the events are this one.** A
  journal entry is a `visit_end` written down, so one fold reads both. A
  visit's last end is the one drawn, whichever life wrote it, and a
  `copy_lost` forgets what lay under its branch, as a resume does. Fed the
  full journal and this life's events together, a visit would count twice.
- **A life is told apart by its `life_start`.** At first the journal wrote
  no mark when a life started, and a killed life was told from the next
  only when that one came in as events. The measurement of lives added the
  mark, and the view reads it, so the journal alone tells them apart.
- **Folding follows state, and so does the glyph.** Running is expanded,
  ended is one line, not visited is the plan line, so the first frame is
  the plan. A visit begun and not ended while nothing runs, which is what a
  kill leaves, is expanded under `○` rather than `●`. A branch, an item or
  an iteration reads `working` while its block runs, which keeps the moment
  between two of its visits from flickering to `○`.
- **A `choice`'s cases are named, not counted.** `○ case 1, default` and
  `○ not taken: default` use the plan's own case names. A count would
  disagree with the plan's `choice of 1 case`, which leaves out the
  default.
- **The counts are over the frame, the cost over every life.** A visit a
  resume ran again counts once, as it last ended, so `failed` matches the
  `✗` a reader can see. The cost adds up every life's, since each one was
  paid. A life that wrote its `run_end` takes the total from there. A killed
  one adds up its outermost ended visits, whose time overstates branches
  that ran together: no clock saw that life end.
- **Only a loop is named `not converged`.** Only a loop ends
  `converged: false` or fails `unconverged`, since a failure that travels up
  becomes `child`. The summary finds them without looking up the node.
- **The glyphs are the TUI's.** `statusIcon` decides `●`, `✓` and `✗` for a
  subagent's row and for a visit's line alike, and the plan's `○` is shared
  with `showPlan`. `–` is the view's own. An ended line names the agent its
  visit ran, and a running one names its subagents, so the journal alone
  and the stream alone draw the same last frame.
- **herdr is left as it is.** A split per subagent named by agent and scope,
  with none for `check`, `ask` or `commit`, belongs with the switch to
  flows, which wires the view into the extension. Until then nothing reads
  the view, and herdr opens its splits as for any run.
- **The width is the caller's.** `showLive` and `showSummary` take it with
  no default, like `truncate`, and cut a line while keeping its indent.

### The shipped flows are the pipelines' work, written as flows

`flows/` holds `build`, `explore`, `split`, `interview` and `build-attended`,
symlinked into `.pi/flows/` like the agents and pipelines, and listed in the
package's `files`. Nothing runs them yet. Each is checked against the shipped
agents by `npm test`, drawn in `docs/reference/flows/`, and walked with
`dryRunFlow`. Their first drafts predate rules the validator now holds, and
writing them for real moved them in the places below.

- **The attended build is `build-attended`.** It sorts beside `build`, and
  names its one difference in the word the run stage already uses: somebody
  is there. A flow's name is a file name, not an id, so the dash costs
  nothing.
- **Ids are CEL identifiers.** `ask-next` became `ask_next`, since a condition
  reads it and a dash is subtraction to CEL.
- **A whole output is read by its bare id.** `input: spec` and
  `commit: message`, as the guide writes them, where the drafts had
  `spec.output` and `message.output`: both check, and one spelling reads
  better than two.
- **The interviewer submits a question or none.** Its output is
  `{ question?: Question }` rather than `{ ready, question? }`. One optional
  field cannot contradict itself, where `ready: false` with no question was a
  card with nothing on it. The choice asks when there is a question, and the
  loop ends when there is none or the card was answered "enough".
- **Six questions end the interview, not the run.** The loop has `on-fail:
  continue`, so reaching its cap ends it `converged: false` and the brief is
  still written, as the interview combinator wrote one at `maxQuestions`.
  Failing the run would throw six answers away.
- **The brief shares the interviewer's subagent with no `output:`.** It needs
  the conversation, which only `memory: flow` keeps. The subagent carries the
  `submit` tool `ask_next` declared, so the brief's prose says the turn calls
  no tool.
- **A pair that reaches its cap goes on to the audit.** `pair` has `on-fail:
  continue`, and the audit reads `work`, where that pair shows
  `converged: false`. The delivery handed the auditor a subtask "reviewed, NOT
  approved" the same way. Failing the build at the first stubborn review would
  leave the auditor, the one reader of the whole, no say. A coder that fails
  past its retry is absorbed the same way, and shows as a failed item.
- **The audit reads `work` in place of `plan`.** From the second round the
  subtasks are what the first audit raised, which the plan never held. `work`
  is what this round did, each item with its task.
- **The coder and the reviewer read `item.text`, not `item`.** The subtask
  arrives as text under its own heading rather than as JSON, and from the
  second round an item is an obligation, `{ id, text }`.
- **The reviewer reads `diff`.** Inside a copy, `diff` is the pair's own change,
  so the review weighs the change itself, not only the coder's summary of it.
- **A failed branch of `explore` or `split` is a failed report.** `on-fail:
  continue` sits on the agent node inside the `map`, not on the `map`, so the
  synthesiser reads a failed branch as `{ ok: false, error }`, as its prompt
  expects. A plan past `max:` still fails the run: that is the flow's bound,
  not a missing report.
- **A disagreement between reports is the synthesiser's to settle.** Its
  definition says to name one and say which report the code supports, and it
  holds `read` and `grep` for that. `explore`'s and `split`'s `answer` said to
  name it "rather than picking one", so one turn carried both rules and the
  model chose which to obey. The flows no longer mention disagreements: the
  rule lives in the definition, which reads whole on its own and reaches every
  flow and workflow that uses the agent.
- **The split planner's prose names the two workers.** `orchestrate` handed
  the planner the workers' descriptions; a flow hands it only its `reads:`, so
  the section says what a scout and a reviewer do.
- **Prose names the heading each read arrives under.** A turn with several
  reads holds several sections, and "the task below" pointed at none of them.
- **The words the pipelines asked for are gone.** "The reviewer approves with
  APPROVED" and `LGTM` were how prose was read; a verdict is a tool call now.
- **The shipped `build` names a script, not a command.** The pipeline shipped
  no check, since `npm test` imposed on a project that has none is worse than
  asking. A path imposes nothing that runs: each project writes its own
  `.pi/checks/test.sh`, and one that has none is refused at the run stage,
  before the first turn, with `check-script-missing` naming the path. An
  unattended build needs something to read "the tests pass" from, and the
  audit alone is an opinion. This repository's own runs `npm test`, and a test
  holds the shipped `build` to it through the run stage.
- **`build` ends with a report, not with its loop.** A run's answer is its
  last root node's output, and `deliver`'s is the loop's whole record as
  JSON: long, hard to read, and read in full by the main session's model. A
  `report` turn after it writes a few lines of prose from the request, the
  diff, the last round's pairs and the audit: what was done, in which files,
  and what is left. It is the `synthesiser`, which already merges several
  reports into one answer and has no tool that writes, so no new agent. The
  price is one more turn per build: 5s and 2.7k tokens of a 1m12s, 47k-token
  run on `ilaas/gemma-4-31b`. It runs only when the build passed; a build
  that fails has no report, and its end line says where and why. It reads
  the two fields of the loop it needs, the pairs and the audit, rather than
  the whole record. The end line loses its `converged`, which only a loop
  gives, and which told nobody anything: a build that passes has converged.
  `build-attended` hands the report to the committer, so what the build left
  undone reaches the commit body and the turn is not spent for nothing.

### A flow run measures itself in its run directory

One directory holds a flow run's state and its exports. The runner writes
the transcripts, since a subagent exports as it closes; `usage.json` is
written by a `measuredRun` subscribed to the run, which the runner never
opens.

- **The measurement stands beside `runFlow`, not inside it.** `runFlow` and
  `resumeFlow` given a `runDir` leave the snapshot, the journal and the
  transcripts. A `measuredRun` opened on the same directory writes
  `usage.json`, and `events.jsonl` with `record: true`. Unplugged, the run
  and its resume are unchanged. An experiment's cell runs a flow with its
  own directory and its own `onEvent` and gets the visits with nothing
  added, and the one writer of `usage.json` stays one. The command that will
  run flows owns the parent session, so it is the one to measure.
- **A taken name takes the first free `~n`.** It is not the life's number.
  A subagent that a timeout replaced in the same visit is a second file in
  the same place, as a later life's is. The name is chosen at the spawn,
  against the disk and the names this life already handed out, so nothing
  already written is overwritten, and a run never resumed keeps clean
  names. `main.jsonl` and `events.jsonl` take the same rule.
- **A delegate's transcript is placed by the flow, not by `delegateTool`.**
  The runner hands the `subagent` tool to an agent naming it, which the flow
  runner did not do before, with a spawn that puts each child in
  `<parent>.children/` beside its parent's files, theirs under them. A tool
  run through the extension keeps its flat layout. The roster is the agents
  the flow names, which the snapshot keeps, so a resume offers the same
  ones.
- **A `visit_end` says what it was.** It carries `node`, `kind` and
  `subagent`. A life killed before its `usage.json` is rebuilt from journal
  entries alone, which had no node or kind. And a subagent a scope keeps
  names only its first visit on its `spawn`, so its `visits` come from the
  ends.
- **Plan order is computed from the ends.** Each visit comes before the
  visits it holds, and visits under one holder come in the order they
  ended. A life folded from the stream and one rebuilt from the journal
  come out the same way, and only branches running together are in end
  order rather than in the order the file writes them.
- **The lives are read from the run directory.** This life is the
  journal's last `life_start` taken after the measurement opened, so a
  measurement whose run never started does not count one. The `usage.json`
  an earlier life wrote gives the lives it lists. Any life between it and
  this one is rebuilt from the journal and marked `partial`: its `run_end`'s
  usage when it wrote one, else what its outermost ended visits cost. A life
  ends `interrupted` when it wrote no end or was stopped. `total` is the
  sum of the lives, wall time included, and so is the top-level `wallMs`.
  With no journal there are no lives, and `visits` and `nodes` stand alone.
- **Each life counts what was measured of it.** This life's usage is its
  subagents', delegates included. A rebuilt one's is its visits', which
  hold no delegate's turns, since a visit's cost is the delta of its own
  subagent's session. The difference is written down rather than filled
  in.
- **`measure/` reads the journal past the flow's door.** `flow/index.ts`
  re-exports the runner, which spawns, and a subagent exports through
  `measure/`: going through the door would close a cycle. `lives.ts` sits in
  `measure/`, and the live view reads it through `measure/`'s door.

### The package exports the flow API, and `/flows` replaces `/pipelines`

The first visible step of the switch: `src/index.ts` exports flows, and pi
lists and plans them. At the time nothing ran a flow from pi, and `/build` and
`/run` still ran pipelines.

- **The root exports the door, minus what nobody outside calls.** Every
  function a script or the extension calls goes out, grouped the way a run
  goes: finding and checking a flow, the ports, running and resuming, drawing.
  The types those functions take and return go with them, the checked nodes
  and `Condition` included, since `CheckedFlow.nodes` names them. The values
  only the validator and the runner call stay in `flow/index.ts`:
  `compileCondition`, `evaluateCondition`, `readSchema`, `showType` and the
  code and kind lists (`FAULT_CODES`, `CONDITION_CODES`, `ANSWER_CODES`,
  `ERROR_KINDS`, `STOPS`) with the value types `CHECK`, `VERDICT`, `LEDGER`
  and `QUESTION`. A caller matching a fault reads its `FaultCode` type. The
  ports are `bashCheck` and `gitPort`, and `Asking` and `Shown` join
  `AskUser`, whose signature names them.
- **A catalogue's flow file says whose it is.** `FoundFlow` is a file with its
  `source`, as an `Agent` has one, so `/flows` gives the origin without
  guessing it back from a path. The snapshot keeps it with the file.
- **Only the user's and the repository's `pipelines/` are read, never the
  package's.** `removedPipelines` refuses each file there with
  `pipeline-format-removed`, whose message names the `flows/` directory beside
  it and links to [From pipelines to flows](guide/from-pipelines.md). The
  package's own pipelines go with the linear format, and a fault only we can
  fix is noise for whoever reads it. It is a function beside
  `loadFlowCatalogue` and not a field of the catalogue: a flow is never
  checked against these files, and a snapshot has none.
- **One line per flow, by name, and a refused file beside the one it would
  lose its name to.** Each line has the name, the source, the bound
  `showBound` writes and the description, in aligned columns; a broken flow
  or a leftover pipeline says `broken`, its faults under it as
  `file at: message`. `.pi/pipelines/build.md` sorts under the `build` flow,
  which is how the shadowing is seen. A count opens the listing, because pi
  prefixes a warning with `Warning:` and would push the first row out of its
  columns.
- **`/flows <name>` ends with what is left under that name.** A person who
  asks for `build` while an old `build.md` sits in `.pi/pipelines/` most
  likely meant that file, so its fault follows the plan.
- **A line is cut to the terminal, and goes on under itself.** pi's
  notification wraps at the terminal's edge back to column one, which broke
  every column of the listing and every level of the plan in the first real
  run. Each line carries the column its text starts at: a description goes
  on under the description, a plan line under its text, cut after a whole
  ` · ` fact when one fits. The width is `process.stdout.columns`, the
  terminal pi draws in; with none, nothing is cut. The working directory is
  left out of paths and the home directory written `~`, since the checked
  flow holds absolute ones.

### `/run` runs flows, and `/build` is `/run build`

The second step of the switch. `/run <flow> <input>` checks the flow at both
stages, runs it in `runs/<timestamp>/` and leaves its answer in the
conversation; `/run resume` carries one on; `/build`, its flags and
`build.json` go from the extension. `/step` and the `subagent` tool still take
pipelines.

- **One command, and the file says the rest.** The interview and the commit
  were what `/build` did beyond `/run`, and both are nodes of a flow now. So
  `/run` takes `--model` and `--timeout` and nothing else: which model you
  hold keys for and how long you will wait are yours, what the work is, its
  checks and its copies are the file's. `--check`, `--worktree` and
  `--pipeline` have no successor on the line.
- **The run stage runs with the terminal's ports.** The card is the `ask`
  port even with nobody there: `somebodyThere` is what keeps a run from
  showing it, and a refusal then says "launched with nobody there" rather than
  blaming a missing port. `bashCheck` and `gitPort` are the other two. The
  refusal outside a repository that `/build` made itself comes from the run
  stage now, naming the node that needs git, so a flow that needs none, such
  as `explore`, runs anywhere.
- **Every run has a run directory, and is measured there.** The live view is
  a `measuredRun` on the run directory, with this session's JSONL, so
  `usage.json` lands beside the journal and a resume adds its life to it.
- **The answer is the last root node's output, then one line on how it
  ended, then the last frame.** The line is `ok`, with `converged` when the
  last node is a loop, and the run directory; a failure says where and why,
  and what `/run resume` would do or why it cannot. The model reads the
  answer and the line. The frame is in the message's details, drawn by the
  renderer at the terminal's width and never sent to the model, since a
  column of glyphs costs its context and says nothing the line does not. A
  typed output is shown as JSON. The message keeps its old `customType`,
  so sessions that hold one still draw it.
- **The widget draws the plan, and pi's own cap would cut it.** pi shows at
  most ten lines of a widget given as lines, and the plan of `build` is taller
  than that while one pair works. The widget is a component instead, painted
  at the width pi gives it, and the plan takes sixteen rows at most: cut above
  and below what runs now, each cut saying how many lines it holds, since the
  finished lines are the ones to lose first. A running visit's subagents hang
  under it with what they are doing, from the pieces the plain widget is
  built of, and the visit line no longer names them twice. Paths leave out the
  working directory and the package, so a shipped agent reads
  `agents/coder.md`.
- **`/run resume` picks the newest run that can go on.** `latestResumable`
  reads each run directory under `runs/` newest first, keeps those whose
  snapshot was taken in this directory, and takes the first `resumePoint`
  accepts. When none would, it says why the newest cannot, since that is the
  one somebody means. It names the run and the visit before it starts, so the
  line is there while the first turn is not. A flow named `resume` would be
  shadowed by it, so `reserved-name` refuses the file.
- **Flags may end the line too, for `/run`.** A known flag written after the
  text used to be read as text, and ran three scouts on the wrong model while
  one of them grepped the repository for the model's name. A line does not
  end on `--model <pattern>` as prose, so `parseFlags` reads `--model` and
  `--timeout` there too; in the middle of the text they stay text. An input
  written as one quoted string is the text inside, and a request ending on
  its own `?` is not given a second stop in the framing line.
- **herdr names a flow's split by agent and home.** A spawn event carries the
  subagent's home, its memory scope's path or its visit's, and the split is
  `coder @ deliver#2/work[1]/pair`: a flow's ids count spawns across the whole
  run, and `coder#3` says nothing of which pair it codes for. A `check`, an
  `ask` or a `commit` spawns nothing, so it opens no split.

### `/step` and the `subagent` tool take flows

The third step of the switch. A stage of `/step` is a flow or an agent, and
the `subagent` tool runs a flow by name. Nothing in the extension reads a
linear pipeline any more.

- **One launch, three callers.** `/run`, a flow stage of `/step` and the
  tool check a flow at both stages, hold it to the terminal and run it under
  the plan's live view, measured in its run directory. That is
  `commands/launch.ts`: `launchable` returns the checked run or the stage
  that refused it with its faults, and `launch` runs it. How a refusal is
  worded and what becomes of the result stays with each caller, since a
  command notifies, a step records and the tool answers a model.
- **The question card is shown during the model's turn.** `ask` cards go
  through `ctx.ui` from the tool's `execute`, as pi's own `question.ts` and
  `questionnaire.ts` examples do, and somebody is there only in pi's
  terminal, `ctx.mode === "tui"`. A call with nobody there gets the run
  stage's refusal, before anything is spawned. Checked in a real pi before it was written down: the
  card came up while the tool's row said `0/1 done`, `enter` answered it and
  the loop went on to its next `ask_next`, and `esc` on the card was the card's
  "that's enough", not pi's interrupt: the brief was written and the turn
  went on. A typed answer through `Other…` went to the card's text box, not
  to pi's editor, and reached the brief.
- **RPC mode is nobody there.** It was first read off `ctx.hasUI`, which
  RPC mode sets too, since its dialogs go to the client. The card is drawn
  with `custom()`, which returns `undefined` there, and a card nobody saw read
  as declined: in `pi --mode rpc` a tool call on a flow whose `ask` has
  `default: true` failed at once with `stopped: the run was stopped at this
  question`. Now the `ask` takes its `default:`, its `enough:` or fails with
  `nobody`, as in `pi -p`. `/run` and `/interview` read the same answer.
  Putting the card through RPC's `select` and `input` instead is a second
  card, and waits for a client that needs it.
- **The tool runs one call at a time.** Two `subagent` calls in one message
  ran side by side, and a flow's cards, its plan above the prompt and its stop
  key are the terminal's, one of each: measured in RPC mode, two flow calls
  started together and one died on the other's run directory. The tool says
  `executionMode: "sequential"`, pi's own knob, rather than a queue of our
  own. pi then runs every call of that message in order, other tools
  included; a model that wants subagents side by side has `parallel`.
- **A codemode script cannot call the tool.** pi 1.0's `codemode` lets the
  model write a script that calls the session's tools, and with the default
  exposure that included `subagent`: measured in a real pi with codemode on, a
  script saw `"subagent" in tools` as `true`. A loop in one script could then
  start any number of runs inside one model turn, past the one-call rule above
  and with none of the tool's rows drawn. The tool says `exposure:
  "model-only"`, pi's term for a tool that orchestrates others: declared to the
  model, never callable from another tool.
- **The widget stops offering `esc` while a card is up.** The first frame of
  a card mid-turn read `esc stops everything` right above `esc That's
  enough`, and the card holds the key. The hint comes back when the card goes.
- **`flow` takes `task`, `model`, `timeoutMs`, `scope` and `herdrAll`, and
  refuses the rest by name.** The file says what runs, so `steps`, `until`,
  `candidates` and the like would be ignored, and a model that believed they
  were used reads a wrong answer as a right one. `lifetime`, `maxDepth`,
  `openInHerdr` and `export` go with them: memory is the file's, a flow run
  always has its run directory, and a field that changes nothing is a
  question the refusal answers once. `mode` is taken when it says `flow`.
- **The tool's flows follow `scope`, like its agents.** A flow names agents
  and project scripts, so a repository's `.pi/flows/` is read only when the
  call asks for `project` or `both`; the package's flows and the user's are
  always there. The commands read `both`, as they always have.
- **The model reads the output and the line on how the run ended.** The
  line names the run directory, and on a failure what `/run resume <run
  directory>` would do. The tool does not resume: a resume is a decision about
  a run that stopped, taken by whoever reads why it stopped. The row draws
  the line and the last frame, the one `/run` draws under its answer.
- **A flow stage runs in the step's folder.** It is the run's directory, so
  a stage that stopped can be carried on with `/run resume <folder>`, and the
  step says so where an agent's failure says where its transcripts are. The
  chain does not learn what a resume finishes: the answer lands in the
  conversation, as `/run`'s does, and a step recorded behind the chain's back
  would be the kind of state the relay exists to keep in view.
- **Flows first, then agents, and a broken flow is refused.** A file of that
  name that does not check is the likeliest thing the person meant, so it is
  not fallen past to an agent of the same name. `--agent` still forces the
  agent.
- **`/step --worktree` goes.** Whether a flow's workers get copies is the
  file's, `copies: true` on a block, as it is for `/run`.
- **The relay writes a step's input itself.** `## Request`, then `## Output
  of step <id>`, as before, but no longer through the pipeline's
  `stepInput`, so the linear pipeline can go without touching the
  extension. A step's kind is `flow` now; an entry an older session wrote as
  `pipeline` still draws.

### The linear pipeline is removed

The fourth step of the switch. `src/pipeline/`, `pipelines/` and its links in
`.pi/pipelines/`, `docs/guide/pipelines.md` and their tests are gone, and the
package's `files` no longer lists `pipelines`. A pipeline written down is a
flow now, and [From pipelines to flows](guide/from-pipelines.md) is where the
old format is still described, beside the flows that replaced it.

- **`deliver`, `pair` and `audit` retire, with `settle` and `resume`.** They
  were pipelines written in TypeScript, and keeping them beside the `build`
  flow would be two implementations of one shape, of which only the flow has
  a journal, restored copies and a script as its check. Nothing else used
  them: `orchestrate`, `plan` and `swarm` stand on the pool and the parser,
  and the flow runner on `review/`, `git/` and `mapConcurrent`, which all
  stay. The generic combinators stay public.
- **`examples/13-concurrent-writers.ts` is rewritten, not kept on `pair`.** It
  was the one caller left, and what it shows, two writers in two copies and
  their patches landed one at a time, is `scratchWorktree`, `run` and `land`,
  the primitives a flow's `copies: true` stands on. It loses the reviewer,
  which the `build` flow has.
- **`Verify` and `commandVerifier` go, and `land` loses `verify`.** Its own
  header said `Verify` stayed for the linear pipeline only, and after it no
  caller passed one. A flow checks the tree with a `check` node after the block, not
  between two patches.
- **The review record's public types narrow to `Obligation` and `Closure`.**
  Those are what a `JournalEntry` names. `Verdict`, `Ledger`, `ReviewRecord`
  and the rest were named by `pair`'s and `audit`'s results, and a type no
  public signature names is off the list.
- **`/flows` keeps its scan of old `pipelines/` directories, with no parser.**
  `removedPipelines` reads file names, never a file's content, so nothing of
  the linear format is needed to refuse it. `definitionDirs` takes no package
  directory for a kind the package does not ship, rather than one that must be
  switched off.
- **`examples/11-build.ts` runs the shipped `build` flow.** `checkFlow`, then
  `checkRun` with `bashCheck`, `gitPort` and nobody there, then `runFlow` in a
  run directory of the target repository. It needs the target's own
  `.pi/checks/test.sh`, and is refused before any model runs without it.
- **The tutorials keep their frames until they are recaptured.** A frame is
  captured, never composed, so each page that shows the linear format says so
  in a note, and its prose stops describing it as current.

## A chain walked by hand

`/run explore …` put its answer in the conversation, and the session picked it
up and started orchestrating: every later command was chosen against a
conclusion the main window had already drawn. That is the documented behaviour
of `/run` working exactly as decided - an exploration is read and then asked
about - and it is the wrong default for the other use, where the main window is
a console and the chain is `explorer → planner → coder → reviewer` advanced one
step at a time. So a knob, not a reversal: `/step`, `/chain`, `/quote`.

- **The output goes to a relay, not to the conversation.** `appendEntry` draws
  the step in the transcript and keeps it out of the model's context, and the
  text waits in module state for the next command. `/quote` is the single door
  into the conversation, taken on purpose, with the same framing `/run` uses -
  pi turns a custom message into a **user** message, and an unattributed report
  in that slot reads as an instruction.
- **A `Result` crosses, never a context.** Which is the lifetime rule already:
  persistent subagents do not share history, you pass `Result`s. Each `/step`
  opens a subagent and closes it, so there is no live session to own between two
  commands and nothing to leak if pi is quit in the middle. Carrying a
  conversation would also have made `--model` per step meaningless.
- **Replaced: the relay writes the same two sections itself**, and a flow
  stage reads them as its `input`; see [`/step` and the `subagent` tool take
  flows](#step-and-the-subagent-tool-take-flows).
  **A step is handed `stepInput`, the same three sections a pipeline's steps
  get.** A chain walked by hand and the same chain written down then send the
  model byte-for-byte the same thing, which is the only way the two can be
  compared. A first step carries nothing and is passed through verbatim, as
  `/run` passes its request: a lone instruction under a `## Request` heading is
  noise.
- **Replaced: a step is a flow or an agent**, flows first, a flow stage in a
  run directory of its own; see [`/step` and the `subagent` tool take
  flows](#step-and-the-subagent-tool-take-flows).
  **A step may be a pipeline or an agent**, resolved in that order, because
  `explore` - a fan-out and a synthesis - is a perfectly good stage of a chain
  and so is a lone `planner`. A name held by both runs the pipeline and says so,
  and `--agent` is there so that the collision is not a dead end. A name held by
  neither is one message, not two listings.
- **A failed step leaves the chain untouched.** It produced no material, and
  recording it would hand the next agent an error message as its input. The
  export stays on disk and the same command can be retried on another model,
  which is the whole reason a step is typed rather than walked.
- **One folder for the chain, one subfolder per step.** A chain walked by hand
  is still a run, and leaves the same trace as one walked by `/run`. What
  `/chain` totals is the sum of the steps' own wall time, not the age of the
  chain: most of a hand-walked chain is spent waiting for a human, and counting
  that as work would be an estimate.
- **`/quote` and not `/share`**, which is what it was called until a real pi
  said so at startup: `share` is one of pi's own twenty-three built-in slash
  commands, and an extension command of that name is dropped from its own
  autocomplete. No test here could have caught it - the fake `pi` a test hands
  the extension has no built-ins to collide with. `quote` is also the better
  word: what lands in the conversation is a quotation, attributed and read as
  one.

## Measurements: time and tokens per subagent

Nothing is estimated, nothing is recomputed by hand: pi already exposes the
numbers, we **collect and attribute** them. The only things the library adds are
**time** (pi does not measure it) and **aggregation per subagent**.

```typescript
type Usage = {
  // time - measured here, on a monotonic clock (performance.now()), not Date.now()
  wallMs: number;        // from spawn to close (includes waiting between two asks)
  busyMs: number;        // time actually spent working (sum of the asks)
  turns: number;

  // tokens & cost - reported by pi, never reconstructed
  input: number;
  output: number;
  cacheRead: number;
  cacheWrite: number;
  cost: number;
  contextTokens?: number;  // current context size (persistent agent)
};
```

Where the numbers come from:

- `session.getSessionStats()` → `{ tokens: { input, output, cacheRead,
  cacheWrite, total }, cost, contextUsage, userMessages, assistantMessages,
  toolCalls, … }`. This is **the** source of truth for tokens and cost.
- `session.getContextUsage()` → context occupancy, to display for persistent
  agents.
- Time is measured around each `session.prompt()`: `busyMs` is the sum of the
  `ask` calls, `wallMs` runs from `spawn` to `close`. On a `"task"` agent the
  two are nearly equal; on a `"workflow"` agent the gap between them **is** the
  interesting information (waiting time vs useful time).

Rules:

- **`getSessionStats()` is cumulative over the session.** A turn's usage is
  therefore the **difference** between two snapshots, taken before and after
  `prompt()`. That is what gives both `subagent.usage` (since spawn) and
  `result.usage` (this turn) without ever recounting a token.
- Counters are clamped at `0`: compaction can walk the totals backwards, and a
  negative usage means nothing.
- **A fan-out aggregates**, it does not average: total tokens, total cost,
  `wallMs` = duration of the fan-out (not the sum of the branches), `busyMs` =
  sum of the branches. The ratio of the two gives the real parallelism - that is
  what we want to see.
- **A failure counts too.** A subagent that crashed after 12k tokens cost 12k
  tokens; its `Usage` is filled in even when `ok: false`.
- **A refused turn does not.** Asking a subagent that has already been stopped
  returns without reaching the session, so `turns` stays where it was: there was
  no request and no answer, and a turn nobody made is the one kind of number
  this section exists to keep out. A failure that *ran* still counts, which is
  the line between the two.
- **Never** estimate tokens by counting characters. If the provider does not
  report them, the field is `0` and we say so.

### Usage does its own arithmetic

`sumUsage` existed, and four places did not use it. The subagent's cumulative
usage added a turn field by field; `usage.json` flattened its total into a
record by naming the fields, leaving two out; the experiment report summed
those records key by key and turned them back into a `Usage` by naming the
fields again; and the hand-walked chain summed its steps in a `reduce` of its
own. A tenth field on `Usage` would have needed six edits and would have
vanished from the experiment table without a test saying so, and the clamp at
zero lived in `deltaUsage` alone.

The arithmetic is `usage.ts`'s: `deltaUsage` for a turn, `accumulate` for a
subagent's life - the one rule `sumUsage` does not have, that a context is a
level replaced by the latest reading rather than a counter summed - and
`sumUsage` for several subagents. Nothing outside the file names the field
list. `usage.json`'s total is a `Usage` with the subagent count beside it, and
its `wallMs` is the run's; the experiment summary is a `sumUsage` over the
cells with their wall times added, because they ran one after another, and its
own `wallMs` field went into the total rather than standing beside it as a
second copy. `asUsage` is gone with the record it read back.

## Session export

Two formats, two uses, both provided by pi (`AgentSession`):

```typescript
await session.exportToHtml(outputPath?);  // → path of the HTML file, readable/shareable
session.exportToJsonl(outputPath?);       // → JSONL of the current branch, replayable
```

Implemented in `src/measure/export.ts`, wired into `spawn` and every workflow
through `exportDir`.

- **An export covering the parent session *and* all its subagents.** An
  orchestration export that lost the subagents' work would be useless. What lands
  in `runs/<timestamp>/`: one `<agent>-<n>.html` / `.jsonl` pair per subagent,
  `main.jsonl` for the parent, and a `usage.json`.
- **`main.html` is not there, and will not be.** pi's HTML renderer is a method
  of `AgentSession`; an extension only ever gets a `ReadonlySessionManager`
  (`ctx.sessionManager`), and `exportFromFile` is not re-exported from the
  package root - the `exports` map blocks a deep import. So we copy the parent's
  JSONL, which pi's own `pi --export <file>` turns into the same HTML on demand.
  Writing our own HTML would break "we reimplement nothing".
- **`main.jsonl` is written from what pi holds when pi has no file yet.**
  pi creates a session's file with its first assistant message, so the
  first command typed in a fresh pi runs in a session whose file does not
  exist, and `--no-session` never writes one. Copying the path alone left
  `main.jsonl` out of those runs, seen in real ones. A run is handed the
  session itself (`MainSession`, the part of `ctx.sessionManager` it reads)
  rather than its path, and reads it when the run ends. pi's file is copied
  when it exists; otherwise the header and the entries pi holds are written
  one JSON object per line, which is all a session file is. Nothing is
  rendered and nothing is added, so this is not a second JSONL writer.
  Waiting for the file was the other way, and a parent that never replies
  never writes one.
- **`exportDir` implies a session directory**, `<exportDir>/.sessions`.
  `SessionManager.inMemory()` persists nothing, and pi answers a request to
  export one with `Cannot export in-memory session to HTML`. Asking for an
  export *is* asking for the session to be kept long enough to export it - one
  decision, not two. `sessionDir` stays available for anyone who wants the
  working files elsewhere. Neither is a default: with no `exportDir`, a subagent
  leaves nothing behind, not in `~/.pi` and not in the working directory.
- **JSONL and HTML are attempted separately.** An in-memory session still yields
  its transcript even though pi refuses to render its page; losing both because
  one is impossible would be a poor trade.
- **Export can be triggered at any time**, not only at the end of a workflow:
  `subagent.export(dir)` works on any live subagent. `close()` exports first and
  disposes after, so the workflow's `finally` - cancellation included - is
  already the "export what was done" path. What is written on the interrupted
  path is what makes this feature worth having.
- **An export never throws.** Every failure is a string in `SessionExport.error`.
  An export is an observer of the run, and an observer that takes the workflow
  down with it is a bug - most of all when it runs on the way out of a crash.
- **`usage.json` is the only artefact we produce ourselves**; we reimplement
  neither pi's HTML nor its JSONL. It is built from the same `TuiSnapshot` the
  TUI draws (`usageReport(snapshot, wallMs)`) - one collected state, two
  consumers - and it carries what pi cannot: time, attribution per subagent, and
  `parallelism` (busy over wall).
- Not to be confused with `pi --export <file>` (CLI, on an existing session
  file): useful when debugging, and the way to render `main.jsonl`.
- **`runs/` gets a `.gitignore` of `*`, written once.** The exports land inside
  the repository the run is working on, and git has no business with them.
  Nobody had hit that here because *this* repository ignores `runs/` in its own
  `.gitignore`: the defect only shows in somebody else's. Two ways it shows, both
  measured in a real run. A delivery given `worktree: true` cannot put its
  subtasks' patches back, because `land` refuses a tree that is not clean and
  `runs/` alone makes it unclean. And `/build` commits with `git add -A`, which
  sweeps a run's transcripts into the user's history.

  Relaxing the landing was the other candidate and a test refused it: a delivery
  lands twice, and the second time it has to recognise the work the first one
  left, which is precisely untracked files. So the export gets out of git's way
  instead of git being told to look away. A `.gitignore` already in `runs/` is
  left alone - the directory is the user's the moment they have said anything
  about it.
- **A run directory is suffixed `-2`, `-3` when its second is taken**, and
  created exclusively. A timestamp to the second alone gave two `subagent` tool
  calls started in the same second one directory, and the second flow died on
  `EEXIST` over `snapshot.json`: a run directory holds one run. A suffix only on
  collision keeps the common name as it was, and reads like the run's branch
  does. Finer timestamps or a random part make a collision rarer, not
  impossible, and the names harder to read.
  The names sort by start with a numeric collation (`newestRunFirst`), which is
  how `/run resume` finds the newest.

## Display: herdr if present, pi TUI otherwise

One event stream, several reporters. The core emits:

```typescript
type SubagentEvent =
  | { type: "spawn";  id: string; agent: string; lifetime: Lifetime }
  | { type: "status"; id: string; status: "working" | "idle" | "blocked" | "done"; task?: string }
  | { type: "text";   id: string; delta: string }
  | { type: "tool";   id: string; name: string; args: unknown }
  | { type: "usage";  id: string; usage: Usage }
  | { type: "close";  id: string; result: Result };
```

A listener that throws is swallowed: a broken reporter must never take a
workflow down with it.

**The task rides on the `"working"` transition**, not on `spawn`: at spawn time
nobody knows yet what the subagent will be asked, and a persistent subagent is
asked several different things over its life. A reporter has no other way to
learn it - and until it did, every collapsed row in the TUI showed a blank task.

### herdr reporter (the default when available)

Implemented in `src/reporters/herdr.ts`, transport in `herdr-client.ts`.
Verified against herdr 0.9.0, protocol 22, by `node scripts/check-herdr.ts` and
by a run with panes open.

**The key idea, because it is not the obvious one.** A herdr pane cannot *host*
an in-process subagent: there is no process and no TTY to attach. So the pane
does not host the subagent - it **displays a stream we write**. We append to a
file and open a pane running `tail -n +1 -f` on it, which takes three calls
because herdr has no single call that does all three:

```typescript
pane.split      { direction: "right", target_pane_id: ours, focus: false }   // → result.pane.pane_id
pane.rename     { pane_id, label: "reviewer#2" }
pane.send_input { pane_id, text: "exec tail -n +1 -f '<log>'", keys: ["enter"] }
```

`target_pane_id` is ours, from `HERDR_PANE_ID`: with none, herdr splits whatever
pane is focused, which can belong to another client.

`exec`, so the pane *is* the stream - closing it closes what it follows - and
the clear that puts the name at the top is written into the log file rather than
run as a command, because the shell is still starting up and writes over
anything printed before it has finished. Measured, as a zsh history warning
sitting on top of a member's first turn.

This also settles the ownership question: a pane carries exactly **one** `agent`
/ `agent_status`, and the main pane's already belongs to herdr's own pi
integration (source `herdr:pi`, installed at
`~/.pi/agent/extensions/herdr-agent-state.ts` - read it, it is the reference
implementation). The panes we open are ours, so we report on those and never
fight for the main one.

**Detection**: `HERDR_ENV === "1"` **and** `HERDR_SOCKET_PATH` **and**
`HERDR_PANE_ID`. All three, or nothing. Missing them is the normal case outside
herdr, so we fall back silently: **never an error, never a noisy warning**.

**Transport**: unix socket, newline-delimited JSON, one connection per request,
resolve on the first `data`, then `destroy()`. First attempt at 500 ms, one
retry at 1500 ms, then give up quietly. Envelope `{ id, method, params }`.

Points that cost time to discover:

- **`PaneAgentState` has no `done`** - it is `idle | working | blocked |
  unknown`, even though the *event* `AgentStatus` enum does have `done`. Our
  `"done"` maps to `idle`, and `pane.release_agent` is what actually retires the
  agent.
- **`pane.split` answers with the `pane_id`** (`result.pane.pane_id`), which is
  the only way to later name, run in, report on, release and close that pane.
- **`agent.start` is not that call**, whatever its name suggests: it starts a
  *recognised* agent in a pane that already exists. `pane.split` is what makes
  panes. See the reversal below.
- **Release before close.** Closing first leaves herdr holding an agent on a
  pane that no longer exists.
- **Every promise chain ends in a `catch`.** A try/catch around the listener is
  not enough: an unhandled rejection escapes it entirely and takes the process
  down. Found by a test, not by reading.
- `seq` is a monotonic ordering field; seed it from the clock like the pi
  integration does, so two processes reporting on one pane do not collide.
- Useful CLI equivalents when debugging: `herdr pane list`, `herdr agent list`,
  `herdr pane read <pane_id> --source visible`. The read sources are
  `visible | recent | recent_unwrapped | detection`, with an underscore - the
  CLI's own `--source recent-unwrapped` is spelt the other way and the socket
  refuses it.

**`openInHerdr` is opt-in, per subagent**, exactly like `lifetime`: a fan-out of
twenty branches must not carpet the screen unless someone asked. The other
regime - **watch everything** - belongs to the reporter, not to the core:
`createHerdrReporter({ all: true })`, toggled for a session with `/herdr on`.
It was also seeded by a `COMBO_HERDR` environment variable, which is gone: the
same rule that removed `COMBO_MODEL` applies to display. Configuration is an
argument or a command, never ambient state. Who gets a pane is a display decision,
and the workflow runs identically either way; putting it on the spawn would have
meant threading a flag through every call site to change what a terminal shows. It travels on
the `spawn` event rather than being read back from the core - a reporter is a
pure observer, it never queries anything.

The split **closes automatically** when the subagent closes, after the final
usage line is written. No orphan panes after a fan-out.

**A board gets a pane, and it is not an agent.** A pane per member shows each
one working and none of them talking: what a member said is in its own pane, and
who it was addressing is only legible where all of them are. So the first thing
anybody says opens one more pane, named `board`, carrying the exchange in order,
and each member's pane keeps its own half of it without repeating its own name.
Nobody works in that pane, so no agent state is ever reported on it and nothing
is released when it closes - a pane that never had an agent has none to give
back, which is now what `finish` checks rather than assumes. It opens only when
the run is watched at all, and closes with the last member.

**Reversed: `agent.start` was never the call that opens a pane.** Under herdr
0.7.3 it took an `argv` and a `split` and did open one. By 0.9.0 it takes a
`kind` and a `pane_id` and starts a recognised agent in an existing pane, so
every request combo sent was answered `invalid_request`. Nothing said so: a
reporter must never throw, so the refusal was swallowed, `paneIdOf` gave
`undefined`, and every call after it was skipped by design. `/herdr on` warned
only about herdr being absent, which it was not. The suite stayed green over a
feature that opened nothing at all, because every test of it asserted on combo's
own idea of the request.

The guard for it is not another test. `node scripts/check-herdr.ts` runs the
reporter through a subagent's whole life against a recording transport, then
holds every request it made to `herdr api schema --json`. Because it validates
what the code sends rather than a copy of it, it cannot drift from the code the
way example payloads would; run against the shape that shipped, it names both
missing fields. `scripts/drive-pi.py` plays the same part for pi: a check a fake
cannot perform, run by hand.

**`/herdr on` asks herdr, and says what it heard.** Being inside herdr was the
only question it used to ask, and it is the wrong one on its own: a herdr that
answers `invalid_request` to everything is indistinguishable from a working one
until a run opens no panes. Nothing in the reporter may say so, because a
reporter that warns is a participant, so the command says it instead.

The probe is harmless because of the pane it names. `pane.split` sent with a
`target_pane_id` no herdr can have separates the two answers: `invalid_request`
when the request itself is wrong, `pane_not_found` when it was understood and
only that pane was missing. Neither opens anything, and if some later herdr
splits anyway, the probe closes what it got. It asks about `pane.split` alone,
because nothing else runs without the pane id it returns;
`scripts/check-herdr.ts` covers the other five, against the schema rather than a
live server. The params come from one function (`splitParams`), so a probe
cannot go on answering yes about a request the reporter no longer makes.

The preference is set whatever the answer. What a user wants for the session is
not herdr's to decide, and a `/herdr on` that silently declined to be on would
be a second way to be surprised by an empty screen.

Two tests asserted that the suite runs outside herdr. It does not, for anyone
developing this in the window it is written for: both passed for the wrong
reason there, and one of them opened a pane in the terminal running `npm test`.
`test/fixtures/no-herdr.ts` takes the markers out of the environment for the
length of a body, so the suite answers the same either way.

The wording of those lines lives in `src/reporters/traffic.ts` and nowhere else.
The console and a herdr pane are read side by side when a run is compared
against another, and a difference in how they say it would read as a difference
in what happened.

**Reversed: a subagent's pane hosts a client, not a file.** The `tail -f`
above was the right answer to a pane that cannot host a subagent, and the
wrong shape for two things asked of it since: to read like pi, and to take a
word for the subagent. A text file can do neither. The split now runs
`pane/main.ts`, a client of the mirror, so what it draws is pi's chat and what
is typed into it reaches the turn in flight. The three herdr calls are the
same three, with a different command typed into the shell; the file, the
one-line tool summary and the usage line the reporter used to write are gone
from it, because the pane draws them from the session itself. Text, tool and
usage events still flow on the bus for every other reporter, and the herdr
reporter reads only `spawn`, `status` and `close` from it: what herdr must be
told, and nothing the pane already knows.

The command names the binary this process runs on (`process.execPath`) and
the client where the package keeps it, both quoted, because the split's shell
may find another `node` or none, and a pane opening on a usage error says
nothing about why. When this process is not node at all, the shell's `node` is
the one candidate left.

The board keeps its file. Nobody works in it and nothing is typed to it, and
its lines are the console's, which is the property worth keeping there. A
member's own pane no longer repeats its half of the traffic: the tool calls
it made are drawn as tool calls, which is how pi would show them. The console
keeps the `⌨` line for a steer, and `traffic.ts` does not carry it, because
the pane draws a steer as the user message it becomes and two displays of one
run must not say it twice.

**A reporter that nobody subscribes reports nothing.** `onEvent` takes a single
listener, so watching in the TUI *and* in herdr means composing them - use
`combineReporters(collector.reporter, createHerdrReporter())`, which drops the
`undefined` that `createHerdrReporter` returns outside herdr. The extension once
passed `openInHerdr` all the way to the `spawn` event with nobody listening, and
nothing failed: no split, no error, no clue.

### The mirror: a live session on a socket, and what a typed word may do

A pane cannot host an in-process subagent, and the `tail -f` above shows one
without letting anybody speak to it. What a pane *can* host is a client: the
mirror (`src/mirror.ts`, wire in `mirror-wire.ts`) registers every live
subagent by id on one unix socket per process, replays `session.messages` to
whoever attaches, then forwards pi's own session events as they come. pi's own
chat components are written against exactly those events, so a client built
from them draws the session the way pi would.

The mirror is **not a reporter**. A reporter reads the stream and never reaches
the session; this one takes `steer` and `abort` from the socket and applies
them. It is a port of the core, beside `ask.ts` and `verify.ts`, and it is
reached only by the act that opens a pane. Registration is a map entry, so
every subagent registers and no socket exists until something asks for its
path: the suite spawns hundreds of fakes and listens on nothing.

**Nothing typed while idle is queued.** Probed, on two pi generations with
the same answer: a `steer` queued on an idle session is delivered with the
next `prompt()` and answered *in place of it* (`user SECOND → user STEERED →
assistant STEERED`); a `followUp` queued while idle is answered *inside* the
next `prompt()`, after the task, so `lastAssistantText` reads the person's
exchange back as the workflow's result. Either silently changes a turn the
workflow believes it composed. So a steer is accepted only while
`session.isStreaming`, refused otherwise with a reason the pane can print, and
`followUp` is not on `SessionPort` at all. Mid-turn a steer lands after the
tool call in flight and the transcript `ask` returns carries it, which is the
feature.

**A steer is an event.** `{ type: "steer", id, text }` goes on the bus when
one went through, so `events.jsonl`, the console and the pane say a person
spoke and where. A run somebody steered is not the run they would have got by
watching, and the record is what tells the two apart afterwards. The line is
`traffic.ts`'s (`⌨ scout#1 ← …`), for the reason every line there is: two
displays of one run must say the same thing.

**The socket path does not ride on the `spawn` event.** The plan had it
there, beside `openInHerdr`; it is a process-wide fact and not a subagent's,
so `mirrorSocket()` answers it once and the same path serves everyone alive.
A reporter asking a module for a path is not a reporter querying the run.

`attached` names the pi package this process resolved
(`import.meta.resolve`), because the events are that pi's and the components
reading them have to be too - the version trap above, seen from the other side.

### The pane is pi's chat, fed by the mirror

`pane/` is the client a herdr split runs: a `TUI` over a `ProcessTerminal`,
a chat of `UserMessageComponent`, `AssistantMessageComponent` and
`ToolExecutionComponent`, an `Editor`, two lines of footer. The chat runs the
switch pi's own interactive mode runs over the same events, which is why it
looks like pi: there was no drawing to invent, only a session to be fed.

It sits at the top level like `extension/`, because it is pi UI code and
imports `@earendil-works/pi-tui`, which `src/` never does: the TUI collector
stays free of it so a snapshot can be tested without a terminal, and so does
this, on `render(width)` of the components themselves.

**It imports pi statically**, from the same tree as `src/`. The plan had it
import the package from the URL `attached` carries, so the components would be
the version that produced the events; but `pane/` and `src/` resolve the bare
specifier the same way from the same checkout, so the URL adds nothing here and
a dynamic import would cost the types. The URL stays on the wire as a fact a
client may compare against its own resolution, and nothing reads it yet.

**pi's editor theme is not exported.** `getEditorTheme()` lives in pi's theme
module and `index.ts` does not re-export it, and `theme` itself is not exported
either, so the pane's editor border is plain dim rather than the theme's
`borderMuted`. The select list inside it is the theme's, through
`getSelectListTheme()`, which is exported.

Verified in a pty against a real subagent: the task, the thinking, each `read`
in its box, the answer; a line typed one second into the turn shown as a user
box, the model's answer to it below, and `Result.output` reading `STEERED` on
the library's side; a line typed after the close answered `done - nobody is
listening`. The frame was drawn with `scripts/frame.py`, which is the check no
fake can do.

### pi TUI reporter (always available)

Implemented across `src/reporters/tui.ts` and `extension/index.ts`.

**The split that makes it testable.** `tui.ts` *collects* - it turns the event
stream into a `TuiSnapshot` and formats strings, with no pi-tui import. The
extension *draws* - it maps tool arguments onto combinators and builds
components. Collection is therefore tested by inspecting a snapshot, never by
scraping a terminal, and the same state would feed a web view or an export
without touching a component.

The rendering is still tested, though: `test/extension.test.ts` captures the
registered tool, calls `renderCall` / `renderResult` with a full `Theme`, and
reads the component's own `render(width)`. That catches the failure that
actually bites - a renderer that throws makes pi fall back to its default
rendering silently, and nobody notices until the demo.

`Theme.fg` throws on an unknown colour, so a partial stub fails on the first
unusual colour rather than on a real defect: `test/fixtures/theme.ts` builds a
complete one. Do not reach for pi's internal `theme` singleton - it is not
exported from the package root.

The extension registers the tool with `renderCall` / `renderResult` (see
`docs/extensions.md`, *Custom Rendering*) and composes with
`@earendil-works/pi-tui` (`Container`, `Text`, `Markdown`, `Spacer`).

Load it with `pi -e extension` (the flag accepts a directory), or `pi install
./extension` to add it to settings.

Display specification - this is the "Claude Code" bar we are aiming at:

- **A dot per subagent above the prompt**, via `ctx.ui.setWidget(key, lines)`
  (`aboveEditor` is the default placement). `●` while it works, `✓` when it
  succeeded, `✗` when it failed, coloured by status; then a dimmed line with
  model, tokens and time. This is the Claude Code shape, and it is the *live*
  view - the tool row below holds the record.
  - The widget **disappears as soon as the work ends**, in a `finally` so a
    thrown workflow does not leave a dead row of dots above the prompt.
  - `widgetRows()` says *what* each line is and applies no colour, so the layout
    is testable without a terminal; the extension paints it.
  - **Events alone are not enough to keep a clock.** `usage.busyMs` only lands
    when a turn ends, so a widget reading it would show `0.0s` for the whole
    wait and then jump to the total. `SubagentSnapshot.startedAt` gives a live
    figure, and the extension repaints on a 250 ms tick - a subagent thinking
    for twenty seconds emits nothing, and a frozen clock reads as a hung agent.
- **Collapsed tool row, compact, one line per subagent**: status icon
  (`⏳` / `✓` / `✗`), agent name, truncated task, last tool called.
- **Streaming**: you see tool calls arrive, not an opaque spinner. Handle
  `isPartial`, call `context.invalidate()` sparingly.
- **In parallel, everything advances at once**: `2/3 done, 1 running`.
- **Expanded view (`app.tools.expand`)**: full task, all tool calls formatted
  (`$ cmd`, `read ~/path:1-10`, `grep /pat/ in ~/path`), final output rendered as
  **Markdown**, usage per subagent.
- **Usage line**: `3 turns 12.4s ↑12k ↓2.1k R8k $0.0412 ctx:34k model` - and,
  for a persistent agent, the **cumulative** usage since its spawn.
- **End-of-workflow summary**: a table with one line per subagent (time, turns,
  tokens, cost), a total at the bottom, and the export directory path if an
  export was requested.
- Use `keyHint("app.tools.expand", …)` rather than hard-coding "Ctrl+O": the
  user's key configuration must be respected.
- Reuse `context.lastComponent` instead of rebuilding the tree every frame.

### A figure on a row is one somebody measured, and a row fits its terminal

Recapturing the tutorials in a real pi showed rows that said more than was
known, or took more room than they had. What each one does now:

- **The summary's time is the clock's, handed in.** It added up the visits
  that had ended, so it read 20s at 39s elapsed and jumped when a long visit
  ended. `livePlan` takes `elapsedMs`, this life's time by the caller's clock,
  and still reads none of its own: the same arguments draw the same frame, and
  the extension hands in `elapsedMs()` of the view at each paint.
- **A line counts what every life spent there.** After a resume, a loop's
  `visit_end` holds the work of the life that wrote it, and the root lines
  did not add up to the summary. Each life keeps the visits it ended, and a
  line adds, for every life, the outermost of them at or under its path. A
  `copy_lost` or a visit run again forgets how it ended, never what it cost.
- **No token figure before a turn has ended.** pi's counters are read when a
  turn ends, so a subagent in its first turn read `↑0 ↓0` for all of it, which
  looks measured. `showTokens` gives nothing while `turns` is `0`, and `↑0 ↓0`
  after a turn, when the zero was read. It is one rule for the widget, the flow
  rows, `formatUsage` and the summary line, and it replaced the list of node
  kinds that had no tokens: a `check`, an `ask` or a `commit` runs no turn.
- **No cost figure when none was reported.** The widget and the flow rows
  already left a zero cost out; `formatUsage` printed `$0.0000`, which reads as
  free. `showCost` leaves it out everywhere, and a table column, which cannot
  leave a cell empty, says `not reported`.
- **A refused file reads as a warning, and only it.** `/flows` notified the
  whole listing as a warning as soon as one file was refused. It notifies at
  `info` now, and colours each line itself, since pi colours a notification
  whole: a refused entry and its faults in the warning colour, the rest as a
  listing.
- **The tool row names what the mode runs.** It named the first of `flow`,
  `agent`, `steps` and `candidates` that was set, so a loop read `loop coder`
  or `loop coder → reviewer` depending on whether the model had also passed
  `agent`. `runsWho` reads the field the mode reads.
- **A notification is fitted before pi draws it.** pi wraps a line where the
  room runs out, and a path is one long word. Every refusal goes through the
  fitting `/flows` already used, now `extension/notice.ts`: paths named from
  the working directory, a cut before a path, and the `Error: ` or `Warning: `
  pi writes before the first line counted. A hang that leaves fewer than 24
  columns goes on four columns in, so a description at 80 columns is not a
  word per line down the right of the terminal.
- **A painted line is cut in columns, last.** A subagent's row under a deep
  visit took two columns more than it had when its activity filled the room,
  because the two spaces before an empty detail sat before a colour code,
  where `trimEnd` cannot see them, and pi exited on `Rendered line exceeds
  terminal width`. The parts are joined only when there is something to join,
  and every line `paintFlow` gives is cut by `truncateToWidth`, which counts
  what the terminal counts: colour codes, tabs and wide characters.
- **A call that came back an error says so.** A call pi refused, `Tool write
  not found`, was listed like one that ran. pi's `tool_execution_end` carries
  `isError`, and `session.ts` reads it into a `tool_error` event with the first
  line pi said; the call is drawn `✗ write notes.txt · Tool write not found`.
  pi does not say whether a call was refused or ran and failed, so neither do
  we: the words are pi's.

### One picture of the run, and readers that do not fold

The TUI collector held the fold of the event stream - who spawned, under whom,
doing what, at what cost - and it was not the only one. The console kept a
`depths` map of its own to indent a delegated subagent, the extension rebuilt
counts and a hand-summed total from the subagents pi handed back in a tool
result, and `setTask` sat on the collector's interface for a caller that no
longer existed, its TSDoc saying the core did not emit the task when `status`
had carried it for some time. A fix to the fold - the launch order of a
fan-out - reached the widget and nothing else.

The fold is one module now, `picture.ts`, named for what it is rather than for
the first thing that read it: `createRunPicture` folds, `RunSnapshot` is what
comes out, and `snapshotFrom` derives the counts and the total from a list of
subagents so the extension gets the same picture back from a serialised
result. Depth is a fact of the fold, decided when a subagent spawns from the
parent the picture had seen, and stored on the snapshot; `treeOrder` only puts
children after their parents. The console reads the depth from a picture of its
own instead of keeping one, `setTask` is gone, and `of(id)` takes its place on
the interface - a read where there was a write-in side door. Formatting stays
in `tui.ts`, with nothing to hold.

What did not move, and why: the herdr reporter keeps a map of the panes it
opened, which is effect state rather than a second copy of the run; the mirror
relays a session's events to a socket and holds nothing; and the pane client
lives in another process, fed by the wire, with three variables. None of them
re-derives what the picture knows.

### The stream on disk

`recordReporter(file)` appends every event as one JSON line. It is a reporter
like the others - the run is identical without it - and it earns its place
because pi's per-subagent JSONL cannot show the **interleaving**, nor carry a
timestamp pi does not have.

- **Verbatim, no filtering, no knob.** A recorder that edits its own record is
  worse than a large file, and the question it will be asked is the one nobody
  planned in advance.
- **`appendFileSync`, not a write stream.** One syscall per event is the cost;
  the run worth reading afterwards is the interrupted one, and a buffered stream
  loses its tail exactly then. Same trade as the export.
- **Every experiment cell gets one**, at `events.jsonl`, with no option to turn
  it off. A cell whose stream was not kept can only be re-run, and a matrix is
  expensive. The day someone needs it off is the day the knob is justified.

### A measured run is the library's

The extension's live view and an experiment's cell assembled the same thing by
hand: a picture, a composition of listeners on it, a clock, and at the end
`usageReport(snapshot, wallMs)` written with `writeUsageReport`. The view added
the parent session's copy and swallowed a failed write; the cell added a
recorder and let a failed write throw. Two adapters at one seam with no name,
each free to drift from the other, and the guide's own example of an export
was a third copy.

`measuredRun({ dir?, record?, listeners?, mainSessionFile? })` in
`src/measure/measured.ts` is the seam: `onEvent` to subscribe, the `picture`,
`elapsedMs()`, and `finish()`, which writes `usage.json` with the time measured
and hands the report back, never throwing, because an export is an observer of
the run. The view is a measured run with a terminal on top and takes its `dir`
up front rather than at `stop()`, since `watched` knows it from the start. The
cell is a measured run with `record: true`, which is where `events.jsonl`
comes from now. A script that wants `usage.json` is one with nothing added,
and that is what the export guide shows.

`run-ui.ts` was four concepts at 224 lines: the view, the report, the herdr
session switch and the widget's paint. The report is the library's now, the
switch is `herdr-switch.ts` because the command sets it and the view reads it
and neither should reach into the other for a boolean, and the counter that
gives escape to a question card is `asking.ts` for the same reason: the card
sets it, the stop key reads it, and `ask-ui.ts` importing from `stop.ts` for
it made the card depend on the stop registry.

The assembly is asserted once, in `test/measured.test.ts`. The tool's two
export tests that proved the parent session's copy and its absence through the
tool are gone; the tool's tests still say where it exports and that it exports
nothing unasked, and the experiment's still say each cell gets its own
directory and its own stream.

### How a subagent stands is read once

The glyph that says how a subagent stands was decided in five places. The
widget rows folded `ok` and `status` into `● / ✓ / ✗`; a helper beside them
folded the same two into `⏳ / ✓ / ✗` for a collapsed text line; the tool's
card and its expanded view each wrote `ok === false ? ✗ : ✓`; the widget's
painter mapped a status to a colour; and the console wrote `✓` on every
`close`, whatever the result said. That last one is the bug the close event's
own comment records having fixed once, for every reporter that read the event
- it survived in the console because the console had a copy of the rule. The
text line, meanwhile, had no caller in production and a thorough test, which
is this repository's own trap in miniature.

`standingOf(snapshot)` is the one fold: the status, with a failure outranking
it. `statusIcon(standing)` and `statusColour(standing)` are the two tables read
off it, and every reader - the widget rows, the card, the summary table, the
painter, the console - draws what they give it. `●` while it lives, because
that is what the widget drew and what the guide shows; the `⏳` the display
specification above kept for a collapsed row went with the row nobody drew.
The console keeps `⏳` on its spawn line, which says something else: born, not
yet working.

### One lookup for a pipeline

**Removed** with the linear format; `checkFlow` is the one lookup of a flow.
See [The linear pipeline is removed](#the-linear-pipeline-is-removed).

The rule that a broken file is refused rather than silently replaced was
implemented three times, with three message shapes: in `findPipeline`, which
the extension never reached for a broken file because two of its callers
checked first; in `choosePipeline`, with a `command:` prefix its unknown-name
sibling never had; and in `resolveTarget`, which searched the catalogue itself
rather than call `findPipeline`, because it needed "absent" to fall through
to an agent instead of throwing. The catalogue was loaded with the same three
flags at three sites, where the roster had `loadRoster` and the paragraph
saying why.

`lookupPipeline(catalogue, name)` answers a pipeline or `undefined`, and throws
on a broken file of that name: absent is a thing a caller may act on, a file
that is right there and does not parse is not, and a caller that falls back on
`undefined` must never fall back past one. `findPipeline` is the lookup plus
"unknown name is an error, with what there is instead". `loadCatalogue` is
`loadRoster`'s sibling. `choosePipeline` is one line; `resolveTarget` calls the
lookup and keeps its own question, which was only ever "pipeline or agent";
`listPipelines` loads through the same door. One message, the library's, for
every command: the `run:` and `step:` prefixes went, as the unknown-name
message had never carried one.

## Stopping a run, and one subagent of it

A run already obeys a `signal`, and that is all it took to call the whole thing
off - as long as somebody held one. Two things were missing, and they are
different: **a signal cannot single a branch out**, because every subagent under
a workflow shares it and a combinator hands out no handles; and **a command had
no signal at all**, because pi's `ctx.signal` is the *agent run's*, which is
`undefined` while no turn is in flight. So `esc` during a `subagent` tool call
already stopped everything, and `esc` during `/run` stopped nothing at all.

`stopSwitch` (`src/stop.ts`) is the pair that fixes both: a signal to give the
workflow, and a `spawn` to give it too. The handles are registered on the way
out of that `spawn`, which is the only place they all pass - delegated children
included. `liveRun` builds one per run and hands the two to the call sites, so
the tool, `/run`, `/build` and `/step` are stoppable by the same act.

**Stopping is a `Subagent` method, not a second signal.** `stop()` aborts the
turn in flight and refuses every later one with `"stopped"`. One-way, because
that is what a person pressing a key means, and distinct from `close()`: the
session is still there to be exported, and whoever opened it still closes it.
The label matters more than it looks - a deadline, a cancelled run and a person
call for different reactions, and before this they all read `aborted`.

Three ways in, because no single one reaches every case:

- **`esc` stops everything, and is listened to rather than consumed.** pi binds
  it to `app.interrupt`; swallowing it would stop the subagents and leave the
  turn running, which hands the model a wall of `stopped` results and every
  freedom to delegate again. Not consuming it means one key with one meaning:
  inside a turn the turn goes too, and during a command pi's own handler finds
  nothing to abort and ours does the work.
- **`ctrl+↑↓` select and `ctrl+del` stops the selection.** The list is already
  on screen, so it is the list a key moves through; the three are bound to
  nothing in pi and are read only while a run is live, through
  `ctx.ui.onTerminalInput`. Registering `escape` as an extension shortcut was
  the other option and is a trap: pi checks extension shortcuts *before* its own
  keybindings and swallows the key whatever the handler does, so it would break
  interrupt, autocomplete cancellation and clearing the editor for the whole
  session.
- **`/stop [<id>|all]`** names one, which no key can. It is not enough on its
  own: pi executes an extension command immediately during a *turn*, but
  processes no submission at all while a slash command of its own is awaiting -
  measured, typing `/stop all` during `/run` ran nothing until the run was over.
  That is the whole reason `ctrl+del` exists.

**Stopping one branch is not stopping the run.** The branch comes back as a
failed `Result` and the workflow decides: a `fanOut` branch dies alone, a
pipeline step that fails ends the pipeline. Saying more than that in the
message would be a promise the extension is not in a position to keep.

**A question card owns `esc` while it is up.** The card had said "esc build
with what you have" since before a run could be stopped from the keyboard, and
the two meanings collided the first time somebody pressed it: measured on
`/build`, `esc` on the first card ended with `interview failed: stopped`. The
card resolved to a submit, the listener stopped every subagent of the run, and
the interviewer that had to write the brief was one of them. So `whileAsking`
in `extension/commands/stop.ts` holds the stop for the length of a question,
free-text box included, and `escape` falls through to pi as before. Held rather
than dropped, because the run's subagents are idle while a question waits: there
is nothing running that the key would have been pressed to call off. The other
way round - making the card's `esc` a cancel, so that both meanings agree - was
rejected, because it throws the answers away, and the reflex to escape out of a
dialog is exactly the moment those answers are worth keeping.

**A flow's card is the interview's card, told more.** `createAskUi` reads the
`Asking` a flow's `ask` hands it, and an `AskUser` called without one draws the
interview's card as before. What the card draws around the answer is
`extension/ui/card.ts`, the same for every form: the header and the visit, the
reads under their names, the question and the help line. What takes the answer
is a list or a text box, and `ask.ts` picks it from the form.

- **A flow's text box is drawn in the card, not by pi's `input`.** pi's box
  shows a title and nothing else, so the reads the person is answering about
  would disappear the moment they chose `Other…`. In the card they stay above
  the question. pi's `custom` takes no signal, so the card calls its own
  `done` when `Asking.signal` aborts, and pi puts its editor back as it does
  for any answer.
- **`esc` in the text box behind `Other…` goes back to the options.** On the
  interview's card it still submits, as it always did. On a flow's card with no
  "enough", declining is the run's stop, and a key pressed to leave a text box
  is not a request to end the run. An empty answer goes back too: it is not a
  typed answer.
- **The help line says what `esc` does on this card.** It gives the label of
  "enough" when the card offers one, and "stop the run" when it does not. A
  card that said "esc cancel" while the key stopped the run would be the
  collision above, again.
- **The card is laid out like pi's own dialogs.** The list sits one column in
  and blank lines separate the parts, where it used to be flush with the border
  and packed. When the help line does not fit, each key gets its own line: at
  44 columns it had wrapped as `esc stop the` / `run`. The interview's card
  gets the same layout. Only its looks change, not what it does.

## One floor under the commands

Five commands launch work - `/build`, `/run`, `/step`, `/swarm`, `/interview` -
and each had written the same shape by hand: choose between the double and the
real thing at every use (`deps.runPipeline ?? runPipeline`, thirty-two times in
ten files), wrap the pre-spawn checks in the same `try`/`catch` that turns a
thrown explanation into a refusal (six copies), then `liveRun`, a footer status,
and a `try`/`finally` that stops the view and writes `usage.json` (five copies,
one of them clearing a status the stop had already cleared). `command.ts` was
called the floor they stood on and held the parser, `refuse` and the roster; the
order things happen in was still every command's own, and so was every way of
getting it slightly wrong.

The floor is three pieces now, each deep and composable, and deliberately not
one template function: the five commands have different shapes - three stops
for `/build`, an editor after the work for `/interview` - and a function taking
parse, target, work and report as parameters would be a configuration engine
with five callers. `resolved(deps)` applies the defaults once and hands back
doubles that are all present, treating a key holding `undefined` as unsaid for
the same reason `override()` does in `pipeline-run.ts`. `checked(ctx, fn)` runs
what must pass before a spawn and gives `undefined` back on a thrown
explanation, the shape every refusal already had. `watched(ctx, deps, { status,
dir, work })` puts the dots up, says what is running, hands the work its live
run, and takes everything down in a `finally` with the time this command
measured - one clock, which is also why a run's `usage.json` now carries the
command's wall time rather than the workflow's own.

Two moves came with it. `command.ts` was mixing four concepts at 253 lines, so
what a command reaches for is `deps.ts` and what was typed is `flags.ts`;
`BuildDeps` is `CommandDeps`, since it never was about one command. And
`pipelineVerifier` left `run-ui.ts`: three callers, none of which paints.

`checkModel` stays in each command's `checked` block rather than in `watched`:
`/build` has to refuse a mistyped model before the interview, and the interview
runs before anything is watched. What a command checks is its own; that it
checks before it spawns is the floor's.

The contract of the floor is asserted once, in `test/command.test.ts` - a
thrown work still throws after the clean-up, the footer and the widget are
cleared whatever happened, `usage.json` lands in the folder given - and the two
commands that each proved the widget goes when the run ends lost those tests.
Their own tests say what each command decides.

### The tool and the committer stand on it too

The floor was built for five commands and two launch sites stayed beside it.
The `subagent` tool opened its own `liveRun`, kept its own clock and wrote its
own `finally`, because it varies four things about the view - a second
reporter, `herdrAll`, the parent session's file, a progress line streamed on
every change - and `watched` let none of them through. And `/build`'s committer
ran through `deps.run` with `ctx.signal`, which the section above says is
`undefined` during a command: the one subagent of a build that neither `esc`
nor `/stop` could reach, with no dots to say it was working. Its footer status
was set and cleared by hand around it.

`Watched` takes `live`, the slice of `LiveRunOptions` a caller may vary, and
`status` became optional because the tool has no footer to write to; who stands
on the floor is a `Watcher`, a `ui` and a `signal`, which pi's command context
and the tool's deps both are. The live run keeps the clock - `elapsedMs()`, and
`stop(dir)` no longer takes a wall time it was handed - so the tool reads the
run's time off the view it ran under rather than keeping a second one. The
committer runs under `watched` with the run's `signal`, `spawn` and `onEvent`,
which is what puts it within reach of `/stop` and on the widget.

The tool's arguments were declared twice: a typebox `Schema` in `index.ts` for
pi, and a hand-kept `Params` in `execute.ts` for the code, each with its own
description of the same twenty-one fields. `Schema` lives in `execute.ts` and
`Params` is `Static<typeof Schema>`. Three of its fields are `enum`s now -
`mode`, `lifetime`, `scope` - so what the code accepts is what the model is
told: `asLifetime` silently took `"session"`, which the description never
named, and `asScope` took anything and fell back in silence; pi refuses what
the schema does not name, and the two validators are gone. What the model
sends is `extension/params.ts`, what the tool does stays `execute.ts`, and the
`switch` over the mode is `perform()`, apart from the wiring: the one file was
mixing three things at 345 lines.

The tool's widget test went the way the commands' did: the floor's contract is
asserted once, and the tool's tests say what the tool decides. The committer
has the test it was owed - the run's signal, not pi's.

### The relay owns the entry a step leaves

`/step` and `/swarm` both end a stage the same way: a numbered subfolder of the
chain for the transcripts, the step recorded in the relay, an entry appended to
the transcript. Each had written the folder name and the entry by hand, and
`/swarm` imported `STEP_ENTRY` and its type from `/step`'s file to do it - one
command depending on another to know what a step of the chain looks like in the
session. The two folders were also named differently: `/step` after the id,
`/swarm` after the agent, so a second swarm of the same member got a folder that
did not say which step it was.

Both now belong to `relay.ts`, which already owned the chain: `stepDir(relay,
id)` names the folder, `STEP_ENTRY` and `entryOf(step)` say what the transcript
gets, and `recordStep` takes the id the step ran under rather than minting one
afterwards, so the folder and the entry cannot disagree. The doors a command has
into the session - `SendMessage`, `AppendEntry`, and the `PipelineDeps` and
`StepDeps` that carry them - are in `deps.ts` with everything else a command
reaches for. What one command file still imports from another is the design
and not a leftover: `/build` opens with `/interview`'s function, and `/quote`
sends the message `/run` sends.

### The relay owns the life of a step

Owning the vocabulary left the sequence to the commands, and `/step` and
`/swarm` still spelled it out alike, eight moves each: continue or start the
chain, mint the id, name the folder, run under the dots, refuse on failure,
record, append the entry, notify. The entry went through
`injected.appendEntry?.(…)` - the door read off the raw dependencies, where
`resolved()` never looked, so a caller that forgot it got a step recorded in
the relay and never drawn, in silence, the very failure `resolved()` guards
against for every other dependency. `stepAnswer` and `pipelineAnswer` were one
framing in two files, each justified by the same paragraph about pi's
user-role slot.

`beginStep(name, runDir)` opens or continues the chain and names the id and
the folder before anything runs; `finishStep(begun, outcome, appendEntry)`
records under that id and appends the entry in the same call, so "recorded but
never drawn" cannot be written. What the commands keep is what each decides:
the flags, the target, the refusal's wording, the notification. The doors are
**required** on `PipelineDeps` and `StepDeps` rather than resolved to a
default: a door that quietly did nothing would be the failure the type exists
to prevent, so pi's is bound once in `sessionDoors` and a test hands a
recorder. `framed(what, asked, output)` is the one framing, in the relay, and
the pipeline's answer is a one-line caller of it.

The mechanism is asserted once, in `test/relay.test.ts`: a begun step opens or
continues the chain, a finished one is recorded under the id it began with and
its entry goes through the door. The two commands' tests keep saying what each
sends and what each draws, because that is the property they exist for.

## pi comes in through one door, and nothing casts it on the way

Every command handler took pi's `ExtensionCommandContext` and handed it on `as
unknown as CommandCtx`, eleven times, and `/stop` did the same into a slice of
its own. Written when pi's context and ours did not line up, the casts had
outlived that: an assignment compiles without them today. What they still did
was keep TypeScript from checking that pi's context has what a command reads,
which is the one check a fake cannot make - and the five slices of pi's UI the
extension read were declared in five files, `notify` three times over,
consistent with one another by luck rather than by construction.

`extension/pi.ts` is where pi comes in, and the only file under `extension/`
that names pi's context types. `Ui` is the slice of `ctx.ui` the extension
reads, declared once; `CommandCtx` is what a command is handed; `RunUi`,
`KeyUi`, `AskUi` and `StopCtx` are picked from `Ui`, narrow so a test's double
stays small, and never redeclared. `PiApi` is the slice of pi's API the
extension registers through, and the test's fake is typed as it rather than
cast `as never`, so a method pi renames fails at compile time. The two wirings
only the door knows - what the tool body is handed, `mainSessionFile` read off
pi's session manager included, and how a command reaches the session - are two
functions, `toolDeps` and `sessionDoors`, and they are tested.

A handler is written against `CommandCtx`. pi hands it the whole context, and
TypeScript checks at every `registerCommand` and at the tool's `execute` that
the whole has what the slice reads. That is the door, and it needs no adapter
function: an identity typed on both sides would be a pass-through, and the
check it would carry already happens at the assignment. The one cast left is in
the test that drives a registered handler with the fake, standing in for the
members pi has and the extension never reads.

### The audit cycle is the audit's

**Removed:** `audit` went with `deliver`; the audit is a node of the `build`
flow, inside its `deliver` loop. See [The linear pipeline is removed](#the-linear-pipeline-is-removed).

`auditOnce` took ten fields and did three things: a pool, a prompt, a turn.
Everything that made a round of audit a cycle - asking, closing the round in
the record, deriving the fixes and attaching the standing check to each, running
them, settling the tree, deciding whether another round was worth paying for -
was fifty lines of `deliver.ts`, and twenty of `deliver`'s tests were about
those lines and nothing else in a delivery. The two pure helpers the cycle used,
`fixesFrom` and `withCheck`, were tested; the caller that had to call them in the
right order with the right check was not, on its own.

`audit(options)` in `audit.ts` is the cycle. It builds the review record with
what a previous run raised, spawns a fresh auditor per round whatever the caller
runs with, reads the verdict, derives and dresses the fixes, and applies the two
rules that end a cycle early: a yes with nothing owed and no failing check
standing, or a round that asked for nothing and closed nothing. What it does not
know is how a fix reaches the tree. It is handed `fix(fixes)`, which runs them
and says what the check is afterwards, and `deliver` hands it the one that runs
a pair and settles the copies - which keeps everything about working copies in
`deliver`, where the next decision moves it. `deliver` reads as plan, work,
audit, settle.

The move surfaced a double count. A fix's result was appended to `tasks` so the
next auditor would read it, and the delivery's usage summed `tasks` and then
every round's `results` again. A delivery with a fix reported the fix's tokens
twice, and the one usage test had no fix in it. The total is planning, the tasks
as audited last - fixes included, once - and each round's review.

`auditOnce` is not exported any more: one turn without a record or a stopping
rule is not an audit in this repository's sense, and it had one caller.

## Asking the user, and touching the world

Two ports, one rule: **the agents produce text, our code performs the act.**

- `src/ask.ts` - `AskUser`, one question at a time. The pi implementation is a
  select card (`extension/ui/ask.ts`), an example uses readline, the tests use a
  scripted array. Returning `undefined` is the **submit**, not a cancel: what was
  already answered still counts, and the brief is still written. `esc` maps to it
  for the same reason.
- `src/verify.ts` - `Verify`, a command we run with `execFile` and no shell. Its
  output is evidence the agents read and cannot argue with.
- `src/git/git.ts` - the git a pipeline may do, as functions. There is no
  `push`, no `reset`, no `rebase`, no `--force`, and no shell: arguments are
  arrays and the commit message is piped to `git commit -F -`, so a message
  containing `rm -rf /` is committed rather than executed.

**Why the committer has no `bash`.** It was the obvious design - give the agent
git and tell it what not to do - and it is exactly what "a prompt is not a
permission boundary" forbids. The agent writes the message, which is what a
model is for; the branch and the commit are ours. Adding a subcommand is a
decision someone takes in a diff, not an argument a model produces at runtime.

**Why the interactive flows are commands, not tools.** **Reversed: a
question card is shown during a model's turn.** pi's own examples call
`ctx.ui.custom()` from a tool's `execute`, and the `subagent` tool now puts a
flow's `ask` cards through `ctx.ui` while the turn waits; a real pi showed
the card answered, and `esc` on it meaning the card's own entry. See
[`/step` and the `subagent` tool take
flows](#step-and-the-subagent-tool-take-flows). What follows is the reason
there were none. A question card owns the terminal until it is answered, and
nobody can answer a question asked inside a model's turn. `/interview` and `/build` are therefore `pi.registerCommand`, and
`/build` stops exactly twice: the brief before any work starts, the commit before
anything reaches history. A refusal at either stop leaves everything where it is
- the brief in the editor, the work in the working tree. Nothing is undone on the
user's behalf. Both stops have since gone: see [`/build` asks
nothing](#build-asks-nothing).

## `/build` asks nothing

**Replaced: `/build` is `/run build`.** The shipped `build` flow still asks
nothing, commits nothing and needs a git repository, now because its copies
need one and the run stage says so; the check is a script the file names,
`.pi/checks/test.sh`, and no longer a flag. See [`/run` runs flows, and
`/build` is `/run build`](#run-runs-flows-and-build-is-run-build).

**This reverses a decision.** `/build` stopped twice, on the brief and on the
commit, and asked for a check when the pipeline named none. Each question made
sense on its own. Together they meant a build could not run without somebody
in front of it, and a build is the work you most want to leave running.

- **The request is the brief.** The interview is `/interview`, a command of its
  own, and what it hands back is text to give `/build`. Opening every build
  with it questioned people who had typed a precise request, and would question
  nobody in a run left alone.
- **Nothing is committed.** The work stays in the working tree, and `git diff`
  is the report. A commit is a decision about the work, taken by whoever read
  it. `extension/commands/commit.ts` went, and the commands' `Git` port shrank
  to `isRepository`. The committer agent and `src/git/git.ts` stay in the
  library, where a script can still use them.
- **The check is a flag or a file.** `--check "npm test"` for one run, `verify:`
  for every run of that pipeline, and the flag wins. With neither, no check
  runs. The flag is split on whitespace and on nothing else, which is not the
  small shell the `verify:` list exists to avoid: no quote, `&&` or `$` is
  interpreted, and every word is an argument. A double-quoted value is the one
  thing the flag parser learnt for it.
- **Resuming says what it picked up.** Typing `/build resume` is already the
  answer, so "Carry on?" became a line.
- **A git repository is still required**, for a new reason: nobody watches the
  run write, and git is how its work gets read and, if need be, undone.
- **`/interview` parses its own flags.** It borrowed `/build`'s parser for
  `--questions`, a flag `/build` no longer has.

## A check the auditor contradicts goes out with the fix

**Removed** with `deliver`: in the `build` flow the auditor reads the check's
report among its `reads:`, and the round ends only on
`tests.output.passed && audit.output.approved`, a condition our code reads. See [The linear pipeline is removed](#the-linear-pipeline-is-removed).

The auditor is handed the check's command and its output, and it can still
write the opposite. Measured on a delivery that worked: four green tests, then
`"Test file has a syntax error causing failure."` and a fix raised for it, and a
round spent rewriting a file that was fine. Audit 2 approved. Every honest
signal was there and none of them stopped the round.

Three ways out, and the middle one is the one taken.

- **Drop the fix.** It needs us to read the auditor's prose for a claim about
  the check, and a false positive there throws away a real remark. Guessing at
  meaning is what the verdict tool exists to avoid.
- **Send the evidence with the work.** Every audit fix carries one line saying
  the check passes on the tree it is about to change, so a worker sent after a
  failure that is not there settles it by reading instead of by rewriting. The
  remark still reaches the worker exactly as the auditor wrote it, which is what
  keeps a real fix from being lost to a heuristic.
- **Say nothing and pay the round.** That is what the measurement cost, and the
  round is not the whole of it: the rewrite lands in the tree and the next audit
  reads a file nobody meant to change.

The audit prompt now states the passing case with the same force as the failing
one, which was already there: a fix is asked for because the code is wrong, not
because something fails. That sentence is not the mechanism, though. Invariant 7
holds here as everywhere - the prompt asks, and the evidence travelling with the
fix is what a worker cannot talk itself out of.

A **failing** check is not attached. It is what the fix is for, and the suite
says so the moment the worker runs it.

## Resuming a build

**Replaced: `/run resume` carries on a flow run from its journal**, and
`/build resume`, `build.json` and `deliver`'s own `resume` are gone, removed
with the linear pipeline (see [The linear pipeline is
removed](#the-linear-pipeline-is-removed)). See [A resume goes as deep as the
journal](#a-resume-goes-as-deep-as-the-journal-and-replays-only-what-did-not-end).

A delivery is long, it costs money and it writes to a working tree. `deliver`
therefore takes `onProgress` and `resume`, and
`src/workflows/deliver/resume.ts` turns the one into the other through
`runs/<timestamp>/build.json`.

- **Only what was approved survives.** A subtask still being argued over left the
  tree in a state nobody signed off on, so it runs again. Approval is the only
  claim from a previous life worth trusting. **Reversed for flows:** every
  visit that ended survives, and a resume starts at the first that did not.
  The next review visit judges the tree anyway, so the sign-off still happens;
  see
  [A resume goes as deep as the journal](#a-resume-goes-as-deep-as-the-journal-and-replays-only-what-did-not-end).
- **The plan is reused, never re-made.** Re-planning would re-split work that is
  already half done on disk, and the plan was paid for.
- **Nothing of the conversation is saved.** Agents are stored by name and
  resolved again; `Result.messages` are dropped. A resumed build re-reads the
  code rather than replaying a transcript - which is also what keeps the file
  small enough to write after every step.
- **A state whose agents no longer exist is refused whole.** Dropping the steps
  that no longer resolve would silently drop work.
- **The audit rounds already spent are spent.** Resuming continues the cycle, it
  does not restart it.
- `onProgress` is a reporting hook, so a listener that throws is swallowed - the
  same rule as the event bus.

### The shape of a message is read where the pi API lives

Two readers in `subagent.ts` knew what a pi message looks like: one walked the
content parts of the last assistant message for its text, the other read its
`stopReason` and `errorMessage` for the failure a turn can end on without
throwing - the trap `AGENTS.md` lists. The event subscription beside them cast
a `SessionEvent` to reach `assistantMessageEvent`, `toolName` and `args`, though
`session.ts` declares that very union. The fake session then had to reproduce
the same shape for those readers to read. Three places knew pi's shape; the
invariant names one.

`lastTurn(messages)` and `streamed(event)` are the two readings, in
`session.ts`: what the last turn said and how it ended, and what a streamed
event means to a listener - a piece of the answer, a tool being called, or
nothing. `subagent.ts` reads a text, an error and two kinds of event, and
holds no cast. The casts themselves did not go: the union keeps a `{ type:
string }` member so that pi's own listener type satisfies the port, and that
member is what stops narrowing; they sit where the shape is known, with the
reason. The fake still builds pi-shaped messages, because it stands in for pi,
and the two readings are asserted against that shape once, in
`test/session.test.ts`.

The port itself did not grow a method. `createDefaultSession` hands back pi's
`AgentSession` as it is, which satisfies the port structurally; a method of our
own would have meant a wrapper proxying every member for the sake of one.
**Reversed** by the next section: the wrapper now does the work.

### The port is a turn

`SessionPort` was a slice of `AgentSession`, and its callers spent their code
turning pi's shape back into what they wanted. `subagent.ts` took
`getSessionStats()` before and after each prompt and subtracted, subscribed to
pi's raw events and parsed them back into text and tool calls, collected the
turn's messages off `message_end`, bridged its signal to `abort()` and removed
the listener again, and read the last message for a failing `stopReason`. The
mirror checked `isStreaming` before every steer. Each fake had to reproduce
all of it: a cumulative counter, pi's event and message shapes, an abort that
cuts. A prototype of a second backend showed what that costs: half of its
adapter rebuilt `AgentSession`'s getters around a backend that has none of
them.

The port now has the shape of what combo asks for:

```typescript
type SessionPort = {
  readonly model?: string;                                  // provider/id
  ask(text: string, options: TurnControl): Promise<Turn>;  // never rejects
  steer(text: string): Promise<"queued" | "idle">;
  watch(listener: (event: SessionEvent) => void): () => void;  // the mirror's
  transcript(): readonly AgentMessage[];                    // the mirror's replay
  exportToHtml?(outputPath: string): Promise<string>;
  exportToJsonl?(outputPath: string): string;
  close(): void;
};
type TurnControl = { signal: AbortSignal; onStreamed?: (event: Streamed) => void };
type Turn = { messages: AgentMessage[]; text: string; error?: string; usage: Usage };
```

`sessionPort(session)` in `session.ts` is the adapter over pi, and every trap
the turn used to handle moved into it with its comment: the abort bridge, the
retry wait that leaves only the signal to say the turn was cut, the failing
`stopReason`, the compaction that makes a turn's messages its events rather
than a slice. A turn never rejects; it comes back with `error`. What stays in
`subagent.ts` is combo's: refusing a turn a stopped subagent was asked, naming
an abort `stopped` or a timeout, and adding time to the bill.

`watch` and `transcript` hand out pi's own events and messages, because the
pane draws them with pi's components, and that is the one place they are
wanted raw. `transcript` stays synchronous: the wire replays it on attach, and
a replay that awaited would let a live event overtake it. A refused turn keeps
the context level where the last turn left it, which is what the session read
then too.

**The delta stays, inside `session.ts`.** The question was whether a turn's
usage could be the sum of the usage on its own messages, which pi reports,
with no counter to subtract. Measured on pi 1.0.2, a persistent session on
`ilaas/qwen-3.6-35b-instruct`: the sum and the delta agreed to the token on a
tool turn, a plain one, an aborted one and a 411k-token one. Then a compaction
cost 5,904 tokens in and 767 out, `getSessionStats()` counted them, and no
`message_end` carried them: pi bills the summary request to the compaction
entry. pi compacts inside `prompt()`, before sending when the context is past
its threshold and after an overflow, so a turn summed from its messages would
drop that request. pi's own source names two more kinds the stats count and no
message carries: a cache warm-up's `usage` entry, off for a subagent, and a
branch summary. `getSessionStats()` is the one total that sees everything pi
billed, so the adapter reads it twice and subtracts, and nothing else does.

`snapshotUsage` went into the adapter, which took the last import of pi's
types out of `usage.ts`. `deltaUsage` stays where it was and public.

The fakes moved onto the port. The test fake and the dry run's scripted
session now script a turn as it comes back, error and per-turn usage included,
and no longer imitate pi's cumulative counters, event shapes or failing
`stopReason`. The pi-shaped fake survives in `test/fixtures/pi-session.ts` for
the two places that are about pi: the adapter's tests, which hold every trap
above, and the mirror's, which forward pi's events. `createDefaultSession` hands back the port, not pi's
session; a caller that wants pi's session builds it and wraps it in
`sessionPort()`, which is public for that.

## The pi API: what you need to know

**combo needs pi 1.0 or later.** The extension refuses an older pi when it
loads (`requirePi`, in `extension/pi.ts`), and `src/session.ts` speaks one model
API: a `ModelRuntime`, with patterns resolved by `resolveCliModel`.

**Reversed:** through 0.80.x combo ran on two pi generations at once. Homebrew
shipped `0.80.6` and npm `0.80.10`, and `0.80.7` had replaced `AuthStorage` +
`ModelRegistry` with a single `ModelRuntime`, so a pi "patch" release broke the
API. `buildRegistry` chose between the two by the **presence of the export**
rather than by version string, and `pane/tui.ts` did the same when 0.86 made
`TUI` an interface and moved the inline class to `TuiMainScreen`. pi 1.0 has
neither `AuthStorage` nor a `TUI` class, so both shims went: the session builds
a `ModelRuntime`, and the pane constructs `TuiMainScreen` itself. Keeping them
meant carrying a model API for a pi nobody tested against.

`requirePi` reads a version string, which the presence rule was there to avoid.
The rule answered a different question: which of two APIs this pi has. There is
one API now, and the question is whether this pi is one combo was written
against, which the version states directly. Without the check, a 0.80.6 host
would load the extension and die at the first spawn on `undefined.create()`,
the failure below. Peer dependencies stay at `*`, as pi's packaging
documentation asks, so npm cannot carry the minimum: the check and the README
do.

The pi an extension gets is the one it is loaded into, even loaded by path from
a checkout (`pi -e extension`): measured with pi 0.80.10 and 1.0.2 as the host,
the extension read the host's `VERSION` whatever `node_modules` held. The pane
and the library run under plain Node, where `node_modules` decides.

`ToolExecutionComponent`, handed no definition, draws a generic box of
`key=value` arguments. pi's interactive mode finds its own tools' renderers by
name but does not export that lookup, so `pane/renderers.ts` names the
definitions through `create*ToolDefinition`, which is what draws
`grep /x/ in a.ts` in the pane.

Standing still was not free either. 0.80.10 pinned an `undici` and a
`brace-expansion` with advisories against them, and only a newer pi fixed them.

This is the failure mode to remember: 158 tests were green while the extension
died on `undefined.create()` in a real pi, because every test injects a fake
`SessionPort` and none of them ever touches pi's real module. **A fake session
cannot tell you the package it stands in for has changed shape.** Anything that
only runs against the real pi has to be exercised against the real pi.

Local reference docs: `node_modules/@earendil-works/pi-coding-agent/docs/`
(read `sdk.md`, `extensions.md`, `tui.md`), examples in `examples/sdk/` and
`examples/extensions/` - in particular `examples/extensions/subagent/`, which we
take inspiration from but **do not copy**: it spawns one process per subagent
and therefore has no notion of lifetime.

Creating an isolated session (see `src/session.ts`, the only place the pi API
lives):

```typescript
import { ModelRuntime, SessionManager, createAgentSession } from "@earendil-works/pi-coding-agent";

const modelRuntime = await ModelRuntime.create();

const { session } = await createAgentSession({
  cwd,
  modelRuntime,
  model: resolveCliModel({ cliModel: pattern, modelRuntime }).model,
  tools: agent.tools,                       // ["read", "grep", "find", "ls"] …
  resourceLoader: new StaticResourceLoader(agent.systemPrompt),
  sessionManager: SessionManager.inMemory(cwd),
});

await session.prompt(task);
const messages = session.messages;
session.dispose();                          // ← always, in a finally
```

Points to watch:

- Built-in tools: `read`, `bash`, `edit`, `write`, `grep`, `find`, `ls`
  (default: `read`, `bash`, `edit`, `write`). `noTools: "all"` disables
  everything.
- A read-only subagent is `tools: ["read", "grep", "find", "ls"]`. **That is the
  recommended default** for exploration agents. The allowlist is genuinely
  enforced by pi - verified: a scout session exposes exactly those four tools.
  A weak model will still *emit* calls to tools it does not have (`edit`,
  `run`); they fail, and the model may retry them in a loop. That is an argument
  for `loop`'s `maxIterations`, not for loosening the allowlist.
- **A prompt is not a permission boundary.** `examples/04-loop.ts` used to hand
  the `coder` its full `edit`/`write` toolset and merely *ask* it not to change
  anything. It edited `src/usage.ts` anyway, twice, in a plain demo run. If a
  subagent must not write, take `write` and `edit` away from it - do not ask it
  nicely. This applies with force to anything shipped in the repository: an
  example must not be able to rewrite the repository it ships in.
- The system prompt goes through the `resourceLoader`, **not** through a
  `systemPrompt` field on `createAgentSession`. We supply our own
  (`StaticResourceLoader`): `DefaultResourceLoader` requires `cwd` and
  `agentDir`, re-reads the disk on every spawn, and loads extensions, skills and
  project trust - non-deterministic context a subagent does not need.
- **`session.prompt()` does not accept an `AbortSignal`.** `PromptOptions` only
  holds `expandPromptTemplates`, `images`, `streamingBehavior`, `source`,
  `preflightResult`. Cancellation goes through `session.abort()`: we bridge the
  signal by hand, and remove the listener after each turn.
- Messages are read from `session.messages`; `session.agent.state.messages` is an
  internal detail.
- A turn can fail **without throwing**: look at the last assistant message's
  `stopReason` (`"error"`, `"aborted"`, and `"length"` for the output limit).
- **Not every provider reports tokens.** Several return a `usage` that is already
  zero at the source; `getSessionStats()` then sums zeros. We display `0`, we
  never estimate it. And what a provider reports changes: one that reported
  nothing at all now reports tokens, so a provider's counters are worth
  re-measuring rather than reading off a list.
- `session.subscribe(…)` is the source of every `SubagentEvent`:
  `message_update`/`text_delta`, `tool_execution_start`, `turn_end`,
  `agent_end`. Never log directly from the core.
- `session.prompt()` in series on one session **is** a persistent subagent. That
  is literally the whole implementation of `lifetime: "workflow"`.
- `session.compact()` exists: it is the way out for a persistent agent whose
  context is swelling.
- Measurements and export are `AgentSession` methods: `getSessionStats()`,
  `getContextUsage()`, `exportToHtml()`, `exportToJsonl()`. **Call them before
  `dispose()`.**
- `session.dispose()` releases the session; an undisposed session leaks.

## Defining an agent

Markdown + frontmatter, compatible with pi's convention
(`~/.pi/agent/agents/*.md`, `.pi/agents/*.md`):

```markdown
---
name: reviewer
description: Reviews code and returns actionable remarks
tools: read, grep, find, ls
skills: diffing             # our own extension: an allowlist, like tools
model: anthropic/claude-sonnet-5
lifetime: workflow          # our own extension: default lifetime
---

You review the code produced and return at most 5 remarks…
```

- `name` and `description` are **mandatory**; a file without them is ignored
  silently (pi's behaviour, we keep it).
- `tools` and `skills` are read the same way: a comma-separated line or a YAML
  sequence, and a field with nothing after it names nothing rather than
  something empty.
- `lifetime` and `openInHerdr` in the frontmatter are only **defaults**: an
  explicit call always wins. Declaring `openInHerdr: true` on an agent is often
  what you want - a scout is worth watching whoever calls it.
- Project agents (`.pi/agents/`) are **repository-controlled** content: loaded
  only on explicit request (`scope: "project" | "both"`), never by default. Do
  not relax that rule "to keep things simple". This repository's own demo agents
  live in `agents/` and are symlinked into `.pi/agents/`, so the extension can
  find them - with an explicit scope, like anyone else's.
- **An empty agent list is almost always a scope problem, not a typo**, so
  `findAgent` says so. Given the old "Loaded agents: none", a model concluded the
  repository had no agent definitions at all and started offering to write some.
  An error message is read by an LLM as often as by a human now; it has to name
  the real cause.
- Agents are rediscovered on every call (hot editing works).

### Every subagent answers in the language it was given

Asked in French where the wall time is measured, `/step scout` answered in
English. No definition asks for English; they inherit it from the prompt around
them, which is English like everything written into this repository. The person
who asked then reads their own repository through a translation they never
asked for.

So the instruction is standing, in `src/language.ts`, appended to every system
prompt beside `situate()`. One place, and it reaches an agent a user wrote
without ever thinking about the question - which a line added to the nine
definitions here would not.

**It is said twice, and the second time is where it lands.** The system prompt
was enough for a single task and stopped being enough as soon as a combinator
framed one: a swarm tells each member which round it is, what is new on the
board and what is still free to take, and all of that arrives in the *same
message* as the goal, which a model weighs far above anything standing behind
it. So `ask()` closes every turn with the rule again, in one sentence - short,
because repeating three would spend a paragraph of context saying what one line
says. Measured on `ilaas/gpt-oss-120b`, a French question put to a swarm of two
over two rounds, counting the board posts that came back in French:

| | standing rule alone | with the closing line |
|---|---|---|
| the wording that disowned only the prompt | 3 of 18 | 9 of 35 |
| the wording that disowns the framing too | 5 of 19 | 62 of 89 |

Neither half carries it, which is why neither can be dropped as a tidy-up. `ask()` is the one funnel every turn
of every workflow goes through, so a combinator cannot forget it and a workflow
somebody else writes gets it for free; the event stream still reports the task
its caller wrote, since a card drawing the line back would show a reader
something they did not write.

**It points at the work, not at the prompt.** The first wording said "the
language of the task you are given", and measured on a small open-weight model
that loses the case worth having: a reviewer whose French goal is wrapped in
English scaffolding ("Review this work.", "It was asked to:") answered in
English, because most of the task really was English. Naming the material
instead - the request, the specification, the report of another agent - and
saying outright that these instructions never decide it, turned the same run
French.

**And it names no language.** The wording that fixed the reviewer did it with an
example: *if the work you were handed is in French, answer in French*. The next
English question came back in French. In a standing instruction an example reads
as the target, and this one is in front of the model on every turn of every
agent, so the rule now says that no example in it decides the language either.
Three turns settle it, and they are the check to re-run on the day the sentence
is touched: an English task answers English, a French task answers French, and
French work inside English scaffolding answers French.

**What must not move is exempted by shape.** A model writing French writes
`PRÊT` for `READY` and `RAS` for `LGTM`: a loop that never converges, a review
nobody can parse, and no error anywhere. The library could list its own
sentinels, but a workflow somebody else writes has its own, so the rule is
written as a shape: a word you were told to answer with, a JSON key, an agent
name, an identifier, a path, anything quoted from code. Measured, on the same
model: a French task asking for `LGTM` alone returns `LGTM`; the router answers
`scout`; the planner answers valid JSON whose `task` strings are French, which
is right, since they are the work the next agent reads.

The interview keeps its own per-turn reminder. It is the one prompt that can
name its exemptions exactly, and it pays the most for losing them.

## The model: an explicit knob at every level

For a long time the model was the one thing invariant 5 did not cover, and it
was measured: agents with no `model:` ran on the operator's
`~/.pi/agent/settings.json` defaults (a model the caller had never named, 402s
inside a session started with a different provider, `thinkingLevel: high` nobody
asked for), while the caller's `--provider`/`--model` never reached a
subagent. An experiment that pinned its repository with a tag while leaving
the model floating was measuring the operator.

The fix is **one option, `model`, at every level, with the nearest override
winning** - the same "an explicit call wins" rule as `lifetime`:

1. the run-time argument: `SpawnOptions.model`, `WorkflowOptions.model`, the
   tool's `model` param, `--model` on `/run` and `/step`;
2. the flow or pipeline file's top-level `model:`;
3. the agent's frontmatter `model:`;
4. pi's own settings, as the last resort - only when nothing was set anywhere.

It is an **override, not a default**: the knob exists to run one workflow
against different LLMs, and a frontmatter model surviving a sweep would make
the experiment measure a mixture. Precedence is resolved once, in `spawn()`,
so an injected fake session observes the *effective* pattern.
### What a shipped agent declares: nothing

None of the nine agents in `agents/` carries a `model:`, and that is a decision
rather than an omission. A definition that ships in a package must not choose its
user's provider: pinning one would override the settings they already made and
break outright for anyone holding no key for it. So a shipped agent falls through
to step 4, which is what makes `pi -e extension` work on a machine we know
nothing about.

The cost is real and worth stating, because it is invisible: **a run with no
`--model` measures the operator.** Verified on 2026-08-01 - `/run explore` with
nothing specified put all four subagents on a model named nowhere in this
repository, read from `~/.pi/agent/settings.json`. An earlier run picked up a
`thinkingLevel: high` nobody asked for the same way.

Which is why the fix is a habit, not a file: **anything whose numbers will be
compared names its model.** `experiment()` takes `models` and refuses to guess,
`--model` exists on `/run` and `/step`, the examples take it in argv, and a
pipeline that belongs to one repository may pin its own. An agent *you* write for
*your* machine is welcome to declare one - it is only what ships that must not.

Three deliberate refusals, so nobody "fixes" them later:

- **No environment variable is read by the library, or anywhere else.** An
  ambient variable is how this hole existed; the examples take `--model` in
  argv instead, and the old `COMBO_MODEL` is gone.
- **The parent session's model is never inherited** by a subagent, even though
  the extension can see it. A subagent floating with whatever the operator's
  TUI happens to be on is the same bug one level up. The model comes from an
  explicit artifact - an argument or a file.
- **No per-step pipeline `model:`** until a real pipeline needs one: agent
  frontmatter already covers "this role runs on X".

An unresolvable pattern still throws at spawn (a workflow on the wrong model
costs more than a lost run). The commands validate `--model` with
`checkModel()` **before** the interview, in the same early block as
`checkPipelineAgents` - a typo costs a second, not a conversation.
`checkModel` reads pi's real model catalogue: only a run inside a real pi
proves it end to end.

### A model pi does not know is refused, even when pi would try it

For a provider it knows and an id it does not list, `resolveCliModel` does not
fail: it builds a custom model and only warns, `Model "nope-model" not found
for provider "ilaas". Using custom model id.` Measured on pi 0.80.10 and 1.0.2
alike: `/run explore x --model ilaas/nope-model` passed `checkModel`, spawned
four subagents and ended `provider: 404 "Unknown model"`. A typo cost a run
after all.

So `resolvePattern` refuses a model that came back with a warning, in
`checkModel` and at spawn alike. The message keeps pi's reason without its
last sentence, which would claim the opposite of what happens, and says where
a model is made known: `Model "nope-model" not found for provider "ilaas". Add
it to pi's models.json to run it.` A pattern pi rejects outright shows pi's
`error` (an unknown model, a bare id ambiguous across providers) rather than a
message of our own.

It reads pi's fields instead of comparing the model with the catalogue, and
the custom model is the only warning `resolveCliModel` returns beside a model:
it parses a `:thinking` suffix strictly. A partial match (`ilaas/qwen-3.6`)
and a suffixed id (`qwen-3.6-35b-instruct:high`) still resolve, with no
warning. If pi ever warns about something else beside a model, that is refused
too, in pi's words: the safer side to err on. The tests run pi's real
`resolveCliModel` over a catalogue held in memory, so they hold these cases to
the pi that is installed.

### A subagent's settings are a named list, and it never warms a cache

`createAgentSession` given no `settingsManager` builds one from
`~/.pi/agent/settings.json` and the project's `.pi/settings.json`, and a
subagent used to get exactly that: every setting the user had, against
invariant 5. It cost little until pi 1.0 turned prompt-cache warming on by
default. A warm-up is a request pi sends by itself before a cache entry expires,
and pi records it as usage that `getSessionStats()` sums, so it lands in a
turn's delta. Measured on a real pi with a model that declares a 20 s cache
lifetime: one turn spent 25 s in a `bash` call, and 2 warm-ups came back in its
usage (2919 input tokens, about 1400 of them warm-ups). With the user's
`cacheWarming: "idle"`, the same persistent subagent sent 2 more during a 25 s
wait between two `ask()` calls. The same run showed the user's
`shellCommandPrefix` running in the subagent's shell.

A subagent now gets `SettingsManager.inMemory(subagentSettings(...))`. The seed
is pi's own reading of the user's files, global and project merged, cut down to
`INHERITED_SETTINGS` in `src/session.ts`, with `cacheWarming: "off"`. That list
is the exception list, and each entry is either the model ladder or a fact or a
consent that a subagent breaks without:

| Kept | Why |
|---|---|
| `defaultProvider`, `defaultModel`, `defaultThinkingLevel`, `modelThinkingLevels` | the last step of the model ladder above |
| `shellPath` | a machine fact: pi looks for bash by itself, and where that fails (Windows with Git Bash outside Program Files and no `bash.exe` on `PATH`) the `bash` tool throws on every call |
| `httpIdleTimeoutMs`, `retry.provider.timeoutMs`, `websocketConnectTimeoutMs` | a fact about the provider: each request carries its own timeout, so a user who raised it for a slow local model would see subagents cut at 300 s |
| `enableInstallTelemetry` | a consent: a user who opted out would see subagents send attribution headers to OpenRouter, NVIDIA and Cloudflare |

What stops being inherited is everything `createAgentSession` and
`AgentSession` read besides those, and each now has pi's default:

| Setting | What a subagent does now |
|---|---|
| `cacheWarming` | off, whatever the user chose |
| `shellCommandPrefix` | none: a prefix sets up the user's own shell, aliases and variables, and a subagent's commands should not depend on it |
| `compaction.*`, `branchSummary.*` | pi's reserve and keep sizes, even for a model with a small window |
| `retry.enabled`, `retry.maxRetries`, `retry.baseDelayMs`, `retry.maxAgentDelayMs`, `retry.provider.maxRetries`, `retry.provider.maxRetryDelayMs` | pi's retry policy |
| `transport`, `steeringMode`, `followUpMode`, `thinkingBudgets` | pi's defaults |
| `images.autoResize`, `images.blockImages` | pi's defaults |

`defaultTools`, `enabledModels` and `theme` are read too but change nothing
here: combo always passes `tools`, never cycles a model, and a theme only
colours an HTML export. `httpProxy` is not a session setting: the host pi
applies it to the whole process. The in-memory manager also means that a
`setModel` on a subagent's session can no longer write the user's default
model back to their file.

Measurements taken before this change include whatever warm-ups pi sent, and
only for a model that declares a cache lifetime and where pi expected the
warm-up to save at least $0.05.

**Reversed in part: a machine fact comes from the user's file alone.** The seed
first came from pi's merged reading, so the repository's `.pi/settings.json`
won for every key on the list, and pi trusts a project by default. A cloned
repository was enough to choose the binary every subagent's `bash` runs.
Reproduced on a real pi: a scratch repository whose `.pi/settings.json` set
`shellPath` to a wrapper script, and a subagent with `tools: ["bash"]` spawned
in it ran its command through that wrapper. `shellPath` is a fact about the
user's machine, which a repository cannot know, so the table above was already
saying where it belongs.

`INHERITED_SETTINGS` now names a layer for each key, and `subagentSettings`
reads them from pi's `SettingsManager` with `getSettings()` and
`getGlobalSettings()`:

| Key | Layer | Why |
|---|---|---|
| `defaultProvider`, `defaultModel`, `defaultThinkingLevel`, `modelThinkingLevels` | global and project merged, project winning | which model works on a repository can be the repository's call, as it is for pi itself, and the user's keys still decide whether it can run |
| `shellPath` | global | where bash is on this machine |
| `httpIdleTimeoutMs`, `retry.provider.timeoutMs`, `websocketConnectTimeoutMs` | global | how slow the user's provider is, a fact about their setup, and a repository could set it low enough to cut every request |
| `enableInstallTelemetry` | global | a consent is the user's to give, and a repository setting it to `true` would override an opt-out |

The same reproduction after the change: the repository's wrapper no longer
runs, and a `shellPath` in the user's own settings still reaches the
subagent's `bash`.


## Experiments: comparing models on the same work

The model knob only pays off when something uses it to compare. `experiment` is
that something: M models × N repetitions of one workflow, each cell in its own
directory with its own measurements, and one table at the end.

It is **a function, not a combinator**. It returns no `Result` and composes with
nothing, because it is a harness placed above a workflow - one that could be
nested inside a workflow would be measuring itself. Everything else here is
nestable on purpose; this one is deliberately not. It is also not a pipeline
kind: a pipeline is data a model must not be able to author, and a matrix over
models is code the operator writes.

A cell is handed a ready-made `WorkflowOptions` and the contract is to **spread
it**. That is what puts every subagent on the cell's model, in the cell's export
directory, and under the cell's collector - a callback that rebuilds those by
hand silently measures something else. It is also why flows need no support
of their own: `runFlow` takes the same `model`, `signal`, `timeoutMs`, `spawn`
and `onEvent`, so `run: (cell) => runFlow(checked, input, { ...cell.options,
runDir: cell.dir })` is the whole integration, as `runPipeline` was before it.

Three rules the arithmetic depends on:

- **Sequential by default.** `concurrency` is 1 unless asked otherwise: two
  cells racing for the same machine measure the contention, not the models.
- **Sums are stored, means are displayed.** `experiment.json` carries totals
  only; the mean wall and mean cost are derived when the table is rendered.
  Averaging averages is how a study starts lying about itself.
- **A failed cell stays in the report, with its usage** - a callback that threw
  included. It spent tokens before it broke, and dropping it would turn "two
  models out of three answered" into a clean comparison of the survivors. Same
  reason `loop` reports `converged` apart from `ok`.

Flag columns are the union of the outcome keys actually seen, so a study
comparing `converged` gets a `converged` column with nothing configured. `error`
never becomes one: a column of distinct sentences compares nothing.

## A board, if there is to be one, is append-only and stamps its own names

Everything else here passes `Result`s between subagents that never meet. A board
is the other arrangement, and it is worth having only if a run that used one can
be read back afterwards exactly as it happened. That is what fixes its shape
before anything is built on top of it.

**Nothing is rewritten and nothing is deleted.** A member can add to the record,
and that is the whole of what it can do to it. The record is the only thing that
says what happened, so a run able to edit it is a run that cannot be
investigated.

**`from` is stamped by the board, never carried in the draft.** The caller says
who is posting. `Draft` has no such field, and a model that sends one anyway is
not believed - there is a test for exactly that, and it is the load-bearing one
of the file. Identity that can be claimed in a parameter is what makes a medium
one where any member can speak as any other, and no prompt repairs that
afterwards. The same rule as "only whoever raised an obligation may close it".

**All three caps have defaults**: 200 posts, 2000 characters each, 50 per
member. Not because those numbers are right, but because a medium with no cap is
bounded by the deadline and nothing else, which is the same reason
`maxIterations` defaults to 5.

Three smaller things, each one a fork that could have gone the other way:

- **A refusal is `{ ok: false, error }`**, the shape `Ledger.close` already
  returns, rather than a post-or-error union a caller can mistake for a post.
- **`to` names a member, not a topic.** A board that cannot tell the two apart
  cannot refuse an address that reaches nobody, and `kind` already carries what
  a topic would have. A board told no members accepts any address: refusing one
  it cannot check would be guessing.
- **A reader is not handed its own posts.** It wrote them, they are already in
  its context, and the point of the cursor is to spend as little of that as
  possible.

**Reading is announced as loudly as posting**, and that came out of running it
rather than out of designing it. Three members dividing one job posted their
claims 1.5, 3.1 and 3.5 seconds in, so the later two could have read the earlier
ones before choosing - and nothing in the record said whether they had. Posts
answer who said what. An investigation asks who *knew* what, knowing comes from
being handed something, so being handed something is an event: `read` carries
the ids delivered and how many were left waiting. The empty read is recorded
too, being the only thing that settles what a member could not have known.

### The medium announces, whoever acts on it

The tool announced. A member's post, its read and its take went onto the bus
from `board-tool.ts`, and the decision above said reading was announced as
loudly as posting. The swarm, meanwhile, read the board itself at the top of
every round to hand each member what it had not seen, and took back what a
member that had stopped was holding, and neither act was on the bus: the
handout that fed every round was in nobody's record, and the console, the board
pane and `events.jsonl` showed the half of the traffic that went through a
tool. The code was behind the decision, not against it.

`src/board/announced.ts` wraps the medium: `announcedBoard` and
`announcedClaims` return the same `Board` and the same `Claims` with every act
on the bus - an
accepted post, every read with the ids it handed and how many it left waiting,
every grant and refusal, and one release per key for a member that is gone. The
swarm wraps what it was given or what it made, once, and hands the wrapped
medium to the tool and to itself. The tool announces nothing any more.

Wrapped at the swarm rather than built in: a caller hands the swarm a board to
read afterwards, and `/swarm --claim` builds its claims with no bus in reach.
A bus baked into `createBoard` would have left both silent. `board.ts` and
`claims.ts` stay the pure data their headers promise.

Paging moved into the board with it. `since(reader, cursor, limit)` hands back
a page, says how many wait, and moves the cursor past what it handed and never
past what it merely looked at - which the tool used to do on its own after the
fact. A `read` event has to say what was handed over, and only the one that
decides the page can say it.

### A member has one cursor, whoever hands it posts

The swarm hands each member what is new at the top of its turn, and the
member's `read` hands it what is new since its last `read`. Each kept its own
cursor, so a `read` after the handout gave the same posts back. In a real pi
(`/swarm --agent debater`, qwen-3.6-35b-instruct), a member handed one post
read the board at once and got that post again. In round four, a `read`
returned posts from rounds two and three. Weak models read on nearly every
turn, so they paid for nearly every post twice, against what the board header
and the debater's definition promise.

`createReader(board, id)` now holds the member's one cursor. The swarm makes one
per member and gives the same reader to the handout and to the member's tool.
A tool built with no reader keeps its own, as before, because outside a swarm
nobody else hands that member anything.

### One result a turn, when a turn is a vote

`16-debate.ts` on gemma-4-31b, twice. Both times one member, in the first round
and while the other two were still working, read an empty board, posted its
vote, read again, and did that seven times, then eleven, in one turn. Each post
was worded a little differently, so the refusal of an exact repeat did not
catch any of them. The next round the other two were handed nine and eleven
posts, most of them that one vote, and both came round to it. A weak model
counts heads, and the flood gave it heads to count.

`SwarmOptions.resultsPerTurn` caps the `result`s one member posts in one turn.
`turnBoard` wraps the board the member's tool posts to, and the swarm starts the
count over before each ask. The refusal tells the member to end its turn: what
the others post reaches it at the top of the next one. The task says the same
thing from the first round on, so there is no reason to wait.

The cap is an option, not a default. `15-swarm.ts` posts one `result` per file
described, several in a turn when a member holds several files, and a cap of
one would refuse work that is right. `/swarm --until agree` sets it to 1, and
so does the debate: there a result is a vote, and the same vote again is not an
argument. A `tell` is never capped, because that is how a member thinks aloud
without answering anyone. The board's own caps still bound it.

`latestVotes` reads only `result`s now, the kind `VOTE_INSTRUCTION` asks for. A
`tell` answering somebody's vote quotes it on its first line often enough, and
it would have counted as the quoter's own. A cap on results would also be
worthless if a vote could still arrive as a `tell`.

### A member's failed turn is asked again once

A member whose turn failed dropped out of the swarm for good. `/swarm --agent
debater --until agree` on qwen-3.6-35b-instruct: one debater's first turn was cut
by the output limit, in a loop repeating one phrase. It dropped out holding no
vote, so `agreed(3)` could not fire, and the three rounds left ran for nothing.

A failed turn is now asked again the next round, and two failed turns in a row
take the member out. The shipped flows follow the same rule: every agent node
is asked at most twice. Asked again in the same round, the member would see the
same board it failed on. The next round hands it what the others said since,
and the round cap still bounds the whole.

### An answer names the post it answers, in public

Debaters on gemma-4-31b wrote `To p2:` and `To the Rust advocate (p2):` at the
top of their `tell`s. They were answering an argument, not writing to a person,
and the board had no way to say so. `to` would have been worse: it is mail, and
the third debater would never have seen the exchange.

A draft now takes `re`, the id of the post it answers. The board refuses an id
that is not on it, as it refuses a `to` that reaches nobody, and records the
answered post's author beside the id. It stamps that name, like `from`, so the
answer's author never supplies it. An answer is read by everybody.
`boardLines(posts, reader)` puts what answers the reader first, under
`Answering you:`: handed ten posts, a weak model reads the first few, and those
should be the ones that are its business. The handout and the tool's `read` both
use it.

`re` is optional. A post that answers nobody is how a member thinks aloud, and
a debate needs that as much as it needs replies. The refusal of an exact repeat
counts `re`, so the same "I disagree" under two posts is two answers.

### A member does not post the same thing twice

`/swarm --agent debater --until agree` in a real pi, on
`qwen-3.6-35b-instruct`: one member posted the same vote, word for word, six
times in a single turn. Its definition says to post one `result` a turn, and a
prompt is not a boundary. Every copy took a slot on each reader's page and told
nobody anything.

So the board refuses a post whose kind, reader and trimmed text match one the
same member already made, and the refusal names the first one's id. Only an
exact match is refused. Deciding that two phrasings say the same thing is a
judgement, and the board does not make judgements. The same words from another
member, under another kind or to another reader are a different post: a vote
repeated by a second member is the agreement a debate is looking for.

## A claim is granted, never announced

A board lets a member say what it is taking, and that turned out not to be
enough in the first run that recorded reading. Three members read within 123ms
of each other, each was handed nothing because nobody had posted yet, and all
three then claimed the same file. Announcing into a medium that was empty when
you looked is a race, and no prompt repairs a race.

So `claims.ts` arbitrates: first to ask holds it, and everyone else is refused
and told who holds it. The refusal is the useful half: "`src/parser.ts` is held
by scout#3 - ask scout#3, or take something else" turns contention into somebody
to talk to, where a silent loss turns it into two members doing one job.

**The keys are a list, not free text**, and that is the part the same run
argued for: one file came back as `console.ts`, `src/reporters/console.ts` and
`I will handle src/reporters/console.ts`. A lease keyed on what a model writes
would have granted all three and arbitrated nothing. A caller that knows what
there is to claim passes the list, and a key that is not on it is refused with
the list - the discipline `delegateTool` already applies to an agent name. A
caller that cannot enumerate the work passes nothing and gets the weaker
behaviour, which is honest rather than convenient.

**How much one member may hold is a knob, and the number came from a run.** With
nothing bounding it, one member of three took all six keys and the other two
spent their turns being refused. Three arms, two repetitions each, same model,
same goal:

| | most held by one member | grants | mean wall | ↑input |
| --- | --- | --- | --- | --- |
| nothing | 3, 4 | 10.5 | 41.2s | 394k |
| a rule in the prompt | 4, 4 | 11.0 | 43.2s | 434k |
| `maxPerMember: 1` | 1, 1 | 14.5 | 53.3s | 646k |

The middle row is the one worth keeping. "Take one thing at a time, and release
it before you take another", written into the member's own definition, changed
nothing at all: four held at once, exactly as with no rule. That is invariant 7
in its own words - a prompt is not a permission boundary - measured for holdings
rather than for tools.

The bound works and the work still goes round: more grants, not fewer, because
members cycle through take, do, release instead of sitting on everything. It is
**not the default** because it is not free: 29% slower and 64% heavier in input
tokens, since taking, releasing and being refused are all calls. A caller with
no contention should not pay for it.

Two smaller ones:

- **Taking what you already hold is granted.** It is not contention, and a
  member told "you cannot have it, you have it" learns nothing from being
  refused.
- **`releaseAll(member)` returns the keys it let go**, because a member that
  dies holding claims leaves work nobody will do and nobody can take. The run
  says which ones rather than leaving a reader to notice the gap.

The same three members, the same job, with `take` on the tool:

```
   +693ms member#3 take console.ts -> granted
  +1063ms member#2 take console.ts -> refused (member#3)
  +1675ms member#2 take herdr.ts   -> granted
  +3081ms member#1 take console.ts -> refused (member#3)
  +3711ms member#1 take record.ts  -> granted
```

All three went for the same file again, which is the point: the models did not
change, the medium did. Both refusals were absorbed on the next call, inside the
same turn, and three different files were described in half the wall time of the
run that collided. Nobody posted and nobody read: arbitration made the
announcing half unnecessary for this job, which is worth knowing before anything
is built on the assumption that members talk.

## A swarm is built to be compared against not having one

Every other combinator decides who does what. A swarm decides none of it, which
makes it the one whose value is an open question rather than a design. So it is
built to lose honestly: **with no board, no claims and one round, `swarm` is
`fanOut`**. That degenerate case is a test, and it is the control arm of any
experiment run with it - the board's worth is the difference between the two, on
the job you actually have, and not an argument.

Three things are decided in the combinator rather than left to a prompt, each
one from a run that happened before it was written:

- **Members remember.** `lifetime` defaults to `"workflow"` here and nowhere
  else. A member that forgets the round before cannot build on what it saw, and
  a swarm of amnesiacs is a fan-out that costs more. `"task"` with more than one
  round is refused outright: a member's id is its name on the board, and a
  task-lifetime member gets a new one every round, so from the second nobody
  would be talking to who they think they are.
- **What is new is handed over, not fetched.** Measured: once something
  arbitrated, members stopped reading the board entirely - they take, they are
  refused, they take something else. Charging them a call to learn what the
  workflow already knows is charging them for its bookkeeping.
- **A member that is gone holds nothing.** Measured: two of three members never
  released what they took and were still holding after they closed. Claims left
  by a member that is gone are work nobody will do and nobody can take, so the
  swarm releases them and the result says which.

`stoppedBy` is separate from `ok` and from `converged`, for the reason `loop`
separates them: a swarm that ran out of rounds did not succeed, it stopped.

## A swarm reaches pi as a step of the chain

`/run` leaves its answer in the conversation, because an exploration is read and
then asked about. A swarm's answer is not one report: it is what every member
said plus the board they said it on, and in a real run that is a screenful. Put
in the conversation it would be an orchestrator's brief that nobody asked for,
so `/swarm` does what `/step` does - recorded in the relay, drawn in the
transcript, out of the model's context until `/quote`.

That reuse is the whole design of the command. The relay gained a third `kind`
and nothing else: `/chain` lists a swarm, `/quote` brings it in, `/step --from`
carries it on, and the entry renderer already knew how to draw a step.

**`--claim` also gives the run something to be finished by.** With things named,
the swarm stops once each of them has been reported on, rather than spending
every round it was allowed. Measured in a real pi, three members and three
files: 3 rounds and 9 turns with the cap alone, 1 round and 3 turns with the
condition, and three of the first run's six posts described a file another
member had already described. A round cap bounds the worst case; it is not a
plan. With nothing named there is nothing to be done with, and the rounds are
all there is.

**`--until agree` is the other way to be finished, for the jobs that do not
split.** Coverage is a test on work: every named thing reported on. A question
put to three members has no named things, so the only end it has is the three of
them saying one thing. The condition reads the roster rather than whoever spoke,
because two of three agreeing is not agreement, and a member that dropped out
therefore never lets it fire - the run spends its rounds and says so.

The vote is a line, `VOTE: <answer>`, and the sentence asking for it is appended
to the goal by the command. Two alternatives were weighed. A JSON shape would be
parsed more exactly and comes back fenced or explained from a small model, where
a line is what one writes correctly on its first turn. Putting the request in an
agent definition instead would have every other run of that agent posting a vote
nobody counts, and would miss any agent a user writes. So the instruction lives
beside the parser that reads it, in `src/board/agreement.ts`: told in one file
and read in another, the two drift the first time either is edited.

Measured in a real pi, three members on one question: 3 rounds and 9 turns with
the cap alone, 1 round and 3 turns with the condition. What they agreed on is
the part worth reading - all three read the repository they stood in and voted
for its language, which is agreement about a fact rather than an argument
anybody won. Copies of one model share its opinion, so a swarm asked to debate
has to be handed its disagreement, and `--claim` is what hands it out: the camps
are leased one owner at a time, and the vote stays free every round.

**An agent that does not name `board` is run anyway**, with a word saying its
copies cannot reach each other. That is a fan-out, which is precisely the arm a
swarm has to be compared against, and refusing it would remove the control from
the one place someone can try it in a second.

## A directory is a module, and its index is the door

`src/` had thirty-nine files at its root, and the ones that formed a module
together said so only in their headers: `scratch.ts` names `worktree.ts`,
`board-tool.ts` sends the reader to `announced.ts`, `settle.ts` says it is
"what `deliver` does with" `land.ts`. A reader who wanted one concept opened
five files, in the order the headers pointed, out of a flat listing that gave
no hint which five.

The rule: **files whose headers name each other sit in one directory, and its
`index.ts` is the only file anything outside imports**. What the index lists is
the module's interface; what it does not list is implementation, and the
directory is what makes that distinction visible in a listing rather than in
prose. Tests are the one exception, and a deliberate one: they reach past every
door, so a helper exported for a test never has to appear on an index.

Two things about the generated reference made this cheap. `npm run docs` follows
`export … from` transitively and names each page after the **declaring** file,
so re-exporting a directory's index from `src/index.ts` changes nothing on the
public surface and moves the module's pages under its directory. And the
reference's toctree is generated, so a page that moves takes its navigation
with it; only the hand-written links in `docs/guide/` have to follow. What a
move cannot lean on is a test: two constants compute a path from their own
depth, `PACKAGE_ROOT` in `builtin.ts` and the pane's entry in
`reporters/herdr.ts`, and nothing offline reads either. That is why
`builtin.ts` stays at the root whatever moves around it.

### The delivery owns its git policy and its saved state

**Removed:** `workflows/deliver/` went whole with the linear pipeline. See
[The linear pipeline is removed](#the-linear-pipeline-is-removed).

The first module to become a directory was the one whose two files pointed the
wrong way. `resume.ts` sat at the root and imported four files from
`workflows/` - `audit`, `deliver`, `pair`, `plan` - while none imported it back:
every arrow left the root and went down. `settle.ts` sat in `workflows/` and
was not a workflow by this repository's own definition - it takes `{ cwd,
worktree, writers, verify }`, no `spawn`, no lifetime, no signal - and its
header said whose it was: "this is what `deliver` does with it". They are also,
with `pair.ts` and `audit.ts`, the most edited files under `src/` after the
barrel and `subagent.ts`, and `deliver` and `pair` change together in most of
those commits.

`src/workflows/deliver/` now holds the combinator, its two stages, `settle.ts`
and `resume.ts`. Its index lists what the barrel and `pipeline-run.ts` reach:
`deliver`, `pair`, `audit`, the saved-build functions, and the two `settle`
types a `DeliverOptions` names. `settling` itself is not on it; the delivery is
its one caller. `plan.ts` stays in `workflows/`, because `orchestrate` plans
too.

### The extension: commands, the terminal, and the floor

`extension/` had twenty-two files at one level, and the suffix was doing the
directory's job: six ended in `-command`, two in `-ui`, and `build.ts`,
`stop.ts`, `commit.ts` and `stage.ts` carried nothing, so a reader could not
tell a command from what it stood on without opening it. `src/stop.ts` and
`extension/stop.ts` were unrelated files with one name, told apart only by
their tests being called `stop.test` and `stop-ui.test`.

The split follows what a file does with pi. `commands/` holds every file that
registers a slash command, plus the two helpers only one command uses -
`commit.ts` for `/build`, `stage.ts` for `/step`. `ui/` holds what paints or
reads the terminal without registering anything: the live view, the question
card, and the two module-level switches a card and a key share. The root keeps
the floor - `pi.ts`, `deps.ts`, `command.ts`, `flags.ts`, `params.ts`,
`relay.ts` - with `index.ts` and `execute.ts`, because the tool is what the
extension *is* and the floor is what every command stands on. The suffixes
went with the move: the path now says what the name used to.

Two type-only cycles between the floor and its leaves went at the same time,
because a directory makes an arrow's direction visible and both pointed up.
`pi.ts` imported `ExecuteDeps` from the tool body to type what it handed the
tool; `ToolDeps` now lives in `pi.ts`, as the slice of pi's context the tool
reads, and `execute.ts` widens it with what a test may replace. `deps.ts` and
`relay.ts` each held one half of `AppendEntry`; the relay owns the entry a step
leaves, so it owns the door it leaves it through as well.

One arrow still crosses from `ui/` into `commands/`: the live view tells `/stop`
which run is on screen. That is the sign that `stop.ts` holds a key and a
command at once, and it is left in view rather than papered over.

### The board is one module with one caller

Five files at the root of `src/` - `board.ts`, `claims.ts`, `announced.ts`,
`board-tool.ts`, `agreement.ts` - had one caller inside the library,
`workflows/swarm.ts`, which imported nine names from four of them. Their headers
already read as one text: `claims.ts` calls itself "pure data, like `ledger.ts`
next door", `board-tool.ts` sends the reader to `announced.ts` for the bus, and
following "a member claims a file" meant four files in the order the headers
pointed. `agreement.ts` is a policy over a `Board`, the way `settle.ts` is a
policy over `land`, and its one caller is the `/swarm` command.

`src/board/` now holds the five, `board-tool.ts` renamed `tool.ts` since the
directory says whose. Its index lists what `swarm` builds and hands out, and
the vote the command reads; `latestVotes` stays behind it, counted only by the
agreement. The one arrow into the module from the core is `events.ts` naming
`Post`, because a board's traffic is an event like any other; a directory makes
that arrow readable where a flat listing hid it.

### Running a pipeline is not a combinator

**Removed:** `src/pipeline/` went with the linear format. The reason it was
not under `workflows/` holds for `src/flow/run/`. See [The linear pipeline is removed](#the-linear-pipeline-is-removed).

`pipeline-run.ts` sat in `workflows/` and was the one file there that depended
on every neighbour: it imported the eight combinators and dispatched on a
step's `kind`, returned a `PipelineRunResult` rather than a `Result`, and
composed with nothing. Every other file in the directory depends on
`options.ts` and `pool.ts` and on nothing sideways. Its own header said what it
was: "our code walks the steps". Meanwhile `pipeline.ts` and `pipeline-load.ts`
sat at the root, two of the three stages of one thing, in a listing that put
`ledger.ts` between them.

`src/pipeline/` now holds the three in the order they are used - `pipeline.ts`
reads the file, `load.ts` finds it, `run.ts` walks it - and `workflows/` holds
only combinators again. `builtin.ts` stays at the root on purpose: its
`PACKAGE_ROOT` is two `dirname`s up from its own file, so a move would point
the shipped `agents/` and `pipelines/` at `src/` and nothing offline would
notice; the layout block now says what it is for, so the next reader does not
try.

### The working copy is a stack, and reads as one

`git-run.ts`, `git.ts`, `worktree.ts`, `scratch.ts` and `land.ts` were a stack
that said so in every header - "`git.ts` covers the repository; this covers the
copies of it", "`worktree.ts` holds the primitives; this holds the one shape
every caller wants" - and sat at the root of `src/` with `ledger.ts` and
`language.ts` between them. `src/git/` now holds the five, `git-run.ts`
renamed `run.ts`, and its index lists what leaves the directory: the git a
pipeline may do, the landing, the scratch copy `pair` opens, and the worktree
primitives. `run.ts` is not on it; running git is the how, and everything
outside asks what.

Two things did not move, each for a reason worth keeping. `verify.ts` is a
port, beside `ask.ts` - its header says so - and seven files depend on it that
have nothing to do with git; putting a port under a mechanism would read the
dependency backwards. And the worktree primitives stay on the public surface
although only `scratch.ts` calls them inside the library, because the
worktree guide teaches them in its first code block; the rule for the surface
is what somebody outside calls, and a guide is somebody outside.

### The review record's door was already in the barrel

`review.ts` said in its header that it was "the one place that joins" the
verdict and the ledger, and the code agreed: `verdictTool`, `declaresVerdict`,
`createLedger` and `openList` had no caller but `review.ts`, and the barrel
exported the types of all three files and the function of one. The interface
was visible everywhere except in the listing, where the three sat among
thirty-nine files with `resume.ts` and `run.ts` between them.

`src/review/` now holds the three, and its index lists what the barrel already
did: `reviewRecord` and the types a result names. Its three callers are the
delivery - `pair`, `audit`, and `resume` for the types - which is why it could
have gone under `workflows/deliver/`; it did not, because a verdict as a tool
call is one of the three tools combo hands a subagent, beside the board's and
the delegation's, and that family reads better from the root. `tool.ts`, the
words those three share, stays at the root for the same reason.

### Measuring is a chain, and `usage.ts` is not a link of it

`export.ts`, `measured.ts`, `experiment.ts` and `experiment-report.ts` were one
chain - a run directory, a run that measures itself, the same run over M models
and N times, the matrix read back - split across the root of `src/` by
alphabetical accident. `src/measure/` now holds the four, `experiment-report.ts`
renamed `report.ts`, and its index lists what the extension's live view, the
examples and `spawn` reach. The chain had one back-edge, `experiment.ts`
importing its report and the report importing `ExperimentOutcome` back; the
outcome now lives with its reader, and the module reads in one direction.

`usage.ts` is not in it, although the word suggests so. It is what `Result`
and every event carry, imported by fifteen files including `result.ts` itself;
a `measure/usage.ts` would have the core reaching into a feature directory for
its own currency. The one arrow that does run from the core into this module,
`spawn` calling `exportSession`, is the honest one: a subagent's transcript is
exported where it is still in hand.

## The public surface: one entry point, grouped as it is learnt

`src/index.ts` is the only door - the examples and the extension import from it,
never from a file inside `src/`. What changed is that it is now **grouped the way
the library is learnt** rather than alphabetically: start here (an agent, a run, a
result), the combinators, watching a run, measuring a run, flows, pipelines, the ports
that touch the world, the pi session. A reader who needs a dozen symbols finds
them in the first section instead of scanning ninety.

Two rules decide whether a symbol belongs on that list at all:

- **A type named by a public option or return value is public.** `HerdrSend` is
  `HerdrOptions.send`, `HerdrEnv` is what `detectHerdr` returns, `CreateSession`
  is `SpawnOptions.createSession` - remove any of them and the option cannot be
  written from outside. This is why the option and result types of every
  combinator stay, even though nothing in this repository names them.
- **Test-only is a reason to stay off it.** Tests reach into `src/` directly, so a
  helper exported for one is not part of the surface. Eleven symbols left on that
  basis: the package's own directory constants (`PACKAGE_ROOT`,
  `BUILTIN_AGENTS_DIR`, `BUILTIN_PIPELINES_DIR`), the herdr transport under
  `detectHerdr` (`createHerdrSend`, `HERDR_SOURCE`), the leaf formatters
  `widgetRows` already composes (`currentActivity`, `detailLine`, `elapsedMs`,
  `widgetLines`) and the id counters (`nextSubagentId`, `resetSubagentIds`).

Both rules are stated at the top of the file, because the next person adding an
export will read that before they read this.

### A value is public when somebody outside calls it

The two rules decided types and test helpers, and said nothing about a value
nobody called. The file was touched in twelve of twenty-five commits, every
time to add, and of 168 exported values 75 were referenced by nothing in the
extension, the examples or the scripts: every prompt builder, the word
constants, the four working-copy layers, the record and the ledger, the board
and its announcing wrappers, the mirror, the event bus, the pipeline directory
constants. The list had become a measurement of feature count. A reader
learning the surface learnt a third of it for nothing, and the generated
reference documented it.

The third rule: **a value is public when somebody outside `src/` calls it** -
the extension, an example, a guide's code block, or the TSDoc of a public option
that names it as the default a caller may wrap (`formatBranches` for
`ReduceOptions.format`, `pickDestination` for `RouteOptions.parse`). Two more
were kept on the same footing although the rule as computed missed them: the
guide for writing a combinator tells its reader to open a `Trail` and to compose
offers with `offerBoth`, and the file's own comment names `createDefaultSession`
as what a caller wraps to reach pi's session. `succeeded` stays as the pendant
of `failed`.

Forty-seven values left: `announcedBoard`, `announcedClaims`, `boardTool`,
`BOARD_TOOL`, `verdictTool`, `declaresVerdict`, `VERDICT_TOOL`, `SUBAGENT_TOOL`,
`reviewRecord`, `createLedger`, `openList`, `latestVotes`, `scriptedAsk`,
`busFor`, `createEventBus`, `experimentReport`, `writeExperimentReport`,
`copyMainSession`, `exportSession`, `applyPatch`, `currentBranch`, `headSha`,
`landable`, `scratchWorktree`, `settling`, `answerInTheirLanguage`,
`inTheLanguageOfTheWork`, `ANSWER_IN_THEIR_LANGUAGE`, `mirrorSocket`,
`registerMirror`, `REFUSED_IDLE`, `loadPipelinesFromDir`, `PIPELINES_DIR`,
`STEP_KINDS`, `abortError`, `loadBuildState`, `BUILD_STATE_FILE`, `skillDirs`,
`accumulate`, `snapshotUsage`, `AUDIT_APPROVAL`, `auditPrompt`, `answerPrompt`,
`briefPrompt`, `questionPrompt`, `remarksPrompt`, `makePlan`. Every one is still
exported by its own file and reached by its tests there; what changed is what a
reader is told they may call. The one reference page whose module kept no public
symbol, `announced`, went with it, and no guide linked to it. The day a script
needs one of these back, the rule says how: show it in a guide, or call it from
an example.

`latestVotes` came back that way. `examples/16-debate.ts` prints each member's
last vote, and it had been written with a parser of its own that stripped
decoration only after the answer: `VOTE: **Rust**` read as `**rust`, so it and
`VOTE: Rust` were two votes, and a debate could spend its rounds on an agreement
it had already reached. The example calls `agreed` and `latestVotes` now, and
the board's door lists the second.

## A default is written after the spread, never before

`pair` and `interview` default their lifetime to `"workflow"` - two agents in a
conversation keep their memory unless the caller says otherwise. That default
used to be written as `{ lifetime: "workflow", ...options }`, which is correct
for a caller that types its options by hand and wrong for every caller that
builds them by merging.

`runPipeline` is such a caller. It laid a step's overrides on top of the run's
options with `lifetime: step.lifetime ?? workflow.lifetime`, and when neither
was set the key still existed, holding `undefined`. Spread over the default,
that `undefined` won: the same `deliver`, run from code, gave its pair one
worker and one reviewer for the whole conversation; run from a pipeline file, it
gave them a fresh pair every round. A reviewer that never remembers its own
remarks is not a detail, and nothing in the output said so - the run simply cost
several times more turns.

Two rules came out of it, and both are now tests. **A default belongs after the
spread**, as `options.x ?? default`, so that an explicit `undefined` reads as
"nobody set this" rather than as a choice. And **a merge must not invent keys**:
`override()` in `pipeline-run.ts` copies an override only when it is defined,
because `{ lifetime: undefined }` and `{}` are the same intent and must become
the same object.

The general form is worth stating, since the next merging caller will be a new
extension command: in this codebase, absent and `undefined` mean the same thing,
and any code that turns the first into the second is a bug even when the types
allow it.

## The pages, split into a guide and a reference

Eleven hand-written pages sat in one flat directory next to the generated
`docs/api/`. They are now `docs/guide/` - task by task, in the order the library
is learnt - and `docs/reference/`, holding the generated API beside the list of
examples. `index.md`, `development.md` and `decisions.md` stay at the root: they
are the way in and the two pages about the repository rather than about using it.

The split is the question a reader arrives with. *How do I make two agents argue
until they agree* and *what does `fanOut` take* are different questions, and a
flat directory answered neither first - it offered fourteen file names, sorted
alphabetically, of which the second was `build.md` and the third `decisions.md`.

`scripts/api-docs.ts` computes its links out of `DOCS_DIR` instead of spelling
them, so the two that leave the generated tree - the design decisions and the
README - follow the next move on their own. That was the actual cost of this one:
the pages moved with `git mv` in a second, and the links took the afternoon.

## The documentation, as a site

The pages were always Markdown in `docs/`, read on GitHub and in the published
tarball. They are now also a Sphinx site - MyST Markdown, the furo theme, built
by `make -C docs html` with `-W`, so a warning fails the build exactly as a
failing test does.

**Sphinx over the same files, not a second copy.** Nothing was written twice: the
site renders the pages that were already there, and the only Sphinx-specific
syntax in them is the toctrees and the cards on the landing page. A page that
reads well in a repository and badly on a site is a page with two audiences and
one author; this way there is one file per subject, whatever is reading it.

**Python is a documentation dependency, and says so.** It lives in
`docs/requirements.txt`, never in `package.json`. `npm test` and `npm run
typecheck` do not reach the directory, and the library still depends on the pi
SDK alone - which is the rule that made this worth stating rather than assuming.

**`guide/` and `reference/`.** Task-oriented pages moved under `guide/`, and the
generated API under `reference/api/` beside `reference/examples.md`. The split is
the question a reader arrives with: *how do I do this* has a different shape from
*what does this export do*, and eleven pages in a flat directory answered neither
first. `scripts/api-docs.ts` computes the links out of `DOCS_DIR` rather than
spelling them, so the next move is one constant.

**The toctrees are the navigation, and the only one.** `docs.json` listed every
page for `test/docs.test.ts` to check reachability; Sphinx needs the same list as
`toctree` entries, and two lists of the same pages disagree the day someone edits
one. The JSON went, and `scripts/doc-links.ts` reads the toctrees instead - so
the offline suite still fails in seconds on a page nobody can reach, and it fails
on the list the site actually uses.

**What `-W` caught on the first build**, and neither the suite nor a reader
would have: ninety-three code blocks that failed to highlight, because `{ … }`
is not TypeScript a lexer accepts, and a dead link at the top of every generated
page. The first is why a signature now elides with `{ /* … */ }` - a comment, so
the block stays lexable - and why a long initialiser keeps its last line: the
bracket it closes. The second is why "Source:" is an absolute URL into the
repository rather than `../../../src/<module>.ts`: that path resolves in a
checkout and in the tarball, and is dead on a site that publishes `docs/` alone.
It is read from `package.json`, so the repository is named once.

**The site is built in CI, and published from `main` alone.**
`.github/workflows/docs.yml` is the first workflow this repository has had, and
it exists because `-W` is only a standard if something enforces it: a build that
runs on one laptop is a build that breaks quietly. Every pull request builds the
site; only `main` deploys it to GitHub Pages. The workflow also regenerates
`docs/reference/api/` and diffs the result, because the site publishes what is
committed - and a reference that no longer matches the source is exactly the
failure the generator was written to prevent, arriving by a different door.

### The type

Three faces, one job each: **EB Garamond** for what is read, **Inter** for what is
navigated - headings, sidebar, tables, cards - and the reader's own monospace for
what is typed. The third is not shipped on purpose: code is read in the face
someone has already chosen for code, and a page that overrides it is arguing
about the wrong thing.

The other two **are** shipped, and that is the decision. A font CDN would tell a
third party who reads this documentation and would leave every page waiting on a
host nobody here controls; system stacks cost nothing and give a different page on
every machine. So the files live in the repository, cut down to the characters
these pages use by `scripts/subset-fonts.py`, from a pinned commit of
`google/fonts` - never `main`, which moves - with the OFL text beside them, as
that licence requires.

**Static instances, not variable.** Measured, subset the same way: variable was
480 kB for three files, static 272 kB for five. A variable font pays for every
weight between 400 and 800 whether or not a stylesheet asks for one, and this one
asks for four weights in total.

**A subset is a silent failure waiting to happen.** A character no shipped face
carries is drawn from whatever the reader has installed - different weight,
different baseline, and visible to them alone. So the script writes
`coverage.json` from the cmap of the files it actually produced, and
`test/fonts.test.ts` fails on a page whose prose needs more. The serif stack also
names Inter before any system face, so a character only one of the two carries
still lands in a face this site ships.

Two details the faces themselves forced. EB Garamond is a sixteenth-century
design with a small x-height, so `article` sets its own size rather than furo's -
16px of it reads a size smaller than 16px of anything drawn for a screen, and
raising the root size would have shrunk nothing but grown the entire chrome. And
its figures are old-style, which is right in a sentence and wrong in a column, so
tables ask for lining and tabular ones.

### The mark

combo had no drawing of its own. It has one now: **three bars, one bracket** -
the bars identical and in the ink, because a fan-out has no favourite branch,
and the bracket in verdigris, because holding the three as one is the claim of
the library. `docs/_static/logo/README.md` holds the palette, the file table and
the two rules that are easy to get wrong: a two-tone drawing needs a file per
ground, and `currentColor` never reaches an SVG referenced as an image.

**The bracket replaced the first mark, three strokes in, one out.** The strokes
read as a generic merge glyph, and their curves closed into a blob below 24px.
Of the directions compared against it, the bracket is the one that states the
library plainest - three equal things made a group by one device - and the only
one legible at 16px with a single geometry, in rectangles only, which is
trysquare's vocabulary. The proposal drew that bracket in brass; here it takes
verdigris, because the accent is the identity and brass is trysquare's.

**The wordmark is geometry, not type.** Circles on a 20-unit x-height at one
stroke width, rather than glyphs outlined from a font. Outlining is the usual
answer, and it costs a vendored typeface, a licence to check and a generator to
run before the lockup can be rebuilt. Drawing it costs a paragraph of
construction notes - and, like an outlined wordmark and unlike a `font-family`,
it cannot fall back silently on a reader who lacks the face.

The neutrals are [trysquare](https://github.com/AI-for-dev/trysquare)'s, and the
site is shaped like its documentation, deliberately: two tools by the same hand,
meant to be read together, cost a reader more when they look unrelated than they
gain by being distinct. What differs is the accent - brass there, verdigris here.

## One package, and what would split it

**The library and the extension ship in the same tarball.** pi lets one package
be both: `exports` answers `import … from "@ai-for-dev/combo"`, and the `pi`
manifest in `package.json` names `extension/index.ts`, which is what
`pi install npm:@ai-for-dev/combo` loads. A real pi says so on startup, listing
the extension with no `-e` on the command line.

Two packages was the other candidate, and nothing pays for it today. The two
halves share every dependency they have, and all of those are packages pi
bundles, so a separate extension package would weigh nothing less. `extension/`
and `pane/` reach into `src/` from twenty files: split them and every internal
change becomes a version bump across a boundary, where a mismatched pair fails
at runtime rather than at the typecheck. What would pay for it is the extension
needing a dependency the library does not, or the two wanting different release
cadences. Neither has happened, and rule 11 covers the rest.

**The published name carries a scope** because `combo` was taken on npm in 2011.

**The packages pi bundles are peer dependencies with a `*` range**, which is
what pi's packaging documentation asks for: an installed copy binds to the pi
it is loaded into, never to a second one of its own. The oldest pi combo
accepts is therefore checked when the extension loads, not declared here.
`@earendil-works/pi-tui` and `typebox` were imported and declared nowhere at
all. In a clone they resolve by transitivity,
and an installed copy is exactly where that stops being true.

## The version comes from the commit titles

Release Please reads the conventional commits landed on `main` and opens a pull
request that is the next release. **Its workflow starts by hand, never on a
merge**, so an ordinary pull request leaves nothing behind it. The price is that
a release takes two runs: the first opens the pull request, and the second,
after it has been merged, tags the release. Running on every push would have
removed that second run, and would have put a release pull request in the way of
every merge. That is the more expensive of the two.

**Publishing is a workflow of its own**, woken by the release event and by a tag
handed to it, and it stages rather than publishes. As a second job beside the release it was reachable only while
that one run existed, and the first release proved why that matters: the tag was
written, the publish failed on a credential, and nothing but re-running that
exact run could finish it. It carries no npm token either. Trusted publishing
authenticates it by the identity GitHub gives the run, which is one fewer secret
to hold and to rotate, and the provenance attestation follows from the same
identity.

What no credential supplies is proof of presence, and npm asks for it on a write
to this package. `npm stage publish` is the shape that admits it: the run builds
and uploads the tarball, and `npm stage approve` finishes the job from a machine
where somebody can answer. The alternative was a token allowed to bypass the
second factor, which is what npm is in the middle of taking away, so the manual
step is not a workaround to be removed later. It is where a release genuinely
stops being automatic.

The cost is a vocabulary. Pull requests are squashed here, so each title becomes
a commit title, and `feat:` or `fix:` has to be written on it rather than
inferred from the diff; `check-pr-title` refuses a title that cannot be parsed.
The titles already in the history are prose, which is why `bootstrap-sha` names
the tip of `main` at the time: the first release reads what comes after it.

`versioning-strategy: always-bump-minor` was in the configuration this one was
taken from, and is not here. It suits a project released as a whole. It does not
suit a package whose readers write `~0.1.0` and would then never be offered a
fix.

## A deadline travels in the abort's reason

pi cannot tell a deadline from a stop: both are `session.abort()`. A subagent
told the two apart only when the deadline was its own `timeoutMs`. A flow
bounds each attempt with a signal of its own and passes it as the turn's
`signal`, so its journal said `timeout` while the subagent's `Result`, and the
`usage.json` built from it, said `aborted`.

The deadline now says what it is in the signal's reason, a `TimeoutError`, as
`AbortSignal.timeout` gives, and `src/deadline.ts` both writes that reason and
reads it back. The subagent reads the reason of whichever signal fired first,
so a caller's abort before the deadline still reads `aborted`, and a person's
stop still reads `stopped`. The other way was to pass the flow's bound as
`timeoutMs`, which would have taken away the flow's own deadline port, the
switch a dry run throws when its script says an attempt timed out.

## A deadline is worded once, as the flow file writes it

One fact was worded three ways: `timed out after 20000ms` in the subagent's
failure and `usage.json`, `no answer within 20000 ms` in a flow's journal, and
`ran past its bound of 20000 ms` for a check. A turn that hit its deadline
read one way in the journal and another in `usage.json`, two files a person
reads side by side.

`src/deadline.ts` now builds the phrase, `timedOutAfter`, and every bound that
fires uses it: the reason a deadline aborts with, an agent node's failure, an
`ask` with no `default:` (which adds that clause after it), a check. The words
are `timed out after`, which the subagent already said. The bound is shown the
way a flow file writes it and the plan prints it, `30m` rather than
`1800000ms`, so the error reads like the `timeout:` a person wrote and the
`timeout 30m by default` the plan showed them. A bound that is not whole
seconds, which only a `timeoutMs` passed from code can give, stays in
milliseconds: the plan's formatter rounds up, which suits a worst case and
would misstate a deadline. That formatter moved out of `src/flow/` into
`src/duration.ts`, because a subagent's deadline exists without any flow.

## A frontmatter lifetime holds inside a workflow

The record said it from the start: a frontmatter `lifetime` is the agent's
default, and an explicit argument wins over it. `spawn` did that. The pool did
not: it passed `options.lifetime ?? "task"` to every spawn, so it chose the
default before it looked at the agent. In a `loop` that named no lifetime, the
shipped `coder` and `reviewer`, both `lifetime: workflow`, got a fresh subagent
every iteration, and the `spawn` events said `task`.

The pool now resolves the lifetime per agent, with the same rule as `spawn`
(`lifetimeOf` in `src/agent.ts`, the one place that writes it): the workflow's
`lifetime` when it names one, then the agent's frontmatter, then `"task"`. This
keeps invariant 6. A `lifetime: workflow` in a definition is persistence asked
for, in the file that defines the agent, and a workflow that wants every agent
fresh says `lifetime: "task"`. The pool closes the subagents it opens either
way: `"task"` after the turn, the rest in `closeAll()`.

Two paths still set the lifetime themselves, on purpose. `run` is one turn and
forces `"task"`. A flow node reads `memory:`: it runs as `"task"` unless it
names a scope, because in a flow the file says which nodes share a subagent,
and a frontmatter default would let an agent definition change how a flow
reads. `interview` and `swarm` keep their own `"workflow"` default, which they
pass as the workflow's lifetime, so it wins over a frontmatter the same way a
caller's would.

## `candidates` beside a field only orchestrate reads is an orchestration

The tool inferred `route` from `candidates` whatever stood beside it, so a call
with `candidates` and `reduceWith` ran a route. The route never read
`reduceWith`, no synthesiser was spawned, and nothing said so. The schema did
not say that `orchestrate` needed an explicit `mode`, and the `task`
description named three of the seven modes that read it.

`candidates` alone is still a route, the cheaper of the two. Beside
`concurrency`, `maxTasks` or `reduceWith`, the fields orchestrate reads and a
route does not, it is an orchestration: a route would drop that field, so the
call has only one reading. The other way was to keep `orchestrate`
explicit-only and say so in the schema. That still let a call that meant an
orchestration run as a route and drop a field.

`extension/params.ts` now keeps one table of which modes read each field.
The descriptions a model reads are written from it, the set of fields that
make an orchestration comes from it, and so does the list `flow` mode accepts.
A test does not take the table on trust: it runs every mode, records which
fields the tool body reads, and checks each description against that.

## A flag written wrong is refused, not read

`parseLeadingFlags` read a switch written `--name=<word>` as on for any word
but `false`. `/step --agent=no explore where usage is measured` forced the
agent, the opposite of what was typed, and refused with `Unknown agent
"explore"` when no agent had that name. A switch now takes `--name`,
`--name=true` or `--name=false`, in any case. Any other value is refused, and
the refusal says what the switch takes. `yes`, `no`, `on`, `off`, `1` and `0`
are refused too. Nothing in this record asked for them, one spelling of each
answer is enough, and a word the parser does not read is better refused than
guessed at. An absent switch still arrives as nothing, so a switch keeps the
three answers `--worktree` was given.

A count went the same way. `/swarm` refused `--members x` in its own words,
and `/interview` dropped `--questions x` without saying so. The reason given
for dropping it was that `0` would skip the interview, but running on the
default of 6 is also a guess, and nobody is told. A count is now a whole
number of at least 1, checked in `flags.ts` for every command that names one.
It is written in digits: `1.5`, `0x10` and `1e2` are all refused.

The parser returns the refusal as `refused` and does not throw. Each command
prefixes it with its own name and shows it through `refuse()`, as a warning,
the same level as its other refusals of what was typed.

## Selecting a subagent with shift+↑↓

pi 1.0 runs its TUI fullscreen by default, and in fullscreen `ctrl+↑↓` jump
between prompts in the transcript. That holds on every platform. pi's
`docs/keybindings.md` says `ctrl+up` is bound only on Windows and WSL, but
`dist/core/keybindings.js` binds `["ctrl+shift+up", "ctrl+up"]` everywhere
else. The fullscreen screen registers its input listener when it is built,
before any extension's, and consumes those keys, so the selection described
under "Stopping a run" never moved: measured with `drive-pi.py`, `ctrl+down`
marked `▸ scout#2` with `--tui-mode regular` and marked nothing in fullscreen.
`esc` was checked the same way and has no such problem. With the transcript
search open, pi's listener takes the first `esc` to close it and the run goes
on. The next `esc` stops the run.

**Selection moved to `shift+↑↓`.** pi binds them to nothing, in either mode,
and its editor does nothing with them either: with a cursor in the middle of
a line, `shift+up` and `shift+down` left it where it was. The other two keys
keep their defaults.

**The four keys are rebindable, through pi's own file.** A key that
cannot be changed is a collision waiting for the next pi release. pi has no
way for an extension to declare a keybinding, so `extension/keys.ts` names
four ids under `combo.` and reads them from `getUserBindings()` of the manager
pi-tui's `getKeybindings()` returns. pi keeps every entry of
`keybindings.json` there, ids it does not know included, and `/reload`
re-reads the file. **That is undocumented**, observed on pi 1.0.2, and it is
the only place in the extension that relies on it. The fallback is silent and
per id: no manager, no `getUserBindings`, or a value that is neither a key nor
a list of keys, and that id keeps its default. A typo in a file must not cost
anyone the key that stops a run. A combo file of its own was the other way,
and it would have been a second place to bind keys, beside the one pi users
already know.

The hints name the keys bound now, read at each paint and written the way pi
writes a key (`shift+up/shift+down`). A key that is unbound is not offered.
In a flow's widget the line goes one key a line when the terminal is too
narrow for it, as a card's help line does. `/stop`'s description names no
key: it is written once, as the extension loads, and would go stale after a
rebind.

Key releases are ignored. pi asks a terminal for the kitty keyboard protocol
with event types, so a terminal that speaks it reports a release after each
press, and both reach an extension's input listener. pi-tui's `parseKey`
reads a release as the key itself, so without the check a press and its
release would move the `▸` twice.

**What replaces this.** The issue drafted for pi asks for either docs that match
the bindings, a way for an extension listener to run before the fullscreen
screen's, or (point c) keybinding definitions an extension declares, which
`keybindings.json` overrides and `/hotkeys` lists. With point c, the four ids
are declared there, read with `getKeys()`, and `keys.ts` keeps only its
defaults.

## An agent can require a tool call, and `tools: []` means none

On `ilaas/qwen-3.6-35b-instruct`, scouts asked where a subagent's wall time is
measured answered without calling a single tool in 5 of 20 runs through
`run()`, and the first scout of `/run explore` did so in 3 of 3. They wrote
prose, pretend `bash` blocks, or "the directory does not exist", and combo
reported each one `ok: true`. The synthesiser then built its answer on paths
nobody had read.

combo cannot judge a text, but whether a turn called a tool is a count pi
keeps: `getSessionStats().toolCalls` counts the calls in the transcript,
compacted ones included. `sessionPort()` takes the delta of that count over
a turn, the same way it takes the tokens, so `Usage.toolCalls` is on every
`Result`, sums in a fan-out and lands in `usage.json`, for every subagent.

**The count alone did not settle anything.** Shown for every agent that had
tools, it marked the synthesiser on every `/run explore` (9 runs of 9 called
no tool), because every shipped definition names read tools and answering
from what it was handed is the synthesiser's job. The same fact meant "did
not do the work" for a scout and "did the work" for a synthesiser, and only
the definition knows which. So the definition says it.

**`mustCallTool: true`.** A turn of an agent that declares it and calls no
tool fails: `ok: false`, the error `called no tool (mustCallTool)`. That is a
failure in the sense invariant 8 already gives one, a turn that did not do
what it was for, and it is decided by our code from pi's count, never from
what the model wrote. In a flow it is the failure kind `no-tool`, and
`retry:` covers it like `schema`: the turn is asked again with the failure
named, and a fan-out's `on-fail: continue` shows the item as failed rather
than handing its guess on. The name follows the frontmatter's own
convention, camelCase like `openInHerdr`, and it names what is checked: a
call. `readsFirst` was the other candidate, and it promised more than a
count can hold, since any tool counts, `submit`, `verdict`, `board` and
`subagent` included.

The displays and `usage.json` say `called no tool` of a declaring agent
only. Nothing is said of the others, whose count stays recorded. A dry run
does not hold a scripted text answer to the key, since the script stands for
the whole turn, reading included; `{ fail: "no-tool" }` scripts the failure,
for an agent that declares it.

**`tools: []` now means no tools.** An empty list used to fall back to the
read-only default, so "no tools" could not be written. A blank `tools:` still
reads as saying nothing and keeps the default, like every other list in a
definition. This is a breaking change for a definition that wrote `tools: []`
and relied on getting read tools. A definition with `mustCallTool: true` and
`tools: []` is refused, since none of its turns could succeed.

**Who declares it.** `scout` locates code, `explorer` splits its reading
across scouts, and `auditor` is told that the reports it is handed are
claims, not evidence: an answer from any of them that read nothing is one
that was not done. `reviewer` does not, since it can judge the diff it is
handed, and in `build` its `verdict` call would meet the key anyway. Nor
does `member`: a swarm member may end a turn with nothing left to take, and
failing that turn would stop a member that was right.

No shipped agent's `tools:` changed. The synthesiser and the committer
answer from what they are handed, and the synthesiser called no tool in any
of the 9 `/run explore` runs measured. `tools: []` would make that a property
of the file. It would also take away the reading the synthesiser's prompt
allows when two reports disagree, so it is a separate decision.
