# Usage and configuration reference

Return to the [English README](../README.md) or [Chinese README](../README.zh-CN.md) for installation. This reference documents the extension-managed Jev path and the separate explicit-spec SDK. The underlying subagent engine comes from Luke Parke's upstream extension; Jev routing and startup capability verification are additions in this fork.

---

### Credential setup

Version `0.10.0` and later read the TypeSafe key from `jevRouting.apiKey` in your private, user-level `~/.pi/subagent.json`. Set the key locally using an editor, preserve unrelated configuration and replace the placeholder in the [routing example](#jev-routing). Do not paste the key into chat, shell commands, source control or task descriptions.

This file stores the credential in plaintext. Restrict access to your user account and protect editor backups and synchronized copies. Same-user processes, including children with filesystem access, may read it; neither profiles nor worktrees provide an OS sandbox. See [SECURITY](SECURITY.md#routing-disclosure-and-credentials).

When upgrading from published npm `0.9.0`, remove `apiKeyEnv` and add `apiKey` with the actual key locally. The old field is rejected even if both fields are present. There is no environment fallback, automatic migration or default key. Published `0.9.0` still requires its older `apiKeyEnv` setup, while `0.10.0` uses the config-file credential contract.

Reload or restart Pi after changing extension code. Subsequent dispatches re-read the config file, so a later key or routing-destination edit takes effect on the next decision without setting an environment variable or restarting the process. Provider credentials for execution models are configured independently using Pi's own authentication mechanisms. Rotate credentials exposed in source, logs or conversation.

---

### Quick usage

These are request objects for the `subagent` tool, not shell commands. New work calls Jev; omit `model` and `fallback_models`.

#### Foreground and named agents

Use an explicit read-only profile for exploration. Single tasks otherwise default to `general`.

```json
{
  "task": "Find all call sites of parseConfig and summarize patterns.",
  "description": "Map parseConfig usage",
  "profile": "explore"
}
```

A named agent supplies a persona and non-model defaults; Jev still chooses the model and tools.

```json
{
  "task": "Review this diff for security issues.",
  "agent": "reviewer",
  "profile": "review"
}
```

#### Parallel tasks and synthesis

Parallel tasks default to `explore`. Optional synthesis starts an additional read-only child and has its own routing selection. Raw worker results remain available if synthesis fails.

```json
{
  "tasks": [
    { "task": "Audit backend error handling.", "description": "Backend audit" },
    { "task": "Audit frontend error handling.", "description": "Frontend audit" }
  ],
  "synthesis": "Merge both audits into one prioritized findings list."
}
```

Omit `synthesis` when you only need the separate worker reports.

#### Background work and collection

Start a background run and keep its returned run ID:

```json
{ "task": "Audit dependency licenses.", "profile": "review", "async": true }
```

Inspect without consuming the deliverable:

```json
{ "action": "status", "id": "<run-id>" }
```

Collect the result with the same tool:

```json
{ "action": "wait", "id": "<run-id>" }
```

Or pass this request to `subagent_wait`, which delegates to the same collection handler:

```json
{ "id": "<run-id>", "timeout_ms": 30000 }
```

An interrupted or timed-out wait does not cancel or consume a still-running task. A real foreground dispatch is registered before local preflight and Jev selection, so a routing/setup timeout or failure remains visible under its run id; use `async: true` when the work must outlive the initiating call. Cancel it explicitly when needed:

```json
{ "action": "cancel", "id": "<run-id>" }
```

#### Inspect a paid plan

A plan runs local preflights and Jev selection without spawning a child or creating a run entry. It can incur selector fees; a later dispatch selects again. It checks the same model/tool/budget/isolation resolution as a launch. Each resolved plan line reports the supplied `difficulty` (or `(unspecified)` when none was given).

```json
{
  "action": "plan",
  "tasks": [
    { "task": "Implement feature A.", "profile": "general", "isolation": "worktree" }
  ]
}
```

#### Task difficulty

`difficulty` is an optional descriptive hint added to Jev's routing context. It never hardcodes a model, reorders your candidates, changes tool permissions, or alters retry/failover behavior.

| Value      | Use for                                                                                                                          |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `simple`   | Bounded, well-specified read-only review; documentation/format checks; local verification.                                       |
| `moderate` | Multi-file analysis; ordinary bug fixes; focused research with known boundaries.                                                 |
| `complex`  | Architecture/design; cross-layer implementation; difficult debugging with an unknown root cause; high-risk or broad changes.     |

```json
{
  "task": "Review the auth diff for security issues.",
  "profile": "review",
  "difficulty": "simple"
}
```

Choose the lowest honest level. `simple` lowers the chance that a simple task is routed to an unnecessarily large model, but it does not guarantee a specific model: candidate descriptions, profile, thinking and Jev's probability ranking still decide. Omit the field when you have no honest signal; older callers stay valid and no default difficulty is inferred. An invalid value is rejected locally before any selector request. `difficulty` is request context only; it is not stored in results, session entries or run snapshots.

#### Structured output

The child must produce a fenced `json:result` block matching the requested schema. After an otherwise successful answer, the parent validates it and allows one repair round. An unresolved schema failure preserves raw output as a partial result rather than discarding paid work. A terminal provider failure does not trigger a repair prompt or publish a structured result, even if its text contains valid JSON; it follows the availability-failure rules instead.

```json
{
  "task": "Audit the auth module.",
  "profile": "review",
  "output_schema": {
    "type": "object",
    "required": ["findings", "risk"],
    "properties": {
      "findings": { "type": "array", "items": { "type": "string" } },
      "risk": { "type": "string", "enum": ["low", "medium", "high"] }
    }
  }
}
```

#### Resume, fork and steering

A fork starts from a branch of the persisted parent conversation. It is single-task only:

```json
{ "task": "Implement the plan we agreed on.", "context": "fork", "profile": "general" }
```

Find a resumable child session ID in status, then start a new invocation:

```json
{ "task": "Continue from your findings and propose a fix plan.", "resume": "<session-id>" }
```

Both operations select again. To guide an existing child instead, steer it; parallel runs use `index` to select the worker:

```json
{ "action": "steer", "id": "<run-id>", "index": 0, "message": "Skip tests; focus on src and wrap up." }
```

#### Budgets and retries

A budget breach requests a final answer and allows the configured grace turns. `timeout_ms` is the absolute task deadline across local preflight, Jev routing, setup, queueing and child execution; status and terminal output distinguish a pre-spawn `timeout (routing)` from child `queued`, `starting`, `running` or `cancelling` timeout. Before any tool starts, a recognized model-availability failure can advance to the next candidate in Jev probability order. All attempts share the selected tools, task deadline and cumulative execution budgets. Switching does not call Jev again.

`max_retries` is the total number of extra child attempts: `0` permits the initial attempt only, `1` permits at most two attempts, and `2` permits at most three. The built-in default is `1`; task, agent, profile and configuration overrides still apply. The list never wraps back to an earlier candidate. Pi's own in-process/provider retries are separate and unchanged, so a child can make multiple provider requests before the extension sees its final failure. See [ranked failover](#probability-ranked-failover) for the failure boundary.

```json
{ "task": "Audit dependencies.", "profile": "review", "max_turns": 15, "grace_turns": 2, "max_retries": 1 }
```

#### Isolated writers

Use separate worktrees for parallel changes:

```json
{
  "tasks": [
    { "task": "Implement feature A.", "profile": "general", "isolation": "worktree" },
    { "task": "Implement feature B.", "profile": "general", "isolation": "worktree" }
  ]
}
```

Inspect each finished patch:

```json
{ "action": "diff", "id": "<run-id>", "index": 0 }
```

Apply it as uncommitted changes in the parent checkout:

```json
{ "action": "apply", "id": "<run-id>", "index": 0 }
```

Discard an unwanted worktree and branch:

```json
{ "action": "discard", "id": "<run-id>", "index": 1 }
```

The `/subagents` overlay provides the same loop: `s` to steer, `a` to apply and `x` to discard.

---

### Side questions (`/btw`)

Ask a side question in Pi:

```text
/btw does this repo have a rate limiter?
```

Or open the question prompt:

```text
/btw
```

`/btw` runs a one-off read-only subagent for _you_, not for the model. It uses the same policy, budget, semaphore and process-lock machinery as any run, but delivers its answer as a custom session entry, which does not participate in LLM context. The main agent keeps working and never sees the question or the answer; useful for checking something mid-task without derailing the conversation or polluting the context window.

---

### Execution runtime

All extension-managed children run through Pi's RPC runtime. Jev selects the
execution model and ordinary locally permitted tools, while the extension keeps
worktree isolation, process locks, depth limits, budgets, startup verification,
steering, resume/fork, and structured-output guarantees in one Pi path.

The low-level SDK and the Pi extension share this same child runtime contract;
there is no second execution runtime or implicit tool augmentation in the task
schema. Unsupported capability combinations are refused rather than silently
ignored.

---

### Profiles

| Profile                      | Tools                                                   | Writes                                      |
| ---------------------------- | ------------------------------------------------------- | ------------------------------------------- |
| `explore` (parallel default) | Jev-selected subset of locally permitted read-only candidates | no project-file writes |
| `review`                     | same as explore                                         | no project-file writes                      |
| `general`                    | Jev chooses from the full available locally permitted catalog | yes for selected write-capable tools; unknown custom tools count as writable |

Jev chooses individual tool names, not a capability bundle. Candidates come from the full available locally permitted catalog, not from the agent file's `tools` defaults and not from the parent's currently active tools. An explicit task `tools` list is an ordinary candidate ceiling, and explore/review keep their read-only rule regardless of what the selector returns. Configured trusted `passthroughTools` are preserved separately under the trust boundary below; they are not selector questions. An empty ordinary selection never means "all tools".

The finalized tool set is passed to the child as Pi's `--tools` allowlist (`--no-tools` for a true empty set). Pi 0.86.0 is the verified baseline for built-in, extension and late-registered tool enforcement; a host that cannot honor that allowlist is refused rather than silently weakened, and the extension does not claim identical behavior on untested older releases. Before the real task prompt is sent, a verified private startup command must confirm the selected model, exact active ordinary tools and registered passthrough definitions. Only the parent-authorized passthrough subset may be inactive; all active names must be allowed. Missing registration, duplicate/malformed evidence, an unexpected active tool, or a source/model/nonce mismatch aborts with a startup diagnostic instead of running with a broader tool set. An older child without passthrough registration proof is refused when the list is nonempty.

Parallel write-capable tasks sharing one checkout are rejected unless each uses `isolation: "worktree"`, distinct `cwd`, or explicit `allow_shared_writes: true`.

---

### Configuration

Defaults can be overridden in `~/.pi/subagent.json`; runtime fields with an env var below also support env overrides (env wins over file):

| Setting                 | Env var                               | Default                               |
| ----------------------- | ------------------------------------- | ------------------------------------- |
| `passthroughTools`      | none (config file only)              | `[]`                                  |
| `maxTasksPerRun`        | `PI_SUBAGENT_MAX_TASKS`               | 8                                     |
| `maxActiveProcesses`    | `PI_SUBAGENT_MAX_ACTIVE`              | 4                                     |
| `maxQueuedTasks`        | `PI_SUBAGENT_MAX_QUEUED`              | 32                                    |
| `maxGlobalActive`       | `PI_SUBAGENT_MAX_GLOBAL_ACTIVE`       | 16                                    |
| `defaultTimeoutMs`      | `PI_SUBAGENT_TIMEOUT_MS`              | 900000                                |
| `maxDepth`              | `PI_SUBAGENT_MAX_DEPTH`               | 2                                     |
| `killGraceMs`           | `PI_SUBAGENT_KILL_GRACE_MS`           | 3000                                  |
| `sessionDir`            | `PI_SUBAGENT_SESSION_DIR`             | `~/.pi/subagent-sessions`             |
| `worktreeDir`           | `PI_SUBAGENT_WORKTREE_DIR`            | `~/.pi/subagent-worktrees`            |
| `lockDir`               | `PI_SUBAGENT_LOCK_DIR`                | `~/.pi/subagent-locks`                |
| `worktreeRetentionDays` | `PI_SUBAGENT_WORKTREE_RETENTION_DAYS` | unused (lifecycle GC)                 |
| `sessionRetentionDays`  | `PI_SUBAGENT_SESSION_RETENTION_DAYS`  | unused (lifecycle GC)                 |
| `lockRetentionDays`     | `PI_SUBAGENT_LOCK_RETENTION_DAYS`     | 7                                     |
| `taskDefaults`          | -                                     | none                                  |
| `jevRouting`            | -                                     | required for new dispatch (see below) |
| `graceTurns`            | `PI_SUBAGENT_GRACE_TURNS`             | 2                                     |
| `stallAfterMs`          | `PI_SUBAGENT_STALL_AFTER_MS`          | 90000                                 |
| `stallKillAfterMs`      | `PI_SUBAGENT_STALL_KILL_AFTER_MS`     | 90000                                 |
| `maxRetries`            | `PI_SUBAGENT_MAX_RETRIES`             | 1                                     |
| `widget`                | `PI_SUBAGENT_WIDGET`                  | `background` (`off` disables)         |
| `notifications`         | `PI_SUBAGENT_NOTIFICATIONS`           | `batched` (`off` disables)            |
| (bin)                   | `PI_SUBAGENT_BIN`                     | auto (`process.execPath` + CLI entry) |

<a id="passthrough-tools"></a>
#### Passthrough tools

Add an optional top-level string array alongside `jevRouting` in the same `~/.pi/subagent.json` file:

```json
{
  "passthroughTools": ["runtime_control", "session_store"]
}
```

These are illustrative names, not presets; use the actual registered names of infrastructure you trust. The default is `[]`. Names are case-sensitive, trimmed and deduplicated in first-occurrence order. The array accepts at most 256 entries, each normalized name at most 256 characters; objects, non-strings, blanks, embedded whitespace/control characters, commas and wildcard patterns are rejected. No environment override, caller field, effect/activation attributes or foreign-extension config discovery exists. Malformed config or unavailable/disallowed names reject new work before paid selection; existing-run management remains available.

The list is **explicit user approval of trusted non-project-writing infrastructure**, including under `explore` and `review`. The host cannot infer a custom tool's effects from its name or registration. Known writers (`bash`, `edit`, `write`), unclassified unsafe builtins and this package's nested-dispatch tools (`subagent`, `subagent_wait`) cannot become infrastructure through list membership. Unlisted custom tools retain conservative writer classification; ordinary writers still require the existing parallel/worktree safeguards. Profiles are tool policy, not an OS sandbox.

Listed definitions must be registered in both parent and child. They are omitted from ordinary selector choices, survive a caller `tools:["read"]` or `tools:[]` ceiling and selector exclusions, and are appended once to the ordinary selected subset. An empty selection can therefore produce a passthrough-only allowlist, never all tools. The original route's `selectedTools` remains the ordinary choice. Plan, single/parallel tasks, resume/fork, `/btw`, optional synthesis and ranked replacements share this boundary; the invocation uses one frozen list, and config edits affect later invocations only.

The child proves registration separately from activity. Each configured definition may be active or host-inactive at startup; the extension does not force activation, wait indefinitely, or assume a group size. Ordinary selected tools must be active, and every active name must be in the final allowlist. Missing registration or unusable proof is a non-transient startup failure, not permission to drop a name, broaden capabilities or try another model. With the list omitted/empty, the existing exact-active startup and ordinary-only SDK behavior is unchanged. The low-level SDK does not discover this config or implicitly trust names; nonempty internal passthrough metadata requires routed attestation.

#### Jev routing

New subagent dispatches are selected by Jev, TypeSafe's structured-decision API, against a dedicated candidate list you maintain in `~/.pi/subagent.json`. Fixed routing is gone: a `modelPolicy` block produces a migration error, and an explicit `model` or `fallback_models` on new work is rejected rather than bypassing selection. Management actions (`status`, `wait`, `cancel`, `steer`, `diff`, `apply`, `discard`) never call the selector and need no credential.

```json
{
  "jevRouting": {
    "selectorModel": "jev-latest",
    "baseUrl": "https://api.typesafe.ai/v1/systemone",
    "apiKey": "<your-typesafe-api-key>",
    "timeoutMs": 15000,
    "models": [
      {
        "model": "<provider/model-id>",
        "description": "<your characteristics notes, Chinese allowed>"
      }
    ]
  }
}
```

| `jevRouting` field | Default / requirement | Purpose |
| --- | --- | --- |
| `baseUrl` | Optional; exact official `https://api.typesafe.ai/v1/systemone` when omitted | Complete normalized HTTPS SystemOne request URL; no implicit path rewriting |
| `apiKey` | Required; private config only | Header-only TypeSafe credential; never put it in URL or routing data |

- `selectorModel` defaults to the stable alias `jev-latest`. Pin an exact version to control which selector version is requested. This does not guarantee deterministic choices; the extension records the version that actually answered.
- `apiKey` is required and has no default. It must be a non-blank string; surrounding whitespace is trimmed and embedded whitespace/control characters are rejected. Store it only in the private config file. The transport uses it only for the Authorization header and does not copy it into prompts, selector JSON bodies, argv, logs, receipts or results. `apiKeyEnv` is rejected with migration guidance; environment variables cannot supply or override the key.
- `baseUrl` is optional and is the complete SystemOne request URL. When omitted, the exact official default `https://api.typesafe.ai/v1/systemone` is used. The URL is canonicalized by the standard URL parser and never receives an implicit path suffix or replacement. It must be an absolute `https://` URL with a hostname, no username/password, query, fragment, whitespace or control characters, and no more than 2048 characters; invalid values reject new routing before HTTP. A custom destination receives the same minimal task/model/tool routing disclosure, so configure it only when that endpoint is trusted. The normalized URL remains in the private frozen config snapshot, not in routing DTOs, prompts, receipts, persisted results or child arguments.
- `timeoutMs` defaults to 15000 and must be an integer between 100 and 600000. It bounds one logical task selection, including all its HTTP requests and queue waits. Parallel workers each have a selection allowance, still capped by their absolute task `timeout_ms` deadline. Deferred synthesis has a separate allowance.
- `models` holds 1 to 255 entries, each with an exact `provider/model-id` and a non-blank description. Those descriptions are what Jev matches against your task, so write them the way you would explain the model to a colleague. `thinking` is optional.

Plan and background-start native usage attachments are limited to 1024 selector HTTP receipts per invocation. Larger requests fail with a request-splitting error; background work has not started at that point. Previously incurred selector tokens remain in the ledger. This bounds an atomic delivery record, not the number of tools Jev may consider within each task.

The candidate list is intersected with the models the local Pi registry reports as available. A configured model Pi cannot resolve is not eligible, and an empty eligible pool fails before any request. Adding a model anywhere else in Pi does not authorize it, and legacy `modelPolicy` entries are never imported automatically. An unknown `jevRouting` field is an error, not a silent default.

Jev receives only the current delegated task text, your model IDs and descriptions, candidate tool names and descriptions, the requested difficulty and the permission/output requirements it needs to choose. It does not receive repository files, conversation history, full system prompts, persona text or tool parameter schemas. Task text and descriptions are user content and may contain sensitive material, so treat what you delegate as disclosure to TypeSafe.

Every new extension-managed dispatch routes through Jev: `task`/`tasks[]`, `action:"plan"`, `/btw`, resume, fork, locally permitted nested dispatch and the optional `synthesis` child. `action:"plan"` calls Jev and runs the same local preflights, returns the resolved model/tool plan and the selector usage, and creates no child or run entry. A later dispatch selects again; there is no cached decision to reuse. If optional synthesis selection fails, the worker plan and its usage stay valid and synthesis is reported as blocked with its diagnostic.

`thinking` is optional and is an opaque Pi thinking-level string. Common values include `off`, `minimal`, `low`, `medium`, `high`, `xhigh`, and `max`, but the package does not remap or restrict model-specific values. Pi receives the value unchanged and decides whether the active model supports it. Resolution order is: explicit task `thinking` > agent frontmatter `thinking` > profile `taskDefaults.<profile>.thinking` > the selected candidate's optional `thinking` > the parent session's thinking level. Jev never chooses a thinking level.

The extension re-reads `jevRouting` on each dispatch and injects non-secret routing guidance into the parent prompt, never the key. Configuration edits reach the next decision without a code change. Missing or invalid `jevRouting.apiKey`, or other invalid routing configuration, rejects new dispatch and plan with a remedy; management stays available.

The mandatory routing above is the Extension dispatch contract. The stable SDK exports (`runTasks` and `runSubagent`) are trusted low-level library APIs: they execute the explicit `TaskSpec` you pass and perform no implicit routing, config discovery or network call. Library callers own model and tool choice, and must not read these SDK calls as Jev enforcement.

#### Probability-ranked failover

Version `0.11.0` adds this behavior. Published `0.10.0` retries the selected model rather than advancing through the ranked candidates.

TypeSafe's Choice response includes a probability for every eligible option. The extension retains that distribution and tries higher-probability candidates first. These values express the selector's preference, not measured model uptime or success rates. The separate `confidence` value belongs to the original answer. A tied maximum keeps Jev's returned choice first; other ties follow configured candidate order. Low or zero probability is not a new exclusion threshold.

Jev selects one task-based tool subset, independent of the first execution model, for every attempt. The initial logical selection may use several HTTP batches; failover adds none. Each candidate still gets its own thinking default under the existing precedence and fresh model/ordinary-active/passthrough-registered startup verification. An attestation mismatch stops the task instead of trying a broader capability set.

Switching requires a settled provider error and conclusive evidence that no tool execution has begun in this invocation. Recognized cases include an explicitly unavailable model, temporary throttling, service overload and identifiable transport failures. Authentication/configuration errors, quota or billing exhaustion, invalid requests, context limits, refusals, poor answers, schema failures, cancellation and exhausted budgets do not trigger a model switch. Recognition uses only the latest completed assistant error's bounded message and documented primitive `diagnostics.error.code`, never ordinary answer text or arbitrary diagnostic details. Authentication, quota and other excluded evidence take precedence over an availability code. Unfamiliar error formats stop conservatively. A tool-start event blocks restart even when no result arrived; missing or malformed protocol evidence is not permission to retry. Historical tool messages in a resumed or forked session are not new execution.

For an eligible failure, the next allowed extension attempt advances immediately to the next candidate; it does not first add a same-model retry. Conclusively pre-work local process failures can retry the same candidate under the same total attempt budget. Attempts also have a local resource ceiling of 255 launches, including infrastructure retries. Expired task deadlines, uncertainty after startup and stalls cannot be used to restart work that may have begun. Pi's internal retries remain enabled or disabled according to your existing Pi settings, which this extension does not change.

Plan and status distinguish the original selector choice from the actual execution model. Results retain bounded attempt metadata, failure categories, output previews and child-session pointers; previews are attributed to the attempt that produced them. A later successful structured answer is not concatenated with failed JSON. If Pi retries within the same child, an empty latest answer stays empty; it never reuses text or JSON from the failed turn. If all attempts fail, the final failure/model/session remain authoritative. Earlier output remains available while its session is referenced on the active branch. Usage accumulates under the same run, without duplicating selector receipts; existing native-accounting limitations for thrown failures still apply.

A selector timeout or invalid response still stops dispatch. Failover uses only the validated ranking from that invocation, never an emergency model or a new selection. New invocations, including resume/fork and synthesis, select afresh. Trusted SDK specs without a ranked route retain their existing explicit fallback behavior.

#### Named agent files

Define reusable subagent personas as markdown files, discovered from the same conventional roots skills use (higher root wins name conflicts):

| Priority | Location                                                                | Scope                       |
| -------- | ----------------------------------------------------------------------- | --------------------------- |
| 1        | `.pi/agents/<name>.md`                                                  | project (authoritative)     |
| 2        | `.agents/agents/<name>.md`                                              | shared cross-tool workspace |
| 3        | `$PI_CODING_AGENT_DIR/agents/<name>.md` (default `~/.pi/agent/agents/`) | global                      |

The markdown body becomes the child's appended system prompt; frontmatter supplies defaults using the same snake_case names as the tool parameters:

```md
---
description: Security-focused code reviewer
# Legacy model/fallback fields are ignored; Jev routing owns model/tool choice.
thinking: high
profile: review
max_turns: 20
spawns: false # or "*", "scout", "[reviewer, scout]"
---

You are a security auditor. Review code for injection flaws, auth issues,
and sensitive data exposure. Report findings with file:line evidence and
severity ratings.

@include shared/review-checklist.md
```

Agent files may also pin a structured contract with `output_schema: {"type": "object", …}` (single-line inline JSON) or `output_schema: @contract.json` (path relative to the agent file).

`spawns:` controls which agents a child of this persona may spawn: `false` disables further nesting (no tool registered in that child), `"*"` (or omit) is unrestricted, and a comma/bracket list is an allowlist (agentless tasks are rejected under an allowlist). The policy is passed to the child via `PI_SUBAGENT_SPAWNS` and enforced on each subsequent spawn.

Body lines that consist solely of `@include relative/path.md` expand that file one level deep (relative to the agent file, same 64KB/symlink guards as `@contract.json`). Missing or rejected includes leave the line verbatim; includes do not recurse.

Invoke with `{ task: "…", agent: "reviewer" }`. The agent file supplies persona, capability, thinking and budget defaults only: model and tool selection stay with Jev routing, and a legacy `model`/`fallback_models` in frontmatter is ignored. An explicit `system_prompt` appends after the persona body. Profiles still enforce capability: `profile: review` filters candidates to read-only tools. Legacy agent `tools` defaults are ignored; an explicit task `tools` list requesting write tools under review fails closed. The agent catalog is advertised in the tool's system-prompt guidelines (session start) and in bare `status` output (live), and file changes are picked up within seconds; no restart needed.

#### Per-profile task defaults

`taskDefaults` in `~/.pi/subagent.json` remains available for non-model fields such as thinking, budgets, and retry counts. Its legacy `model` and `fallbackModels` fields are ignored; model and tool routing belong only to `jevRouting`. A profile `thinking` value overrides the selected candidate's optional `thinking` default. Invalid fields are dropped field-by-field.

Notes on behavior:

- `timeout_ms` is the absolute task deadline: local preflight, Jev selection, setup, queue time and runtime all count against it, and selection cannot reset it. Timed-out tasks report `state: "timeout"` with `timeoutPhase: "queued"|"starting"|"running"` so agents can retry capacity issues without confusing them for task failures.
- Budget stops (`max_turns`, `max_cost`) trigger a **graceful wrap-up**: the child is steered to produce its final answer NOW and allowed `graceTurns` more turns before SIGTERM. Results end as `partial` with `wrappedUp: true` when the child concluded in time. `graceTurns: 0` restores immediate stops.
- A **stall watchdog** flags children with no protocol activity for `stallAfterMs` (a liveness probe distinguishes quiet-but-thinking from dead), then kills after `stallKillAfterMs` more silence. A stall is not proof that no tool ran and does not authorize ranked failover.
- **Ranked failover** follows the [availability and pre-tool rules](#probability-ranked-failover), bounded by total `maxRetries`, candidate exhaustion, cumulative budgets and the original deadline. A selector failure has no emergency fallback. Task-quality, cancellation and budget failures never trigger model reselection.
- `context: "fork"` starts a single child from a real branched copy of the parent conversation (`--fork` on the parent's session file). It requires a persisted parent session, cannot combine with `resume`, and is rejected for parallel fanout (context duplication × N is a cost bug, not a feature).
- **Structured output** (`output_schema`): the contract is appended to the child's system prompt; the final message must end with a fenced `json:result` block. Validation runs parent-side against a dependency-free JSON-Schema subset (type/properties/required/items/enum/const; unknown keywords are ignored, never rejected). An otherwise successful but invalid answer gets **one steer-based repair round**; still-invalid results end `partial` with `structuredError` and raw text retained. Failed provider attempts neither repair nor publish structured output. Validated successful parallel results feed the `synthesis` child as clean JSON instead of prose.
- **Arg repair**: double-encoded task text (literal `\n` / `\"` escapes from LLM re-encoding) is conservatively de-mangled once at validation time. Identifier fields and paths are never touched. Protocol streams truncated after useful assistant output also end as `partial`.
- Aborting a `wait` returns immediately without cancelling the background run.
- Child processes are launched via the same Node runtime + CLI entry as the parent when possible (`PI_SUBAGENT_BIN` overrides). Bare `pi` on PATH is only a logged last resort.
- Direct resume is exclusive **across processes** via durable locks under `lockDir`. Lost runs block resume until startup orphan reconciliation kills (or confirms dead) the recorded child process group.
- `maxGlobalActive` bounds concurrent children across every Pi parent process on the machine (in addition to the per-session semaphore).
- Nested children at the depth ceiling do not re-register the subagent tool; only top-level parents run maintenance/orphan reclaim/worktree GC.
- Preserved worktrees live under `worktreeDir` (durable, not `/tmp`) and are garbage-collected on startup by **lifecycle**, not wall-clock retention: once a run is over (not live, past a 1h concurrency race guard), the worktree's unique work is archived as one applyable patch under `<repo-container>/_patches/` and the directory is reclaimed immediately. Branches holding commits that exist on no other ref are never deleted. `diff`/`apply`/`discard` transparently fall back to the archived patch when the directory is already gone. Live runs are never swept: the current session's live worktrees plus any worktree recorded on a running run record (concurrent Pi processes) are shielded machine-wide.
- Startup GC sweeps **every** repo container under `worktreeDir`, not just the current checkout's, so repos you stop visiting are still reclaimed. A container whose base repo no longer exists is kept and reported, never deleted; its worktrees' object stores lived inside the deleted repo, so unique work cannot be distinguished from a pristine checkout, let alone archived. Empty containers (no worktrees, no archived patches) are removed.
- Child session transcripts are likewise distilled on lifecycle: when a run is over and nothing on the parent branch references its session, the transcript is reduced to a small `.digest.json` (task, final output, model, usage, turn/tool/error counts) and the raw `.jsonl` is deleted. Resume needs the transcript, so anything referenced or busy machine-wide is kept.
- `keep_background: true` on a task keeps processes the child intentionally backgrounded (e.g. dev servers) alive after a clean exit.
- `include_wip: true` (with `isolation: "worktree"`) seeds the worktree with the parent checkout's uncommitted changes so the child sees your dirty baseline. `diff`/`apply` subtract that baseline when clean, else report the combined delta with an explicit `[includes parent WIP]` warning.

---

### Using the runner as a library

Import the stable public SDK from the package root or the explicit `/sdk` subpath; do not reach into `src/*` internals (those paths are not part of the supported contract):

```ts
import {
  runTasks,
  runSubagent,
  ChildRunner,
  WorktreeManager,
  Semaphore,
  ProcessLockManager,
  addUsage,
  normalizeUsage,
  emptyUsage,
  type TaskSpec,
  type TaskResult,
  type RunState,
  type UsageStats,
} from "@cr1ms0n/pi-subagent/sdk";
```

The package root is an alias for the same SDK: `import { runTasks } from "@cr1ms0n/pi-subagent"`.

The Extension dispatch path routes every new task through Jev, so callers never pass a model to it. The low-level SDK is the opposite contract: it is explicit-spec and performs no implicit routing, config discovery or network call, so embedding code supplies the model (and tool list) it resolved itself. This placeholder is illustrative only and configures nothing:

```ts
const task: TaskSpec = {
  task: "Audit src/ for unsafe parsing",
  profile: "explore",
  model: "<provider/model-id resolved by your embedding code>",
  timeoutMs: 10 * 60_000,
};
```

Prefer `runTasks()` for multi-task / worktree orchestration (same path the extension and pi-workflows use). `runSubagent()` runs a single child process directly without the extension host, but durable coordination is **opt-in**. Pass both `locks` (a `ProcessLockManager`) and a stable `runId` if you want global concurrency slots and orphan reclaim to see the child. Without those options no durable run record is written, so a parent restart cannot reclassify the process and nested children vanish from reconcile. There is intentionally no implicit default lock manager; embedding code that needs durability must construct and share one.

The Pi extension entry is unchanged: package `pi.extensions` still points at `./extensions/subagent.ts`.

---

### Cost accounting

`status`, `/subagent-cost`, and the `/subagents` overlay header show separate **root**, **subagent**, **routing**, and **combined** totals based on provider-reported usage. On Pi builds after v0.80.10, delivered runs also report their total usage natively on the tool result ([pi#6671](https://github.com/earendil-works/pi/pull/6671)), so Pi's own footer, `/session`, and RPC totals include subagent spend exactly once per run. Older Pi hosts ignore the field. Nested usage reported by a child's tool results (e.g. grandchild subagents) folds into the run's totals and budgets. The extension footer stays terse (running/ready counts only). Delivery and replay do not double count runs. See [docs/COST-ACCOUNTING.md](COST-ACCOUNTING.md).

Jev selection is billed separately from execution. TypeSafe reports tokens, not currency, so the ledger shows routing tokens as their own category, counts each selector request once by its request ID (including plan and pre-spawn failures), and marks routing cost as **unreported** rather than free. Numeric dollar totals exclude unreported routing spend, and `max_cost` caps provider-reported execution cost only; it does not cap TypeSafe charges. Route metadata (original selected model, ranked probabilities, selected tools, locally added controls, selector version, confidence, outcome, latency) travels with the run alongside usage. Actual attempt models are recorded separately; advancing through the ranking does not create another selector receipt.

---

### Engine contract

The [architecture contract](ARCHITECTURE.md) owns the complete lifecycle, persistence, permission and delivery invariants. The [security model](SECURITY.md) explains the limits of tool profiles and worktree isolation. For source layout and checks available in a fresh checkout, see [development](DEVELOPMENT.md).
