# Design

This document owns the maintained architecture and security model. Code is authoritative for exact protocol shapes and limits.

## Invariants

1. Pi owns the system prompt, current branch, compaction, active tools, tool execution, model selection, and cancellation.
2. Except for organization-trusted Claude Code managed policy, the main provider's Claude process may propose Pi tools but may not execute file, shell, web, MCP, plugin, hook, browser, agent, or memory capabilities invisibly. The separately invoked visible Pi web-search tool is the narrowly scoped exception described below.
3. Provider failures are errors, never successful-looking assistant text.
4. Subscription use is not reported as API billing.
5. External interfaces are capability-checked and mismatches fail closed.
6. Private request state is removed before success is published.

## Request and transcript transport

`streamSimple(model, context, options)` receives Pi's prepared context and applies Pi's logical `before_provider_request` replacement when present. The provider does not parse session files or rebuild Pi state.

That same function is registered twice, with Pi and with Pi-AI's API registry. Pi's own `registerProvider` populates its model runtime, which serves the user's turns; Pi-AI's own `completeSimple` and `streamSimple` resolve the model's api in Pi-AI's registry instead, and Pi's provider composer falls back to that registry as a last resort, so a miss there throws into a loop Pi does not await, which ends the process. Side requests use the same cwd routing rules as user turns: the caller transports its prompt and tools and executes proposals itself, but a tool-bearing direct Agent must provide a recognized cwd declaration or registered ID. The registration is shared by every instance in the process and is withdrawn only when no session is left to serve, and each session start re-asserts it, because a reload anywhere in the process clears that registry for everyone. It is also best effort: Pi-AI declares the entrypoint that carries the registry temporary, so the module is loaded rather than imported and its absence costs only these side requests. Importing it would make its removal a failure to load the extension at all, before any notice could report why.

Validated initialization is the transport's response boundary: capabilities are known and no content has been published. The provider announces it to Pi's `after_provider_response` observers with a synthetic success status and no headers, because the headless protocol exposes no HTTP response. The observer completes before body events are mapped; a failing observer fails the request.

Observer waiting is bounded by caller cancellation, process failure, and the existing total deadline. That deadline starts at Claude launch and remains active through ordered response processing even after the child exits; it does not include payload hooks or preparation. A stalled observer therefore cannot hold private files or an image lease indefinitely. The provider cannot cancel the observer's own work, but observes late rejections and publishes no late content. Process termination and safe cleanup are still awaited before lifecycle metrics finalize.

Each request serializes the effective system prompt, messages, and active tools into a versioned semantic transcript. Separate append-stable records preserve model-visible text, reasoning, tool calls, tool results, and images while omitting operational metadata, provider error messages, usage fields, signatures, and UI-only details. Tool-result `isError` state is preserved so historical failures remain meaningful. Historical tool results pair by `toolCallId`.

Pi owns the context window. The provider refuses only a system prompt too large for the window Pi configures, before any private state exists, and words that failure so it does not read as a context overflow, since compacting history cannot shrink a system prompt. Nothing else is gated before launch, and no output is reserved: on every model the aliases serve, the API accepts input plus the output limit beyond the window and stops generation at the window. A transcript too large for the window is refused by Claude Code, or by the API, as "Prompt is too long" without billing anything, and Pi reads that as an overflow, compacts and retries. Generation that reaches the window ends with `model_context_window_exceeded`, which Claude Code answers like an output limit and the provider hands off as Pi's `length` stop, which Pi also compacts and retries. Claude Code estimates a request at about three characters per token before sending it, and refuses one near the window on that estimate alone. This transport's JSON transcript runs denser than that, measured at 2.5 to 2.9 bytes per token, so the window itself usually binds first and Pi's own compaction threshold is reached before either refusal. The doctor names a window Claude Code reports serving that stops matching the configured one, because Pi places that threshold by the configured value.

At `message_end`, the extension prefixes recognized Claude context-overflow errors with `context_length_exceeded` so Pi can compact. It leaves unrelated errors unchanged.

Literal at signs are JSON Unicode-escaped because Claude expands `@path` syntax. Only provider-generated, validated images remain attachment references. They are quoted absolute paths: Claude runs in Pi's session directory, where a relative reference would resolve against the project, and the quotes keep a temporary root containing spaces in one reference. Claude Code's quoted form cannot contain a double quote, so an image request from a temporary directory whose path has one fails with `image_path`. The session image store checks the physical temporary root before creating its directory, so that failure writes no image and the store works again once the root is corrected. Their reference list follows the append-stable transcript blocks. Each Pi session stores image bytes under private, content-addressed paths that stay identical across its requests; a later request reattaches every image still in its effective context, including images Claude has already answered. Claude Code narrates image reads ahead of the transcript, so stable paths are required for cache reuse. A newly added image can still change that prefix. Images remain subject to the existing 20-image count and aggregate byte limits on every request. This conservative count aligns with [Anthropic's documented limit of 20 images per claude.ai message](https://platform.claude.com/docs/en/build-with-claude/vision); it is not a verified Claude Code print-mode limit.

Requests ask for summarized thinking display. Claude 5 and Opus 4.7 or newer otherwise return thinking blocks whose text is empty, so Pi would show a signature and nothing else while the turn still paid for the reasoning. The summaries are replayed in later transcripts as ordinary data, which grows the transcript by that text; thinking signatures still are not replayed, so the model does not resume its earlier reasoning either way. The option carrying this is hidden from Claude Code's `--help`, so preflight cannot scrape it. The documented minimum supported version covers it instead; the doctor reports that minimum but does not enforce it. `PI_CLAUDE_CODE_PROVIDER_THINKING_DISPLAY` chooses another display or sends none.

Resending the complete current context keeps branches, compaction, reloads, and provider handoff Pi-authoritative without a second session store. It also adds framing tokens, cannot replay thinking signatures, and is not wire-equivalent to the Messages API. Claude controls prompt-cache keys; the transport chooses only the placement and one-hour TTL of its single history breakpoint, described under compatibility below.

## What Claude Code adds on its own

The provider controls what it sends; it does not control everything the model sees. Claude Code adds content of its own that no documented flag removes, and the isolation guarantees above should not be read as covering it. It describes Claude Code at the verified baseline, run with this project's exact argument vector and environment, and `npm run capture:claude-breakpoints` rechecks the request-shape claims whenever that baseline moves.

The CLI prepends a billing header block, carrying its own exact version, ahead of everything else in the system array, where it is never marked for caching. Upgrading Claude Code therefore invalidates every cached prefix, which is expected but worth knowing before reading a cold start as a regression.

In print mode the CLI also prepends its own identity line, `You are a Claude agent, built on Anthropic's Claude Agent SDK.`, directly to the system prompt supplied through `--system-prompt-file`, with no separating newline. The branch is selected by non-interactivity rather than by the Agent SDK, so it applies on this project's documented path; `--append-system-prompt` only exchanges it for a different identity line. Nothing is appended after the supplied prompt.

The CLI also injects a block of its own environment context. It states Claude's process working directory as the model's primary working directory and says whether that directory is a git repository. Neither `--setting-sources ""`, the pinned settings object, nor `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` suppresses it. `--bare` does, but it never reads a subscription login. A request therefore carries context the provider never supplied and cannot remove.

Where the block lands depends on the model family, and can move between releases. The Claude 5 aliases receive it as a trailing `role: "system"` message after the transcript, the last thing the model reads, while Haiku 4.5 receives it at the head of the first user message. That difference decides whether prompt caching works, as the compatibility section below explains.

The model treats the block as fact, which is why Claude runs in Pi's session directory (see process lifecycle below). The block then agrees with the working directory Pi's system prompt states, and it stays identical across requests. Run anywhere else, such as the private request directory, the model proposes Pi tool calls into that directory and can describe a git project as not being a repository.

Starting in the project has effects of its own, which the provider bounds:

- **Startup Git collection.** Claude Code collects git status, recent log and user name at startup for its default system prompt, which `--system-prompt-file` replaces. That status, run in a repository with a configured clean filter, executes the filter. The provider sets `CLAUDE_CODE_DISABLE_GIT_INSTRUCTIONS=1`, which stops the collection without changing what reaches the model, and `npm run capture:claude-breakpoints` fails if a release runs the filter again.
- **What still runs.** With that setting, Claude Code still runs read-only `git ls-files` and `git remote` probes and an `rg --files` count of the project, and reads the project's `.claude/settings*.json`. None of the project's customizations, hooks or MCP servers take effect, and nothing is written to the project.
- **Cost.** The file count runs on every launch, so very large repositories add measurable latency to each tool round trip.

These are observed behaviors, not guarantees; a Claude Code release can change them.

Claude Code can further append turn nudges to user turns and tool results. Those are conditional and should be treated as possible rather than certain.

## Tool proposal boundary

Active Pi schemas are sorted into an ephemeral MCP catalog. Names of at most 48 characters made only of the characters Claude Code keeps unchanged in an MCP tool name (letters, digits, `_`, and `-`) are preserved; other names receive deterministic request-local aliases. Removed historical tools receive non-callable labels.

Response mapping resolves exact advertised transport names first. As a deliberate exception for an intermittent model naming error ([#14](https://github.com/chem/pi-claude-code-provider/issues/14)), it also accepts exactly `bash`, `read`, `edit`, and `write` when `mcp__pi__<name>` maps to that same active lowercase Pi name. A same-name Pi override remains the implementation that executes. Capitalized Claude Code names, other bare names, and generated aliases receive no recovery. IDs and arguments are preserved: this does not translate `file_path`, edit fields, timeout units, or any other Claude Code inputs into Pi's schema. Pi remains responsible for schema validation and execution. The whitelist and resolver are confined to the stream mapper; removing that fallback restores strict transport-name matching without changing catalogs, initialization, or transcript serialization.

The Claude child runs in `dontAsk` mode with local tools disabled. A proposal-only MCP server implements `initialize` and `tools/list`; any `tools/call` writes a violation marker and returns an error. The provider waits for catalog readiness, maps complete known proposals back to Pi, terminates Claude, verifies cleanup and violation state, removes private transport files, and only then publishes the Pi `toolUse` result.

On POSIX, a handoff accepts a SIGTERM or SIGKILL exit only when the supervisor recorded sending that signal to its owned process group. This includes cleanup escalating after the SIGTERM grace period; proposal validation and successful cleanup still precede publication. Unrecorded signal exits remain failures.

Tool names outside the advertised mapping and the four-name exception, malformed arguments, execution attempts, private transport paths, unexpected exits, caller cancellation, and cleanup failures fail the request. Initialization still requires the exact advertised MCP inventory; the response-name exception does not admit Claude Code's built-in tools.

Private-path detection recursively inspects argument values. Literal references are rejected even inside command strings; complete string values are also resolved against the request's selected working directory and checked for equality with or descent from a private directory. Native path rules handle dot segments and repeated separators, with separator and case normalization on Windows and case normalization on macOS, whose volumes are case-insensitive by default; POSIX backslashes remain literal characters. Each private directory is checked under both spellings of the temporary root, the resolved one it was created at and the one `TMPDIR` or `TEMP` names, because that root is commonly an alias: `/var/folders` links into `/private/var` on macOS, and Windows `TEMP` can hold an 8.3 short name. This is a heuristic for avoiding misdirected Pi tools, not a filesystem sandbox: beyond that root alias it follows no symlinks, and it does not evaluate shell expressions. Arguments are checked without being rewritten.

## Process and storage lifecycle

The doctor's proposal-bridge probe uses a marked private directory and records its child PID. A valid handshake is healthy only after a clean exit and confirmed tree cleanup. If termination fails, doctor reports unknown liveness and retains the directory. POSIX stale recovery reclaims it once the owner and child tree are gone; Windows requires manual cleanup, as for other retained runtime state.

For an unknown session, prompt-derived working directory resolution trusts the caller to preserve Pi's cwd declaration. This is an integration contract, not an authorization check on arbitrary caller text: a direct Agent can author the same marker, and a caller-controlled prompt can select another existing directory. Known sessions remain bound to their registered directory. Pi passes no authenticated cwd to the provider callback.

On POSIX, Claude runs as a detached process-group leader and cleanup targets the group. On Windows, Claude is a hidden, non-detached child and forced cleanup invokes `%SystemRoot%\System32\taskkill.exe /PID <owned-pid> /T /F`. The PID comes only from the retained request child; the provider never terminates by image name or process enumeration.

System prompts, transcripts, catalogs, and markers live in randomized per-request temporary directories. Images live in a separate private directory for the Pi session so their paths stay stable. A request takes its lease on that directory when its session is resolved, not when it reaches preparation, so a session ending in between cannot fail a request already placed on it — which a request borrowing another live session's store would otherwise see. The image directory is removed at session shutdown when no request still holds a lease on it. Shutdown deliberately does not wait for one that does: Pi emits `session_shutdown` before it aborts the turn, and awaits the handler with no timeout of its own, so a request still holding a lease has not been told to stop and waiting for it would hold Pi open for the rest of the turn. The directory and its recorded paths are kept instead, so that request keeps writing to the paths it was already given, and the last lease to be released removes the directory, because no second shutdown arrives to do it. When the host exits before that release runs, the exit reaper described below removes it instead. An unknown child-process liveness leaves it for stale-directory recovery instead, for the same reason the request directory is retained. POSIX uses mode 0700 directories and mode 0600 files. Windows relies on the per-user temporary root's ACL. Generated image names are content-addressed, bounded, and independent of user filenames.

**Claude's own working directory is the one Pi states for that request.** The extension records each session's directory from its own `session_start` context, not from Pi's process directory, which a resumed or imported session can differ from. Those records are process-wide and keyed by Pi's session id, because neither this extension's instances nor its module evaluations map one-to-one onto sessions. Pi re-runs every extension factory for a new, resumed, forked or cloned session, clears its extension cache on reload or a change of working directory, which re-evaluates the module, and lets a host hold several live sessions that share one set of instances. Pi's model runtime keeps only the most recently registered provider, so state held in one instance's closure would serve every session, and a map in module scope would hide one evaluation's sessions from another's. Each session's image store and rate-limit notifier are recorded with it, for the same reason. Pi also binds extensions twice for one session in its RPC mode, so the `session_start` handler is idempotent throughout.

A request first checks the ID Pi sends against sessions this extension registered. A matching ID supplies the cwd; a conflicting prompt declaration fails. For an unregistered ID or no ID, the provider accepts a cwd declared in the original Pi system prompt and borrows private image and notification state from a live session, preferring one in that directory. Pi supplies a `<cwd>` section in its transcript system messages, which the provider replays before routing; a direct or upstream caller may instead append a `Current working directory:` line after its own prompt and project context. The parser ignores sections inside tagged project context and takes Pi's last eligible declaration. It does not treat a generic `Working directory:` line in a direct Agent prompt as Pi metadata.

Without either cwd signal, a tool-bearing request fails with `working_directory`, even if only one session is registered: an unregistered child may have another cwd. `PI_CLAUDE_CODE_PROVIDER_BORROW_SOLE_DIRECTORY=on` explicitly restores that ambiguous sole-session borrow for compatibility. A tool-free markerless request borrows the newest registered session's directory, including when several sessions are live; this keeps Pi compaction and branch summaries working but does not prove which session owns the request. Metrics distinguish `registered`, `prompt`, `single` (explicit compatibility borrow), and `oneshot` (tool-free borrow), and the doctor names the last source.

The session is selected before `before_provider_request` runs so that a concurrent session change cannot move the request. If that hook adds tools to a tool-free borrowed request, the request fails before launch, even when only one session is registered. A hook cannot supply cwd provenance retroactively.

**Validation before any launch.** Each request resolves that directory once and validates it before any private state exists. A request fails with `working_directory` without launching when no Pi session has started, the path is not absolute, it cannot be read (for example because it no longer exists), or it is not a directory. A directory that disappears between validation and spawn fails the same way.

**Borrowing limits.** A prompt-routed request borrows only private state; Claude runs in its declared cwd. Tool-free summaries and the explicit compatibility mode borrow a live session's cwd without proving it belongs to the caller. No other tool-bearing request may borrow a cwd.

**Private-directory exceptions.** Only web search and the doctor's bridge probe still run in private directories; neither has Pi tools to misdirect.

A broken pipe on Claude's stdin is evidence, not the failure: a child that closed its input is exiting, so the supervisor records it and lets the exit status, Claude's own terminal record, and its stderr explain the request, naming the closed input only in the exit message. Other stdin, stdout, and stderr errors still fail at once.

Cancellation is checked before launch, so an already-cancelled provider or web-search request starts no Claude process; after launch, cancellation terminates the owned process. Normal success, failure, timeout, and cancellation remove private request state. A rejected tree terminator leaves process liveness unknown even after the leader exits: descendants can still own its process group or inherited pipes. Supervision rejects an outstanding wait promptly, quiesces the retained handle, and preserves private request state and leased images. A diagnostic raised after successful tree termination does not require retention. Provider requests and web search launch Claude through one shared lifecycle in `src/claude-process.ts`, so this ordering is implemented once.

**An exit reaper finishes what the host's exit cuts short.** Hosts exit as soon as a session is disposed. Pi awaits `session_shutdown`, then its dispose aborts the turn, then it calls `process.exit`, and a pi-subagents runner exits the same way once its run stops. The abort sends Claude SIGTERM synchronously, but directory removal and lease release follow an `await` that never resumes, so an in-flight request's private state would otherwise wait for a later POSIX stale pass or, on Windows, stay. Shutdown still does not wait for the request (see the image store above). Instead, `src/runtime-directories.ts` records every private directory this process creates in a process-global registry, and a single `exit` listener reclaims synchronously whatever asynchronous cleanup did not. A child whose termination was not yet confirmed is forced down first: SIGKILL to its POSIX process group, or `taskkill /PID <owned-pid> /T /F` on Windows unless the retained handle shows the leader already exited. Only then is the directory removed. macOS refuses to signal a group whose members are all zombies, which is the usual state at exit because the abort's SIGTERM already ended Claude and the exiting host never reaps it; on that `EPERM` the reaper lists the group's members with `ps` and proceeds only when every member is a zombie. Ordinary termination applies the same rule, because Claude can exit on its own just before the provider stops it, as after an output limit, and until Node reaps it the group is all zombies; there the wait continues until the reap shows the group gone. State already marked liveness-unknown is kept, and so is a child the reaper cannot confirm gone; while any such child may live, image stores are kept too, which is stale recovery's rule. The supervisor reports each termination's outcome to the registry, so every path that stops Claude feeds it without the provider or web search tracking anything themselves. The registry is process-global because pi-subagents runners and Pi's extension-cache resets evaluate this module more than once in one process. The reaper never throws and does no asynchronous work. It covers `process.exit` and the signals the host handles; a crash, SIGKILL, or a signal the host leaves to Node's default bypasses it.

POSIX stale recovery inspects all matching temporary directories before making at most 256 deletion attempts per pass. It removes only old, same-user, package-marked directories whose owner, recorded child, and child process group are gone; only an `ESRCH` probe establishes absence. A potentially live provider child or group also protects its owner's image stores. Incomplete request-ownership inspection retains images for that pass and reports an aggregate failure. Retained and ineligible directories do not consume the deletion budget, so larger stale image populations drain over successive startups. Inspection cost grows with the number of matching directories; there is no persistent scan state. Windows does not perform automatic stale recovery because Node provides no equivalent ownership check.

**Mid-response recovery is a failure, not a source of content.** Claude Code recovers from an API failure that arrives after a response has started streaming in three ways: it replays the request, it continues from the partial it kept, or it re-requests without streaming. A fallback-model chain from managed settings can also move an overloaded request to another model. None of them can be published here. Pi's assistant events are append-only and the provider has already streamed the blocks that arrived, so rewriting them to reuse a recovered attempt breaks Pi's own frame encoder and reducer; and whatever Claude produces after the interruption answers Claude Code's internal continuation prompts, which Pi never saw. The mapper therefore stops at the first recovery signal — an `api_retry` record, a `model_fallback` record, a synthetic assistant error record, a synthetic user turn, or an assistant message whose id is not the open stream's — and fails the request with `stream_interrupted`. A second `message_start` with none of those signals ahead of it is treated the same way: every one observed was Claude Code starting another request on its own, so an unrecognized path is retried rather than reported as protocol drift that loses the turn. The message wording matters: for a transient cause it carries the fixed phrase Pi's retry classifier matches, so Pi drops the failed message and re-runs the turn from unchanged context in a fresh process, while categories repeating cannot clear, such as a billing error, are worded so Pi stops. The replayed transcript is a prompt-cache hit, so only the output tokens are produced twice. Signals are not read before the response starts, where a retry is ordinary, nor after a tool-use or output-limit stop, where the provider is already terminating Claude.

**A refusal is final.** Fable, Opus 5.5, and Opus 5 run safety classifiers that can stop a response with a `refusal` stop reason. By default Claude Code then re-runs the request on a fallback model, for example Opus 4.8 for a cybersecurity flag, which the provider can never publish: the refused blocks are already in Pi's append-only events, and the served model would change mid-response. The provider therefore pins Claude Code's documented `switchModelsOnFlag: false` setting, under which a print-mode request ends with the refusal instead. Claude Code can still retry once on the same model behind a synthetic prompt, and a managed policy can turn the switch back on, so the mapper fails at the first refusal signal: the `refusal` stop reason, a `model_refusal_fallback` or `model_refusal_no_fallback` record, a refusal error record, or a refusal result. The error reads "The model refused to complete the request", with the category and model when Claude Code reports them as plain tokens, mirroring Pi's own Anthropic provider. It must not match Pi's retry classifier, because the same turn would be refused again. Claude Code's explanatory text is therefore never repeated, and each detail is kept only while Pi's own `isRetryableAssistantError` still reads the message as final. Refusal records are ignored after a tool-use or output-limit handoff, where they belong to Claude Code's next internal turn.

**An output-limit stop is handed off like a tool call.** A response that ends with `max_tokens` or `model_context_window_exceeded` is complete as far as Pi is concerned, but Claude Code answers it with a synthetic "Output token limit hit" turn and another message of its own. The provider latches the stop, terminates Claude through the tool handoff's path, and publishes Pi's ordinary `length` stop carrying what the model produced. Once either handoff is latched, stream events still in the pipe are ignored rather than mapped or rejected, because they belong to a turn Pi did not ask for and the outcome must not depend on whether termination wins that race.

Protocol outcomes are published through `ClaudeEventMapper`, whose failure and completion gates are idempotent; setup failures that occur before a mapper exists are published directly by the provider. The request finalizer deliberately does not publish a second terminal event: it owns abort-listener disposal, last-chance safe directory cleanup, and exactly-once metrics. This keeps protocol mapping separate from process and storage finalization while preserving the rule that success follows cleanup. The doctor stores the last lifecycle to finalize, which can be an older request finishing after a newer one; a published failure does not imply its cleanup or metrics are already complete. An internal recorder dependency lets deterministic tests observe each request's own finalization without changing that production ordering.

Host crashes, SIGKILL, and signals the host does not handle bypass even the exit reaper, and their private state waits for POSIX stale recovery. On Windows, `taskkill` cannot reconstruct descendants after their root exits, so normal Claude shutdown is trusted to close its children. `taskkill /T` can also fail on a descendant that was already exiting as it walked the tree; once the owned child has closed, a failure whose named PIDs all no longer exist counts as completed cleanup. Those PIDs are read as digits, which survive localization, and nothing is signalled again, because the closed child's PID may be reused. Name-based termination is deliberately prohibited because it could kill unrelated sessions.

## Web search

`pi_claude_code_provider_web_search` is a visible Pi tool backed by a separate Claude process restricted to WebSearch and WebFetch. Initialization must confirm the exact tool inventory, `dontAsk`, no unexpected MCP server or customization, and subscription-backed authentication.

The outer Pi tool invocation is visible in Pi; Claude's inner WebSearch and WebFetch calls are intentionally restricted but are not individual Pi tool executions.

The query is supplied through a private generated file. Output is validated and bounded; truncated full output may be retained in a session-scoped private file that is removed at shutdown. The main provider never gains invisible web access.

Registration is optional. Pi puts an active tool's snippet and guideline in the system prompt of whatever model is selected, not only this provider's, and every call spends Claude subscription capacity. So `PI_CLAUDE_CODE_PROVIDER_WEB_SEARCH=off` leaves the tool unregistered, and the guideline tells models to prefer any other available web-search tool. An unrecognized value also leaves it unregistered, because someone who set the variable most likely meant to opt out; registration is decided once per provider instance, and the doctor reports the outcome.

## Security and privacy

The package trusts the installed Pi and Claude executables, Node, the operating system, and the user's account. It defends against malformed protocol records, unexpected capabilities, unsafe file references, tool-name confusion, private-path disclosure, child-process leaks, oversized data, and accidental diagnostic content leakage.

Claude children receive an allowlisted environment for authentication (including a relocated `CLAUDE_CONFIG_DIR`), locale, proxies and a proxy's `NODE_EXTRA_CA_CERTS` bundle, shell discovery, and temporary storage. Pinned variables disable automatic memory, nonessential traffic, Claude Code's startup Git status collection, and its own compaction: given a replayed transcript near the window, Claude Code otherwise starts compacting history that Pi owns before refusing it. API keys, `CLAUDE_CODE_OAUTH_TOKEN`, alternate routing, hooks, plugins, and arbitrary parent variables are not forwarded, and a caller that adds one of those to a child's environment fails rather than silently overriding the allowlist. That allowlist is POSIX-shaped: on Windows the child additionally receives the variables Node's process spawner injects into any replacement environment, among them `SystemRoot`, `USERPROFILE` and `TEMP`, which this package neither lists nor grants. `APPDATA` is outside that set and is not forwarded. Explicit settings suppress unmanaged user and project Claude customizations, including Claude Code's built-in `agents-md` and `telemetry` plugins, and initialization verifies the resulting inventory, rejecting any plugin that still loads. The explicit telemetry pin supplements `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1`. The same object pins `switchModelsOnFlag: false`, so a safety-classifier flag ends with the refusal instead of a switch to another model (see "A refusal is final" above). Administrator-managed Claude Code settings, hooks, and MCP policy are an organization-trusted boundary: the package cannot suppress them or prevent their startup effects before validation.

Concretely, the provider passes `--setting-sources ""`, which drops the user, project and local setting sources, and supplies a fixed object through `--settings` in their place. A user who has configured Claude Code will reasonably expect otherwise, so the consequence is worth stating: their own model, effort and per-model settings have no effect on Pi requests. That is Invariant 1 in practice, not an oversight. The alias-resolution environment overrides are excluded from the allowlist for the same reason. Because both are true, alias resolution for this provider's child is fixed by Claude Code's own defaults, which is what makes the model versions the doctor reports accurate rather than a guess.

Diagnostics and optional metrics contain bounded system, version, size, usage, and lifecycle facts. The doctor and its report also name the provider copy Pi loaded, its version and home-redacted install directory, because Pi can keep a project-local install ahead of a newer global one and loads npm and Git installs side by side. They exclude prompts, messages, tool arguments and results, queries, output, request stderr, credentials, and temporary paths. A failed bridge handshake can include a bounded, path-sanitized startup diagnostic derived from bridge stderr. Request error messages can likewise carry a bounded tail of Claude Code's stderr with the private request directory replaced; for web search that message reaches the model as a tool result. Sanitized paths outside home and temporary roots may remain, so reports must be inspected before sharing.

Model-proposed file and shell operations execute only through Pi. That is narrower than saying Claude has no local effects: the startup operations described under "What Claude Code adds on its own" run before any proposal, and administrator policy or a later CLI release could add others. The package does not sandbox Pi, protect against a compromised local executable, make untrusted prompts safe, or prove undocumented server behavior is absent. Model-visible content is sent to Anthropic and inherits ordinary Claude Code confidentiality risks.

**Fast mode is deliberately not exposed.** Claude Code supports it in print mode through the `--settings` object this provider already passes, and it could be offered as a separate Opus alias without weakening isolation. It is not, because fast mode is paid for only from usage credits, at rates well above plan usage and even when plan limits have room, while this provider reports zero cost for every request and cannot tell how a subscription request was billed (Invariant 4). Claude Code also falls back to standard speed and still reports success, so a control reflecting what was requested would not be truthful. A feature that quietly spends money is the opposite of the fewest possible surprises.

## Compatibility and performance

Machine-readable verified versions and the model family each alias must serve live in `src/compatibility.ts`; procedures live in [DEVELOPING.md](DEVELOPING.md). Version metadata is advisory, while protocol and capability mismatches fail at runtime.

Full transcript serialization is required by the stateless design. Stable record boundaries preserve cacheable prefixes, so performance changes must not rewrite unchanged history or weaken validation and cleanup ordering.

Append-stable serialization is necessary but not sufficient. By default the CLI appends a changing token reminder of its own after the transcript; because this provider starts a fresh print-mode process per request and replays the whole transcript, the next request inserts its new turn ahead of that appended content, so the previous request is no longer a prefix of the current one and reuse collapses entirely. The stateless replay design is what converts a Claude-side append into a total loss of caching. The provider therefore pins an undocumented settings key to suppress that content, on upstream maintainer guidance. Because the key is undocumented, the paid cache gate rather than the setting is the contract: see [DEVELOPING.md](DEVELOPING.md#prompt-caching) for the measurement procedure and the conditions for changing it.

The CLI also places its final `cache_control` marker on the content it appends rather than on the transcript, leaving the replayed history with no cacheable entry of its own, and no setting changes that. The transport therefore sets that breakpoint itself, on the last history block and ahead of the generated attachment references. The breakpoint uses a one-hour TTL because the API requires breakpoints in longest-TTL-first order and Claude Code's own markers after it are one-hour; that doubles the cache-write rate over the five-minute default, a cost accepted because the ordering leaves no choice. It is also the last of the four breakpoints the API allows on every Claude 5 alias, so a Claude Code release that adds one of its own would fail every request. `PI_CLAUDE_CODE_PROVIDER_TRANSCRIPT_BREAKPOINT=off` drops it at the cost of prompt caching, and the provider recognizes the API's limit rejection and names that setting. A request on which Pi asks for no cache retention carries no breakpoint either: Pi sets that on its one-shot summaries, whose prompt is unique per compaction, so the entry could only ever be written and never read, at the long TTL's doubled rate. Haiku 4.5 receives the same environment content ahead of the transcript rather than after it, so it caches only while that content is identical across requests. Pi's session working directory and the image paths are stable; request-private files remain in their own exclusively created directories. Account-specific per-session content or changed attachment narration can still disturb that prefix, so `npm run test:paid:cache-haiku` and the Sonnet and Haiku image-cache stages gate reuse separately.

A request can append only about twenty records beyond the history the previous request cached. Each breakpoint looks back about twenty block positions for a prior entry. The API collapses runs of consecutive `tool_use` or `tool_result` blocks to one position, but that exemption does not apply here, because the transcript gives every Pi message, each tool result included, its own plain text block. Past the ceiling the loss of reuse is total and silent. Pi makes a new request after every tool round trip, so the records appended between requests are usually one assistant message plus one per parallel tool result, and only a step with about nineteen or more parallel tool calls reaches the ceiling. An intermediate breakpoint would exceed the four-breakpoint budget. Sending each run of consecutive tool results as one block would lift the ceiling, but that serialization change would need the paid cache gate.
