# Security model

SpecPi provides scope monitoring, an explicit harness improvement loop and an optional Jev advisor, and installs eight pinned packages as its default base. Extensions run as trusted code with Pi's permissions. Scope is not an OS sandbox or a general command guard. Use OS isolation for hostile code.

## Scope monitoring

The human declares project-relative paths. Interactive `write` and `edit` calls outside those paths require a decision; headless calls are recorded as pending. Git snapshots detect changes from other tools after execution. Paths, snapshots, and displayed findings are bounded, and uncertainty remains visible until an explicit recheck. Ignored files, subprocess activity, timing races, and changes outside the observed project can escape snapshot coverage. A clean report does not prove that no other mutation occurred.

Scope records use Pi's current session branch. Restoring a branch does not create new authority or accept pending drift. An improvement contract can supply paths only through the human's `/scope task` command.

## Capability requests

Optional tool groups ship withdrawn. The `request_capability` tool lets the model name a withdrawn group instead of silently working around it; it never grants one. Activation requires a human decision, either accepted in the moment or recorded earlier as a standing grant, so the tool is not offered in headless sessions or when no requestable group is installed, refuses if called without an interactive user, and a declined prompt leaves the session unchanged. Activation is additive, offers only the named group's tools, and lasts for the session: it writes no startup preference, and the next session starts from the saved preference as before. Delegation is not requestable, because its own package requires a human command to bind a model. Acceptance offers the tools; it is not a per-call permission, and the Permission System continues to govern what those tools may do.

`/capability allow <name>` records a standing grant so that group is offered without a prompt, and `/capability ask <name>` returns it to prompting. Recording either requires an interactive human command, and a grant that fires still announces itself in the session. Grants live in `<agent-dir>/specpi/capabilities/settings.json`, written atomically with owner-only permissions; links and irregular files are refused, and an unreadable file, a foreign schema or an unknown capability name reads as no grant. A standing grant removes the prompt, not the requirement for a human: headless sessions are still refused, so a granted capability cannot be activated by an unattended run. It also does not offer the group at startup — the model must still ask, and the group stays withdrawn until it does.

## Background jobs

The `background` tool starts a shell command and returns at once; the job reports its exit code and last lines of output as a follow-up message when it ends. It runs through Pi's own local shell runner with the session's `shellPath` and `shellCommandPrefix`, so the shell, working directory and process-tree cleanup on Pi's exit match the `bash` tool. It adds no new authority, and every start fails closed:

- **Permission parity.** The installer records `shellTools.background = { commandArgument: "command" }` in the permission system's global config, so its full `bash` stack — command decomposition, wrapper flooring, path and outside-directory checks, and every `bash:` rule — applies to the command. The entry is merged into the existing file, backed up and rolled back with the rest of the transaction, left alone when the file carries comments, and removed at uninstall only if it is still SpecPi's. Before every start the extension reads the effective mapping (global, then a trusted project's, which could remap it) and refuses unless it gates `command`. An unreadable config refuses too.
- **Command guard.** `specpi-lancet-guard` gates the `background` tool's command exactly as it gates `bash`, so background jobs run under it with no extra condition. The retired `specpi-jev-guard` checked `bash`, `powershell`, `write` and `edit` only, so while a `jev-guard` command is present -- someone who still installs that package themselves -- background jobs stay off unless its saved state resolves off, read the way it reads it: its default on, the global file, then a trusted project's. Its session-only switch lives in its own memory and cannot be read, so save that choice with `--global`.
- **Interactive only.** Headless sessions are not offered the tool and are refused if they call it; there is nobody to report back to. A Pi that cannot list its commands is refused, because the guard and permission system cannot be detected.
- **Bounded.** At most four jobs run at once. Each log is capped at 8 MiB under `<agent-dir>/specpi/background/`, owner-only, and deleted when the session ends; directories left by a crashed session are pruned after a day. A completion message carries at most 40 lines and 4,000 characters of output with control characters removed. Output is untrusted text, like any tool result.
- **Session-owned.** `/jobs stop` and session end kill each job's process tree; Pi's own exit handling also kills tracked children. Process cleanup is best effort: a program that detaches itself from its parent, or a daemon a job starts, can outlive it, exactly as with `bash`.

Eval runs overwrite the permission config with their `yoloMode` opt-in, which drops the mapping, so the tool is never offered there and eval rows keep their measured tool surface.

## Jev advisor

The advisor is the only part of SpecPi that sends anything off this machine, and it ships off. With the master switch off there is no network call, no key read, no consent read and no prompt injection: the harness behaves exactly as it did before the extension existed.

What is sent: a summary object of at most 1 KB per call to `openrouter.ai`, which is where Jev is published and the default route, or to `api.typesafe.ai` when `JEV_BACKEND=typesafe` selects the direct API. Either way, over HTTPS. It carries tool names, byte counts, relative paths, entry counts, short descriptions and a bounded sample of lines from the material being judged. `sanitize.mjs` builds the wire state and sanitizes question text. Known credential, token, JWT, email and URL patterns are redacted, including quoted credential values and multiline private-key blocks before sampling; workspace paths become relative and other absolute paths are replaced. Source and cluster questions use opaque IDs, with their descriptions inside the state budget. Required fields have explicit allocations: samples may be shortened, but missing objectives, missing samples or candidate lists that cannot fit cause local abstention, not an evidence-free classifier call. Redaction reduces exposure; it does not guarantee that a model-authored string carries nothing sensitive.

Earlier versions of this document said file contents and command output were "refused outright". That was never accurate and is corrected here. Deciding whether a tool result is spent, or whether a fetched page is addressing the agent, cannot be done from byte counts alone, so `outline()` has always sent a sample of the result's own lines: up to six from the head and two from the tail, each collapsed to at most 80 characters and passed through the same redaction as everything else. As of 0.27.0 it also samples up to four lines evenly spaced through the middle, because an instruction planted in a fetched page is rarely in its first six lines and a digest that could never contain one would ask system 7 a question its own state made unanswerable. So the accurate statement is: **a bounded, redacted sample of at most twelve short lines per result, inside a 1 KB total budget** &mdash; not the whole content, and not none of it. What the sample cannot include is anything beyond that budget. The ledger records offered and retained sample counts and whether required evidence was present; that does not establish that the sample contains every relevant fact.

Consent is separate from the switch. The first time any system would transmit, an interactive dialog names the endpoint, the shape of the data and the byte budget. Without an existing valid grant, a noninteractive run cannot start transmitting. A previously granted consent may be reused without a UI. The grant lives in `<agent-dir>/specpi/jev/consent.json` and is bound to the destination origin, including its transport scheme. Transport rejects redirects rather than forwarding evidence to an unapproved destination. Schema 2 requires renewed consent because the earlier dialog incorrectly promised never to send file contents or command output. It now explicitly discloses sampled content, the current request/objective, best-effort redaction and any non-HTTPS endpoint override. `/jev forget` revokes it.

Every advisor request attempt appends one line to `<agent-dir>/specpi/jev/transmissions.jsonl`: timestamp, system, question keys, state byte count, a SHA-256 of the sanitized state and questions, latency, the gated outcome, each answer's value, confidence and distribution, and code-defined labels for the local signals that led to the call. Local evidence refusals are marked `sent: false`, consume no call budget, and are counted separately from calls. Effect tags distinguish shortening, warnings, recorded assessments, notifications, capability proposals and queued steering; queueing does not prove the model consumed a message. The payload itself is never written. `/jev ledger` reads it back, so "only digests are sent" is checkable rather than asserted. The ledger is user data and survives uninstall.

The advisor also keeps `<agent-dir>/specpi/jev/usage.json`, the running call count for the current session, written atomically with owner-only permissions. It exists because the ledger has no session boundary in it, so nothing outside the advisor's own process could say what *this* session had spent, and the count is what SpecPi Chat shows beside each budget. It is the one file in this layer meant to be read by another process, and it is deliberately the least interesting one: counts per system, the budgets they are counted against, a session identifier and two timestamps. No state, no questions, no answers, not even the ledger's digests. It is written only while the master switch is on, so a layer nobody has enabled leaves no trace of having been installed, and the last session's counts are kept rather than deleted at shutdown, because "this has never run" and "the session that just ended spent its whole budget" are different facts and a reader should be able to tell them apart. One file serves the directory, as the settings file does, so where several Pi sessions share an agent directory it describes whichever wrote to it last; it carries a session identifier, two timestamps and an active flag so a reader can say which rather than having to assume.

The advisor holds no authority. It never grants a capability, never calls a tool and never allows one. Its only blocking action is to refuse a capability-gap report that confidently carries a secret value or names a specific person or machine, which asks the model to rewrite its own text and discards nothing. A report that only mentions such things is not refused. Failure is silent, not closed: a timeout, HTTP error, missing key, missing consent, an exhausted per-system or session budget, or ungated confidence produces no advice, and the existing code path runs unchanged. Tool-result retention can shorten a large observational result before it is appended; it never shortens a write, edit, shell or error result. Shell success does not establish that a command is read-only or safe to repeat. Other replacements still suggest re-reading; preserving an immutable original result is not implemented.

Task context comes from the workflow owner's validated active contract, not a rendered heading or a search through stored conversations. A bounded current-request fallback is refreshed when the task changes and cleared on branch/session changes. Delegation ranking adapts `packet.jobs[]` independently, uses each declared mode, and preserves the exact source multiset and ungated positions.

Gap triage runs only after local collection is enabled. Reports bind their task and session before any awaited consent or root lookup; collection is rechecked before transmission. Locked writes recheck task authority and strip the advisory assessment if Jev was disabled while waiting, without preventing otherwise authorized local collection. A bounded shortlist comes from sanitized wishlist observations, never from Pi state. A persisted `assessment` is a Jev opinion based on the reporting model's account, **not independent evidence**. It cannot change the original impact or suggested fix, canonical identity, priority, qualification, selection or retirement. Suggested cluster matches require a human merge decision. Model-authored report arguments cannot inject these assessments.

One of the five systems can add fixed advice to model context, and it is off by default like everything else here.

**Progress detection, formerly system 5, has been withdrawn.** It was the only system that could steer the model mid-session, and replayed over recorded runs its stuck verdict did not predict failure. Nothing in the layer now appends a message to the conversation. Schema 6 settings drop its switch, budget and `progressNudge` mode.

The command guard is **not** part of this layer. It is a separate pinned package, `specpi-lancet-guard`, with its own gate, its own configuration file, its own commands and a fail-closed posture, and it has its own section below. It is named here because a reader who took "the advisor holds no authority" to cover everything SpecPi installs would be wrong about exactly one component. It sends nothing anywhere: it scores commands locally. `/jev` does not mention it and the advisor imports nothing from it.

**System 7, untrusted-content classification**, prepends a fixed warning line to externally fetched content that confidently reads as instructions addressed to an agent. It applies to web and browser tool results and to shell commands whose output is a fetched document -- `curl`, `wget`, `gh api`, HTTPie and PowerShell's web cmdlets -- and never to other shell output or to file reads. That shell exception is decided by the command, locally, before anything is sent; the result is then sampled and redacted like any other. It is defence in depth and explicitly not a control: it never blocks, it has no authority over what the model then does, and it should not be relied on to stop prompt injection. Its value is that untrusted content in a fetched page is currently owned by nothing at all.

An earlier version of this document said its false-positive rate was measurable on tier 5 of the eval suite. It was not: no task produced a web or browser result, so the system was never called. The shell exception is what made it measurable. Replayed over 737 shell fetches in recorded Terminal-Bench runs (`scripts/jev-untrusted-replay.mjs`), it flagged 18, and all but one were text genuinely addressed to an AI agent -- benchmark task prompts and dataset rows that were themselves system prompts. None of those runs contained a planted injection, so this measures how often it fires on real content, not how often it catches an attack; that rests on synthetic cases, where a subtly phrased injection sits at the 0.85 bar.

Settings live in `<agent-dir>/specpi/jev/settings.json`, written atomically with owner-only permissions. Links, irregular files, oversize files, an unrecognised schema and unparseable contents all read as off. The one exception is the layer's own previous schemas: an older file is migrated forward rather than read as unrecognised, because "collapse to all-off" is a rule for corrupt input and applying it to our own earlier version would silently disable a layer the user had switched on. Keys for withdrawn systems -- compaction guidance in schema 5, progress detection in schema 6 -- are dropped rather than migrated, and two more keys are dropped for a different reason. A `schema: 2` file kept a `guard` pair and a `schema: 3` file kept `systems.guard`; schema 4 keeps neither, because whether the command guard runs is not this file's answer to hold. The package keeps its own switch in its own configuration, and a stale copy here could only ever disagree with it. Call budgets are per system under a session total, so one busy system cannot exhaust the allowance of the others and leave them dead for the rest of the session with event ordering deciding which one won. Turning the layer on or off writes that preference, because a switch that forgets is not a switch; `--session` is how a one-off change is kept out of the file. The key is resolved the way Pi resolves every provider credential, in Pi's own documented order: the `openrouter` entry that `/login openrouter` writes to `<agent-dir>/auth.json`, then the environment variable. That is a deliberate, narrow exception to the rule that Pi state is never read, and it is bounded to exactly one question &mdash; the one `api_key` entry for the one provider this layer calls. Be precise about when: the key's *value* is read only on a call path the master switch already gates, but whether an entry *exists* is checked whenever status is displayed, including with the layer off, because that is the answer someone needs in order to configure it in the first place. A presence check reports a source name and nothing else. An `oauth` entry is ignored rather than unwrapped, because Pi refreshes those under its own lock and a second reader would race a rotation. The value is returned by a single function straight into a request header; nothing else in the layer receives it, and status output, the ledger and the Chat panel carry source names only. Pi trust decisions, sessions, missions and history are never read. `JEV_KEY_SOURCE=environment` restricts resolution to the environment, which is what this repository's own eval scripts set so a measured run cannot silently bill a developer's personal account.

## Command policy and the LANCET guard

`specpi-lancet-guard` is pinned in the base set and ships **off**. SpecPi writes nothing into its configuration: the package's own default is off, and only its own `/lancet-guard on` (this session) or `/lancet-guard on --global` (saved in `<home>/.pi/lancet-guard.json`) turns it on. An installer run never arms or disarms it, so a saved choice survives every update. `@gotgenes/pi-permission-system` stays pinned and decides every tool call whether the guard is on or off; the guard is a second, independent check: either one can block a call, so the guard never allows what the permission system would refuse.

When on, local rules copied from `specpi-jev-guard` decide first -- hard-deny patterns block, provably read-only commands and the user's allow globs pass -- and LANCET Nano v0.4.3, a 111M-parameter CodeT5+ 220M classifier, scores the rest of `bash`, `powershell` and `background` commands in-process on the CPU through ONNX Runtime, reading Bash and PowerShell alike, so a PowerShell command it does not flag runs. `risky` asks (or blocks, after `/lancet-guard mode block` saves `"risky": "block"`), `review` -- the band Nano is unsure about -- asks, `not_flagged` runs, and what LANCET cannot read (input over 8,192 bytes) asks. Longer commands are read in full, as overlapping 512-token windows, but never cleared: a `not_flagged` verdict on one asks, because harmless lines in front of a risky command pull it under the threshold once it spans a second window. Writes and edits to protected or out-of-workspace paths ask without a model call. With no UI, an ask becomes a block unless `uncertain` is `allow`. A missing, damaged or unloadable model blocks what the rules leave open: switching the guard on accepts that, and `/lancet-guard on` is refused until the model is installed. `not_flagged` is not a safety claim; on lancet-bench-2-next, LANCET's 3,204-command release benchmark, the classifier alone caught 78.0% of risky commands and stopped 12.4% of safe ones, against 95.1% and 21.8% for Jev, and credentials are among its weakest areas. Updating the guard to a release that pins a new model leaves the old model unused: a guard saved as on fails closed until `/lancet-guard setup` fetches the new one.

Commands never leave the machine and are never executed by the guard. The only network access is `/lancet-guard setup`, run by a human: it downloads one ZIP from the pinned `TannerMidd/LANCET-model` GitHub release into a private staging directory, capped at its pinned size, and refuses it unless its SHA-256 matches the digest compiled into the package. Only then does the package's own small ZIP reader extract the six model files, each inflated to its pinned size and refused unless its own SHA-256 matches. GitHub is the only host involved, so it works where Hugging Face is blocked. The digests are checked again on every load. The package's one production dependency, `onnxruntime-node` 1.30.0, is pinned exactly; its install script would fetch CUDA libraries on Linux x64, so the installer sets `ONNXRUNTIME_NODE_INSTALL=skip`.

The guard publishes a per-session counter as a Pi status item under the key `lancet-guard`: the number of commands LANCET scored, the number blocked once there are any, and the last verdict. SpecPi Chat renders that line in its footer and does nothing else with it. The key is absent whenever the guard is off or its `auditDisplay` is. Every decision is also written to Pi's own session file by the package.

`specpi-jev-guard`, which this replaced, is no longer pinned. An update retires its package entry like any other SpecPi-added pin and leaves its downloaded bytes and `jev-guard.json` in place. When that file had it switched on, the update says so, because retiring a gate someone had armed is not a change to make quietly. The LANCET guard does not inherit that setting; it starts off.

## Improvement authority and evidence

Collection is off by default. Enabling it permits sanitized gap observations, not implementation. The observation tool is offered only where a report could land: collection on, or undecided in an interactive session, where its first report asks for consent. The improvement skill is kept out of the model's skill list and is named by path only in the `/harness-improvement` kickoff. Only an exact human `/harness-improvement` selection authorizes a wishlist-sourced change. The selected contract is bound to the gap, source checkout, session, and selection generation. Web access tools (`web_search`, `source_check`, `fetch_content`, `get_search_content`) are hidden until a human offers them, through `/webaccess on` or by accepting a `request_capability` prompt.

Retirement requires source registry integration, unchanged verification policy, a matching contract, bounded source snapshots, `npm run check`, and closed registered validators. Receipts distinguish machine-observed gates from model-reported acceptance evidence. Stale selections, changed source, missing evidence, and failed checks reject retirement. A validator proves only the behavior it exercises; the human remains responsible for accepting the result. The loop never commits, publishes, or installs a resulting change automatically.

Wishlist records remain local under `<agent-dir>/specpi/`. Sanitization and salted identifiers reduce exposure but do not guarantee anonymity. Evidence supplied by a model can still be sensitive. Reports, archives, outcomes, and issue drafts are local; external sharing requires explicit human action. Pi's own model requests and retention are governed by Pi and the chosen provider.

## Installer and migration

Package acquisition requests exact npm dependency saves through the child process environment and verifies the installed top-level versions before completing the transaction. A mismatch fails the operation and triggers managed-state rollback. Upstream transitive dependency ranges remain outside this pinning guarantee.

`plan` is read-only. Installation, updates, and removal require confirmation or `--yes`. SpecPi manages its two first-party extension families, improvement skill, manifest, AGENTS marker block, and the eight package entries listed in `templates/settings.json`. Normal install/update runs `pi install npm:<name>@<pin>` for every default package. Only package entries and one `defaultTools` entry are merged; existing resource filters and unrelated settings are retained. `--skip-package-install` allows a core-only install or preserves an already configured base during update, skipping Chromium setup too. Confirmed install/update otherwise invokes the installed Browser QA Node bin with bounded subprocesses to download Chromium and check readiness. `--skip-browser-install` skips that setup only; doctor still checks readiness. SpecPi never passes `--with-deps`, auto-installs OS libraries, or installs/removes Bun. No new shell profile integration is installed.

Managed configuration and files are locked and backed up before mutation; first-party writes are atomic and checksum-tracked. Failure restores the saved configuration and first-party files. Package acquisition runs upstream package-manager scripts and may leave downloads, dependency changes, or external script effects even after configuration rollback. Those effects, including downloaded browser-cache bytes, are outside SpecPi's transaction and survive uninstall. Updates require `--force` before replacing modified retained resources. Retired resources are backed up before deactivation; pre-install files are restored where ownership records identify them. Old runtime directories are moved into backups without inspecting their contents. Backups and private evidence remain after uninstall and can contain sensitive local material.

Install and update, including core-only runs, add `+codemode` to `defaultTools` so Pi's built-in codemode tool is on. The installer first reads `pi --version` and adds nothing below Pi 0.99, which would read the entry as a plain tool list and drop its default tools. It also adds nothing when `defaultTools` already names codemode in any form (including `-codemode`), is empty (no built-in tools), is not an array, or when `extensions` contains `-builtin:codemode`. The manifest records only whether `defaultTools` existed before; uninstall removes `+codemode` and deletes the key only if SpecPi created it. Codemode runs model-written JavaScript in Pi's QuickJS sandbox, whose only capability is calling the session's tools. Each nested call goes through the same `tool_call` and `tool_result` handlers as a direct call, so `/scope`, the permission system and the LANCET guard see and can block it; an isolated Pi 1.0.0 session with a faux provider confirmed a nested `bash` call reaches `tool_call` with a `<parent>/<n>` id and that a block rejects it inside the script. Codemode is itself one more tool to these gates; under the permission system it falls to the `"*"` fallback unless a rule names it.

Legacy migration restores only recorded settings ownership, preserves differing user values, and removes only the SpecPi shell marker block before applying the new base. The installer does not enumerate or modify authentication, provider credential stores, trust, sessions, missions, history, or unrelated private evidence. Normal updates deliberately reapply the default package pins, retaining the original entry for removal. Uninstall restores only package entries that still match the last installed value; user edits are preserved. Downloaded packages and tools are not deleted. Resources that SpecPi never owned require separate human management.

## Browser QA package

`packages/browser-qa` is an independently released Node-native Pi package extracted from the retired browser tools. The installer acquires the immutable `specpi-browser-qa@0.3.0` release through Pi; its source and dependencies are not bundled in SpecPi's npm artifact. Its fourteen tools ship withdrawn: until a human runs `/browser on`, saves `/browser startup on`, or accepts a `request_capability` prompt, Pi is offered no browser tool and the model cannot launch a browser at all. Delegation ships off on the same basis, so neither package's tools reach a default session. Its explicit setup downloads Playwright Chromium, and its doctor performs offline rendering, image-comparison and accessibility smoke checks. It uses package-local dependency resolution and the standard Playwright browser cache rather than SpecPi's retired managed runtime. No Pi settings, credentials or personal profiles are migrated. The ephemeral context is not OS/network isolation; pages can reach localhost and private networks. See the package's [security documentation](packages/browser-qa/SECURITY.md) for artifact retention, best-effort redaction, permission and cleanup limits. This is QA tooling, not general-browser feature parity. BetterWright remains optional/manual. Normal updates restore recorded pre-existing entries and remove only unchanged SpecPi-added BetterWright entries; modified entries, personal browsers, profiles, cookies, and user-owned tools are not migrated or deleted.

## VS Code frontend

SpecPi Chat is a separate VSIX, retained alongside the npm harness. It launches Pi only in a trusted filesystem workspace and communicates over local RPC. Provider credentials stay with Pi. Sign-in is delegated rather than implemented: Chat can start the configured Pi executable in a user-visible VS Code terminal, without RPC flags or a session, and Pi alone prompts, runs any OAuth flow, and writes `auth.json`. Chat sends no input to that terminal, runs no Pi auth subcommand, and accepts no credential over RPC; it reads `auth.json` for one question only &mdash; whether an `openrouter` api_key entry is present &mdash; and the host sends the webview a boolean per source and never a value, so the panel reports where a key would come from without being a place one could leak from; it observes only the terminal's closure, and restarts the connection so Pi re-resolves its catalogue. Its sign-in prompt is derived from Pi's own available-model list and missing-credential error text. The frontend manages only its own workspace-storage conversation catalog and sessions for user-directed history, branching, and export; it does not import unrelated Pi histories. Attachments, images, webview messages, and file navigation are validated and bounded. Rendered model and tool output is untrusted text, never executable HTML.

Approval replies are tied to the current conversation, client, and request ID. Disconnect, cancellation, and expiry never grant permission. Multiline package approval context is shown in the dialog body; requests exceeding the display budget are cancelled rather than approved against incomplete context. The Permissions editor can change only the package's documented global/project `config.json`, with bounded UTF-8/schema validation, explicit native confirmation, connection/scope binding, stale-revision checks, private backups, and atomic replacement. It reads no permission logs, trust decisions, credentials, or agent frontmatter. Symlinks, hardlinks, and special files are refused. Backups remain local and may contain sensitive policy; another same-user process can still race filesystem operations. Saving does not claim runtime activation: the user can restart the selected chat and inspect `/permission-system show`. Global/project changes may affect other chats as upstream reloads them; session approvals are cleared by restart. Delegated agent activity comes from a read-only, connection-local widget the delegation package publishes; Chat registers no tools or commands of its own for it. The VSIX does not install packages or duplicate their enforcement; only Permission System configuration has an explicit editing UI. TUI-only custom components are outside Pi's RPC rendering support.

The optional **Destructive guard** preset replaces the complete global configuration draft; it never merges with the old global rules or options. The user can preview/edit the replacement or undo it without a write. Saving uses the same native confirmation, bound global destination, backup, and atomic replacement as ordinary settings edits. The preset asks by default, adds explicit destructive-command denials, disables YOLO and logging, and clears the authorizer chain. It performs no project-policy or agent-definition inspection and no additional directory enumeration. This is a configuration template, not independent enforcement or a claim that global settings override every runtime scope. Project/per-agent policies and session approvals remain upstream-controlled, including 32.0.2's ordered-map merging behavior. Case-sensitive patterns can block benign uses and miss scripts, alternative executable spellings, or unconfigured shell tools. Restart after saving and inspect effective policy; the template is not a sandbox.

## Upstream package boundary

The eight packages add their own extensions, tools, prompts, skills, network connections, filesystem operations, and subprocesses under their upstream defaults. They are not confined by the improvement loop's selection requirement. Permission System owns tool policies; SpecPi does not inject a duplicate guard or claim its coverage. Delegation starts real Pi child sessions in the same process tree, restricted to a frozen source snapshot and three read-only tools; experiments run the user's own `git` and create worktrees on disk. Neither is an OS sandbox. See [delegation](packages/delegation/SECURITY.md) and [experiments](packages/experiments/SECURITY.md). `pi-lens` and `pi-background-tasks` (including its Anthropic provider wrapper) are no longer part of the default base. Normal updates remove only unchanged entries originally added by SpecPi; pre-existing or modified entries and downloaded bytes remain. `--skip-package-install` preserves the old base. Restart Pi and Chat connections to unload retired extensions. Independently retained installations remain trusted upstream code; retained Pi Lens can still apply configured formatting/autofixes. Web access and usage reporting can contact services and use credentials through their upstream implementations. SpecPi's local-only wishlist collection policy does not describe all activity of those packages.

Top-level versions are pinned; upstream transitive dependency ranges are not frozen by SpecPi. `doctor` reads configured pins and installed package metadata, then invokes the installed Browser QA bin for real offline rendering, pixel-comparison, and accessibility checks when the managed base includes it. Doctor never downloads a browser, and missing Chromium or OS libraries fail with recovery guidance. It does not validate provider access or browser readiness for independently retained BetterWright. Core-only installations do not run browser checks. `check:base` acquires the packages, checks combined resource loading and Chat RPC startup, and exercises upstream approval, denial, and cancellation with synthetic context. It uses temporary home/configuration directories and sends no model prompt. The base check also exercises Node-only Chromium setup and offline Browser QA readiness; authenticated services, OS isolation, and every upstream tool's behavior remain outside that check. See [THIRD_PARTY.md](THIRD_PARTY.md) for sources and compatibility limits.

Manifests, source checkouts, Pi, dependencies, and the operating system are trusted inputs. Validation rejects unexpected managed resource paths and symlinked managed paths; it does not claim protection against another process changing files during an operation. Repository and release checks use isolated temporary Pi directories.
