# Changelog

## v3.0.1 — 2026-08-13 — patch: a peek no longer hides a recipient's aged undelivered mail

<!--
  VERSION: 3.0.1. PATCH — one shipped-in-3.0.0 silent-non-delivery fix, its wording, and a
  changelog cleanup; nothing else. The 3.0.0 undelivered-mail work keyed the DRAIN's
  since-escape on the OBSERVED axis (`seq`, stamped by a non-consuming peek) while the purge
  exemption and wake detector keyed on the DELIVERED axis (`read_by_session`), so a peek
  flipped an aged undelivered message out of the recipient's own drain. Routes all three
  through one NEVER_DRAINED_SQL SSOT so they cannot disagree again.
-->

A patch for a silent-non-delivery path shipped in 3.0.0.

### Fixed — a peek no longer removes aged undelivered mail from a drain (#198 follow-up)

3.0.0 made a pending drain return undelivered mail regardless of `since`, but keyed
"undelivered" on `seq IS NULL` (never *observed*). A `peek` — which every watcher does
(Sentinel, the dashboard, `relay watch`) — stamps `seq` without delivering, so peeking an
aged undelivered message lapsed that escape and the `since` window then hid it from the
recipient's own default drain (recoverable only via `since='all'`). The drain's escape now
keys on `read_by_session IS NULL` (NOT DRAINED — the same `NEVER_DRAINED_SQL` SSOT the purge
exemption and the wake detector already use), so a peek can no longer remove mail from
reach. Undelivered means not drained, not "not looked at."

### Docs

- The `get_messages` `peek=true` description no longer claims it "suppresses the
  read-side-effect entirely" — it suppresses the read-MARK but still stamps the observation
  cursor (`seq`), like any first view. It now says what it does.

## v3.0.0 — 2026-08-12 — BREAKING: `since` strict-cursor contract withdrawn; wake-coverage detector; undelivered-mail durability

<!--
  VERSION: 3.0.0. Twenty-two merges since 2.25.0. MAJOR because a PUBLISHED public contract
  changed: 2.25.0's `since` field described "only return messages after this time" (a
  strict cursor), and #198 makes a pending drain return never-observed mail regardless of
  `since` — a consumer coded to that sentence now receives rows created BEFORE their
  cursor. Backward-incompatible whatever the intent, so a major bump, not a minor: the
  version number is a claim, and a MINOR release that tells users to audit for a
  behavioural break would contradict its own notes. The change itself is correct — the old
  promise could only be kept by leaving mail permanently undeliverable — so we label it
  honestly rather than shipping it quietly. Also ships a new default-on feature (the
  wake-coverage detector) and a retention change (undelivered mail held to a 30-day grace).
-->

Twenty-two merges since 2.25.0. **This is a major release for one reason: a published
`since` contract changed.** Around that break, wake-coverage detection is now durable and
undelivered mail is no longer silently deleted at seven days — plus a pending-predicate
SSOT unification, guard hardening, and a security-advisory + dependency sweep.

### ⚠ BREAKING — the `since` strict-cursor promise is withdrawn (#198/#199)

2.25.0 documented `since` as "only return messages after this time." **That is no longer
true for a pending drain.** A pending `get_messages` drain now returns mail that was never
observed by anyone **regardless of `since`** — `since` bounds only already-OBSERVED
history. A consumer that used `since` as a strict pagination cursor on a pending drain
will now receive messages whose `created_at` predates the cursor (the never-delivered
ones). **Audit any pending-drain `since`-cursor logic before upgrading.** History reads
(`all` / `read` / `resolved`) keep the plain `created_at >= since` bound and are
unaffected, as is `peek=true`.

Why it changed, and why it could not stay: the old promise could only be honoured by
leaving never-observed mail permanently undeliverable — a message older than the caller's
window was never returned, never seq'd, and eventually purged (silent non-delivery).
Delivering owed mail and honouring a strict cursor are mutually exclusive here; we chose
delivery and are versioning the break honestly rather than shipping it quietly under a
minor. No tool signature changed.

### Messaging

- **One pending predicate, everywhere (#56/#187/#191).** Wake and drain can no longer
  disagree about what "pending" means — `resolved_at IS NULL AND read_by_session IS NULL`
  is a single SSOT (`pendingGlobalClause`) routed through every surface (drain, wake,
  `stop-check.sh`, `getInboxSummary`, and the purge exemption below).

### Durability — undelivered mail is an obligation, not history (#201/#203)

- **Undelivered mail is no longer deleted at 7 days.** A message **no recipient has
  drained** (`pendingGlobalClause` — the SAME predicate the wake detector uses as its
  candidate set, so the two cannot drift) is an undelivered *obligation*, and the 7-day
  transient purge now **exempts** it, holding it to a bounded **30-day** operational-tier
  grace (`RELAY_UNDELIVERED_GRACE_DAYS`, default 30; `0` disables the extension but never
  the announcement). This includes mail that was *peeked but never drained* — the wake
  regression the detector reports on.
- **Deadletter announcement.** When an undelivered obligation is finally dropped at the
  grace horizon, a `deadletter:` line is written to **stderr** (recipient, age, id,
  effective grace) — the only record that survives the row's deletion. Operators who
  parse logs will see this new message type. It is unconditional (fires at any bound,
  including `grace=0`), emitted after the drop commits.

### Wake-coverage detector (new, default-on) (#60/#200/#202)

- A periodic daemon sweep classifies each agent with mail stuck past the **effective
  threshold** (`RELAY_WAKE_BOUND_MS` + `RELAY_WAKE_ANTI_FLAP_MARGIN_MS`, default 24h+24h =
  48h) into **three verdicts** — *covered* (draining since the mail arrived; not
  reported), *uncovered* (drained before but not since — a possible wake-path regression,
  reported as an alarm), and *unobservable* (no MCP drain recorded for the identity —
  reported, but **never** as uncovered). It emits to **stderr** (REPORT-FIRST — a human
  decides), never mutates the DB, and is default-on (`RELAY_WAKE_DETECTOR=0` to disable).
- New durable **`agents.last_drain_at`** column (schema auto-migrates; a one-time
  best-effort backfill from retained drain events). Reserved-for-forward correctness: an
  identity dark past event retention at migration reads *unobservable* until it next
  drains.

### Guards + hardening

- Auth-generation drift guard now sees SQL hoisted to module-scope consts and refuses
  concatenated/cross-module SQL at `prepare()`/`exec()` (#57/#59). Secret-register guard
  wired into a gate with a filesystem-driven wiring meta-check (#61). AST guards pinned to
  `typescript-legacy` behind a loud-fail parse gate, unblocking a future TS7 bump (#174).
- `relay backup`/`restore` work on the wasm driver (#171/#190); the schema migration
  chain is single-sourced across both init paths (#171/#188); a golden-snapshot guard
  locks the 37 tools' emitted input schemas as an external contract (#186).

### Dependencies + CI

- Security-advisory sweep: `js-yaml` / `undici` / `brace-expansion` HIGH + `hono` MED
  override pins (#175), `nanoid` 3.3.16→3.3.18 HIGH (#184); an extension `npm-audit`-high
  CI gate with override-retirement discipline (#176). ESM-safe `createRequire` in the
  `getDb()` native fallback (#170); the pre-publish gate builds `dist/` before vitest so
  it passes from a fresh clone (#169).

## v2.25.0 — 2026-08-05 — ADR-0011 message disposition + read-receipts; a security-advisory sweep; dependency majors

<!--
  VERSION: 2.25.0. Sixteen merges since 2.24.0. Additive feature release (ADR-0011,
  schema v24 auto-migrates) plus a security-advisory sweep and dependency majors —
  no breaking change, so a MINOR bump. This is the RELEASE PR: content is final and
  package.json is bumped to 2.25.0 — ready to cut once merged + published.
-->

Sixteen merges since 2.24.0. The headline is **ADR-0011 — message disposition +
read-receipts** (schema v24, the new `get_outstanding` tool, now **37 MCP tools**).
Around it: a security-advisory sweep that unblocked the whole merge queue — three new
HIGH advisories had entered the tree and red-failed the audit gate on every open PR —
the dependency majors that arrived with it (uuid 14, TypeScript 6, `@types/bcryptjs`
3), and docs/tooling hardening that keeps the release honest (a de-versioned README
masthead behind a drift-guard, a locked SSRF regression, extension-tree parity).

### Messaging

- **ADR-0011 — message disposition + read-receipts (#127).** Messages now carry a
  disposition and a write-once `read_at`; the new **`get_outstanding`** tool
  (sender-scoped, auth-gated) surfaces what a recipient has not yet acted on.
  Schema **v24** — additive, auto-migrates from any prior version. Tool count is now
  **37** (36 + `get_outstanding`). Resolution and read-state are orthogonal.

### Security

- **Three new HIGH advisories cleared — they had silently blocked the entire merge
  queue (#159).** The advisory DB moved after 2.24.0 published: `fast-uri`
  (host-confusion, GHSA-7p8r) — which our own `overrides.fast-uri: 3.1.4` **pinned to
  the vulnerable version** — **`ajv`**, flagged high because our **direct** `ajv@8.20.0`
  dependency pulls that same vulnerable `fast-uri` (it declares `fast-uri ^3.0.1`) — and
  `ip-address` (SSRF / trust-boundary bypass). The audit-high step runs inside the required Test
  jobs, so under strict branch protection every open PR failed CI identically until
  this landed. Fix: `fast-uri` override → 3.1.5, `ip-address` → ^10.4.0. **Measured:
  published-2.24.0 *consumers* were never exposed** — npm overrides do not travel to a
  package's dependents, so consumers resolved the patched versions through normal
  range resolution; the pin only ever affected our own dev/CI. A security pin without
  an expiry re-check had become the sole hazard.
- **SSRF alternate-spelling rejection locked (#167).** Test-only regression covering
  the leading-zero / hex / decimal-integer / IPv4-mapped-IPv6 / NAT64 spellings
  (`012.0.0.1`, `127.1`, `0x7f.0.0.1`, `2130706433`, `[::ffff:…]`, `[64:ff9b::…]`) —
  all already fail closed via WHATWG-URL canonicalization; this locks it against a
  silent refactor regression.
- **Publish-gate hygiene (#156/#157, #158).** The prebuild-guard test no longer leaks
  an inherited `RELAY_ALLOW_PROD_BUILD` into its spawned guard (a false green), and the
  gate's stale "audit-level=moderate" header comment is corrected to the actual `high`
  threshold it enforces.

### Dependencies

- **uuid 11 → 14 (#80).** We call only `v4`; clean bump; kills the moderate uuid<14 line.
- **TypeScript → 6.0.3 (#164), NOT 7.x.** 6.0.3 is the latest *classic-compiler* major.
  TypeScript 7.0 is the native "Corsa" rewrite and **drops the classic JS compiler API**
  (`ts.SyntaxKind`, `ts.ScriptTarget`, `ts.createSourceFile`) that our AST-walk guard
  scripts consume — so 7.x is a deliberate guard-port task, not a dependency bump.
  Recorded so the next bump does not re-learn it.
- **`@types/bcryptjs` → 3.0.0 (#165)**, dev types only. **Extension-tree parity:**
  `fast-uri` → 3.1.5 and `ip-address` → 10.4.0 in `extensions/vscode` (#161/#162),
  matching the root security fix; plus `brace-expansion` / `linkify-it` / `postcss`
  (#104/#110/#135) and CI `actions/checkout` + `setup-node` → 7 (#68/#96).

### Docs & tooling

- **README masthead de-versioned + retired badge fixed + drift-guarded (#166).** The
  masthead hard-coded "v2.22" and had gone stale while shipping 2.24.0 — and the public
  Glama listing mirrors the README. Per-release facts now live in the CHANGELOG; the
  masthead states the (test-guarded, accurate) tool count and points there. The retired
  shields.io VSCode-Marketplace version badge is replaced with a static one that always
  renders. A new pre-publish **masthead version-guard** fails the gate if a bold
  `**vX.Y**` version marker ever reappears in the first 15 README lines.

## v2.24.0 — 2026-07-28 — Security hardening: operator auth, dashboard content isolation, orchestration integrity

<!--
  VERSION: 2.24.0, locked by Maxime this session. Strict semver would make this a
  MAJOR (3.0.0) because #142 is a breaking change, but a major bump signals a
  rewrite to anyone browsing npm and this is a hardening release, not that — so
  2.24.0 WITH the loud upgrade note below. This is the RELEASE PR: content is final
  and package.json is bumped to 2.24.0 — ready to cut once merged + published.
-->

Sixteen merges since 2.23.0. The spine is a security cluster: operator power can no
longer be inferred from being on the box, the dashboard no longer hands message
content to unauthenticated callers, the orchestrator no longer loses work to a dead
agent or a stale daemon, and the diagnostic traffic recorder no longer captures live
tokens. The rest harden the tooling around the release — one shared structural
detector for the must-call guards, a network deadline that covers the whole
exchange, a build-time guard against an accidental production-tree deploy, a
config-path fix so `init` and the daemon agree, and cacheable snapshot output for
the fleet board.

### ⚠️ BREAKING — the operator dashboard now requires a secret (one action needed on upgrade)

**Required step — existing installs AND fresh installs: run `relay init`.**
Operator actions in the dashboard — kill/wake an agent, send a message, focus a
terminal, set status, change the theme, the keyring — return **401 until a
dashboard secret exists**, and **a fresh install ships with no secret, so it has
no working operator actions at all until you run it.** `relay init` generates one
(printed once, preserved on re-run, added to legacy installs automatically). This
is a required setup step, not a footnote — without it the dashboard's operator
controls are inert. **If your dashboard starts refusing actions right after
upgrading, this is the reason — not a regression.**

Why it changed (**#142**, ADR-0006): before this, in the default loopback
deployment, **any local process that could reach the dashboard port could invoke
operator powers** — kill or wake agents, send messages as you, focus terminals —
with no operator credential. Worse, the shared `http_secret` that *every* HTTP
agent already holds was accepted as operator authorization, so a mere agent
transport credential silently granted operator power. Power was derived from
network position plus a transport secret, not from an operator identity. ADR-0006
("location is not a principal") closes it: operator endpoints now require a
**verified dashboard secret regardless of network position**, and the
`http_secret`→operator escalation is removed at every site.

> **◆ FLAGGED FOR MAXIME — a possible SECOND breaking change for this release, not yet decided.**
> Node 20 reached **end-of-life on 2026-04-30** (Node.js release schedule) and no
> longer receives security patches; Node 22 is supported to 2027-04-30. The pending
> better-sqlite3 major (**#81**, 11.x → 13.0.1) declares `engines: node >=22` and,
> verified in CI this session, failed **only** the Node 20 cell (Node 22 + macOS
> passed). Taking it moves this package's own `engines` from `>=20` to `>=22` and
> drops Node 20 from the support matrix — itself a breaking change that belongs in
> these notes. The slot is left ready here rather than bolted on afterward. **Not
> decided.**

### Security

- **The dashboard no longer serves message content to unauthenticated callers (#141).**
  In the default loopback deployment, the dashboard's snapshot endpoint served
  **decrypted previews of every agent's messages and tasks — plus each agent's
  session and process identifiers — to anyone who could reach the port, with no
  password.** That defeated the per-agent token isolation `get_messages` enforces,
  and where at-rest encryption was not configured (the default) it exposed
  plaintext. The unauthenticated snapshot is now built from an explicit
  **allowlist** of non-sensitive identity/presence fields; message and task
  content and process locators are **absent by construction** (a first denylist
  attempt still shipped the raw content column and was caught and replaced).
  Authenticated operators still get full previews.
- **Operator power now requires operator authentication, everywhere (#142).** (See
  the BREAKING note above.) The `http_secret`→operator-auth escalation is removed
  at all three sites; a single resolver decides operator auth from the dashboard
  secret and nothing else.
- **Tasks are never handed to a dead agent (#141).** `post_task_auto` could route a
  task to a **registered-but-closed** agent; the task then dead-ended forever
  while the caller was told `routed=true` — **silent work loss on the core
  orchestration path**. Routing now targets only agents with a live session, and a
  task orphaned by an agent vanishing between post and accept is **loudly
  requeued**, never silently redirected.
- **A stale daemon after an upgrade is surfaced, not silently kept (#141).** After
  `npm update`, the already-running daemon kept serving every HTTP client on the
  **old code until reboot** (launchd does not restart on a code change), with no
  warning — the "merged but never takes effect" trap at the daemon layer. `relay
  init` now **loudly flags a version drift** with the exact remedy, and a new
  operator-invoked **`relay restart`** (with `--dry-run` to preview) applies it. It
  never auto-restarts — bouncing a shared daemon mid-session is the operator's
  call.
- **Token redaction — the traffic recorder no longer leaks secrets (#145).** The
  diagnostic traffic recorder could write agent tokens (and other secret values)
  verbatim into its capture. A central **redact-by-value registry** now scrubs known
  secret values from recorded traffic, so a captured exchange cannot expose a live
  credential.

### Fixed

- **Tether: `claude wake` now actually submits the prompt (#129, Tether v0.7.0).**
  The VS Code wake typed the prompt into the terminal but **never submitted it**
  (the newline rode in the same chunk and did not register as Enter), so a woken
  agent sat with an unsent prompt and looked like it had ignored the wake. It now
  types, settles briefly, then sends the Enter as a separate keystroke.
- **The server no longer closes an idle keep-alive socket first (#149).** The HTTP
  server's keep-alive timeout sat below the idle window a pooling client may hold a
  connection, so under load the server could close a socket the client was about to
  reuse — surfacing as an intermittent `ECONNRESET` on the next request. The
  timeout is raised (to 61s, with the header-read timeout above it) so the client
  always closes first. Affects raw/proxy/non-undici HTTP clients regardless of
  platform (LOW).
- **Network calls bound the WHOLE exchange, not just the headers (#155).** The
  request deadline capped the header-read phase but left the response body/exchange
  unbounded, so a peer that accepted the request and then stalled mid-body could
  hang the caller indefinitely. The deadline now covers the entire request→response
  exchange, so a stalled peer fails fast instead of hanging.
- **`relay init` writes config to the daemon's RESOLVED path (#153, P1).** `init`
  wrote `config.json` to the flat default location instead of the per-instance path
  the running daemon actually resolves, so a configured setting could silently land
  where the daemon never reads it. `init` now writes to the resolved path, so config
  and daemon agree.
- **Tether surfaces a no-delivery condition to a human, not an 8-second blip (#154,
  Tether #3).** When a wake or message failed to deliver, Tether showed only a
  transient ~8s status-bar note that was easy to miss; the condition is now surfaced
  persistently, so a human actually sees that delivery did not happen.

### Internal / hardening

- **Dashboard theme values validated against a CSS-color grammar (#147):** the
  dashboard theme endpoint now rejects theme token values that are not well-formed
  CSS colors, closing a stored-value injection vector on an operator-settable field
  (LOW).
- **From-source native-build CI gate (#148):** CI now compiles `better-sqlite3`
  from source on the matrix, closing a prebuilt-only blind spot — a dependency that
  breaks on a from-source build (or drops a Node version) now reds CI instead of
  passing on a cached prebuilt.
- **Liveness probe caches evicted on every state change (#140):** the five writers
  that change an agent's session/anchor now evict the probe caches unconditionally,
  so a stale cached verdict cannot outlive the change.
- **Backup test asserts the harm, not a proxy (#144):** the backup test asserts the
  atomic-swap **harm** directly instead of a poll-count proxy (ADR-0015).
- **Docs (#146):** ADR-0007 (control-plane target architecture) committed to
  version control.
- **Must-call guards share one structural detector (#151):** the guards that
  assert a required call is present now use a single AST-based helper, closing
  comment-, alias-, and class-field-shaped evasions a per-guard textual check
  missed. One detector, so the guards cannot drift apart.
- **Prebuild guard against an accidental production-tree build (#156):** an npm
  `prebuild` hook refuses `npm run build` in the live-serving tree — where a build
  silently overwrites the `dist` the daemon and the whole stdio fleet load — unless
  `RELAY_ALLOW_PROD_BUILD=1`. Positive-match only, so CI checkouts and throwaway
  worktrees build freely; repo/dev tooling only (excluded from the npm package).
- **Stable snapshot body + ETag/Date caching for `/api/snapshot` (#152):** the
  dashboard snapshot endpoint now emits a stable body with `ETag`/`Date`, so a
  polling consumer (the lumen fleet board) can cache and conditionally re-fetch
  (304) instead of re-pulling the full snapshot every tick. `last_alive` is excluded
  from the ETag input so a liveness-only heartbeat does not churn the cache.

## v2.23.0 — 2026-07-25 — Config-clobber closed, orchestration-anchor safety, DX + a HIGH dependency advisory

> _**Backfilled 2026-07-27.** 2.23.0 shipped without release notes; this entry was
> reconstructed after the fact from the first-parent commit log between v2.22.0
> (`1f64ec8`) and the **published** 2.23.0 commit (`22c92db` — the npm `gitHead`
> for 2.23.0, which is two commits past the `2.22.0 → 2.23.0` version bump and so
> includes #138 and #139) — it is **not** contemporaneous. The PRs are named per
> item; for the precise diff see those commits._

The load-bearing one is a **shipping defect that anyone who ran the test suite hit.**

- **`npm test` no longer rewrites your real Claude config (#125).** **Anyone who
  cloned this repo and ran `npm test` had their real `~/.claude.json` and
  `~/.claude/settings.json` silently rewritten to point at their checkout — no
  warning, no error.** On dev machines it quietly pointed every agent at an
  unmerged build for nine days; run from a path containing a space it wrote a
  percent-encoded path (`%20`) that does not exist, so it failed mutely rather than
  loudly. **If you ran this repo's test suite before 2.23.0, check your
  `~/.claude.json` MCP entries.** Fixed in four layers, each with a proven-failing
  negative control: HOME sandboxes for the three offending test environments, a
  chokepoint guard that throws on any real-user-config write under test, a
  suite-wide tripwire hashing the real config files, and the `%20` path fix.
- **The user-config tripwire compares only relay-owned regions (#139).** A
  follow-up to #125: that suite-wide tripwire's whole-file byte comparison
  false-tripped on unrelated user edits to `~/.claude.json`, so it now compares only
  the relay-owned config regions — a legitimate user edit no longer reds the suite.
- **Security — postcss HIGH advisory patched (#134, GHSA-r28c-9q8g-f849).** A
  path-traversal advisory (CVSS 7.5) landed on `postcss <= 8.5.17`. It reaches this
  project only as a **dev-only transitive dependency** (vitest → vite → postcss;
  production deps are empty), so runtime exposure is nil — but it reds the `npm
  audit --audit-level=high` gate. Patched lockfile-only (postcss 8.5.15 → 8.5.23),
  no production or SDK change.
- **Anchor takeover is CAS-guarded; no silent auto-takeover (#136, ADR-0012).**
  `force` re-anchoring uses a compare-and-swap plus a dead-anchor diagnostic
  instead of an unconditional takeover, so two agents racing for the same identity
  anchor cannot silently clobber each other.
- **Publish gate hardened (#138).** The pre-publish gate refuses to publish from
  the wrong branch and treats a **missing CI run as a failure**, not a pass —
  closing a stacked-PR hole where an empty CI rollup read as clean/mergeable.
- **`relay resolve` — a CLI ack for MCP-mute sessions (#133).** A session that
  cannot call the MCP tool (e.g. a plain terminal) can still acknowledge/resolve
  messages from the command line.
- **`relay send` stdout hygiene (#130).** Usage/help now goes to stderr and `relay
  send` prints a clean message id on stdout, so a script can capture the id without
  parsing noise.
- **Tether wakes routed by observed agent-state; wake-stacking ended (#126).** A
  "landed" gate stops the extension from stacking repeated wakes on an agent that
  has already been woken and is working.
- **Stop-hook is a read-only wake (#124).** The stop-check hook never consumes mail
  it cannot prove it delivered — a silence-as-failure guard so a wake can never
  quietly eat a message.

## v2.22.0 — 2026-07-22 — ADR-0005: relay DX hardening (curl/script-caller friction)

Five fixes for callers driving the relay over raw HTTP (no MCP client wired in), from a real friction report. The load-bearing one is a **safe self-serve orphan cleanup**.

- **Safe orphan cleanup (#4) — `abandon_registration` (caller-initiated ONLY; no auto-GC).** A caller who registered but lost the `agent_token` before ever authenticating (e.g. a truncated `curl` capture) orphaned a row it couldn't unregister (unregister needs the lost token) — and had to reach for the *destructive operator endpoint*. Now: register returns a one-time, name-scoped **`registration_recovery`** handle; `abandon_registration(name, handle)` self-cleans the orphan without a token. **Safe by construction** via the establishment keystone — an `established_at` marker (stamped at token auth, credential recovery, and spawn provisioning) means abandon can **only ever** remove a row that has NEVER become a legitimate identity, so a working agent self-excludes and can never be reached (the handle proves the caller is the registrant + covers the just-registered race). **An automatic orphan GC was designed, adversarially audited five rounds, and CUT by final ruling:** abandonment is undecidable from row state — a slow-spawned child, an idle recovered agent, and a genuinely abandoned registration are byte-identical in the data, and each audit round found a different legitimate identity being reaped. The rule that came out of it: *do not attach an irreversible action to a predicate that cannot be decided by observation.* Un-abandoned rows persist harmlessly (session-less, never established, name reclaimable via unregister or `relay recover`); `tests/v2-22-0-no-auto-gc.test.ts` proves the GC is gone, not disabled. Architect-designed (ADR-0005); the keystone invariant is adversarially tested.
- **Reliable token capture (#2).** `agent_token` is now the **first** field of the register response, so a truncated `head -c` read still captures it (it was buried 4th).
- **Plain JSON for one-shots (#3).** A non-streaming one-shot `POST /mcp` now returns `application/json`, not an `event: message\ndata: {…}` SSE frame — curl/script callers can `JSON.parse` the body directly. The stateful/streaming (Tether SSE) path is unchanged.
- **`send_message` accepts `message` (#5).** The MCP tool now takes `content` **or** its alias `message` (exactly one; both-with-different-values rejected), matching the REST endpoint + the agent-team `SendMessage` — one send vocabulary across every surface.
- **HTTP one-shot recipe (#1).** New `docs/http-one-shot.md`: register → capture → send → self-clean over `curl`.

- **The v2.0 30-day dead-agent purge is also CUT (same ruling, worse instance).** `purgeOldRecords` deleted any agent row with `last_seen` older than 30 days — no principal asking, establishment not even checked, so a working token-authed agent that went idle 31 days was deleted and its name freed for anyone to claim (the v2.14.0 reserved-name exemption had already documented that freeing a name reopens the bootstrap-claim window — the rationale applied to every name, not just reserved ones). Found by codex's re-audit of the orphan-GC cut as the last remaining autonomous agent-row deletion. The purge tick now deletes messages/tasks/logs/events — records with retention windows — and **no agent row, ever**. Deliberate pruning stays available via `relay purge-agents` (dry-run by default, `--apply` + audit).

Schema v22 (`first_authed_at` + `registration_recovery_hash`/`_expires_at`, additive). 36 MCP tools (adds `abandon_registration`). Native + wasm parity.

## v2.21.0 — 2026-07-22 — ADR-0002: agent class/flare topology

Agents can now declare a **coordination class** — a coarse "flare" of their posture in the team — and `discover_agents` gains a **`view='topology'`** that renders the live team grouped by class. A fresh agent can also learn its peers on session start (opt-in).

- **`class` — a self-declared, immutable coordination posture (schema v21).** One of `orchestrator | builder | advisory | auditor | transient`, declared at `register_agent` and immutable thereafter (the `managed`/`host_id` precedent). It is a THIRD axis, orthogonal to `role` (free-text label) and `capabilities` (what an agent does) — kept deliberately coarse so it never collapses into capability. Undeclared/legacy rows read as `unclassified`; `bridge` is reserved for a future federation node. Single source of truth: **`src/agent-class.ts`**.
- **`discover_agents view='topology'`** (default `view='list'` is 100% back-compat) groups the live team by class, flat within each `{name, role, class, status}`. **Two independent exclusions:** dead/terminal agents (liveness verdict) and the `transient` + `unclassified` classes are omitted from the who's-who (with honest excluded-counts — no silent truncation). The `class` field is surfaced through the discovery projection (the gate that silently drops `managed`/`visibility`), so it actually reaches clients.
- **Opt-in SessionStart onboarding map** (`RELAY_ONBOARD_TOPOLOGY=1`, default OFF) — a compact "team by class" roster delivered to a freshly-started agent via `check-relay.sh`. Off by default so existing installs' session output is unchanged.
- **Taxonomy drift guard** (`scripts/agent-class-guard.mjs`, wired into the pre-publish gate) — a TS-AST walk that rejects a class-value branch (equality/switch) or a parallel class vocabulary (an array of ≥2 class ids) defined anywhere outside `src/agent-class.ts`, mirroring the cli-profile guard. It prevents a taxonomy re-fork; an adversarial negative-fixture test proves it fails on a synthetic re-fork. Native + wasm parity. No new tool (a `view` param on `discover_agents`); not security-critical.

## v2.20.1 — 2026-07-22 — Verified-token cache on the explicit-caller auth path

ADR-0003 made the O(N) **token-only** auth scan O(1). This extends the same verified-token cache to the **explicit-caller** path (`enforceAuth` for tools that name their caller — `send_message.from`, `get_messages.agent_name`, … — the orchestration hot path), which still bcrypt-verified on every call.

- **Cache short-circuit, impersonation-gated.** The explicit-caller branch now consults the same verified-token cache first; a hit skips the per-call bcrypt. A hit is honored **only when the cached verdict belongs to the claimed caller** — a valid token for agent X can **never** authenticate a claim `from: Y`, even with a warm cache (the name-match gate). A miss falls through to the unchanged `authenticateAgent` flow, so every revoked / recovery / rotation-grace / legacy-can't-actor-stamp behavior is preserved, then re-populates the cache.
- **One cache layer, one invalidation.** Both auth paths now route through the same `verifiedTokenCacheGet` / store-or-heal helpers (a security argument, not just DRY — two paths that must invalidate identically shouldn't be maintained separately). The generation counter still gives instant revocation on both: a revoked/rotated token → generation moved → cache miss → the correct error. bcrypt remains the sole verifier.
- **Deterministic O(1) migration.** The explicit-path miss now lazily **self-heals** the caller's lookup digest too. Since every agent makes explicit-path calls, the whole fleet's `token_lookup` populates on first send — so token-only calls also go O(1), closing the NULL-digest gap left by ADR-0003 (which only self-healed on the token-only path). Native + wasm parity. No new tools, no API change.

## v2.20.0 — 2026-07-22 — ADR-0003: O(1) token auth (indexed HMAC locator + verified-token cache)

Token-only tool calls no longer scan every agent with bcrypt. Two O(N) linear bcrypt scans (`resolveCallerByToken` for the dispatcher, `checkToken` for `health_check`) are replaced by an O(1) indexed lookup plus a verified-token cache — bcrypt stays the sole verifier throughout.

- **O(1) HMAC locator (schema v20).** New `agents.token_lookup` / `previous_token_lookup` columns hold `HMAC-SHA256(lookup_key, token)` (hex), indexed by `idx_agents_token_lookup`. Auth resolves the candidate row via a single indexed SELECT, then **bcrypt confirms** — the digest only narrows candidates, never authorizes (a digest collision is rejected by the bcrypt check). The lookup key is a dedicated HKDF subkey of the encryption keyring when one is configured (key-separated from `http_secret` + the record key; rotates with the keyring), or a persisted per-instance secret (`<instance>/token-lookup.key`, 0600) in plaintext mode.
- **Verified-token cache (`src/auth-cache.ts`).** Per-process, LRU-bounded, TTL-capped cache of positive verdicts keyed on the digest — never the plaintext token, never the bcrypt hash. A hit skips the locator + bcrypt.
- **Invalidation via a global `auth_meta.generation` counter — the correctness mechanism.** Every mutation that can change a token's validity bumps it; a cache entry is served only when its stamped generation still matches, giving **instant revocation** regardless of TTL. The classic trap is covered: `revoke_token` keeps `token_hash` for forensics (the token still bcrypt-matches), so validity is keyed on generation, not hash presence — a revoked token is denied on its very next call. A build-time drift guard (`scripts/auth-gen-guard.mjs`, with an adversarial negative-fixture test) fails the build if any token/auth mutator omits its bump.
- **Zero-lockout migration.** `token_lookup` can't be backfilled from a bcrypt hash, so legacy rows stay NULL and authenticate via an O(N) fallback that **lazily self-heals** the digest on first authenticated call — each agent goes O(1) thereafter. Native + wasm driver parity covered.
- No new tools, no API changes; auth semantics (active / rotation_grace / revoked / recovery / legacy) are byte-for-byte preserved.

## v2.19.0 — 2026-07-22 — Liveness derivation: presence that stops lying

Fixes the presence **lie**: a rate-limited-but-alive agent (e.g. an agent mid-audit) reported `status=offline` because presence was still derived from `last_seen` **age**, and the liveness verdict anchored **only** on a registered `agent_pid`. Consumers misread "offline" as "dead."

> **⚠ Presence-semantics contract change** (minor-version bump). `computeStatus` (the age→status mapper) is **removed**; the coarse `status` is now PID/verdict-derived. A programmatic register with **no** `host_id`/`agent_pid` anchor now reads `unknown` (was age-based `online`) — real hook/Tether registrations stamp an anchor and read `online`. The `status` enum drops `stale` (`online | offline | unknown`). External snapshot readers keying on `status` should treat `unknown` as "no liveness signal," never as offline/dead.

- **Security — `fast-uri` HIGH advisory pinned out.** Pins `overrides.fast-uri = 3.1.4` to clear **GHSA-4c8g-83qw-93j6** and **GHSA-v2hh-gcrm-f6hx** (CVSS 7.5 each — URI host-confusion via failed IDN canonicalization + literal-backslash authority delimiter; both affect fast-uri `3.0.0`–`3.1.3`). It reaches us transitively via `ajv@8.20.0`, which declares `fast-uri ^3.0.1`, so `3.1.4` is in-range and non-breaking (the MCP SDK is untouched — `npm audit fix` was rejected because it force-upgrades the SDK to a breaking major). Repo-wide `npm audit (high+)` in CI now passes.

- **The coarse `status` is now derived from the liveness VERDICT, not `last_seen` age.** `alive → online`, `dead → offline`, `unknown → unknown`. A live agent **never** reads `offline`; `last_seen` is pure telemetry. (The verdict-based `agent_status` already dropped age in v2.15.0; this finishes the job for the last age-based surface. The `status` enum is now `online | offline | unknown`, was `online | stale | offline`.)
- **The liveness verdict gained an argv-scan fallback.** When an agent has no (or a stale) `agent_pid`, the verdict also confirms *alive* by finding a live process on **this host** that advertises `RELAY_AGENT_NAME="<name>"` in its argv — the agent's **own** process. Both-side anchored + a **literal** substring search (not a `pgrep -f` regex), so `foo` never matches `foobar`/`foo-x`, and a name metacharacter can't inject. Host-scoped + behind the 5 s probe cache.
- **`host_shell_pids` is deliberately NOT probed** as a fallback: the Tether ancestry chain includes the terminal/shell, which outlive the agent — probing it would false-read a crashed-agent-in-an-open-terminal *alive* (the v2.13.0 §3 contract). The argv scan fixes the same case via the agent's own process, without that false-alive.
- On-read (no periodic sweep). Remote/federated agents stay `unknown` (a daemon can't PID-probe another host) — a cross-host heartbeat is a parked follow-on.
- Acceptance test reproduces the exact failing shape (agent_pid null, process alive with the name in argv) and asserts the surface shows **alive**, not offline.

## v2.18.0 — 2026-07-21 — Sentinel: `relay watch` — autowake for any terminal

Sentinel is the relay's built-in autowake for terminal agents **not** in VS Code/Tether (iTerm2 personas, plain terminals, remote sessions) — the poll/marker-based sibling of Tether's push-based wake. It's the shipped replacement for the hand-armed inbox-watcher bash loop. Pairs with Tether: **Tether = push-wake in VS Code; Sentinel = wake anywhere.**

- **`relay watch <agent> [--interval S] [--once] [--json]`** — stands watch over `<agent>`'s inbox and prints a wake line when new mail arrives, so a harness Monitor (or you) can nudge the agent to read it. It consumes the sanctioned cheap primitive `peekMailboxVersion` **in-process** (a single indexed `COUNT`, not a raw `sqlite`-every-Ns scan loop); the wake signal is `total_unread_count` **rising** (`last_seq` is unreliable before first observation — v2.3.0 Codex HIGH #2). `--once` checks + exits (scripts/smoke); `--json` emits machine-consumable wake lines.
- **Event-driven when `RELAY_FILESYSTEM_MARKERS=1`:** waits on the daemon-written delivery marker (`~/.bot-relay/marker/<agent>.touch`) via `fs.watch` (near-zero idle cost), with a slow fallback re-check so a **dropped** fs event is never a silent permanent miss (the marker is a HINT, not a queue). Falls back to bounded polling otherwise — no busy-spin.
- **Local-trust auth** (filesystem authority, like `mint-token`/`recover`): reads the **ACTIVE per-instance DB** directly — no token. It resolves the instance DB exactly as the daemon does (`resolveInstanceDbPath`), **never** the legacy `~/.bot-relay/relay.db` (the stdio-legacy-DB-split trap), so the marker and the count read describe the same live instance. A token-authenticated `--remote` path is documented forward-compat, not built (YAGNI).
- **`check-relay.sh` JSON hardening (fast-follow).** The 2 latent register-JSON bugs left in the Claude hook when the same fixes landed in the Codex hook: (1) a hostile `RELAY_TERMINAL_TITLE` (quote/backslash/newline) was raw-interpolated into the `register_agent` payload → malformed JSON → the whole register + mail delivery failed; now validated against the server allowlist and **dropped** if it doesn't match. (2) the capabilities→JSON `awk` used `next` (which skips the whole record, dropping the closing `]`) + an index-based separator; an **empty token** (a double/leading/trailing comma) triggered malformed caps JSON → register failed; now `continue` + a count-based separator → always valid JSON. Both are byte-parity with the Codex hook and proven end-to-end against a real daemon.

## v2.17.1 — 2026-07-21 — WakeSpec reconciliation + transient-send retro quick wins

Fast relay patch. Corrects the 2.17.0 interim WakeSpec placeholders to the Tether extension's **proven** wake behavior (so the data-driven Tether 0.6.0 reads real values), plus three one-line-send DX wins from the transient-send retro.

- **WakeSpec reconciled** (`src/agent-cli-profiles.ts`): codex `wakeText` is now **byte-identical** to the extension's `DEFAULT_CODEX_WAKE_TEXT` (kept in sync by a drift-guard test that reads the extension source), `submitMethod: "sendSequence"` + a new **`submitDelayMs: 150`** match the tuned `codexAdapter`, and claude wakes by typing `inbox` inline (`sendText`, 0 ms). The interim-placeholder marker is **removed** — these are the real values now.
- **`relay send <to> <content>`** — one-line send. Resolves the sender's token (`$RELAY_AGENT_TOKEN` → per-instance vault → `--mint-if-missing`) and POSTs `/api/send-message`; `--from NAME` sets the sender (default `$RELAY_AGENT_NAME`). **Never sends a bad credential** — it sends `from_agent_token` (the impersonation gate still applies), and it refuses **locally** (exit 2, no POST) on every path when no token resolves OR the vault token does not authenticate against the DB (stale/missing/mismatched) — the mismatched credential is never handed to the daemon.
- **`mint-token --json`** verified to emit **only** JSON on stdout (the daemon advisory + init logs already go to stderr); locked with a regression test.
- **`/api/send-message` accepts `message` as an alias for `content`** — one send vocabulary across the MCP `send_message` tool, the agent-team `SendMessage`, and the HTTP endpoint. Exactly one is required; sending **both with different values is rejected** (no silent precedence).

## v2.17.0 — 2026-07-20 — LLM-agnostic parity (P1–P4)

The LLM-agnostic parity arc (`audit-findings/llm-agnostic-parity-scope-brief.md`; P0 = the Codex Tether handshake + cold-start launcher, 2.16.3/2.16.4). Bundled into one minor across the relay phases (P1, P3, P2); the Tether extension change (P4) ships separately as a Tether VSIX minor.

> **Interim data note (honest scope).** The profile registry's `wake` values (`src/agent-cli-profiles.ts`) are **interim placeholders** in 2.17.0 — populated in P3 to freeze the shape. **No shipped relay code consumes them** (Tether 0.5.0 does not read the registry), so this is documented-interim data, not silent-wrong data. They are reconciled to the extension's proven wake behavior (tuned codex `wakeText` + 150 ms submit delay + correct `submitMethod`) in **2.17.1**, and the data-driven Tether **0.6.0** reads only the corrected registry.

**P1 — Codex hook generation.**
- **`relay generate-hooks --codex`** — emits a `~/.codex/config.toml` fragment with a **register-only** SessionStart hook (points at `hooks/codex/codex-session-start.sh`). Reconciled to the current no-poller model: Codex wakes via Tether + `bin/codex-relay`, so there is **no Stop-hook poll loop** (the `codex-stop.sh` poller was removed in 2.16.4). The output documents the cold-start launcher + the MCP-server requirement inline.
- **`--all`** emits both the Claude JSON and Codex TOML, each in its own labeled section. Default (no flag) and `--full` are unchanged — Claude Code, back-compat.
- **Honest hook-model mapping** (no forced symmetry): Claude Code uses SessionStart + PostToolUse + Stop; Codex uses a single register-only SessionStart hook (its wake is Tether-driven, not a hook poll loop). PostToolUse/Stop have no Codex analog by design — documented in `--help`.
- Parser-based tests (`tests/v2-1-cli-tooling.test.ts` 8b–8e): `--codex` and both `--all` payloads are TOML/JSON-parsed and asserted register-only; TOML basic-string escaping covers all control chars (round-trip regression).

**P3 — CLI-profile registry (registry-first, ahead of P2 spawn).**
- **`src/agent-cli-profiles.ts`** — one declarative registry (claude + codex; schema extensible = new CLI is one entry) carrying each CLI's `processPattern`, `hookInstall`, `launch` (consumed by P2 spawn) and `wake` (consumed by P4 Tether). The shared machine-GUID / `host_id` derivation is deliberately **not** per-profile (the federation-safety invariant).
- **`relay cli-profiles [--json]`** — prints the registry (human summary or JSON for cross-boundary reads).
- **Refactors** `generate-hooks` (output byte-identical) and `liveness.ts:DEFAULT_AGENT_PATTERN` to read the registry — every hardcoded `claude|codex` branch in those consumers is gone.
- **Drift guards** (vitest + pre-publish share one TS-AST walk, `scripts/cli-profile-guard.mjs`): reject hardcoded claude/codex *decision logic* in `src/` outside the registry — id-equality on **either** operand, switch/case on a CLI id, or a regex alternation of the ids — with planted regressions for each dangerous form (incl. the three regex bypasses codex's audit found). A bash-mirror guard also asserts the `_vault-helpers.sh` PID-finder pattern's CLI tokens track the registry (kept a mirror, not read on the per-hook hot path).

**P2 — LLM-agnostic spawn (registry-driven).**
- **`spawn_agent` gains a `cli` field** (`SpawnAgentSchema`): `"claude"` (default — every existing call is byte-identical) or `"codex"`, validated against the profile registry at the MCP boundary (an unknown CLI is rejected). The drivers resolve the launch **strategy** from the registry — no hardcoded `claude|codex` branch (the P3 drift guard stays clean).
- **`LaunchSpec` extended (additive):** `strategy: "binary" | "launcher"` + `launcherScript`. `binary` runs the CLI directly (Claude, unchanged); `launcher` runs a repo POSIX launcher that self-registers the relay handshake then execs the CLI. Codex → `bin/codex-relay`, so a **spawned** Codex gets the 2.16.4 cold-start `host_shell_pids` handshake **at launch**.
- **macOS** (`bin/spawn-agent.sh`): one hardened path, now generic — when the driver sets `RELAY_SPAWN_LAUNCHER` it runs that launcher instead of `claude`; otherwise the claude line is **byte-identical** to before. All identity/cwd hardening (allowlist, control-char + symlink-root defense, osascript escaping, vault hydration) is CLI-agnostic and reused; the launcher path gets its own defense-in-depth validation (absolute, no metachars, resolves to an executable **within the repo `bin/`**). There is no CLI literal in the bash — it branches on launcher-presence.
- **Linux**: `exec claude` generalized to the registry launch target; Codex runs `bin/codex-relay` (POSIX) across every emulator, name/cwd POSIX-quote-escaped. **Windows**: a launcher-strategy CLI throws a clear "not supported on Windows (codex-relay is POSIX)" error — Claude spawn is unchanged.
- **Security**: no escaping relaxation anywhere. The launcher path is **fully canonicalized** (the whole symlink chain of the final target, portable pure-bash — a symlink physically inside `bin/` pointing outside it now resolves to its real target and is rejected; the codex P2 audit hole) and must be a **regular executable** under the canonical `bin/`; the canonical path is what runs. New adversarial coverage — codex-launcher path escapes hostile name/cwd (defense-in-depth); the real-bash suite rejects launcher path-escape / metachars / non-executable / traversal / **symlink-out / symlink-cycle** while still enforcing name/cwd hardening on the codex path. Back-compat asserted byte-identical across the env matrix.
- **Tracked follow-ups (documented limitations, non-blocking):** (1) **Windows codex-spawn parity** — `codex-relay` is POSIX, so Windows codex-spawn errors clearly; a PowerShell codex-relay equivalent is the follow-up to close full cross-platform parity. (2) **Codex-spawn kickstart** — a spawned codex is not handed a startup prompt (codex's positional-prompt CLI contract is not guessed on the security path); it registers + Tether-wakes. Both slated post-P4.

## v2.16.4 — 2026-07-20 — Codex cold-start launcher (autowake at pure launch)

Closes the last gap in the Codex autowake story. The P0 handshake (2.16.3) works, but Codex runs the SessionStart hook's `register_agent` at the **first turn**, not at idle launch — so a freshly-summoned Codex has no `host_shell_pids` until you take a turn, and Tether can't PID-bind it until then ("summon → nothing happens until you talk to it").

- **`bin/codex-relay <agent-name>`** — a cold-start launcher that pre-registers the Tether handshake (`host_shell_pids` + `host_id`) **from the shell**, before exec'ing Codex. Because it runs as a child of the launching shell, its ancestry (`relay_pid_chain`) includes the VS Code `Terminal.processId`, so `host_shell_pids` is populated at **pure launch** → Tether binds + wakes immediately, zero manual turn. Uses the SAME shared helpers as the hook (`hooks/_vault-helpers.sh`, no drift). Cause-independent — it sidesteps Codex's hook-lifecycle timing entirely. Generalizes to any summoned Codex (the agent name is the first arg); brand-new agents register auth-free at first launch and the launcher vaults the minted token.
- **The wrapper→hook handoff (no double-register collision, no force).** The launch register is **non-force** — a genuinely-live same-name session correctly rejects it (duplicate-session protection intact). On success the launcher captures the registered `session_id` and exports it as `RELAY_LAUNCH_SESSION`; Codex's `SessionStart` hook then **skips its own register only when that marker equals its row's current `session_id`** (proof *this* launch registered *this* row — never on DB-state alone), and **otherwise registers normally with `host_shell_pids`** (the plain-`codex` / failed-launch fallback). The launcher **unsets any inherited `RELAY_LAUNCH_SESSION` at entry** and re-exports it only on its own successful register, so a stale/leaked marker can't cause a wrong skip.
- **Liveness via the stdio server.** The launcher does **not** send `agent_pid` (the contract requires the exact Codex process, which only exists after `exec`). Codex's stdio MCP server stamps the exact detected Codex process on startup (`src/transport/stdio.ts`, v2.13.0) — the universal capture point.
- **Time-bounded — never stalls the launch.** The pre-register is a bounded **synchronous** health + register (`--connect-timeout 1 --max-time 2`, ~3s worst case), then Codex `exec`s regardless of daemon state: down/hung daemon → prompt `exec`, no marker.
- **Local-trust note.** The marker is the row's `session_id` (exposed by unauthenticated `discover_agents`), so a hostile same-user process could forge a skip. Accepted under the relay's existing local-trust boundary — a same-user attacker already reads the `0600` token vault and can act as any agent — so no cryptographic launch-nonce was added. Documented in `docs/agents/codex-autowake.md`.
- **Tested** (`tests/v2-16-4-codex-coldstart-launch-register.test.ts`, real launcher + hook vs a real daemon): wrapper marker == row `session_id` + `host_shell_pids` contains the launching PID; marker-match → hook skips; marker-mismatch / no-marker / cross-agent marker → hook registers WITH `host_shell_pids`; a real duplicate-live collision + inherited marker → cleared, live row untouched; daemon down/hung → prompt bounded `exec`; `-c` identity override + arg forwarding.
- Doc: `docs/agents/codex-autowake.md` gains a "Cold-start" section (handoff + trust boundary) + the launcher alias. JSON-separator hardening in `bin/codex-relay` + `codex-session-start.sh` (count-based, so a leading invalid capability can't emit invalid JSON).

## v2.16.3 — 2026-07-20 — Tether wakes Codex terminals (P0 LLM-agnostic parity)

Restores the reported break — **"Tether stopped waking Codex."** It was a registration-contract gap, not wake logic. Tether binds a VS Code terminal to an agent by **PID**, host-scoped (`pid-binding.ts` needs the agent's `host_shell_pids` + a `host_id` matching this machine). The Claude hook (`check-relay.sh`) has always sent them; the Codex `SessionStart` hook sent **only** `agent_pid` — its comment (*"Tether is Claude/VSCode-only"*) was frozen from before Tether went LLM-agnostic (v0.4.0). So Tether abstained on Codex and fell back to fragile terminal-name matching, which broke when a workspace move / alias change renamed the terminal.

- **Shared helpers.** `relay_machine_guid()` + `relay_pid_chain()` moved out of `hooks/check-relay.sh` into the shared `hooks/_vault-helpers.sh` (byte-identical behavior for Claude — no inline copy), so the Codex hook can report the **same** handshake.
- **Codex handshake.** `hooks/codex/codex-session-start.sh` now sends `host_shell_pids` + `host_id` + `terminal_title_ref` on register — byte-parity with the Claude hook, so a Codex agent's `host_id` agrees with the extension reader (`extensions/vscode/src/host-identity.ts`). Tether can now PID-bind and wake Codex terminals **token-free** (the extension does the idle waiting, not the model).
- **No poll loop.** This is a **register-only** `SessionStart` change. The token-burning `Stop`-hook keep-alive poller (which re-prompted the model every ~90s and blocked terminal input) stays **removed** — Tether does the waking; Codex spends tokens only when there is actually mail. The now-orphaned `hooks/codex/codex-stop.sh` script is deleted (zero references remained; git history retains it).
- **Hardened `terminal_title_ref`.** The Codex hook validates the title against the server's `TERMINAL_TITLE_REF_PATTERN` allowlist (`[A-Za-z0-9_.- ]`) and drops it if it fails, so a hostile title (quote / backslash / newline / JSON fragment) can never malform the register payload — or be rejected server-side — and take the handshake down with it. The `SessionStart` context also no longer references a Stop-hook waker (it names Tether).
- **Test.** `tests/v2-16-3-codex-tether-handshake.test.ts` drives the shipped Codex hook against a real daemon and asserts the register payload lands `host_shell_pids` + `host_id`, and that the Codex `host_id` is byte-identical to the Claude hook's (shared-helper, no drift).
- **Doc.** `docs/agents/codex-autowake.md` rewritten to the Tether-wake model (no poll loop).
- Ships as a hooks-only relay patch — no Tether VSIX republish (hooks travel with the npm package). Scope is P0 only; broader LLM-agnostic parity (spawn, hook generation, profile registry) is P1–P4 in `audit-findings/llm-agnostic-parity-scope-brief.md`.

## v2.16.2 — 2026-07-13 — Publish-gate fix for Node 24 / npm 11 (carries 2.16.1)

A release-tooling patch: **2.16.1 never reached npm** because the pre-publish gate failed on Node 24 / npm 11, so 2.16.2 is what ships and it **carries all of 2.16.1** (stable mint-once-reuse, below) plus this fix.

- **The block.** The extension's VSIX-contents drift guard (`v0-1-4-vsix-contents.test.ts`) shells out to `vsce ls` / `vsce package`, which internally run `npm list --production --parseable --depth=99999`. Under npm 11 that scan **exits 1** on a false-positive `ELSPROBLEMS` (the `qs`/`form-data` `overrides` mark `call-bind-apply-helpers` / `get-intrinsic` "invalid" though the installed versions satisfy their ranges). npm 20/22 accept it, so CI stayed green while the gate — which Maxime runs on Node 24 — went red on all 11 assertions.
- **The fix.** Both vsce invocations now pass `--no-dependencies`, skipping the dependency scan. The extension is esbuild-bundled, so runtime deps never ship in the VSIX — the packaged file list and byte ceiling the guard asserts are unchanged; only the spurious scan is gone. Verified passing on Node 24 / npm 11.
- **Known follow-up (flagged, not in this patch):** the *relay* pre-publish gate runs this *extension* VSIX test, so a Tether-tooling failure can block a relay npm publish. Decoupling the two gates (or making the shared gate Node-version-robust) is a tracked follow-up.

## v2.16.1 — 2026-07-13 — Stable agent tokens (the relay half of the durable autowake fix)

The recurring "autowake stops working after a relaunch" bug had two halves. The relay half: a launcher running `relay mint-token --force` on **every** relaunch rotated the agent's token, invalidating any holder of the old one. `relay mint-token` now defaults to **stable mint-once-reuse**:

- **No agent yet** → mint a token **and write it to the vault** (closing a long-standing gap where the CLI minted a token but never wrote the vault the SessionStart hook reads).
- **Agent exists and its vault token still authenticates** → **reuse it** — the `token_hash` is byte-stable, no rotation, no churn.
- **Agent exists but its vault token can't authenticate** (missing / stale / mismatched) → **refuse**, with guidance (`--force` to rotate deliberately, or `relay recover`). It never silently rotates — that would invalidate a possibly-live token and mask a stale/compromised credential.

`--force` still rotates on demand for the genuine "I want a new token" case (and now also writes the vault). Paired with the Tether 0.5.0 vault-first token read, this ends the manual "Set Agent Token" babysitting for good.

## v2.16.0 — 2026-07-10 — One-command install (`relay init`), the adoption gate

Getting a working relay used to take ~6 disjoint manual steps across three files plus a hand-authored launchd plist. `relay init` is now the single idempotent macOS install path: a stranger runs one command and gets a working relay + autowake loop, and re-running is always safe.

### What `relay init` now does (all reconcile-not-clobber, safe to re-run)

1. **`~/.bot-relay/config.json`** — reconciles: PRESERVES an existing `http_secret`, `instance_id`, and any operator edits; adds only missing defaults. Records a `default_agent_name` (`--agent NAME`) the SessionStart hook falls back to.
2. **`~/.claude.json`** — deep-merges the `bot-relay` stdio `mcpServers` entry (absolute path), preserving other servers.
3. **`~/.claude/settings.json`** — deep-merges the SessionStart hook, deduped by command path, preserving unrelated hooks.
4. **macOS launchd** — installs + bootstraps a `KeepAlive` daemon plist, but SKIPS if `:3777` is already served by any relay (collision-safe, label-agnostic — never double-loads an existing supervisor).

Opt-outs: `--config-only`, `--skip-hooks`, `--skip-daemon`, `--skip-mcp`. Deep-merges are structural + atomic (tmp+rename+`.bak`); a second run is a strict no-op.

### Token-safety — the load-bearing invariant

`relay init` is **token-blind by construction**: it imports no token/db module and never mints, rotates, registers, recovers, or writes/deletes a token or touches the agents token-hash column / the vault. So init / deploy / bounce can never desync a live agent's credential. Agent identity is established by the already-token-safe SessionStart hook on first launch (vault-first read; register captures the minted token → writes the vault). Guarded by a source+compiled token-blind scan and a regression that seeds a matching vault, runs init twice, and asserts both the DB hash and the plaintext vault stay byte-stable + authenticating (a negative control proves the assertion catches a rotate).

### Local installs are secret-free — a deliberate local-trust default

A default local install OMITS `http_secret`: the daemon treats a `127.0.0.1` bind as local-only (a transport secret would 401 the SessionStart hook's register and break the loop). **Be clear-eyed about what this means: the default trusts every process on the same machine.** With no secret, any local process can reach the loopback HTTP surface, including the dashboard operator endpoints (`/api/snapshot` read, `set-status`, `wake-agent`, `kill-agent`, `operator-identity`). Per-agent tokens + gate-11 from-verification protect only agent **identity** — you still can't send or act *as* another agent without its token — but they do **not** gate the whole surface.

This is a reasonable boundary for a single-user machine: a local process already has OS-level access to the same data (PIDs via `ps`, the DB file). It is enforced conservatively — non-loopback binds are still **refused** without a secret, and the loopback-literal Host check still blocks DNS-rebind / malicious-webpage access on every route.

**On a shared machine, or any non-loopback / team / remote setup, set a secret:** `relay init --secret <strong-random>` (or export `RELAY_DASHBOARD_SECRET`) to gate the surface. The daemon *requires* a secret to bind to any non-loopback host.

macOS-first: Linux/Windows daemon supervision is "coming" (init prints manual guidance there, not gated). D2: the config `default_agent_name` never overrides an explicit `RELAY_AGENT_NAME` or spawn manifest.

## v2.15.2 — 2026-07-07 — Signal teardown stops re-polluting terminal states

v2.15.1 was a one-time cleanup of stale stored terminal states. This closes the source that *re-creates* them: the stdio signal handler. A `SIGHUP`/`SIGINT`/`SIGTERM` is delivered to the agent's **MCP-server process**, not to the agent itself (the agent — `claude`/`codex`, tracked by `agent_pid` — can survive a terminal reflow / editor reload and relaunch its MCP server). The pre-v2.15.2 handler stamped a terminal `agent_status` on that signal, which stuck on two surfaces and phantom-closed a surviving agent:

- **`getAgents`** — `deriveAgentStatus` R1 makes a stored `offline` win even over a confirmed-alive probe, so a signal-stamped `offline` never self-heals.
- **dashboard** — `deriveDashboardState` returned `closed` from the signal stamp alone, *and* the stamp was never cleared on re-register, so a fresh session with no probe-able anchor (a valid case: no agent ancestor matched / cross-host) read `closed` from the stale stamp.

### The fix (liveness governs; a signal is context, not a verdict)

1. **New signal-only teardown** (`endAgentSessionOnSignal`) replaces `closeAgentSession`/`markAgentOffline` on the stdio signal path. It **clears the liveness anchor** (`agent_pid`/`agent_pid_start`/`last_alive`), writes **no sticky terminal status** (`agent_status = 'idle'` — a neutral, derivable value: with the anchor cleared `deriveAgentStatus` returns `unknown`, and if the agent comes back it reads active), stamps `signal_received_at`/`signal_kind` for dashboard forensics only, and releases the session. There is **no `markAgentOffline` fallback** — on failure it logs + audits and writes no terminal status. `markAgentOffline` is untouched for spawn's deliberate offline-pre-register.
2. **Dashboard liveness gate** — `deriveDashboardState` now suppresses the signal-derived `closed` when the liveness probe is `alive`. An explicit `unregister` (a deliberate delete) is **not** gated and still wins.
3. **Session-scoped stamp** — `registerAgent` clears `signal_received_at`/`signal_kind` in the same UPDATE that rotates `session_id`, so a signal stamp can never outlive its session. Long-term forensics live in `audit_log` (`stdio.session_ended_on_signal`).

A genuine operator `set_status('offline')` is unaffected — that declaration path (R1) still wins over a live probe.

### Also

- **HTTP `/health` now reports `uptime_seconds`** (monotonic, `process.uptime()`-based, not wall-clock), so a follow-on Tether can detect a silent daemon restart under a still-open connection. The detector must compare with a strict `<` (uptime is non-decreasing within one process).
- **CI now runs the VSCode extension unit suite** (`extensions/vscode` `npm run test:unit`) in GitHub Actions — previously it ran only locally, which is how the v0.4.1 health-poll status-check gap slipped past CI.

## v2.15.1 — 2026-07-06 — Clear stale stored terminal states

v2.15.0 honors a stored `offline`/`closed` as a declaration (rules R1/R4), which is correct going forward — but agents carrying a *leftover* terminal state from before the presence fix (the stdio signal handlers, force-rotation, and the spawn offline-pre-register all STORE `offline`/`closed`, sometimes on a transient signal the agent survived) read that stale value instead of their live state, and an `offline`-polluted live agent never self-corrects. This is a one-time cleanup so the live verdict governs again.

### The migration

A **version-guarded (18→19), data-only, run-exactly-once, idempotent** migration (`migrateSchemaToV2_19`) clears the stale terminal states:

```
UPDATE agents SET agent_status = 'idle'
WHERE agent_status IN ('closed','abandoned','stale')
   OR (agent_status = 'offline' AND session_id IS NULL);
```

**Narrowed contract (deliberate + documented).** `set_status` can only produce `offline` (its enum excludes `closed`/`abandoned`/`stale`), so those three are *always* cleared. For `offline`: a **session-present** row is a genuine `set_status` declaration and is **preserved** (R1 keeps honoring it); a **session-NULL** `offline` has ambiguous provenance — it could be signal/rotation pollution *or* a genuine-but-sessionless `set_status`, so it is **intentionally cleared to `idle`** and the live presence verdict governs (dead→closed, alive→idle, no-anchor→unknown). This can never produce a false-alive (`idle` only ever comes from a real positive probe, never from clearing the column); any genuine offline intent is trivially re-assertable via `set_status`; and because the migration runs exactly once (guarded on the stored `schema_version < 19`), all future operator/dashboard offline-sets — on any row, sessionless or not — persist normally. It is **not** a read-path mutation (init/deploy path only).

### Tests

`tests/v2-15-1-stale-status-cleanup.test.ts` — the five contract regressions: (a) a session-present `set_status('offline')` is preserved; (b) a session-NULL `offline` (signal-derived *or* genuine-but-sessionless) is intentionally wiped → verdict-governed, asserted as designed; (c) `closed`/`abandoned`/`stale` are always cleared; (d) runs exactly once — a post-migration sessionless `offline` persists (not re-wiped) — and is idempotent; (e) no false-alive — a signal-marked row reads its real probe (dead→closed, alive→idle). Schema version 18→19; no new column, no new tool.

## v2.15.0 — 2026-07-06 — Presence you can trust: the `unknown` state

The relay could mislabel a live-but-quiet agent as `closed`/`offline` and callers would relaunch it (rotating its token). Root cause: `deriveAgentStatus` guessed death from `last_seen` age whenever it had no liveness data, and "no data" was indistinguishable from "confirmed dead". This rebuild makes the misread impossible: **staleness alone can never produce a terminal state again.**

### The model

- **Three-way liveness verdict** (`alive` | `dead` | `unknown`), computed by a live same-host process probe: `unknown` = no probe-able anchor (`agent_pid` absent or cross-host) or; `dead` = a POSITIVE dead signal (`agent_pid` present + process confirmed gone / PID reused); `alive` = process confirmed up.
- **`deriveAgentStatus` is now age-free** — a pure function of `(stored_status, verdict)`. Precedence: an explicit `offline` declaration wins (R1); a live process is authoritative even with days-stale `last_seen` — the rate-limited-but-alive case (R2); a dead probe → `closed` (R3); a clean-SIGINT-close declaration → `closed` (R4); otherwise **`unknown`** (R5). Age is not an input.
- **New `unknown` agent_status + `liveness` field.** `agent_status` gains `unknown`; a new `liveness: alive|dead|unknown` is the **field of record** for any dead-vs-unknown decision (close/relaunch/purge). The back-compat `alive` boolean stays, but `alive===false` means "not confirmed alive", NOT "dead" — a drift-guard test fails if any in-repo runtime path reads the lossy `.alive` bool to drive a terminal decision. The dashboard state machine + HTML, `get_standup`, and `health_check` (new `agent_count_unknown`) all surface `unknown` distinctly from `closed`.

### report_liveness (new tool) + self-heal

- **`report_liveness`** (tool #35) — a narrow, metadata-only presence self-report that restamps ONLY the liveness anchor (`agent_pid` + start-time, and fills `host_id` when NULL) via the sanctioned `setAgentLivenessAnchor`. It does NOT rotate `session_id`, bump `last_seen`, or touch the read cursor — so an old/existing session that predates the anchor mechanism becomes probe-able WITHOUT a re-register (which would re-surface already-read mail).
- **PostToolUse self-heal.** The PostToolUse hook computes the agent's PID + start-time and calls `report_liveness` when the stored anchor doesn't match the current process (PID changed, or a readable start-time differs/fills) — gated on a real mismatch so it's a one-shot with zero steady-state churn, and it never downgrades a present start to null when the current one is transiently unreadable. `relay_agent_pid`/`relay_pid_start` moved to `_vault-helpers.sh`, shared by all three hooks.

### Reads are pure

- The positive liveness cache moved from the `last_alive` DB column to in-memory (mirroring the negative cache), so `getAgents`/`health_check`/`get_standup` compute the verdict from a live probe + in-memory TTL caches and write **nothing** to the DB on read. `deriveAgentStatus`'s narrow-dead rule matches `isAgentProcessAlive`: kill-ESRCH or start-time mismatch → dead; kill-ok + start-time missing/unreadable → alive (never dead).

### Tests

`tests/v2-15-0-presence-unknown.test.ts` (NEW): the six named regression cells end-to-end (rate-limited→idle, pre-anchor-live→unknown, crash→closed, clean-close→closed, declared-offline→offline, start-missing→alive), the `report_liveness` self-heal invariants (session/last_seen/message-plane untouched, stale-start restamp flips dead→alive, host_id COALESCE no-churn), read-path purity (rows byte-identical after repeated reads), `agent_count_unknown`, and the `.alive`-drift guard. The full `deriveAgentStatus` stored×verdict matrix is pinned in `tests/v2-13-0-presence-liveness.test.ts`. No schema change (derivation-only; `agent_pid` already existed).

## v2.14.1 — 2026-07-01 — Every agent captures its liveness anchor on launch

v2.13.0's presence probe was **live but inert** in production: `agent_pid` (the process the probe checks) was only populated by the stdio MCP server's ancestry walk, but live agents register via the SessionStart hook over HTTP — which sent `host_shell_pids` but **not** `agent_pid` — and the HTTP daemon can't self-detect (its own ancestry is launchd). So `agent_pid` stayed NULL and presence fell back to the age-based chain. This makes the probe real for every agent.

### Fixed — the SessionStart hooks now capture `agent_pid`

- Both hooks (`check-relay.sh` for Claude, `codex/codex-session-start.sh` for Codex) now compute the agent's own process id + start-time and send them in `register_agent` (which already persists them). So every hook-registered agent — summoned, spawned, or relaunched — captures its liveness anchor on first launch, HTTP or stdio, zero-token.
- The agent process is identified by its **`comm` (executable basename), not its argv/path** — critical, because the relay can live under a directory named `Claude` (e.g. `.../Claude AI/bot-relay-mcp`), so an argv/path match would false-hit any process launched from there (including the hook itself). Node/bun/deno-hosted CLIs are matched via the runtime `comm`; the relay's own node is excluded by its entrypoint. Extensible via `RELAY_AGENT_PROCESS_PATTERN`. Unidentified → omitted → age-based fallback (graceful).
- `start-time` is read with `LC_ALL=C` in **both** the hook and the relay's probe (`src/liveness.ts`), so the captured token matches the daemon's probe-time read byte-for-byte across the user-shell/launchd locale boundary (a drift would false-read a live agent dead).
- The **stdio-server matcher (`src/liveness.ts`) is brought to the same identity discipline**: `findAgentProcess` now matches process **identity** — the executable basename (via `comm`, else argv[0]), or for a runtime-hosted CLI the hosted script's basename — never the raw `ps command=` string. The old full-command regex would stamp any non-agent ancestor whose *path* merely contained `claude`/`codex` (e.g. a wrapper under `.../Claude AI/...`) as the agent, reading a dead agent alive while that ancestor lived. `DEFAULT_AGENT_PATTERN` / `RELAY_AGENT_PROCESS_PATTERN` are now exact-basename matchers, and the self-exclude is narrowed to the entrypoint only (an agent launched from a `bot-relay` checkout is no longer wrongly excluded).

### Fixed — spawned agents capture their PIDs

- `spawn_agent` pre-registered the child parent-side as a live row, so the child's own SessionStart hook re-register would trip the name-collision guard (`NAME_COLLISION_ACTIVE`) — the child could never fill its `host_shell_pids`/`agent_pid`. The parent now marks the pre-registered row **offline** (after the driver launches, so failure rollbacks still work), so the child re-registers freely, the presence derivation resets `offline → idle`, and it captures its PIDs. The token stays valid and mail still delivers to the offline row.
- `check-relay.sh`'s `SKIP_REGISTER` liveness gate now treats a row as skippable only when it is live **and already carries `host_shell_pids`** — a freshly-provisioned/spawned row with empty PIDs re-registers on its first hook run to fill them.

### Tests

`tests/v2-14-1-spawned-agent-pid.test.ts` (NEW) runs the shipped hook against a real daemon: an offline (spawn-preregister-state) row → the hook registers and fills `agent_pid` + `host_shell_pids` (a live ancestor process); a populated-live row → the hook skips; and the hook-captured start-time equals the relay's probe (`LC_ALL=C` parity). `tests/spawn.test.ts` pins the offline pre-register (session cleared, mail still delivered), and the driver-failure rollback tests confirm the offline transition doesn't break the session-scoped rollback. No schema change; no new tool.

## v2.14.0 — 2026-06-30 — Reserved-name / impersonation protection

The local daemon is intentionally auth-free for zero-config onboarding. An agent that already holds a token is protected from impersonation by the token↔identity binding (v2.12.0: a `from`/actor field authenticates only against the caller's own token — so you cannot `send_message(from=someone-else)` without their token). The remaining hole was the **register side**: `register_agent`'s auth-free bootstrap path let any caller **claim** a persona/sentinel name that had no live row and mint its token — becoming that identity. This closes it.

### Added — reserved names

- A **reserved-name set** that cannot be self-registered via the bootstrap path. It is the relay's own sentinels (hardcoded: `system`, the sender sentinel `send_message` uses for relay-authored messages) **plus** an operator-configured list via the **`RELAY_RESERVED_NAMES`** env var (comma-separated). Matching is case-insensitive. Private persona names live only in the operator's env — never in the codebase.
- `register_agent` now **rejects** a bootstrap (no existing row) attempt to claim a reserved name, with a clear error pointing at the provisioning path. A reserved name must be **provisioned by an operator** first via `relay mint-token <name>` (filesystem-authorized, bypasses the dispatcher); after that, the normal token requirement on re-register / send-`from` protects it.
- Reserved names are **exempt from the 30-day dead-agent purge** — their row + token hash is their protection, so they never age out and reopen the claim window.

### Changed — legacy grace can't authenticate an actor

- Under `RELAY_ALLOW_LEGACY=1`, the legacy-bootstrap path authenticates a token-less pre-v1.7 row without proving identity. It is now **rejected on actor-stamping tools** (`send_message`, `broadcast`, `post_*`, etc.) — grace remains only for the no-actor bootstrap path, so a legacy `from` can never impersonate. A legacy agent mints a token (re-register / `relay mint-token`) before using actor tools.

### Notes

- This is **Part 1** (the security fast-track) of the transient-identity governance. The transient onboarding model (auto-named `tmp-*` agents that may send *to* a persona but never *as* one) and the `rename`/`promote` tool are the queued relay feature — a documented follow-on, not in this release.
- No schema change; no new MCP tool.

### Tests

`tests/v2-14-0-reserved-name-protection.test.ts` (NEW): reserved-set resolution (hardcoded `system` + env, case-insensitive); bootstrap-registering `system` or an env-reserved name → rejected (and no row recorded) while an ordinary name still self-registers; the operator provisioning path (`mint-token` → the persona acts with its token); send-`from` requires the actor's token (pinning the v2.12.0 protection + the provisioned-persona case); the legacy-grace actor rejection; and the purge exemption (a stale reserved name survives, a stale ordinary agent is purged).

## v2.13.0 — 2026-06-30 — Presence liveness (alive-and-idle vs. closed)

The relay couldn't reliably tell an **alive-but-idle** agent from a **closed** one. `last_seen` only advances on activity (observation isn't liveness, v1.3), so an agent sitting open and waiting aged out to `stale → offline → closed`/`abandoned` — and orchestrators misread live idle agents as dead. This adds a **positive** liveness signal so "is this agent awake and waiting?" has a trustworthy answer.

### Added — same-host agent-process liveness probe + `last_alive`

- New additive columns **`agents.last_alive`** (ISO timestamp of the most recent positive liveness confirmation; distinct from `last_seen` activity), **`agents.agent_pid`** (the agent's OWN process id), and **`agents.agent_pid_start`** (a start-time token guarding PID reuse).
- On a presence read (`discover_agents`, `get_standup`, `health_check`), the relay lazily probes **the agent's own process** with a `process.kill(pid, 0)` signal-0 check **(same-host only)**. If that process is alive, the agent is open — even while idle — and `last_alive` is stamped. Crucially this probes the **agent process**, not the `host_shell_pids` ancestry chain: the chain's shell/terminal ancestors outlive the agent, so probing them would keep a dead agent "alive" while its terminal stays open. `host_shell_pids` is left untouched for its terminal-binding purpose.
- The probe is host-scoped by the OS machine GUID (`host_id`) — a PID is only probed when the agent shares the relay's host, so it can never false-match an unrelated process elsewhere. `EPERM` (cross-user) counts as alive; `ESRCH` as dead. A recycled PID (different `agent_pid_start`) reads dead. Two caches (positive `last_alive` + an in-memory negative cache) suppress re-probing on rapid successive reads.
- **The capture point is universal**, not Claude-specific: the relay's **stdio MCP server** — which every stdio agent spawns as a child — walks its own ancestry once at startup (a single `ps`, zero-token, no loop) to identify the agent CLI. Identification matches **claude + codex** out of the box and is **extensible via `RELAY_AGENT_PROCESS_PATTERN`** for other CLIs; an unrecognized binary falls back to age-based presence (safe). Managed/script agents self-report `agent_pid` via `register_agent`.
- New module computing this host's machine GUID with the exact OS source the SessionStart hook + extension use (macOS `IOPlatformUUID` / Linux `/etc/machine-id` / Windows `MachineGuid`), so host comparison is byte-identical.

### Changed — both presence derivations honor liveness

- `deriveAgentStatus` (the `discover_agents`/`health_check` surface) and `deriveDashboardState` (the dashboard) now take the liveness signal: a **fresh** `last_alive` overrides the age-based `offline`/`abandoned` promotions and a **stale stored `closed`** from a prior session, so an alive-and-idle agent reads `idle`/`waiting` instead of `closed`. A genuine current-session teardown wins: `closeAgentSession`/`markAgentOffline` clear `last_alive` + `agent_pid` in the same atomic update, so the probe has nothing to restamp and the close is never masked.
- Presence derivation follows a **canonical precedence table** (single source of truth, exercised cell-by-cell in tests): (1) an explicit current-session `offline` declaration (`set_status('offline')`, force token rotation) wins over liveness — a deliberate "unavailable" isn't un-declared by a live process; (2) a fresh liveness signal surfaces the active state, overriding age-derived `stale`/`offline`/`abandoned` and a stale prior-session `closed`; (3) with no liveness, the unchanged age + stored-terminal chain. The `alive` boolean requires a fresh confirmation **and** an active surfaced status, so it stays consistent with `agent_status`.
- **Resume is clean:** an existing-row re-registration resets a TERMINAL lifecycle state (`offline`/`closed`/`abandoned`/`stale`) to `idle` — a relaunched agent (or the next valid registration after a force token rotation) comes back **available**, not stuck offline. Active declared states (`working`/`blocked`/`waiting_user`) are preserved across the session rotation as current intent. This reset is what lets a stored `offline` be read unambiguously as a current-session declaration.
- New query surface: `discover_agents` + `get_standup` agents carry **`last_alive`** (ISO) and **`alive`** (boolean — the trustworthy "awake right now?"); `health_check` adds **`agent_count_alive`**.
- Freshness window tunable via `RELAY_AGENT_ALIVE_WINDOW_SEC` (default 120s).

### Schema

- Migrates **v17 → v18** — three additive `NULL`-default columns (`last_alive`, `agent_pid`, `agent_pid_start`). Zero data migration: with no liveness signal an agent derives exactly as before (pure age-based), so behavior is byte-identical until a probe populates it.

### Notes

- This ships the relay-side foundation. A Tether-extension liveness **heartbeat** (cross-host + crash-without-clean-exit coverage, with its own auth model) is a planned follow-on; the relay-side probe covers same-host fleets today.

### Tests

`tests/v2-13-0-presence-liveness.test.ts` (NEW): the headline regression (idle agent with a live process reads `alive`/`idle`, not `closed`/`offline`, even with a stale stored `closed`); a current-session close clears the anchor so liveness can't mask it; we probe `agent_pid` not the chain (a dead agent with a live shell ancestor reads dead); the negative cache (dead rows aren't re-probed in-window); **the agnostic proof** — a non-Claude (codex) agent reads alive-when-idle, and the ancestry walk identifies codex as readily as claude; the PID-reuse start-time guard; cross-host fallback; back-compat (no signal → byte-identical age-based); `process.kill` errno handling; and the machine-GUID parsers.

## v2.12.0 — 2026-06-29 — Pending vs. history (a permanent "resolved" plane)

A fresh agent session calling `get_messages(status="pending")` could re-surface messages it had already handled in a **prior** session, as if newly actionable. This is a direct consequence of the intentional **session-scoped** read model (a new terminal re-sees previously-read mail so a handover never drops *unfinished* work, v2.0 final #6): a new `session_id` makes every prior-session-read message look pending again — correct for unfinished work, wrong for already-handled items.

This release adds a session-**independent** "resolved" plane on top of the existing session-scoped read plane. "Resolved" ≠ "read": read is a per-session observation; resolved is a permanent "handled, archive it." The change is purely **additive** — zero behavior change until an agent opts into resolving.

### Added — `resolve_messages` (tool #34) + `ack` on `get_messages`

- New tool **`resolve_messages(agent_name, message_ids[])`** permanently resolves specific messages (partial handling — "I did these, not those"). Recipient-scoped: the dispatcher binds the caller's token to `agent_name`, and the DB layer additionally scopes the `UPDATE` by `to_agent`, so an agent can only resolve its **own** mail. Idempotent — re-resolving, unknown ids, or ids addressed to another agent are no-ops (reflected in the returned counts). Own-mailbox primitive; no extra capability required.
- **`get_messages` gains `ack` (default `false`).** `get_messages(status="pending", ack=true)` drains **and** resolves the returned set in the **same transaction** as the per-session read-mark, so the next poll — even from a fresh session — won't re-flood. Only the pending drain path resolves; browsing history never does. The drain transaction is promoted to `BEGIN IMMEDIATE` so concurrent acking drains serialize and neither double-counts nor drops.

### Changed — the `pending` filter honors resolution

- `status="pending"` now also requires `resolved_at IS NULL`. This single clause keeps already-handled mail out of the cross-session action queue **permanently**, while genuinely unfinished work still re-surfaces across sessions (handover safety preserved). The same clause is mirrored in `get_messages_summary` so the cheap preview agrees with the mutating drain.
- New `status` values: **`history`** (alias of `all` — the full durable record) and **`resolved`** (only acked mail). `read`, `all`, `lane`, `since`, and sequence assignment are unchanged.

### Schema

- Migrates **v16 → v17** — one additive `NULL`-default column `messages.resolved_at`. Existing rows stay `NULL` (unresolved) → every currently-pending message stays pending; zero data migration. No backfill: the live re-flood self-heals once agents adopt `ack=true`, and stale mail ages out via the existing 7-day purge.

### Tests

`tests/v2-12-0-pending-vs-history.test.ts` (NEW) — 7 contract groups: the cross-session re-flood regression (ack-drain in S1 ⇒ S2 pending empty; **without** ack S2 still re-sees them, proving the fix is the resolve and not a change to read semantics); resolved-absent-from-pending / present-in-history; `ack` resolves exactly the returned set (read-mark + resolve move together); `resolve_messages` id + recipient scoping; **adversarial authz** (one agent's token cannot resolve another's mail — verified through the real dispatcher); back-compat byte-identical response when `ack=false`; and idempotent, no-double-count resolution. Tool count moves 33 → 34.

## v2.11.1 — 2026-06-25 — Republish from cleaned sources

No functional or API change. Genericized examples and removed non-public references from docs, changelog, code comments, and tests; `dist` rebuilt from the cleaned sources. Protocol unchanged.

## v2.11.0 — 2026-06-19 — Tether PID-handshake: refresh on relaunch (GAP 1)

Completes the Tether v0.3 PID-handshake so it works for **existing** agents, not just freshly-named agents. v0.3.2 made a token-free agent fetch its binding and autowake — but only an agent whose row was *inserted* with a live PID chain. A long-lived agent (a pre-existing row, e.g. `build-agent`) re-launched with an **empty `host_shell_pids`/`host_id`**, so Tether had nothing to bind to → no autowake. This release closes that gap end-to-end.

### Fixed — `host_id` now refreshes on an authenticated re-register

`register_agent` previously treated `host_id` (the OS machine GUID) as **immutable** (INSERT-only), so a row created without a handshake could never have its `host_id` populated on relaunch. It is now **session-refreshable** with the same semantics as `host_shell_pids`: provided → overwrite, omitted → preserve. (`host_shell_pids`, `session_id`, and `terminal_title_ref` already refreshed on re-register since v0.3.0.)

- **Security unchanged + still load-bearing.** The re-register path is gated by `enforceAuth`: an `active` row can only be re-registered by the **token-holder**. A wrong/missing-token re-register of an existing name is rejected and writes neither PIDs nor `host_id` (regression-tested). A `host_id` refresh is therefore always the *owner* declaring its current machine — never a cross-agent overwrite.
- **No schema change** — `host_id` column already existed (schema stays **v16**); `protocol_version` stays `2.4.0`.

### Fixed — SessionStart hook re-registers on relaunch (not just first launch)

`hooks/check-relay.sh` previously skipped `register_agent` entirely whenever a token was present **and** the agent row already existed (the Phase 4j spawn-handoff optimization). For a relaunched agent that meant its live PID chain was **never re-sent**, so the handshake fields stayed stale/empty. The skip is now **liveness-scoped**: it fires only when the existing row holds a **fresh, live session** (true spawn handoff / concurrent-terminal guard). When the row is **offline or stale** (a genuine relaunch), the hook re-registers, refreshing `host_shell_pids` + `host_id` and repopulating `session_id`. (This also repopulates an empty `session_id` that could glitch inbox reads.)

- **Known follow-up:** a *spawned* agent still doesn't capture PIDs on its first launch (the parent pre-registers a live row without the child's PIDs, and the child hook then sees a live session → skips). Tracked as a candidate GAP 2; not required for "an existing builder wakes."

### Tests

`tests/pid-handshake.test.ts`: the "host_id is IMMUTABLE" contract is replaced by "host_id refreshes when re-reported, preserves when omitted"; new coverage for populating an initially-empty `host_id` (the `build-agent` case), `session_id` rotation on re-register, and an end-to-end **authenticated** re-register refreshing both `host_shell_pids` and `host_id` via the owner's token. The wrong-token-rejected governance test is retained unchanged.

`tests/v2-11-0-hook-liveness-register.test.ts` (NEW): load-bearing coverage of the **shipped `check-relay.sh`** liveness gate — invokes the real hook as a subprocess against a real HTTP daemon and asserts register-was/wasn't-called via deterministic signals (session_id rotation + host_shell_pids overwrite, plus a host_id refresh assertion where a machine GUID is derivable). L1 (fresh+live → skip), L2 (stale → re-register), L3 (offline/`session_id` NULL → re-register, the build-agent case). Verified load-bearing by negative control: reverting `SKIP_REGISTER` to its old unconditional-skip form fails L2 + L3.

## v2.10.0 — 2026-06-15 — Coordination + safety

### Added — capability-routed messaging (`post_to_capability`)

New MCP tool **`post_to_capability`** (tool #31): a sender tags an FYI/coordination message by a single domain/capability and the relay fans it out to the CURRENT owner(s) of that capability via the normal inbox — no manual channel join, no named recipient. Use case: an ad-hoc agent tags a finding by domain → the owning agent (e.g. an agent that owns the "support" capability) picks it up automatically on its next `get_messages`. This implements principle #1 — capability routing over named recipients.

- **Routing** reuses the proven `agent_capabilities` index + exact-string match (same contract as `post_task_auto`), but fans out to ALL current owners (an FYI should reach every agent that owns the domain, not just the least-loaded one) and delivers into the existing `messages` inbox — recipients drain via the normal `get_messages`, so there is **zero new read path** (Tether and the SessionStart hook work unchanged).
- **Action-vs-FYI line stays intact (machine-enforceable).** Action-required completions remain point-to-point completion reports (`send_message`); capability-routed topics are the FYI/coordination lane only. A new nullable `messages.routed_capability` column + a `get_messages` `lane` filter (`all` | `direct` | `capability`) keep the two lanes distinguishable, so an action item is never lost in FYI noise.
- **No current owner** → `routed_to: []` and nothing is stored (fire-and-forget to current owners; NOT queued-until-owner, which is task semantics).
- New webhook event **`message.capability_routed`** (fired once per fan-out, carrying the capability + recipients).
- Schema migrates **v13 → v14** — one additive `NULL`-default column; zero data migration.

### Added — schema-gated task completion (`register_task_schema` / `task_schema_get`)

Server-enforced JSON Schema validation of a task's completion result, gating the `accepted → completed` transition. Kills the root cause of the 2026-06-09 false-completion incidents — an agent marking a task complete with NO proof — at the protocol layer, for **any** agent (not just Claude hooks).

- A requester attaches a schema to a task via `post_task`'s new optional `schema_id`. On `update_task(action='complete', result=…)` the relay parses `result` as JSON and validates it against the registered schema **before** the CAS UPDATE. Non-conforming → **rejected** (`RESULT_SCHEMA_VIOLATION`), the task stays `accepted`, no `task.completed` webhook. **Opt-in + backward-compatible:** a NULL `schema_id` completes exactly as today.
- **Validator hardening (ajv):** strict mode; **no remote `$ref` / no `$data`** (no fetch/SSRF surface); the schema document is meta-validated + ref-scanned **before** ajv compiles it (a registered schema is code-generated from, so it is an attack surface); compiled validators are cached per id.
- **Rollout-safe:** `RELAY_SCHEMA_GATING = warn (default — shadow: log + allow) | enforce | off`. Ships in **warn** so a brand-new gate can't wrongly reject legit completions; flip to `enforce` after the dual-audit window.
- New tools **`register_task_schema`** (immutable; authz-restricted to the `manage_schemas` capability — a registered schema is compiled) + **`task_schema_get`** (open read). Built-in schemas `ship_pong_v1` / `audit_verdict_v1` / `merge_ready_v1` auto-registered on init.
- Schema migrates **v14 → v15** — new `task_schemas` table + additive nullable `tasks.schema_id`; zero data migration.
- Follow-up (deferred to keep v1 focused): `post_task_auto` `schema_id` threading (v1 gates via `post_task`).

## v2.9.1 — 2026-06-10 — HOTFIX: autowake recursion (Tether 0.2.0 + HTTP daemon crash)

Fixes the deployment-blocking RangeError that surfaced when Tether 0.2.0 connected to the v2.9.0 HTTP daemon. The bug class is the same shape as v2.5-R1 (InMemoryTransport-style tests passing while real HTTP transport breaks); the v2.9.0 autowake test was real but exercised the wrong close path, so the recursion went undetected through v2.5–v2.9.0 and only triggered on a live VS Code + Tether deployment.

### The bug

`src/transport/http.ts:1504-1510` (pre-v2.9.1) set the stateful-init transport's onclose handler to:

```ts
transport.onclose = () => {
  const id = transport.sessionId;
  if (id) sessions.delete(id);
  server.close().catch(() => {});   // <— recursion
};
```

`Server.close()` inherits from `Protocol.close` at `node_modules/@modelcontextprotocol/sdk/dist/esm/shared/protocol.js:500-502`:

```js
async close() {
  await this._transport?.close();
}
```

So calling `server.close()` from inside `transport.onclose` re-enters `transport.close()` → fires `this.onclose?.()` → handler runs again → `server.close()` again → infinite recursion → `RangeError: Maximum call stack size exceeded` surfacing at `node_modules/@modelcontextprotocol/sdk/dist/esm/server/webStandardStreamableHttp.js:639` (the closing brace of `async close()`; V8 reports async rejection at the declaration end).

The recursion fires from two real-deployment paths that the v2.5-v2.9.0 test suite never exercised:
- **Session reaper** at `src/transport/http.ts:1443` — `sess.transport.close()` on idle sessions (60s default interval, 5min idle timeout).
- **DELETE /mcp endpoint** at `src/transport/http.ts:1691` — `await session.transport.close()` when an MCP client (Tether 0.2.0) sends a session-termination request.

Tether's `maxRetries: 20` reconnection (`extensions/vscode/src/extension.ts:506-511`) amplifies — each crash → another connect → another crash. The live repro produced ~530 RangeError lines.

### The fix

Remove the `server.close()` call from `transport.onclose`. Keep the `sessions.delete(id)` so the daemon's sessions Map doesn't grow unbounded. Subscription cleanup is preserved: the v2.5.0 Tether Phase 1 Part S hook at `src/server.ts:262-272` already patches `server.onclose` to run `unsubscribeAllForServer(server)`, and the SDK fires `server.onclose` from the transport-close path WITHOUT going through `Protocol.close()`. So unsubscribe still happens; we just stop the recursive `server.close()` self-call.

One-line code change at `src/transport/http.ts:1504-1527` (replacing the recursive call with a multi-line comment explaining why). Zero behavior change for any non-crashing path.

### Regression test

`tests/v2-9-1-autowake-server-close-recursion.test.ts` exercises the exact path: spawn daemon, connect via `StreamableHTTPClientTransport` with Tether's `reconnectionOptions`, register + subscribe, then `DELETE /mcp` to trigger server-side `transport.close()`. Asserts daemon stderr contains zero `RangeError` / `Maximum call stack` / `unhandledRejection` lines, and `/health` still responds.

Empirically verified during this release: without the fix, the test fails with the same `RangeError: Maximum call stack size exceeded` from the repro. With the fix, it passes. The test catches the bug class on every future PR.

### Also tracked in this release

- `tests/v2-9-0-autowake-headless-push-smoke.test.ts` — the v2.9.0 autowake push-delivery test was untracked in git (a Codex audit caught this). Committed to main alongside the v2.9.1 fix so it runs in CI from now on.

### Out of scope

- The Tether extension itself — no changes. The Tether-side reconnectionOptions + autoInjectInbox behavior already shipped on the marketplace and is unchanged.
- Other transport-close call sites (session reaper, DELETE /mcp, stateless one-shot) — none of these had buggy handlers; the bug was specifically the stateful-init handler calling `server.close()`. All other call sites already correctly let the SDK + the Part S hook handle subscription cleanup.
- v2.9.0 docs — `docs/ambient-wake.md` + `roles/auto-poll-loop-template.md` stay unchanged. The operator-facing contract is the same; v2.9.1 only fixes the underlying recursion that prevented path α (Tether) from actually working end-to-end.

### Pre-publish gate

Expected 20/20 PASS — fix is a one-line removal + comment, lockfile regenerated, no other changes.

### References

- Spec: the ambient-wake design spec (still authoritative for the operator pattern; v2.9.1 is the implementation fix that makes path α actually deliver).
- Origin: a live VS Code + Tether 0.2.0 repro 2026-06-10.
- Root-cause confirmation: empirical via the regression test in this release.

## v2.9.0 — 2026-06-08 — Ambient-wake operator guide + auto-poll-loop template

Added: ambient-wake operator guide + auto-poll-loop template + corrected stale source comments about seq-assignment timing (comment-only, no behavior change).

v2.9.0 is a docs-and-template release that turns the v2.3.0 Phase 4s ambient-wake protocol into a usable operator pattern. The relay-side primitives (`peek_inbox_version`, mailbox/epoch, filesystem marker, `/api/wake-agent`, MCP resource subscription) shipped in v2.3.0 + were verified by a Codex audit of the v2.9.0 spec. v2.9.0 documents how operators wire those primitives into a real workflow so builders wake themselves on relay mail and surface only decisions to the human.

### What's in scope

- **`docs/ambient-wake.md`** — appended a v2.9.0 "Operator setup" section covering the two MVP paths:
  - **(α) Tether** for VS Code (already-shipped extension's MCP-resource subscription + `autoInjectInbox` keystroke injection — zero idle cost, push-based).
  - **(β) `/loop`** for iTerm2 / Terminal.app / tmux / SSH (Claude Code v2.1.71+ `/loop` + `ScheduleWakeup` self-paced peek polling — cheap when idle, scales by cadence choice).
  - Decision-gating discipline (what surfaces to the human operator vs flows agent→agent automatically).
  - Stretch paths (B `fs.watch` sidecar; D `TeammateIdle` hook + `Monitor` tool — gated on Claude Code v2.1.98+).
  - Measured per-tick `peek_inbox_version` cost from a real call against the live daemon (replaces the spec's estimates).
- **`roles/auto-poll-loop-template.md`** — canonical `/loop` recipe for path β. Drop-in for builders who want self-pacing without operator nudges. Includes cadence tuning table, drift detection (epoch handling), failure modes, and a copy-paste quick-start.
- **`roles/README.md`** — references the new template alongside the existing role templates.

### Binding constraint preserved

Claude Code v2.1.89's **Auto-mode is OFF the table.** It triggers per-tool-call token burn that defeats the cheap-polling premise. Path (β) uses `ScheduleWakeup` self-paced, NOT Auto-mode. Any future revision suggesting Auto-mode as the harness should be rejected at audit. Recorded in `docs/ambient-wake.md` + the role template.

### Out of scope (explicitly NOT in v2.9.0)

- No daemon/src changes. Pre-publish gate stays 20/20 — the v2.8.0 R1 lockfile-version-sync guard self-applies to this docs release.
- No tests/ changes. The existing v2.3.0 contract test at `tests/v2-3-0-ambient-wake.test.ts` continues to cover C.1-C.6.
- Tether `autoInjectInbox` default flip stays opt-in for v2.9.0; recommended as a separate Tether arc after real-world burn-in (see `docs/ambient-wake.md` § "Recommended next").
- Stretch paths P4 (`fs.watch` sidecar) + P5 (TeammateIdle hook) NOT in this release. P5 is gated on a `claude --version` check anyway.
- The cross-terminal smoke (two-terminal end-to-end wake verification) is methodology + operator-execution territory — documented in `docs/ambient-wake.md` § "Cross-terminal smoke — pending operator validation"; the operator runs it post-publish before declaring the arc complete.

### Compatibility

- Backwards-compatible. Existing operators on Tether keep working unchanged (Tether's `autoInjectInbox` still defaults to false; opt-in to flip per agent).
- Existing role templates (`builder.md`, `worker-loop.md`, etc.) unaffected; the new template composes with them (e.g., a worker-loop agent can run the auto-poll harness alongside its task-pull loop).
- v2.3.0 `peek_inbox_version` contract unchanged. The wake signal is still `total_unread_count` (not `last_seq`) — Codex HIGH #2 patch at `src/db.ts:3014-3021` remains authoritative.

### Pre-publish gate

20/20 PASS — same shape as v2.8.0 R1. The lockfile-version-sync guard (step #9, added in v2.8.0 R1) passes cleanly because the lockfile was regenerated post-version-bump per the documented rebase discipline.

### References

- Spec: an ambient-wake design spec.
- Primitives doc: `docs/ambient-wake.md` (existing v2.3.0 protocol doc; v2.9.0 appends operator setup).
- Role template: `roles/auto-poll-loop-template.md`.
- Origin: ambient-wake workflow requirement (agents run in the background and surface decisions only when needed).

## v2.8.0 — 2026-05-26 — Dashboard daemon precursor: 5-state state machine + SIGHUP + decay broadcaster + wire-emit sites

Closes the dashboard ambiguity called out in the 2026-05-21 dashboard recon: an iTerm2 tab-close looked identical to "alive but slow" looked identical to "waiting normally" looked identical to "stale and needs attention". v2.8 fixes the DAEMON side end-to-end; v2.9 (separate arc) will wire the new state model into the actual dashboard UI rendering.

All architectural calls locked during the 2026-05-25 v2.8 dashboard state-machine design review.

### Consolidated state machine — 5 operator-meaningful states

New module `src/agent-state-machine.ts` exposes `deriveDashboardState(inputs, now, thresholds)` returning one of `'active' | 'pending' | 'waiting' | 'stale' | 'closed'`. Pure function — no I/O, no side effects. Precedence top-wins: `closed > stale > pending > active > waiting`.

| State | Derivation rule |
|---|---|
| `closed` | `signal_received_at != null` OR `unregistered_at != null` OR `last_seen` older than `RELAY_SESSION_TIMEOUT_SEC` (default 30 min) |
| `stale` | was-active inside `RELAY_STATE_WAS_ACTIVE_WINDOW_SEC` (1 hr) AND quiet for `RELAY_STATE_STALE_WINDOW_SEC` (5 min) AND (pendingCount > 0 OR `RELAY_STATE_RECENT_DISPATCH_SEC` window has lastDispatched event) |
| `pending` | `pendingCount > 0` (caller pre-filters by `RELAY_STATE_PENDING_WINDOW_SEC`, 60s) |
| `active` | `last_seen` within `RELAY_STATE_ACTIVE_WINDOW_SEC` (30s) |
| `waiting` | default catch-all |

All six thresholds env-tunable via the listed vars. `resolveThresholdsFromEnv()` falls back to defaults on any invalid value (non-numeric, negative, NaN).

The legacy `deriveAgentStatus` in `src/db.ts:131-159` stays untouched — it returns the pre-v2.8 8-state union (`agent_status`) which existing consumers still read. v2.9 will migrate the dashboard UI; v2.8 only emits the new state on the WebSocket wire and stamps the new DB columns.

### Schema migration — three new NULL-default columns (v12 → v13)

`migrateSchemaToV2_11` (idempotent additive) adds:

- `agents.signal_received_at INTEGER NULL` — epoch ms stamped by `closeAgentSession` when SIGHUP / SIGINT / SIGTERM fires.
- `agents.signal_kind TEXT NULL` — `'SIGHUP' | 'SIGINT' | 'SIGTERM'`. Operator visibility into which path closed the terminal (iTerm2 tab-close vs ctrl-C vs OS shutdown).
- `agents.last_dispatched_at INTEGER NULL` — epoch ms of the most recent dispatch-relevant event. Powers the `stale` derivation by distinguishing "was given work to do recently" from "just idle".

All three columns are NULL on existing rows; their absence routes cleanly through the v2.7 fallback paths in `deriveDashboardState` and existing consumers. Schema version 12 → 13. Migration registered in `applyMigration` at both [12, 13] pair AND finalize chain.

### SIGHUP wired alongside existing SIGINT/SIGTERM

`installAutoUnregister` at `src/transport/stdio.ts:151-181` now listens for SIGHUP in addition to SIGINT and SIGTERM. iTerm2 / Terminal.app tab-close and ssh-disconnect both deliver SIGHUP — pre-v2.8 the relay would never mark the agent closed in those paths and the dashboard saw a perpetually-alive ghost. Exit-code semantics preserved per POSIX: SIGINT→130, SIGTERM→143, SIGHUP→129.

`closeAgentSession` at `src/db.ts:1740-1779` extended to accept an optional `signalKind` parameter; when set, it ALSO stamps the new `signal_received_at` + `signal_kind` columns. Non-signal close paths (explicit unregister, dispatcher-driven) preserve their pre-v2.8 column-set (NULL signal columns).

Real-process integration test at `tests/v2-8-sighup-handler.test.ts` — spawns the built `dist/index.js`, sends `process.kill(pid, 'SIGHUP')`, verifies the row's `signal_kind = 'SIGHUP'` + `signal_received_at` populated + exit code 129. Same coverage for SIGINT (regression) + SIGTERM (regression). Test path matches the shipped path: real OS signal, no mocks.

### Decay-tick broadcaster

New module `src/dashboard-state-broadcaster.ts` with class `DashboardStateBroadcaster`. Runs in the HTTP daemon process only; stdio agents skip it.

- `setInterval` at `RELAY_DECAY_TICK_MS` (default 30000ms).
- Per tick: for each registered agent compute `deriveDashboardState`, compare to `lastBroadcastedState[name]` (in-memory Map dedup), fire `broadcastDashboardEvent({ event: 'agent.status_changed', entity_id: name, kind: newState })` ONLY on transitions.
- Steady-state: no events. Wire stays quiet while nothing is changing.
- Lifecycle: `start()` and `stop()` are idempotent. The class accepts `scheduler` + `now` + `getAgents` + `broadcast` for DI so unit tests use fake clocks/timers without `vi.useFakeTimers()` global pollution.
- Opt-out via `RELAY_DECAY_TICK_DISABLED=1` (test rigs, smoke).

Token cost: zero. Pure Node setInterval + in-process Map + WS push. No LLM involvement anywhere.

### Wire-emit sites

`broadcastDashboardEvent` calls added at two previously-missing mutation sites so the dashboard sees activity instantly instead of waiting up to a 30s decay tick:

- `src/tools/identity.ts:188-204` — `register_agent` now fires `agent.state_changed` with `kind = 'registered'` (first-mint) or `'reregistered'` (re-register against existing row). Metadata-only per the v2.2.0 H4 audit — no token, no caps payload.
- `src/transport/stdio.ts:118-130` — `performAutoUnregister` (signal handler) fires `agent.state_changed` with `kind = <signal name>` immediately after `closeAgentSession` succeeds. Pre-v2.8 the close was only visible via the decay broadcaster's next tick (up to 30s latency); now the dashboard sees it instantly.

Existing emit sites (send_message, post_task, update_task, set_status, unregister_agent) ALREADY routed through `fireWebhooks → emitDashboardBroadcast` at `src/webhooks.ts:110-145`, no changes needed there.

`DashboardEvent` union at `src/transport/websocket.ts:66-86` extended with `'agent.status_changed'` for the decay broadcaster's transitions. The pre-v2.8 events (`agent.state_changed`, `message.sent`, `task.transitioned`, `channel.posted`, `dashboard.theme_changed`) all preserved verbatim.

### Tests

`tests/v2-8-agent-state-machine.test.ts` — 32 tests (M1-M32). Every state, every transition, every boundary. Custom-threshold injection. Env-var resolution with fallback on invalid values. Tests assert the exact contract, not a proxy: exact `toBe('closed')` style assertions, not proxy checks.

`tests/v2-8-sighup-handler.test.ts` — 3 tests (SH1-SH3). Real OS signals (SIGHUP / SIGINT / SIGTERM) against a forked dist/index.js. Verifies DB columns populated correctly + exit codes match POSIX convention. Skipped on win32.

`tests/v2-8-decay-broadcaster.test.ts` — 19 tests (D1-D19). Mock-clock + ManualScheduler + FakeBroadcastSink. Dedup correctness (no event when state unchanged). All five transitions reachable. Lifecycle (start/stop idempotency). Env resolution.

`tests/v2-8-wire-emit-sites.test.ts` — 6 tests (WE1-WE6). Real HTTP transport + WebSocket round-trip. Asserts each mutation site fires the expected broadcast event with expected `kind` tag.

Brief projected ~40-50 new tests. Honest count: 60. Each test pins a distinct contract — no padding.

### Pre-publish gate

New step `extension_state_machine_smoke` runs the v2-8 test files specifically as a fast feedback loop before the full root vitest. Pre-publish `--full` gate count goes from 18/18 to 19/19.

### Environment variables added

| Var | Default | Affects |
|---|---|---|
| `RELAY_STATE_ACTIVE_WINDOW_SEC` | 30 | `deriveDashboardState` active window |
| `RELAY_STATE_PENDING_WINDOW_SEC` | 60 | pending-message age filter (caller-side) |
| `RELAY_STATE_STALE_WINDOW_SEC` | 300 | stale promotion quiet threshold |
| `RELAY_STATE_WAS_ACTIVE_WINDOW_SEC` | 3600 | stale "was active in last hour" guard |
| `RELAY_SESSION_TIMEOUT_SEC` | 1800 | closed-via-session-timeout cutoff |
| `RELAY_STATE_RECENT_DISPATCH_SEC` | 600 | stale "recently dispatched" alternative trigger |
| `RELAY_DECAY_TICK_MS` | 30000 | decay broadcaster interval |
| `RELAY_DECAY_TICK_DISABLED` | unset | set to "1" to disable decay broadcaster (test rigs) |

### Compatibility

- Schema migration v12 → v13 is additive (ALTER TABLE ADD COLUMN). No DROP, no rename, no breaking change. Rollback to v2.7.x leaves the three columns as unused NULL fields.
- `closeAgentSession` parameter list extended with optional `signalKind = null`. Existing callers without the parameter preserve the pre-v2.8 column-set (signal columns stay NULL).
- `DashboardEvent` union additive. Existing dashboard clients that switch on the pre-v2.8 event names ignore the new `agent.status_changed` events without error.
- v2.9 dashboard UI work depends on this release shipping first — it consumes the new WS events + new DB columns.

### Out of scope (explicitly NOT in v2.8)

- Dashboard UI rendering changes — v2.9 arc. Theme system stays untouched (locked 2026-05-25).
- Multi-machine federation visibility — v3 federation arc.
- Heartbeat protocol changes — current keepalive model adequate.
- Per-agent state notifications via webhook — Phase 4 territory if anyone asks.
- Operator alerts (toast, sound, email) — v2.10+ if operator pull surfaces this.

### References

- Locked during the v2.8 dashboard state-machine design review.
- Auditor: external Codex audit

## v2.7.4 — 2026-06-03 — spawn-agent.sh kickstart apostrophe-quoting fix (HIGH)

Closes a spawn-agent kickstart-quoting bug. Spawn-flow hardening; no schema, no API, no runtime behavior changes. Bug 2 (length truncation) is NOT in scope — that's a separate arc requiring token pre-mint + message ordering work.

### The bug

`bin/spawn-agent.sh` wrapped the kickstart prompt via `printf '%q'`, which emits values in `$'...'` ANSI-C-quoted form. Inside `$'...'`, a literal apostrophe terminates the string — any `RELAY_SPAWN_KICKSTART` value (or the default kickstart) containing a `'` produced an unbalanced quoting that wedged the spawned subshell at `quote>` continuation forever.

The default kickstart contained three offending substrings (`status='all'`, `since='session_start'`, `since='1h'`), so first-attempt spawns wedged. Operator workaround until this lands: always pass an apostrophe-free `RELAY_SPAWN_KICKSTART="..."` override.

### Fix — Option C (defense in depth)

**(A) Sanitize the default kickstart** (`bin/spawn-agent.sh:289`).

`status='all'` → `status=all`, `since='session_start'` → `since=session_start`, `since='1h'` → `since=1h`. MCP tool calls accept unquoted enum values in prompt text; the LLM interprets the call the same way. Safety net so even if the helper below is reverted in the future, the default never wedges.

**(B) Replace `printf '%q'` with `shell_escape_double` for the kickstart** (`bin/spawn-agent.sh:194-218, 313`).

New helper `shell_escape_double` (paired with the existing `applescript_escape`) emits values in plain double-quoted shell context `"..."` with backslash escapes for `\`, `$`, `"`, and backtick. Apostrophes are literal inside `"..."`, so kickstarts with arbitrary apostrophe content no longer wedge.

Scope is intentionally narrow: only `Q_KICKSTART` switches off `printf '%q'`. The other `Q_*` variables (`NAME`, `ROLE`, `CAPS`, `CWD`, `PERM`, `EFFORT`, `DISPLAY`) stay on `printf '%q'` — they're tight-allowlisted (no apostrophes possible) and the bug only manifests for the free-form kickstart.

### Tests

Five new K-tests in `tests/spawn-integration.test.ts`:

- **(K1)** default kickstart contains no apostrophes — safety net for Option A.
- **(K2)** default-kickstart CMD parses cleanly under `bash -n` (would have failed pre-fix).
- **(K3)** `RELAY_SPAWN_KICKSTART` with apostrophes parses cleanly (root-cause fix for Option B).
- **(K4)** `RELAY_SPAWN_KICKSTART` with every shell-special char (`'`, `"`, `\`, `$`, backtick) parses cleanly.
- **(K5)** round-trip: execute the assembled CMD with a `claude` stub, verify the kickstart literal content (including apostrophes) reaches the final argv intact.

One existing test (`spawn-agent.sh — v2.1.4 brief_file_path (I10) → default spawn with valid brief path embeds the pointer sentence in KICKSTART`) updated to match the new escape form — the brief-pointer backticks now appear as `\\\`<path>\\\`` in the dry-run CMD (the helper escapes backticks to prevent command substitution inside `"..."`); when bash actually parses CMD the escapes collapse so claude still sees `` `<brief>` ``.

### Compatibility

- No schema migration. No DB changes. No API changes. No env var changes.
- `bin/spawn-agent.sh` operators on Node 20+ have no action required — `RELAY_SPAWN_KICKSTART` override values that used to wedge now work without modification, and overrides that didn't contain apostrophes continue to work verbatim.
- The apostrophe restriction no longer applies.

### Out of scope (explicitly NOT in v2.7.4)

- Bug 2 (length truncation past ~1000-1024 chars via AppleScript `do script` buffer limit) — separate arc; requires `--initial-message` flag + token pre-mint + message ordering guarantees.
- Other `Q_*` variables continuing to use `printf '%q'` — they're tight-allowlisted and don't trigger the bug.
- Linux/Windows spawn drivers — they use a different code path (`src/spawn/drivers/{linux,windows}.ts`); Bug 1 was specific to the macOS bash-script path.

### References

- Note: the prior working label predated the v2.7.3 CVE hotfix that took the v2.7.3 slot; this fix lands as v2.7.4.
- Bug discovery: 2026-05-21 during v2.7.2 dispatch attempts

## v2.7.3 — 2026-06-02 — SECURITY HOTFIX: vitest@4.1.8 (CVSS 9.8) + qs/ws transitives + Node 18 drop

Narrow security-only release. Closes one CRITICAL + two moderate npm advisories that surfaced after v2.7.2 shipped. Drops Node 18 support (14 months past EOL) because vitest@4 — the only version with the CVE fix — requires Node 20+.

### ⚠️ BREAKING: minimum Node version is now Node 20

`engines.node` bumped from `">=18.0.0"` to `">=20.0.0"`. CI matrix bumped from `['18', '20', '22']` to `['20', '22']`.

Reasoning:

- Node 18 went **EOL on 2025-04-30** — 14 months past end-of-life as of this release. Active + Maintenance LTS coverage is over.
- vitest@4 (the only version that fixes the CRITICAL CVSS 9.8 advisory below) imports `node:util.styleText`, which only exists on Node 20.12+. There is no vitest@3.x patched release; the upstream fix is semver-major.
- Anyone still on Node 18 is running an unpatched runtime regardless; v2.7.2 (also affected by the CVE) remains available on npm for legacy consumers who can't upgrade. They are in the same security position as if v2.7.3 didn't exist for them.

Operators on Node 20+ have no action required beyond `npm install`.

### Security

- **[CRITICAL CVSS 9.8] `vitest <4.1.0` (GHSA-5xrq-8626-4rwp).** "When Vitest UI server is listening, arbitrary file can be read and executed." Bot-relay-mcp does not run the Vitest UI server in production, but the package was a transitive in the test/dev surface; advisory triggers the pre-publish `npm audit (high+)` gate regardless of usage path. Upgrade: `vitest@^4.1.8` (semver-major from `^3.2.0`).
- **[moderate] `qs 6.11.1 - 6.15.1` (GHSA-q8mj-m7cp-5q26).** `qs.stringify` DoS via TypeError on null/undefined entries in comma-format arrays with `encodeValuesOnly`. Transitive via Express. Auto-fix via `npm audit fix`.
- **[moderate] `ws 8.0.0 - 8.20.0` (GHSA-58qx-3vcg-4xpx).** Uninitialized memory disclosure. Direct dep. Auto-fix via `npm audit fix`.

Post-fix: `npm audit --audit-level=high` reports **0 vulnerabilities**.

### Why standalone

The v2.8 dashboard-daemon-precursor arc was mid-flight on `hotfix/v2.8-dashboard-state-machine` when these advisories landed in the npm advisory DB (between v2.8 R1 push 2026-05-26 and R2 push 2026-06-02). Per the 2026-06-02 lock: ship the security hotfix as a standalone arc off `main` so npm users get the CVE patch immediately, then rebase v2.8 R2 on top of clean main — Codex audits each arc on narrow scope rather than juggling a wider mixed scope.

### Compatibility

- **Breaking:** Node 18 no longer supported. See the BREAKING section above.
- vitest 4 reporter / runner internals changed but the project's `vitest.config.ts` defaults work as-is; full suite passes under Node 20+ with no test rewrites.
- Production runtime untouched. `dist/` build is byte-identical to v2.7.2 at the JS level (the security advisories are dev-surface; runtime deps unaffected at the user-facing layer).

### Tests

No tests added or modified. Full v2.7.2 test suite passes verbatim under the new vitest:
- Root vitest (sequential `--pool=forks --no-file-parallelism`): **1353 / 125 files PASS** under vitest@4.1.8 on Node v24.13.0 locally.
- Pre-publish `--full` gate: **18/19 PASS**. The single expected FAIL is the GitHub CI green-gate, which inspects `main`'s CI status — `main` is red from the exact advisories this PR closes (chicken-and-egg). Resolves automatically on merge. The actionable pre-push check is BRANCH CI at HEAD, which is CI run 26805501670: Node 20 + Node 22 + 25-tool smoke ALL PASS.

### DEFERRED-LOCAL

Only Node 24 available locally; Node 20 + 22 verified via the GitHub Actions CI matrix at `.github/workflows/ci.yml`. The completion report cites the CI run number with Node 20 + 22 + 25-tool smoke all green.

### References

- Dispatched 2026-06-02 after a Codex R1 audit caught a vitest critical on the v2.8 R2 push.
- Option A (drop Node 18) locked 2026-06-02 after a scope-expansion request.

## v2.7.2 — 2026-05-21 — spawn_agent identity-recovery defense-in-depth

Closes the silent "default" failure mode when `mcp__bot-relay__spawn_agent` produces a child terminal whose SessionStart hook never sees `RELAY_AGENT_NAME`. Origin: a bug report (2026-05-13).

### Reframe of the failure mode

The original bug report quoted the hook as printing `$RELAY_AGENT_NAME is empty — this terminal isn't registered as a named agent.` — **that string does not exist anywhere in the codebase**. `grep -rn` returns zero matches in `src/`, `hooks/`, `bin/`, and `tests/`. The user-visible symptom (identity unresolved, queued kickstart never surfaces) is real; the failure mode is silent, not loud:

- `hooks/check-relay.sh:40` — `AGENT_NAME="${RELAY_AGENT_NAME:-default}"` falls back to the literal string `default` with no warning when the env var is unset or empty.
- The hook then registers an agent named `default` over HTTP (`hooks/check-relay.sh:259-271`).
- Mail queued for the *intended* name (e.g. `build-agent`) dead-letters — the SELECT at `hooks/check-relay.sh:315-321` only matches `to_agent = 'default'` once that's the registered name.
- From the operator's perspective the terminal "looks unregistered"; from the hook's perspective everything succeeded at exit 0.

The script-side CMD is structurally correct (`bin/spawn-agent.sh:201` exports come first, before the conditional vault hydration and before `claude`). A new contract test (`tests/spawn-integration.test.ts` — v2.7.2 block) executes the dry-run CMD in a clean subshell with a `claude` stub and confirms `RELAY_AGENT_NAME` reaches both the immediate claude env AND a subprocess forked by claude. The test passes on this commit — the propagation contract holds end-to-end inside the script's scope. The remaining attack surface is between `osascript … write text` and the live `claude` binary's hook-subprocess env policy, which is opaque from outside.

### Defense-in-depth fix — spawn-manifest fallback

Rather than guessing which downstream layer drops the env var, this release converts the silent-default-fallback into a recoverable + loud-failing one:

- **`bin/spawn-agent.sh:218-232`** — after sourcing `_vault-helpers.sh`, the script drops a `<instanceDir>/agents/<name>.spawn-manifest` file beside the per-instance token vault. Manifest format is key=value (`name`, `role`, `spawn_pid`, `spawned_at` ISO8601 UTC). Write is atomic (tmp+rename), chmod 0600, owner-only readable. Write failure is non-fatal — typed-env transport remains primary and a warning lands on stderr.
- **`hooks/check-relay.sh:42-50` + `:60-85`** — when env-derived `AGENT_NAME` is unset OR equals literal `default`, the hook scans the per-instance `agents/` dir for *.spawn-manifest files modified within 60s. If **exactly one** fresh manifest exists and its filename matches its `name=` content, the hook recovers identity from it, emits a stderr breadcrumb, and deletes the manifest so it can't be re-used by a later unrelated terminal. Multiple-candidate ambiguity → fall through to `default` (loud warning beats wrong identity). Opt-out: `RELAY_DISABLE_MANIFEST_FALLBACK=1`.
- **`hooks/_vault-helpers.sh:104-256`** — four new helpers (`resolve_relay_spawn_manifest_path`, `write_relay_spawn_manifest`, `find_fresh_relay_spawn_manifest`, `delete_relay_spawn_manifest`). Single source of truth, sourced by both the hook and the launcher script. The freshness check uses `stat -f %m` / `stat -c %Y` (macOS + Linux) for real-second granularity — an earlier `find -mmin -N` impl was rejected during test development because `-mmin` can only express integer minutes and silently let 30s-old manifests through a 10s window.

### Caveat carried into this ship

This is a defense-in-depth ship, not a root-cause kill. If the actual failure is in `claude`'s hook-subprocess env policy (e.g. claude strips `RELAY_*` before forking the hook), the manifest fallback recovers identity AND emits a stderr breadcrumb that surfaces the underlying issue. Operators who observe the recovery breadcrumb in production should report it — it's an instrumentation handle on the opaque layer.

### Tests

- New file `tests/v2-7-2-spawn-manifest.test.ts` — 22 tests covering manifest write atomicity, chmod, malformed-input rejection, find-fresh ambiguity handling, mtime windowing (including the sub-minute precision the `-mmin` impl couldn't deliver), delete idempotence, hook integration (5 cases including opt-out + explicit-name-wins), and drift guard for both shipped-helper definition and consumer-side shadow-redefinition.
- New `tests/spawn-integration.test.ts` block — 4 contract tests asserting `RELAY_AGENT_NAME` reaches both `claude` env AND its subprocess env, for plain/hyphenated/dotted name shapes (`worker1`, `build-agent`, `pod.alpha`).

Pre-publish gate adds 26 tests on top of v2.7.1's count. Existing 75 spawn-driver tests + 41 spawn-integration tests + 15 hook-contract tests pass unchanged — no regression in the adjacent surface.

### Cross-references

PR opened against post-v2.7.1-merge `main` (HEAD: `d516cfd`). Branch: `hotfix/v2.7.2-spawn-name-propagation`. Tether VSCode extension v0.1.3 (`d516cfd`) and v2.7.1 R1 (`189a170`) both shipped first.

## v2.7.1 — 2026-05-13 — SECURITY HOTFIX

Two CRITICAL findings + one HIGH + three LOW/MED that surfaced in a security review AFTER the PR #31 Codex audit had shipped v2.7.0 cleanly. The PR #31 audit was scoped to the four release commits; these findings live in PRE-EXISTING code that would have shipped via `npm publish` if the publish step had run before the security review landed.

All fixes ship in one PR. PR opened against `main` post-v2.7.0-merge; no v2.7.0 tag was created and no npm publish ran, so v2.7.1 is the first public npm artifact in the v2.7 series.

### Security

- **[CRITICAL] `expand_capabilities` now requires `admin` capability** (`a92d22a`). Pre-fix the dispatcher's `TOOL_CAPABILITY` map at `src/auth.ts` omitted `expand_capabilities` entirely. Unmapped tools fall through to "no capability required", so any authenticated agent — even one with the default `{user}` cap set — could call `expand_capabilities` on themselves to add `admin`, `manage_others`, `rotate_others`, then `revoke_token` or `rotate_token_admin` any peer. ~3 calls to compromise a relay. Origin: a security review.

- **[CRITICAL] Plaintext agent tokens no longer logged to stderr on `register_agent`** (`1a56beb`). Pre-fix `src/tools/identity.ts:150-153` interpolated the freshly-minted token at info level on every registration; every stderr-capturing surface (terminal scrollback, CI logs, journald, Docker logs, log aggregators, SaaS observability) ended up with the token in cleartext. The "shown ONCE in the API response" model was broken from day one. Fix has two layers: (a) strip the token from the log line entirely (agent name + "(shown once in the tool response; not logged)" hint preserve the operational signal); (b) `redactSecrets` defense-in-depth scrubber in `src/logger.ts` applied to every log line at every level, matching `RELAY_AGENT_TOKEN=<value>`, `Authorization: Bearer <value>`, `Authorization: <other-scheme>`, `X-Agent-Token: <value>` (case-insensitive), and JSON-ish `"<key>": "<value>"` for token / agent_token / recovery_token / secret / http_secret / webhook_secret / password keys. Origin: an external security review.

- **[HIGH] Dashboard `send_message` now requires `from_agent_token` for any `from` field that names a registered agent** (`aa75924`). Pre-fix `src/transport/http.ts` treated `from_agent_token` as optional defense-in-depth (Option (a) audit-only). Whoever held the dashboard's HTTP secret OR an authenticated session cookie could POST send_message with `from=<victim>` and impersonate any registered agent across the relay's full message + task surface. Option A was selected: make from_agent_token REQUIRED when from-agent has a stored token_hash. Missing → 403 AUTH_FAILED; mismatch → 403 AUTH_FAILED; correct → success. System-message senders (e.g., `dashboard-system`) need their own row + token; no "no token = system" fallback. Origin: an external security review.

### Fixes

- **[MED] Walked all handlers that do "fetch → JS-filter → mutate"; confirmed v2.7.0's `get_messages` fix was complete.** No code change in v2.7.1 — the walk confirmed no other sites match the silent-data-loss class. Sibling read paths (`getMessagesSummary`, `peekMailboxVersion`, `getMessagesInWindow`, `getTasksInWindow`) all push `since` into SQL OR are pure reads with no mutation. One adjacent fidelity gap noted for v2.8: `src/tools/standup.ts:159` fetches the audit log with a hard LIMIT 500 then filters in JS by `created_at >= sinceIso` — pure read, no data loss, but the window may be incomplete on high-traffic relays. Deferred. Origin: codex security review.

### Operational

- **[LOW] Remaining `[broadcast-trace]` info-level lines downgraded to debug** (`962aff7`). v2.7.0 cleaned up the per-event lines; v2.7.1 downgrades `subscribe added` (once-per-subscriber-lifetime) and `GET /mcp (SSE stream open attempt)` (once-per-session). Only `fanout enter` stays at info per codex PR #31 audit pt 4 — it's the load-bearing per-event observability summary. Operators chasing live correlation can set `RELAY_LOG_LEVEL=debug`.

- **[LOW] `SECURITY.md` + `architecture.md` now included in npm tarball** (this commit). Pre-fix `package.json` `files` array omitted both, so `npm install bot-relay-mcp` didn't ship the security disclosure or architecture-overview docs.

### Cross-references

PR opened against post-v2.7.0-merge `main` (HEAD: `afcf6ef`). Branch: `hotfix/v2.7.1-security`. Tether VSCode extension v0.1.3 (HIGH F10: SecretStorage migration) ships as a separate hotfix branch; the marketplace re-publish does NOT block this npm release.

## v2.7.0 — 2026-05-13 — Tether-ready cross-process inbox notifications + externally-flagged P1 correctness fix

The v2.7.0 release closes the cross-process notification gap between stdio MCP terminals and the HTTP daemon that ships with `bot-relay-mcp`, lands an externally-flagged P1 correctness bug in `get_messages`, and pairs with the [Tether VSCode extension v0.1.2](https://marketplace.visualstudio.com/items?itemName=lumiere-ventures.bot-relay-tether) on the marketplace.

The Tether v0.1.2 marketplace listing assumes the daemon has Phase 5 keepalive — anyone running `npm install bot-relay-mcp@latest` pre-v2.7.0 would silently degrade to ~2.5min Electron-fetch idle disconnects. v2.7.0 closes that gap end-to-end.

### Phase 3 — durable outbox + cross-process notification delivery

Pre-v2.7.0 the inbox-changed event bus (`src/inbox-events.ts`) was a module-local Node `EventEmitter`. A message sent from a stdio MCP terminal (one OS process) emitted ONLY inside that stdio process and never woke a subscriber connected to the separate HTTP daemon. Live Tether smokes proved the gap — `[broadcast-trace] event emit` never fired in the daemon log when stdio writers committed.

Fix is the durable-outbox pattern:
- **New table `inbox_events`** (schema v11 → v12, idempotent migration). `CREATE TABLE inbox_events (id INTEGER PRIMARY KEY AUTOINCREMENT, agent_name TEXT NOT NULL, reason TEXT CHECK (...) NOT NULL, created_at TEXT NOT NULL, source_pid INTEGER)` with two indexes on `id` and `(agent_name, id)`. Per-row retention drops outbox rows older than `RELAY_OUTBOX_RETENTION_DAYS` (default 7d).
- **Producer wiring** at all 4 mark-as-read / send-message / broadcast sites in `src/db.ts` (`sendMessage` system + normal paths, `getMessages` drain, `broadcastMessage`). Each producer INSERTs an outbox row inside the same SQLite transaction as the underlying write and threads `lastInsertRowid` into `emitInboxChanged` so subscribers can dedup by event id.
- **Cross-process tail** at `src/outbox-tail.ts` — runs ONLY in the HTTP daemon. Polls `inbox_events` for rows past an in-memory cursor (`MAX(id)` at startup, no replay of historical rows) and dispatches each to the existing `mcp-subscriptions` broadcaster. `PRAGMA data_version` cheap-skip means quiet periods cost one uint64 read per tick. Batched at 500 rows with `setImmediate` re-tick on backlog. Configurable via `RELAY_OUTBOX_POLL_MS` (default 100ms; 0 disables).
- **Broadcaster refactor** at `src/mcp-subscriptions.ts` — extracted `broadcastInboxChange(agentName, reason, eventId, source)` as the single fan-out path called by BOTH the in-process bus AND the cross-process tail. `lastBroadcastIdByUri` Map dedups by event id so when sender + subscriber share a process the subscriber receives exactly one notification (bus fires first; tail catches up and is deduped).

Schema migration is the structural fix to a latent bug across every prior bump: `initSchema`'s `INSERT OR IGNORE` on the single `schema_info` row was a no-op for existing DBs, and the subsequent `UPDATE` only refreshed `last_migrated_at` — never the `version` field. `advanceSchemaVersionIfBehind` helper at `src/db.ts` now runs at the end of both init chains AND from each `applyMigration` case, so the recorded version stays in sync with the actual DB shape.

Cite: commits `dacc121` (Phase 3a schema + producer), `a093ee9` (Phase 3b tail + broadcaster), `0b0387c` (Phase 3d cross-process tests), `818823f` (Phase 3 R1 schema_info version advancement, P1 from codex audit).

### Phase 4a — reaper skips sessions with active SSE GET stream

The HTTP daemon's session reaper (`src/transport/http.ts`) closes sessions whose `lastSeen` is older than `RELAY_HTTP_SESSION_IDLE_SECONDS` (default 300s). Pre-v2.7.0, `lastSeen` was bumped only in `handleSessionRequest` — which fires ONCE per long-lived SSE GET stream open. After 5min of MCP idle (subscriber listening, no POSTs), the reaper culled the session and the SSE stream dropped.

Fix: each `HttpSession` now carries an `openGetStreams: number` counter. The GET handler increments before `handleRequest` and a single `res.once("close", ...)` listener decrements with a `Math.max(0, n-1)` floor that survives reconnect-races where a newer GET stream is live before the older one closes. Reaper skips any session with `openGetStreams > 0` regardless of `lastSeen` age.

`shouldReapSession(session, now, idleMs)` pure predicate exported from `src/transport/http.ts` for unit tests. New env seams `RELAY_HTTP_REAPER_INTERVAL_MS` (default 60_000) and `RELAY_HTTP_REAPER_TEST_MODE=1` (or `NODE_ENV=test`) relax the 30s floor on `RELAY_HTTP_SESSION_IDLE_SECONDS` for fast integration tests; production keeps the 30s floor.

Cite: commit `f957c29`.

### Phase 5 — SSE keepalive comment frames (fixes Electron-fetch idle timeouts)

After Phase 4a fixed the server-side cull, post-Phase-4 Tether smoke revealed a SECOND disconnect class at ~2.5min where the daemon log went silent (no reap, no close, no error) but the extension reported `SSE stream disconnected: TypeError: terminated`. Root cause: VS Code's Electron-based `fetch` implementation has its own idle-stream timeout on `response.body` reads, separate from Node's and undici's.

Fix: daemon-side periodic `:keepalive\n\n` SSE comment frame writes. Per the SSE spec, lines beginning with `:` are comments that clients MUST ignore — invisible to the MCP message stream but visible to the fetch runtime as recent byte activity. Default interval 20s gives ~3-4× margin against typical 60-90s intermediary idle thresholds AND ~7× margin against Electron's observed ~2.5min cull. Configurable via `RELAY_SSE_KEEPALIVE_MS` (default 20_000; 0 disables).

`setupSseKeepalive(res, intervalMs)` exported from `src/transport/http.ts` as a pure helper with idempotent cleanup, self-cancelling on `writableEnded` / write-throws / `res.close`. Wired into the GET `/mcp` handler immediately after the Phase 4a `openGetStreams` increment.

Cite: commit `6ec32d6`.

### Externally-flagged P1 — `get_messages` filter-after-mark silent data loss

An external review (2026-05-11) surfaced a correctness bug not caught across 4+ audit cycles: `src/tools/messaging.ts` ran `filterBySince(rows, sinceIso)` in JS at line 159 AFTER `src/db.ts` `getMessages` had already SELECTed pending rows + UPDATEed their `read_by_session` field in the same SQL call. Net effect: a message older than the caller's `since` bound was consumed (marked-read for this session) even though the caller never saw it in the response. The message never resurfaced in subsequent `status='pending'` calls from the same session because `read_by_session != null` excluded it.

Fix:
- `getMessages(agentName, status, limit, peek, sinceIso)` — new 5th parameter. SQL SELECT now stitches `AND created_at >= ?` BEFORE the mark-as-read transaction. Only rows that survive the filter get marked read.
- `handleGetMessages` (`src/tools/messaging.ts`) — passes `sinceIso` into the DB call; obsolete `filterBySince` removed entirely.
- `sampleGetMessagesConsistency` (`src/transport/consistency-probe.ts`) — takes `sinceIso` so its SUPERSET SQL query mirrors the same bound (else every since-narrower-than-all call emits false-positive divergence warnings post-fix).

Regression test at `tests/v2-7-0-get-messages-filter-after-mark.test.ts` (3 cases): the exact scenario the external review described, idempotency within the same window, and the `peek=true` no-mutation control. Pre-fix code fails the first assertion ("old message is missing after a since='15m' call"); post-fix all 3 pass. Full investigation + walk-analogous results at `docs/v2.7.0-get-messages-filter-after-mark.md`.

Walk-analogous (verified by reading source): `getMessagesSummary` already applies `since` as SQL filter and is pure-read, no mutation — unaffected. `peekMailboxVersion` returns aggregate counts, no since filter, no mutation — unaffected. The bug was specific to `getMessages` + `filterBySince`.

### Phase 2 broadcast-trace log cleanup

Phase 2 of the cross-process investigation added six `[broadcast-trace]` log lines at info level for live correlation during the Tether smoke. v2.7.0 downgrades the per-event lines to debug (visible under `RELAY_LOG_LEVEL=debug`) while keeping the per-session and per-fanout summary lines at info for production observability.

Downgraded (debug):
- `src/inbox-events.ts` event emit (fires per send/read/broadcast)
- `src/mcp-subscriptions.ts` `dedup-skip` (per duplicate observed)
- `src/mcp-subscriptions.ts` `notifying` + `notify accepted` (per subscriber per fanout)
- `src/server.ts` `resources/subscribe`/`unsubscribe RPC arrived`

Kept at info:
- `src/mcp-subscriptions.ts` `fanout enter` (semantic "broadcast happening" — once per emit with full context)
- `src/mcp-subscriptions.ts` `subscribe added` (once per subscribe — semantic "subscriber registered")
- `src/transport/http.ts` `GET /mcp (SSE stream open attempt)` (once per SSE session — verifies long-lived stream actually opened)

`tests/v2-6-tether-broadcast-trace.test.ts` updated to set `RELAY_LOG_LEVEL=debug` so it still asserts the full chain.

### Tool count drift fix (externally-flagged)

README + adjacent docs claimed 25 tools. Actual count is **30 MCP tools** (verified via `grep -c "name:" src/server.ts`). Drift accumulated across v2.2 (`get_messages_summary`, `expand_capabilities`, `get_standup`), v2.2.1 (`set_dashboard_theme`), v2.3.0 (`peek_inbox_version`) — none of these landings updated the quoted count. Updated:
- Root `README.md` v2.1 status block → v2.7 status + "30 MCP tools"
- Root `README.md` Layer 2 Managed Agents tool count
- Root `README.md` Tether section now leads with marketplace install
- Roadmap entry for v2.7 added

Long-term external-review recommendation (out of v2.7.0 scope): tool registry should be the single source of truth that generates docs/MCP-defs/auth-bundle-maps/smoke-test-expectations. Deferred to a v2.8 hygiene round.

### Extension Tether v0.1.2 (shipped to marketplace separately, paired with this release)

The marketplace listing assumes daemon v2.7.0+:
- Constructor passes `reconnectionOptions: { maxRetries: 20, ... }` to `StreamableHTTPClientTransport` (SDK default is 2 retries — too aggressive for a long-running editor extension).
- Retry-budget-exhausted status bar text now reads `Tether: error — run "Tether: Reconnect to Relay"` so the existing `botRelayTether.reconnect` palette command is discoverable.
- Drift guard `tests/v2-7-tether-reconnection-options.test.ts` pins the contract against both `src/` and the compiled `out/extension.js` that ships in the VSIX.

Cite: commits `270acad` (Phase 4b extension wiring), `8b6f984` (VSIX bump + CHANGELOG).

### fast-uri vulnerability patched

GHSA-q3j6-qgpj-74h6 (path traversal) + GHSA-v39h-62p7-jpjc (host confusion) — both high-severity. Patched via `npm audit fix` (semver-compatible transitive bump: fast-uri 3.1.0 → 3.1.2). Lockfile-only change.

### CI hygiene during the release arc

Three small CI hotfixes landed in PR #29 to unblock auto-merge:
- `9843d69` — `.github/workflows/ci.yml` compiles `extensions/vscode` before vitest (drift-out tests need `out/extension.js`).
- `5772f66` — fast-uri vuln patched directly into the branch instead of routing through dependabot PR #30 (same fix, faster to merge).
- `992769c` — `scripts/pre-publish-check.sh` now compiles AND emits the extension via `npm run compile` BEFORE the main vitest step (was running `tsc -p . --noEmit` AFTER vitest — never produced `out/`).

### Schema migration safety

Schema v11 → v12 migration is additive (CREATE TABLE IF NOT EXISTS + new indexes) and idempotent. Existing v11 DBs are upgraded in place by the next daemon start. The Phase 3 R1 `advanceSchemaVersionIfBehind` helper retroactively fixes a latent bug across every prior bump — DBs that were structurally upgraded under any prior `migrateSchemaToV2_X` but whose recorded `schema_info.version` stayed at the old number will sync to current on next start (logged as `[schema] advanced schema_info.version N → M`).

### Cross-references

PR #29 squash commit: `05e2b40`.

## v2.6.3 — 2026-05-08 — port-flake hardening (dynamic-port allocation) + dependabot vuln patches

Two hardening items closed in one patch — known port-collision flake class across integration tests, plus 5 medium-severity dependabot alerts on `package-lock.json` (all auto-patched cleanly via `npm audit fix` — no breaking changes).

> Naming convention switch: the version label now matches the npm publish version directly. v2.6.3 here = npm v2.6.3, the version label matches the published version going forward.

### Port-flake hardening

Pre-v2.6.3 several integration tests hardcoded specific ports in the 39413-39988 range to spawn isolated `node dist/index.js` HTTP daemons. When a prior gate iteration crashed mid-run and left a stale process holding the port, the next run failed with port-already-in-use. Caught manually via `lsof -i :<port>` during v2.6.0 publish-prep, then called out by codex on v2.6.4 R0/R2 audits as a known-flake class. Plus two tests using `40000 + Math.floor(Math.random() * N)` which is non-deterministic but still has a small collision probability against the same stale-process pattern.

- **`tests/_helpers/port.ts`** *(new)* — `getFreePort()` async helper. Binds a throwaway `net.createServer()` to port 0 (kernel-assigned), reads the bound port from `server.address()`, closes the server, returns the port for the caller to use. Race window between close + caller-bind is microseconds; standard Node port-claim idiom (used by `get-port` npm + most Node test rigs). No new npm dependency added — Node `net` is stdlib.
- **6 test files converted** to use `getFreePort()` in place of hardcoded literals:
  - `tests/v2-6-1-token-store.test.ts` — 4 sites (39413, 39414, 39415, 39416)
  - `tests/v2-6-2-recovery-flow.test.ts` — 3 sites (39420, 39421, 39422)
  - `tests/v2-6-2-spawn-to-ready.test.ts` — 1 site (39430)
  - `tests/v2-6-4-hook-token-extraction.test.ts` — 3 sites (39450, 39451, 39452)
  - `tests/chaos.test.ts` — 1 site (was `40000 + Math.random()*10000`)
  - `tests/v2-4-2-tty-guard.test.ts` — 1 site (was `40000 + Math.random()*20000`)
- **Verification** — each affected test file run individually post-conversion; all pass.

#### Walked analogous surfaces — out of scope this round

- `tests/v2-1-cli-tooling.test.ts:34` (`RELAY_HTTP_PORT: "39988"`) — uses synchronous `spawnSync` inside `runRelay()` with the env baked at file scope. Conversion would require restructuring (`runRelay` → async, or per-test `beforeEach` allocation) and the CLI tests don't actually bind the daemon for most subcommands — the port is only consequential for `relay test` / `relay doctor` paths. Not on the known-flaky list. Documented as a v2.6.4+ residual; can revisit if it ever flakes.

### Dependabot vuln patches

`gh api .../dependabot/alerts` enumerated 5 open medium-severity alerts (initial estimate was 3 — actual is 5). All cleanly resolvable via `npm audit fix` (semver-compatible patch upgrades; no `package.json` change, only `package-lock.json`).

| # | package | from → to | scope | advisory |
|---|---|---|---|---|
| 7 | hono | 4.12.14 → 4.12.18 | runtime (transitive) | bodyLimit() bypass for chunked requests |
| 6 | hono | 4.12.14 → 4.12.18 | runtime (transitive) | hono/jsx unvalidated tag names → HTML injection |
| 5 | ip-address | 10.1.0 → 10.2.0 | runtime (transitive via express-rate-limit) | XSS in Address6 HTML-emitting methods |
| 4 | uuid | 11.1.0 → 11.1.1 | runtime | missing buffer bounds check in v3/v5/v6 with `buf` |
| 2 | postcss | 8.5.9 → 8.5.14 | development | XSS via unescaped `</style>` in Stringify |

Pre-fix: 5 moderate, 0 high. Post-fix: 0 vulnerabilities. `npm audit` (high+ threshold) gate-step continues to pass; this patch lowers the residual surface to zero.

Express-rate-limit also bumped (8.3.2 → 8.5.1) as a side-effect of the ip-address upgrade — within its declared semver range, no API change.

### Pre-publish gate

13/13 PASS. Test count unchanged (port hardening is an existing-test refactor, no new tests).

## v2.6.2 — 2026-05-06 — check-relay.sh agent_token capture regex fix (SSE-escape) + recovery-flow parsing

> An earlier working label was "v2.6.4"; this publishes as v2.6.2 (npm semver patch on v2.6.1). The label is a tracking convention, not a semver promise.

Closes a first-spawn bug a builder agent hit on 2026-05-06: first spawn against a v2.6.1-LIVE daemon registered correctly in DB but the per-instance vault file was never written. The cumulative v2.6.1 R1-R3 arc was supposed to close this exact failure mode, yet here it was on the first post-publish spawn.

### Root cause

The daemon serves MCP-over-HTTP responses in SSE-wrapped + JSON-stringified format. Concrete shape (verified via curl against the live `:3777` daemon):

```
event: message
data: {"result":{"content":[{"type":"text","text":"{
  \"status\": \"ok\",
  \"version\": \"2.6.1\",
  \"agent_token\": \"<token>\",
  ...
}"}]},"jsonrpc":"2.0","id":1}
```

Three things matter:
1. Outer envelope is SSE (`event: message\ndata: {...}`), not plain JSON. Server requires `Accept: application/json, text/event-stream` (forcing application/json only returns 406).
2. The `"text"` field's content is a JSON-stringified inner object — quotes are escaped as `\"`.
3. Inner JSON is pretty-printed: `\"agent_token\":` is followed by a SPACE then `\"<value>\"`.

Five `grep` patterns in `hooks/check-relay.sh` expected the un-escaped, un-spaced shape (`"agent_token":"..."`) and silently never matched against the actual bytes. The bug was strictly bigger than the originally-flagged two sites:

| line | pattern | failure mode pre-fix |
|---|---|---|
| 134 | `"auth_error":true` | health_check error detection silently never fires — operator never sees stale-token recovery prompts |
| 137 | `"auth_state":"<state>"` | auth_state extraction returns empty — recovery flow never enters the `recovery_pending` branch even when admin issued a recovery token |
| 164 | `"recovery_completed":true` | recovery branch never extracts the new token from response |
| 165 | `"agent_token":"<token>"` (recovery) | recovery never persisted to vault |
| 271 | `"agent_token":"<token>"` (register) | first-spawn never persisted to vault — the first-spawn symptom |

All five were in the SAME file with the SAME shape gap; the original report named the latter two but the **walk-analogous-surfaces** discipline surfaced the other three pre-fix. Without addressing all five, the recovery flow would have remained dead even after the originally-flagged sites were patched.

### Fix

`hooks/check-relay.sh` — all 5 patterns updated to match the actual SSE-escaped + space-after-colon shape. Token-shape charset narrowed from `[^"]*` to `[A-Za-z0-9_=.-]+` (mirrors `src/token-store.ts:67` `TOKEN_SHAPE_RE`) for defense-in-depth — tightening also defends against any future change in escaping that would otherwise pass-through corrupt bytes.

Verified against the live `:3777` daemon: pre-fix patterns match 0 times, post-fix patterns match 1 time (the actual bytes). Walked analogous surfaces in `scripts/migrate-existing-tokens-to-vault.sh`, `bin/spawn-agent.sh`, `src/cli/recover.ts`, `hooks/post-tool-use-check.sh`, `hooks/stop-check.sh` — none of them parse `agent_token` from MCP HTTP responses; only `check-relay.sh` does.

### Why v2.6.2 SR-D didn't catch this

`tests/v2-6-2-spawn-to-ready.test.ts` test SR-D claimed end-to-end coverage of "spawn → vault written → first MCP call authenticates" — but it parsed `register_agent`'s response with TS `JSON.parse()` (native; handles SSE+escape correctly) and wrote the vault from TS test code, bypassing the bash hook's grep/sed extraction entirely. Test path did NOT match shipped path. SR-D is kept as-is for its different coverage value (the prelude → stdio MCP env-inheritance path); v2.6.4's new test file replaces SR-D as the regression guard for the bash hook's response parsing.

SR-D's docstring updated to call out the divergence + point at the v2.6.4 test as the actual coverage.

### Tests — `tests/v2-6-4-hook-token-extraction.test.ts` *(new, 3 tests)*

Test path matches shipped path: invokes the actual `hooks/check-relay.sh` as a subprocess against a real `node dist/index.js` HTTP daemon. Vault is written by the hook (NOT by test code), so a regression in any of the 5 patterns surfaces immediately.

- **(T1)** first-spawn — hook calls `register_agent` via HTTP, parses SSE-wrapped response with the new pattern, writes vault. Pre-fix this would have failed at the vault-exists assertion (the first-spawn symptom).
- **(T2)** re-spawn — hook with valid env-token + agent already in DB → `SKIP_REGISTER` branch, vault content unchanged.
- **(T3)** recovery flow — admin revokes target with `issue_recovery=true` → operator sets `RELAY_RECOVERY_TOKEN` + invokes hook → hook detects `auth_error` (line 134) + reads `auth_state=recovery_pending` (line 137) → recovery branch fires, parses `recovery_completed:true` (line 164), extracts `NEW_TOKEN` (line 165), writes fresh vault. Exercises ALL 5 patched patterns end-to-end.

### v2.6.4 R1 — codex residual: tightened agent_token extraction regex to `{8,128}` length bounds

The R0 Codex SHIP verdict flagged that the extraction patterns at L174 + L292 used `[A-Za-z0-9_=.-]+` (any length) while `write_relay_token_to_vault` validates the same charset with `{8,128}` length bounds (mirrors `src/token-store.ts:67` `TOKEN_SHAPE_RE`). Contract inconsistency — extraction would match a malformed short token, then vault-write would reject it. R1 aligns the extraction regex to the same length bounds so the hook refuses what vault-write would refuse, before the round-trip.

- `hooks/check-relay.sh:174` (recovery) and `:292` (register) — `+` → `{8,128}` in BOTH the `grep -oE` pattern AND the `sed -E` substitution. Lines 140 (`auth_error`), 143 (`auth_state`), 173 (`recovery_completed`) are NOT agent_token extractions and stay as-is.
- `tests/v2-6-4-hook-token-extraction.test.ts` — new **(T4)** test exercises the actual shipped regex bytes by reading `hooks/check-relay.sh` at test load, extracting the L292 grep + sed via JS regex, then piping synthetic SSE fixtures through `bash` with those exact bytes. No regex re-implementation in TS — drift between this test and the shipped hook surfaces as a real failure. Cases: drift guard (assert pattern contains `{8,128}`), 3-char token rejected, 7-char rejected, 8-char accepted (boundary), valid 43-char accepted.
- Port-flake hardening (codex residual #2) — queued separately for v2.6.5; out of v2.6.4 R1 scope.

### Pre-publish gate

R0: 13/13 PASS. Test count: 1258 → 1261 (+3 from initial regression suite). R1: 13/13 PASS. Test count: 1261 → 1262 (+1 from T4 contract pin).

## v2.6.1 — 2026-05-06 — drift-grep guard extension to scan tests/

> An earlier working label was "v2.6.3"; this publishes as v2.6.1 (npm semver patch on v2.6.0). The label is a tracking convention, not a semver promise.

Small follow-up to the v2.6.0 publish-prep regression caught at the bumped state. `tests/v2-2-0-full-dashboard-smoke.test.ts:112` had a hardcoded `"2.5.0"` literal that slipped past the existing `src/`-only drift-grep guard. The `--full` gate caught it during publish-prep, but only at the bumped state — too late to be caught early in iteration. v2.6.3 adds a guard step so the regression is caught immediately on any future bump.

### Scope

- **`scripts/pre-publish-check.sh`** — new step `tests_drift_guard()` that reads the CURRENT package.json version and greps `tests/` for that exact literal as a quoted string. Fails the gate if any hits remain after excluding lines marked `// ALLOWLIST:` and comment-only lines (`//` / `*` prefix). Strategy is selective: only the CURRENT version triggers, not arbitrary X.Y.Z patterns. This keeps legitimate older-version fixtures (e.g. `"2.3.0"` in traffic-replay tests, `"2.4.0"` in protocol assertions / instance-metadata test data) un-flagged without forcing ALLOWLIST comments across ~50 lines.
- **Per-line allowlist** — any line ending in `// ALLOWLIST: <reason>` (or `# ALLOWLIST: <reason>`) is exempt. For the rare case where a test legitimately needs to assert against the current version literal (e.g. testing a migration whose expected output is "current version was X"). Use sparingly with a justification.
- **Error message** points the operator at the `package.json`-read pattern (mirror of `src/version.ts`):
  ```typescript
  const __pkg = path.resolve(path.dirname(fileURLToPath(import.meta.url)), "..", "package.json");
  const EXPECTED_VERSION = JSON.parse(fs.readFileSync(__pkg, "utf-8")).version;
  ```
- **`tests/v2-6-3-drift-guard-tests-scan.test.ts`** *(new, 5 tests)* — regression coverage that exercises the actual gate function via bash + awk extraction (test path matches shipped path; not a TS reimplementation). Cases: D1 baseline at HEAD → exit 0; D2 planted current-version literal → exit 1 + file:line + remediation; D3 ALLOWLIST-marked literal → exit 0; D4 older-version literal (selective scan) → exit 0; D5 literal in `//` / `*` comment → exit 0. Plant fixtures use `.fixture.ts` extension (not `.test.ts`) so vitest does not auto-load them; `try/finally` cleanup so a crashed test never leaks a fixture file.
- **`scripts/pre-publish-check.sh` header comment** updated to document the new step (5a in the gate-step list).

### Walked analogous surfaces

Per the walk-analogous-surfaces discipline: checked whether other directories also need scanning. Findings:

- **`hooks/`, `scripts/`, `bin/`** — at HEAD, zero current-version literals (`grep -rnE "[\"']2\.6\.0[\"']" hooks/ scripts/ bin/` returns nothing). These directories are bash; the v2.6.0 bump-test bug was specifically a TS test asserting `expect(...).toBe("X.Y.Z")` against a fresh server's reported version. The bash hooks read versions only via the daemon's HTTP `/health` endpoint or via the `relay` CLI, not via hardcoded literals. No regression found, no scan added.
- **`docs/`** — markdown documentation legitimately mentions versions in changelogs, migration notes, and version-specific instructions. Scanning would produce mostly-false-positives. Not added.
- **`extensions/vscode/`** — has its own `package.json` and version. Out of scope for v2.6.3; treat as a separate package if it ever needs the same guard.

The new `tests_drift_guard` is targeted; expansion to other dirs is deferred until a real bug surfaces there.

### Verification (3 cases)

- Case 1 (HEAD, no planted): gate green, step output `No tests/ drift — no hardcoded "2.6.0" literals (current package version)`.
- Case 2 (planted hostile literal): gate exits 1, stderr cites `tests/_planted_drift_test.test.ts:7:const HOSTILE_CURRENT_VERSION_LITERAL = "2.6.0";` plus the remediation block.
- Case 3 (planted removed): gate green again, identical to Case 1.

All three cases pinned in the regression test file (D1/D2/D3 plus D4/D5 selective-scan + comment-exclusion edge cases).

### Pre-publish gate

12-step gate at HEAD (with new step): all PASS. Test count: 1253 → 1258 (+5 from new regression test).

## v2.6.0 — 2026-05-05 — `relay mint-token` CLI + spawn-flow self-bootstrap + cross-platform parity

Adds an operator-side credential issuance path for external CLI agents whose safety monitors block the `register_agent` → use-returned-token sequence in a single response (Codex was the canonical case as of 2026-04-27). Pure-additive — zero behavior change for existing users; existing `register_agent` flow is unchanged.

### `relay mint-token <name>`

New CLI subcommand. Mints an agent token directly via filesystem access to the per-instance DB and prints the plaintext token ONCE to stdout. The operator exports it as `RELAY_AGENT_TOKEN` and launches the external CLI; the agent then authenticates on its first MCP call without ever invoking `register_agent` itself.

```bash
relay mint-token codex --role builder --capabilities build,test,audit
relay mint-token codex --force                      # rotate token, caps + role preserved
relay mint-token NAME --json                            # structured output for scripts
relay mint-token NAME --description "human-readable"    # discoverable in dashboard
```

- **`src/cli/mint-token.ts`** *(new)* — argv parsing, daemon-running stderr advisory, audit-log entry on every invocation (success or refused), human-readable + `--json` output formats. Mirrors `src/cli/recover.ts`'s shape end-to-end.
- **`src/db.ts`** — new sanctioned mutation helper `mintAgentToken(name, role, capabilities, opts)`. INSERT path mirrors `registerAgent`'s 13-column shape exactly so a minted-but-never-registered row is auth-layer-indistinguishable from a registered row. UPDATE path (under `--force`) rotates token only: caps and role preserved, `session_id` cleared, `agent_status='offline'`, auth-side fields (`previous_token_hash`, `rotation_grace_expires_at`, `recovery_token_hash`, `revoked_at`) zeroed.
- **`src/auth.ts`** — `BCRYPT_ROUNDS` is now exported (still `10`) so the regression test can pin the cost factor without duplicating the constant.
- **`bin/relay`** — `mint-token` wired into `SUBCOMMANDS` and `printHelp`. CLI count: 13 → 14.
- **`docs/agents/external-cli-setup.md`** *(new)* — full setup walkthrough: how the env-token pattern works, worked examples for Codex + Cursor (best-effort), token storage best practices (`direnv`, macOS keychain, rotation hygiene), audit-trail shape, cross-platform notes.
- **`README.md`** — new "External-CLI Token Mint (v2.6)" section after Lost-Token Recovery, cross-linking to `docs/agents/external-cli-setup.md`.

### Tests — `tests/v2-6-mint-token.test.ts` *(new)*

Spawns `node bin/relay mint-token …` against a throwaway DB and asserts on:

1. mint on new agent → row created, token-hash valid, plaintext token printed once.
2. mint on existing agent without `--force` → exit 2, clear stderr error pointing at `--force`, original token still authenticates.
3. mint on existing agent with `--force` → token rotates, old token no longer authenticates, new token does, caps and role preserved.
4. token shape is 43-char base64url (matches the existing `generateToken` convention).
5. bcrypt hash carries cost factor 10 (parses `$2a$10$…` prefix; pinned via the exported `BCRYPT_ROUNDS`).
6. `--json` flag emits parseable JSON with `success`, `token`, `name`, `agent_id`, `created`, `force`, `env_block` fields.
7. `audit_log` entry written on every mint (`tool='agent.token_minted'`), including the refused path.
8. `audit_log` entry written on the REFUSED path (existing-without-`--force`) too.
9. `--db-path` pointing at a non-existent directory → exit 2 with clean error.
10. `--help` prints usage, leaves DB alone.
11. Missing `<name>` argument → exit 1 with usage.
12. **(R1 regression)** `--json` still emits the daemon-running advisory to stderr — listener-backed test binds a real TCP listener and asserts both `stdout` parses as JSON and `stderr` contains the verbatim phrase `Daemon currently running on 127.0.0.1:<port>`. Closes the Codex R1 audit finding (P2 #2) (the prior `!args.json` exclusion suppressed the Item 3.7 safety signal for scripted callers).
13. **(R1 regression)** Human-readable mode also emits the advisory to stderr when a listener is bound. Pinned exact phrases verified against `src/cli/mint-token.ts:217-222`.

### Pre-publish gate

One new gate step (R1 add): **npm pack contents check** — `scripts/pre-publish-check.sh` now greps `npm pack --dry-run --json` output for every doc cross-linked from the CLI runtime (currently `docs/agents/external-cli-setup.md`) and fails if any are absent. Closes the Codex R1 audit finding (P2 #1): pre-R1, the `package.json.files` array omitted `docs/`, so the tarball lacked the doc despite `src/cli/mint-token.ts:317` and `README.md:654` both linking to it. The mint-token mutation itself routes through the sanctioned helper `mintAgentToken` in `src/db.ts`, so the existing sanctioned-helper guard passes unchanged. CLI smoke (`scripts/smoke-25-tools.sh`) is unchanged in scope (mint-token tested via dedicated integration test, not the smoke harness).

### Out of scope (deferred)

- Bulk operations (`relay rotate-all-tokens`, `relay list-stale-tokens`, `relay agents --filter ...`) — queued as v2.7.
- Cross-machine identity, namespace prefixing, hardware-backed token storage, GUI mint flow — Tether Cloud / Pro territory.

### v2.6.2 R2 — codex PATCH-THEN-SHIP (universal powershell.exe gate, generalize R1)

R1 verdict: PATCH-THEN-SHIP (Codex audit). R1 gated `cmd.exe` on `powershell.exe` presence — but wt.exe ALSO routes through `powershell.exe -NoExit -Command "..."` (v2.6.2 wrapped the inner shell to give the prelude a script context). Codex caught the same self-contradiction one sub-driver up: selecting wt without powershell creates a terminal that opens and immediately fails. Same bug class, second instance.

R2 generalizes: ALL three Windows sub-drivers require `powershell.exe`. Single universal gate.

- **`src/spawn/drivers/windows.ts:isAvailable`** — replaced the cmd-specific gate with a universal one. Verbatim:
  ```typescript
  function isAvailable(ctx: DriverContext, sub: WindowsSubDriver): boolean {
    // Universal: every sub-driver delegates the vault prelude to powershell.
    if (!ctx.hasBinary("powershell.exe")) return false;
    return ctx.hasBinary(BINARY_FOR[sub]);
  }
  ```
  JSDoc cites both Codex findings (cmd gap → wt gap → R2 generalization). The auto-fallback chain effective shape: wt|powershell|cmd, all predicated on powershell.exe.
- **`tests/spawn-drivers.test.ts`** — R1 auto-fallback chain test reshaped + 1 NEW wt-without-powershell rejection test mirroring R1's cmd pattern. The new chain test enumerates 4 positive cases (powershell present + each combo) + 4 negative cases (powershell missing). Other 9 wt-only test contexts updated to include powershell.exe (legitimate wt-with-powershell case). 74 → 75 spawn-drivers tests.
- **No functional regression on Windows boxes without powershell.exe** — daemon-side R2/R3 stdio-only fallback in `src/server.ts:resolveToken` still authenticates the agent on first MCP call from the per-instance vault. Operators on those boxes (rare; powershell.exe ships with every Windows since Vista) lose the operator-visible `printenv RELAY_AGENT_TOKEN` UX win, not the auth itself.

Generalization rule (added to discipline list): **when fixing a bug class, scope the fix to ALL similarly-shaped surfaces, not just the one flagged.** Walking the analogous code paths is part of declaring scope complete. R2 codifies this after R1 missed the wt path. (My v2.6.1 R3 work caught this rule for the resolveToken auth-oracle bug class via the parallel `resolveTokenForHealthCheck` instance — but didn't apply it when the v2.6.2 R0 → R1 change landed. Discipline: walk every similarly-shaped surface before writing the fix.)

### v2.6.2 R1 — codex PATCH-THEN-SHIP (Windows cmd.exe contract + revoke_token vault scrub)

R0 verdict: PATCH-THEN-SHIP (Codex audit). Two small fixes — one Codex P2 contract gap, one ergonomic call.

- **`src/spawn/drivers/windows.ts:pickSubDriver` — cmd.exe gated on powershell.exe presence.** Codex caught a self-contradiction in v2.6.2: `pickSubDriver` would auto-fall-through wt → ps → cmd, but the cmd sub-driver branch invokes `powershell.exe -NoExit -Command "..."` for the vault prelude. If cmd was selected because powershell was missing, the cmd window would open and immediately fail (`'powershell.exe' is not recognized`) before claude could run. R1 closes by gating cmd selectability on powershell.exe availability — both via auto-fallback AND via `RELAY_TERMINAL_APP=cmd` override. Effective auto-chain: wt → powershell → (cmd skipped). cmd remains selectable when an operator explicitly chooses it AND powershell.exe is present. Daemon-side R2/R3 stdio fallback in `src/server.ts:resolveToken` is still the universal safety net for any path where the launching shell can't run.
- **`src/tools/identity.ts:handleRevokeToken` — vault scrubbed alongside DB revoke.** Design decision recorded 2026-05-05. The security boundary already held via the auth_state check at `src/server.ts:870-878` (revoked / recovery_pending tokens refuse), so a stale token on disk was harmless. R1 aligns the mental model: revoke_token now leaves NO credential on disk for the agent. New `FileTokenStore.deleteSync()` in `src/token-store.ts` mirrors the `readSync ↔ read` pair — sync to avoid cascading the handler to async. Best-effort try/catch swallows IO errors so a vault-scrub failure can't roll back a successful DB revoke. Idempotent: ENOENT is a clean no-op (legitimate revocation of an agent that never wrote a vault entry).
- **Tests** — 2 NEW spawn-drivers tests (cmd-without-powershell rejection + auto-fallback chain assertion). 2 NEW recovery-flow tests (revoke without recovery also scrubs vault + best-effort no-op when vault missing). Previous recovery-flow vault assertion flipped from "vault is NOT auto-scrubbed" → "vault scrubbed by revoke_token". 1248 → 1252 tests.
- **Docs** — `revoke_token` tool description (`src/server.ts:592-598`) extended with the v2.6.2 R1 vault-scrub line. CHANGELOG entry above. Q7 hash bumped in `tests/v2-4-4-tool-description-quality.test.ts`.

### v2.6.2 — Windows FIX 1 closure + spawn-to-ready integration + final security boundary closure

The last piece in the v2.6.0 cumulative arc before publish. Three audit rounds (R1 → R2 → R3) progressively closed the resolveToken auth-oracle bug class; v2.6.2 closes the cross-platform parity gap deferred from R1 and adds the missing shipped-path integration tests that hid the original v2.6.0 dropped-token bug. After v2.6.2 ships green, `npm publish` makes the cumulative v2.6.0 (mint-token + spawn-flow self-bootstrap + cross-platform FIX 1 + integration coverage) live on npm.

#### v2.6.1 R2 — gate vault fallback to stdio transport (Codex audit)

R1 verdict: REJECT. R1's `resolveCallerNameForVault` honored `args.agent_name`, which let any unauthenticated HTTP caller name an agent and the daemon would obligingly read that name's vault file — turning the local file vault into a network-reachable auth oracle.

- **`src/server.ts:resolveToken`** — added `if (ctx.transport !== "stdio") return null;` gate before the vault read. HTTP requests fall through to `AUTH_FAILED`. `resolveCallerNameForVault` rewrites: drops the `args.agent_name` path entirely and uses only `process.env.RELAY_AGENT_NAME`, with explicit `MUST NOT honor args.agent_name` JSDoc.
- **Tests** — `tests/v2-6-1-token-store.test.ts` test 17 rewritten as a real stdio MCP subprocess test (positive path); test 17b NEW negative HTTP test pins the security boundary (HTTP daemon refuses vault fallback even with a valid vault file present).
- **Docs** — `SECURITY.md` precedence list extended with the gated 4th entry; `docs/agents/local-identity.md` lifecycle table adds an "HTTP daemon launched with `RELAY_AGENT_TOKEN` in its env" row; spawn_agent description updated for stdio-only scope (Q7 hash bump in `tests/v2-4-4-tool-description-quality.test.ts`).

#### v2.6.1 R3 — generalize stdio-only gate to env-token + close adjacent health_check oracle (Codex audit)

R2 verdict: REJECT. R2 closed only the vault half. Item 3 of the precedence chain (`process.env.RELAY_AGENT_TOKEN`) was still ungated — an HTTP daemon launched with `RELAY_AGENT_TOKEN` in its env would let any unauthenticated HTTP caller authenticate against the daemon's own env token. Same bug class as R1's vault, one layer up.

- **`src/server.ts:resolveToken`** — restructured. The chain is now split by trust origin: items 1-2 (`args.agent_token`, `X-Agent-Token` header) caller-presented and accepted on any transport; items 3-4 (env, vault) daemon-side and gated COLLECTIVELY on `ctx.transport === "stdio"` with a single boundary check. JSDoc fully enumerates the trust-origin split.
- **`src/tools/status.ts:resolveTokenForHealthCheck`** — adjacent instance found and fixed during R3 implementation per "close the full bug class" discipline. Same shape: read `RELAY_AGENT_TOKEN` env without transport gate. health_check is no-auth, but if a token validates it returns `agent_name` + `auth_state`, so an HTTP caller with no creds could probe what agent the daemon's env belongs to (information disclosure of the same bug class). Surfaced explicitly in the R3 completion report.
- **Tests** — `tests/v2-6-1-token-store.test.ts` test 17c NEW: HTTP daemon respawned with `RELAY_AGENT_TOKEN=<real-token>` in env, attack `get_messages` with `args.agent_name` and no caller-presented token must `AUTH_FAILED`; control with `X-Agent-Token` succeeds.
- **Docs** — `SECURITY.md` precedence list restructured with explicit caller-presented vs daemon-side split + general rule; `docs/agents/local-identity.md` bootstrap §2 expanded; spawn_agent description updated; Q7 hash bumped.

#### v2.6.2 — cross-platform FIX 1 closure (Windows) + spawn-to-ready integration

The Windows launching-shell prelude was deferred from R1 due to PowerShell escaping rigor. v2.6.2 closes it, plus the integration coverage that would have caught the original v2.6.0 dropped-token bug.

- **`src/spawn/drivers/windows.ts:buildVaultPreludePowerShell(agentName)`** *(new)* — single-line PowerShell snippet that mirrors the macOS bash + Linux ts shapes: pre-resolves the absolute vault path on the parent (Node) side, embeds as a literal in the launched PowerShell, runs `Test-Path` + `Get-Content -Raw` + shape regex `^[A-Za-z0-9_=.-]{8,128}$` + `$env:RELAY_AGENT_TOKEN = ...`. Single source of truth shared across all 3 sub-drivers — drift-free.
  - **wt.exe** sub-driver: changed args from `["-d", cwd, "claude"]` to `["-d", cwd, "powershell.exe", "-NoExit", "-Command", inner]` so the prelude has a script context. Pre-v2.6.2 wt.exe ran `claude` directly and had no shell context for prelude injection.
  - **powershell.exe** sub-driver: prepends prelude to the existing `-Command` inner string.
  - **cmd.exe** sub-driver: delegates the inner shell to `powershell.exe -NoExit -Command "..."` via `cd /D <cwd> && powershell.exe -NoExit -Command "<psInner>"`. cmd.exe lacks native `Get-Content` / regex match, so Option A — single PowerShell source of truth — is cleanest. The cmd `/K` window stays open after powershell exits.
- **`tests/spawn-drivers.test.ts` — +5 new prelude tests + 5 updated existing tests** for the new wt/cmd shapes. New `v2.6.2 — Windows FIX 1 vault prelude` describe block asserts: each sub-driver embeds the prelude (4 signature regex), prelude omitted when agent name fails the FileTokenStore allowlist (defense-in-depth), cross-driver invariant (all 3 sub-drivers embed the SAME prelude byte-pattern). 67 → 72 tests.
- **`tests/v2-6-2-spawn-to-ready.test.ts`** *(new, 4 tests)* — establishes the missing shipped-path test that hid the v2.6.0 dropped-token bug. Three vault-state scenarios (SR-A: no vault → prelude no-ops, SR-B: valid vault → prelude exports, SR-C: corrupted vault → prelude refuses to export) plus SR-D bonus end-to-end: register agent via HTTP daemon, capture token, write vault, run spawn-agent.sh DRY_RUN, exec captured prelude, then run `node dist/index.js` stdio MCP subprocess inheriting the prelude's env and assert `health_check` returns `token_validated:true` + `agent_name` + `auth_state:active`. macOS-gated; Windows / Linux equivalents covered by spawn-drivers.test.ts string-construction.
- **`tests/v2-6-2-hook-contracts.test.ts`** *(new, 15 tests)* — shipped-path contract tests for all 3 hook scripts (`check-relay.sh`, `post-tool-use-check.sh`, `stop-check.sh`). Closes the "untested code path in shipped flow" class. Each hook tested for: empty stdin / no DB → exit 0 silent; valid agent name + DB present + pending message → correct stdout shape (plain-text for SessionStart, single-line JSON with right `hookEventName` for PostToolUse / Stop); invalid agent name → silent exit 0 with stderr warning; daemon-down → graceful degrade; vault hydration path. Cross-hook invariants assert no partial JSON / stack traces ever leak to stdout.
- **`tests/v2-6-2-recovery-flow.test.ts`** *(new, 1 test)* — end-to-end recovery cycle: register admin (capabilities=["admin"]) + target → write target token to vault → admin `revoke_token({issue_recovery: true})` → original token now `AUTH_FAILED` → `register_agent` with `recovery_token` → fresh `agent_token` minted → write to vault → authenticated call with new token succeeds. Confirms `recovery_token` is single-use. **Surfaces a finding**: the spec's step 3 ("Verify vault is now scrubbed") doesn't match the implementation — `revoke_token` does NOT touch the vault; only the `relay recover` CLI scrubs it. The original token still fails auth post-revoke because the daemon's state check refuses `recovery_pending`, not because the vault is gone. Whether to add `vault.delete()` to `revoke_token` is a v2.6.3 / v2.7 design decision.
- **`tests/v2-6-1-token-store.test.ts` test 17d NEW** — health_check HTTP-negative regression: daemon respawned with `RELAY_AGENT_TOKEN=<real-token>` in env, attack `health_check` with no token must NOT include `agent_name` / `auth_state` in response (the v2.6.1 R3 codex non-blocking note about the parallel resolver in `src/tools/status.ts`). Pins the resolveTokenForHealthCheck fix end-to-end. Control with `X-Agent-Token` returns `token_validated:true` + agent metadata.
- **Docs** — `docs/agents/local-identity.md` adds a "Launching-shell vault prelude (cross-platform parity per v2.6.2)" subsection with platform table (macOS bash / Linux ts / Windows PowerShell) and Windows verification path. `SECURITY.md` extends pin-test references to include `:17b` (vault, R2), `:17c` (env-token, R3), `:17d` (health_check, v2.6.2). README + spawn_agent description unchanged from R3.

Test count delta v2.6.0 baseline → v2.6.2 STAGE: +26 tests across 5 files (5 new prelude tests in spawn-drivers + 4 vault-state tests in v2-6-2-spawn-to-ready + 15 hook-contract tests + 1 recovery-flow + 1 health_check oracle).

Pre-publish gate: 12/12 PASS (tsc + vitest 670+ tests + audit + build + extension TS + drift + sanctioned-helper + ip-classifier + 25-tool/CLI smoke + GitHub CI + npm pack + split-brain DB).



R0 verdict: REJECT. Codex caught the foundational gap: v2.6.1 R0 wrote the vault but never CONSUMED it on the daemon side. The SessionStart hook's `export RELAY_AGENT_TOKEN` only mutates the hook subprocess; the already-running stdio MCP server forked with whatever env it had at exec time, and never re-reads that env. So a builder agent's first MCP call still failed `AUTH_FAILED` even with a vault file on disk.

R1 closes the gap with two complementary fixes:

- **`src/server.ts` resolveToken vault fallback (FIX 2 — load-bearing, every platform).** Adds a 4th fallback after env. When env / args / X-Agent-Token are all empty, resolve the caller's name (from `args.agent_name` for HTTP or `process.env.RELAY_AGENT_NAME` for stdio) and sync-read the vault file. Sync `fs.readFileSync` of a single-line file is microseconds — the auth path stays sync. New `FileTokenStore.readSync` mirrors the async `read` exactly. This is the foundation that makes spawn-to-ready zero-touch on every platform: even on Windows where the launching shell can't easily hydrate env (deferred to v2.6.2 / future), the daemon still authenticates from the vault.

- **`bin/spawn-agent.sh` + `src/spawn/drivers/linux.ts` launching-shell vault prelude (FIX 1 — macOS / Linux UX win).** Before `exec claude`, the launcher reads the vault file into `RELAY_AGENT_TOKEN` so claude (and its stdio MCP server) inherit the token at fork time. Pre-resolves the absolute vault path on the parent side (no path-resolution in the launched shell — just `head -n 1 | tr -d '[:space:]' | grep -Eq <shape>`). Operator-visible `RELAY_AGENT_TOKEN` in `printenv` + non-MCP tooling (curl scripts) get the token for free. Windows is deferred — daemon-side fallback (FIX 2) handles every Windows sub-driver correctly, and a PowerShell/cmd prelude with full escaping rigor is its own audit cycle.

- **`hooks/_vault-helpers.sh` (NEW) — single source of truth for bash mirrors (FIX 3 — test contract closure).** v2.6.1 R0 inlined the bash helper functions in three hooks with a SHA byte-identity pre-publish check. Codex correctly flagged this as drift-blind: the test inlined a SECOND copy in vitest, and edits to the shipped hook would silently pass. R1 consolidates: hooks/_vault-helpers.sh defines `resolve_relay_db_path`, `resolve_relay_token_path`, `read_relay_token_from_vault`, `write_relay_token_to_vault` once. The 3 hooks + `scripts/migrate-existing-tokens-to-vault.sh` + `tests/v2-6-1-token-store.test.ts` all `source` the same file. Drift is now physically impossible. Removed the SHA-identity guard in the gate (it was solving the byte-identity problem that no longer exists).

- **`src/server.ts:374` corrected description.** The pre-R1 spawn_agent doc string claimed "already exported into the spawned child's `RELAY_AGENT_TOKEN` env" — false post-v2.6.1 R0. Updated to describe the vault-write-then-launcher-prelude flow + the daemon-side fallback. Triggers a tool-description hash bump in the v2.4.4 quality gate (Q7).

- **`docs/agents/local-identity.md` rewrite.** Lifecycle table replaced with a two-layer description (launching shell hydration + daemon-side fallback) + an updated scenarios table covering: parent-MCP-spawn, direct-spawn-agent.sh-without-pre-mint, re-spawn, bare claude start, lost vault, revoked token, recovery flow.

- **Tests — `tests/v2-6-1-token-store.test.ts` (was 15, now 19, +4 R1).** Test 12 now sources the actual `hooks/_vault-helpers.sh` instead of an inline copy. Test 12b adds the bash-write → TS-read direction for full bidirectional contract. Test 14b is a drift guard asserting all 3 hooks + migration script reference the helper file. Test 16 captures the spawn-agent.sh CMD via DRY_RUN=1, runs the prelude in a tmp shell, and asserts `RELAY_AGENT_TOKEN` in the spawned shell's env equals the vault content (FIX 1 e2e). Test 17 spawns an isolated HTTP daemon, registers an agent, writes the captured token to the vault, then makes a get_messages call WITHOUT supplying token via env / header / args.agent_token — auth must succeed via FIX 2 vault fallback (proves the foundation).

- **`tests/v2-4-5-stdio-per-instance-db.test.ts` updated.** The byte-identity Q-identity test is replaced with a "all 3 hooks source the helper, no inline copies" test — same drift discipline, structurally cleaner. The parameterized parity tests still extract `resolve_relay_db_path` and run it under controlled env, but now read it from `_vault-helpers.sh`.

- **`tests/spawn-drivers.test.ts` + `tests/v2-1-spawn-token-passthrough.test.ts` regex tightening.** The Linux launch-command pattern shifted from `^cd '...' && exec claude$` to `<vault prelude> cd '...' && exec claude$` — the end-anchor still pins the tail. The token-leak guard now compares against the actual minted token rather than a permissive shape regex (the prelude legitimately contains the public shape regex string).

Pre-publish gate: 12/12 PASS (tsc, vitest 1220 tests, npm audit, build, extension TS compile, drift guard, sanctioned-helper guard, ip-classifier guard, 25-tool smoke, GitHub CI green-gate, npm pack contents, split-brain warn).

#### Windows scope (deferred, judgment call flagged for audit)

Codex warned: "If you hit any architectural fork (e.g., Linux/Windows driver path requires more than the bash snippet), STOP-AND-SURFACE." Windows requires PowerShell + cmd snippets in `src/spawn/drivers/windows.ts` — non-trivial escaping inside `-Command "..."` strings. v2.6.1 R1 ships **macOS + Linux launcher preludes** AND **daemon-side vault fallback (FIX 2) for every platform**. Windows correctness is preserved: every Windows sub-driver (wt.exe / powershell.exe / cmd.exe) goes through `resolveToken` on every MCP call, which falls back to the vault when env is empty. The only thing Windows operators DON'T get is operator-visible `RELAY_AGENT_TOKEN` in `printenv` — the daemon authenticates without it. Full Windows launcher prelude queued as v2.6.2 follow-up (or its own dual-audit round).

### v2.6.1 stack — spawn-flow self-bootstrap (TokenStore + FileTokenStore)

Closes the latent v2.1 Phase 4j bug exposed 2026-05-04 when a child terminal spawned via `bin/spawn-agent.sh` without a pre-minted token: the SessionStart hook called `register_agent` over HTTP, the relay correctly minted a fresh `agent_token` and returned it in the response body, **and the script discarded the response** — 3-min spawn-to-broken-state. Latent since v2.1 Phase 4j; hidden until now because every prior spawn used the 5th-arg-token plumbing in `spawn-agent.sh`. Fix is a single elegant primitive — a per-instance file vault — that solves both first-spawn bootstrap AND respawn persistence with zero operator mediation, working across macOS / Linux / Windows from day one. Built on the canonical credential-helper pattern (Docker / Git / gh / AWS / kubectl all use this shape), with a pluggable `TokenStore` interface so v2.9+ can plug in OS-specific helpers (Keychain / Credential Manager / libsecret / 1Password / Vault) without breaking changes.

- **`src/token-store.ts`** *(new)* — `TokenStore` interface + `FileTokenStore` impl. Per-instance scoped path at `<instanceDir>/agents/<name>.token`; chmod 0o600 file, 0o700 parent; atomic write via tmp-file + rename so concurrent spawns of the same name converge cleanly; token shape validated on read against `/^[A-Za-z0-9_=.-]{8,128}$/` (mirrors the spawn-agent.sh allowlist + the bash hook helpers).
- **`hooks/check-relay.sh`** — vault-first bootstrap. If `RELAY_AGENT_TOKEN` is unset in env but a vault file exists, hydrate. On register_agent over HTTP, parse `agent_token` from the response body and persist to vault + export. Recovery flow also writes the new token to vault + exports inline (operator no longer needs manual paste).
- **`hooks/post-tool-use-check.sh` + `hooks/stop-check.sh`** — byte-identical mirrors of the new `resolve_relay_token_path` / `read_relay_token_from_vault` / `write_relay_token_to_vault` helpers + same vault-hydration block as check-relay.sh. The three hooks read from one vault.
- **`bin/spawn-agent.sh`** — 5th-arg-token slot dropped. `BRIEF_PATH` moved from positional 6 to positional 5. `RELAY_AGENT_TOKEN` no longer exported into the child shell via the AppleScript embedding; the hook hydrates from disk instead.
- **`src/spawn/dispatcher.ts` + `src/spawn/types.ts` + `src/spawn/drivers/*.ts`** — `token` parameter removed from `buildCommand` / `buildSpawnCommand` / `spawnAgent` (or kept as a `_legacyTokenSlot` placeholder for positional API stability with existing tests). macOS no longer pushes the token as arg 5; Linux/Windows no longer set `RELAY_AGENT_TOKEN` in `buildChildEnv`. The child env actively drops any inherited `RELAY_AGENT_TOKEN` from the parent's `RELAY_*` glob — child agent identity is exclusively vault-resolved.
- **`src/tools/spawn.ts`** — `handleSpawnAgent` keeps parent-side `registerAgent` (race-safe + name-collision UX), captures `plaintext_token`, **writes to the vault before driver dispatch** (instead of passing through the env channel). Driver failure rolls back BOTH the agent row AND the vault entry so a future spawn can succeed cleanly.
- **`src/cli/recover.ts`** — recovery now also scrubs the vault entry. After `relay recover <name>`, the next `register_agent` writes a fresh vault file via the SessionStart hook; no manual operator step.
- **`scripts/migrate-existing-tokens-to-vault.sh`** *(new)* — one-shot migration for operators who had `RELAY_AGENT_TOKEN` baked into `~/.zshrc` / `~/.bashrc` / `~/.envrc`. Reads env, writes vault, prints "you can unset the env var now."
- **`docs/agents/local-identity.md`** *(new)* — full lifecycle doc: vault model, bootstrap → persistence → recovery, cross-platform story, pluggable credential-helper interface, operator hygiene, audit trail.
- **`SECURITY.md`** — new "Local agent token storage (v2.6.1)" subsection in the threat-model section. Same threat-model bar as `~/.aws/credentials` / `~/.ssh/id_rsa` / `~/.config/gh/hosts.yml`.

#### Tests — `tests/v2-6-1-token-store.test.ts` *(new)*

- Unit: `FileTokenStore` read / write / delete on a tmp dir; round-trip preserves token; missing file returns null; malformed content (whitespace, length out of bounds, disallowed chars) rejected on read.
- POSIX perms: file 0o600, parent dir 0o700 after write (assertion skipped on Windows).
- Atomic write: concurrent writes never produce a half-file readable by `read()`.
- Integration: `handleSpawnAgent` writes the vault before driver dispatch; `RELAY_SPAWN_DRY_RUN=1` captures the constructed command and the vault file is verified on disk; rollback on driver failure scrubs both the agent row AND the vault file.
- Recovery: `relay recover` deletes the vault entry alongside the agent row.

`tests/v2-1-spawn-token-passthrough.test.ts` rewritten as legacy-rejection: assertions flipped to verify the env-var passthrough path is gone — `cmd.env.RELAY_AGENT_TOKEN` is now `undefined` for every driver, the macOS arg vector no longer carries the token, and `handleSpawnAgent` writes the vault file before dispatch.

#### Cross-platform parity

- macOS / Linux: POSIX file (chmod 0o600), parent dir 0o700. Verified by tests.
- Windows: NTFS file at `%USERPROFILE%\.bot-relay\instances\<id>\agents\<name>.token`. The `chmod` calls are best-effort no-ops; the parent dir under `%USERPROFILE%` already inherits a user-restricted ACL from the Windows profile defaults. Same threat model as `~/.aws/credentials` on POSIX. Documented in `SECURITY.md` + `docs/agents/local-identity.md`. v2.9+ Windows Credential Manager helper (pluggable via `TokenStore`) will move beyond profile-dir defaults.

## v2.5.0 — 2026-04-29 — Tether Phase 1 (MCP resource subscriptions + VSCode extension)

Introduces **Tether** — the operator-awareness layer for bot-relay-mcp. Free, bundled. The strict architectural line between FREE Tether Local (this repo) and FUTURE Tether Cloud / Pro (paid, separate repo) is documented in `docs/tether-roadmap.md` — features that need a server, third-party API, or cross-machine sync do NOT belong here.

Phase 1 ships three pieces:

### Part S — MCP resource subscriptions

The `relay://inbox/<agent_name>` resource is now subscribable per the MCP spec. Any MCP-aware client (Claude Code CLI, VSCode panel, Cursor, Cline, etc.) subscribes via `subscribe` and receives `notifications/resources/updated` pushes when the agent's inbox mutates.

- **`src/mcp-resources.ts`** — extended with dynamic per-agent `relay://inbox/<name>` descriptors (one entry in `listResources()` per currently-registered agent), and a `buildInboxSnapshot(agentName)` helper that surfaces `pending_count`, `total_count`, `last_message_at`, `last_message_from`, `last_message_priority`, and a 200-char `last_message_preview` with a `last_message_truncated` flag.
- **`src/mcp-subscriptions.ts`** *(new)* — central per-Server-instance subscription registry. `subscribe(uri, server)` adds, `unsubscribe(uri, server)` removes, `unsubscribeAllForServer(server)` cleans up on transport close, and an internal listener on the inbox event bus fans out `server.sendResourceUpdated({uri})` to every subscribed Server.
- **`src/inbox-events.ts`** *(new)* — module-singleton EventEmitter so DB write paths and the subscription registry share a stream without circular imports. Three event reasons: `message_received`, `broadcast_received`, `message_read`.
- **`src/db.ts`** — `sendMessage` (both system + normal branches), `broadcastMessage` (per-recipient fan-out), and `getMessages` (drain pending → read) now emit inbox events via `emitInboxChanged`. Emission lives AFTER the SQL commit so a subscriber's re-fetch sees fresh state.
- **`src/server.ts`** — `resources: { subscribe: true }` capability declared; `SubscribeRequestSchema` + `UnsubscribeRequestSchema` handlers wired; `server.onclose` patched to drop subscriptions on transport tear-down.

### Part E — VSCode extension v0.1

`extensions/vscode/` ships the first IDE consumer of the new subscription primitive. Activation is `onStartupFinished`. Status bar shows `Tether: <pending_count> | last <relative time>` with severity coloring (gray ≤0, yellow 1-3, red 4+); click opens a read-only webview with the last message preview; optional `autoInjectInbox` writes `inbox\n` to the integrated terminal whose name matches the agent's name. Settings live under `bot-relay.tether.*` with VSCode setting > env var > default precedence. Marketplace publish steps are documented for the maintainer to run when ready.

### Part D — documentation

- **`docs/tether-roadmap.md`** *(new)* — canonical free-vs-paid architectural line. Phase 2 features explicitly named out-of-scope-of-this-repo.
- **`extensions/vscode/README.md`** *(new)* — install + config + troubleshooting for the extension.
- vsce publish checklist + manual smoke verification before publish (developer-facing, not shipped).
- **`README.md`** — new "Tether (v2.5)" section pointing at the bundled extension + the roadmap doc.

### Tests

- **`tests/v2-5-mcp-subscriptions.test.ts`** *(new)* — 8 cases via real `Client` ↔ real `Server` over the SDK's `InMemoryTransport`. Every assertion answers "if the contract drifted, would this fail?" with yes. Covers single-subscriber receive, unsubscribe stops, multi-subscriber fan-out, cross-agent isolation (subscriber for X must NOT see events for Y), `readResource` snapshot shape, drain-fires-message_read, server-close cleanup, and broadcast fan-out per recipient.
- **`tests/v2-5-tether-format.test.ts`** *(new)* — 6 unit cases for the VSCode extension's pure-function format helpers (status-bar text, severity buckets, relative-time formatter, toast text). Full UX integration (booting a real VSCode instance via `vscode-test`) is heavyweight + CI-only — surfaced as a manual checklist instead.

### Protocol

`protocol_version` stays at `2.4.0` — adding subscription support is additive (the resources capability flag flips from `{}` to `{ subscribe: true }`), the existing tools/prompts/resources surface is byte-identical for non-subscribing clients.

**Total after v2.5.0: 1180 tests pass** (1166 from v2.4.5 + 14 new across Parts S + E).

### R1 patch round (Codex audit, 2026-04-29)

R0 verdict: REJECT. Codex caught a foundation-level gap and 3 mechanical bugs. R0's tests asserted the in-process subscription pipeline (InMemoryTransport) but the SHIPPED path is HTTP, and the HTTP daemon was strictly stateless — every GET /mcp returned 405. The VSCode extension's `StreamableHTTPClientTransport` requires a stateful SSE GET to receive `notifications/resources/updated` frames; the whole Tether UX was structurally broken in the actual deployment. The test path must match the shipped path — an InMemoryTransport test is a proxy, not a contract.

Decision: **Option A** — make the HTTP daemon support stateful sessions alongside the existing stateless POST. Aligns with v2.3 federation direction (which needs stateful anyway), works for every HTTP MCP client (Cursor, Cline, future tooling), additive (existing stateless POST clients unchanged).

#### Architectural fix — stateful HTTP MCP sessions

- **`src/transport/http.ts`** — added a per-port session map keyed by SDK-generated UUID. POST `/mcp` now branches on three signals:
  - Request carries `mcp-session-id` header → route to existing session's transport. 404 if session unknown.
  - Body has `method:'initialize'` AND `Accept: text/event-stream` → mint new stateful session via `sessionIdGenerator: () => randomUUID()`. The transport assigns the id, we register the entry in the map AFTER `handleRequest` returns, response carries the id back per spec.
  - Otherwise → original stateless one-shot transport (current smoke + curl behaviour preserved verbatim).

  GET `/mcp` now opens the SSE stream against an existing session. Missing `mcp-session-id` → 400; unknown session → 404. DELETE `/mcp` is the canonical client-initiated session close — `transport.close()` triggers our `onclose` hook which removes the entry from the map and drops every subscription that session held via `unsubscribeAllForServer(server)` from the v2.5.0 R0 plumbing.

  Memory bound: `RELAY_HTTP_MAX_SESSIONS` (default 64) caps the map size; new sessions over the cap return 503. `RELAY_HTTP_SESSION_IDLE_SECONDS` (default 300) defines how long a session can sit idle before the 60-second reaper sweeps it. The reaper interval is `unref`'d so it never blocks process shutdown.

- **`tests/http.test.ts`** — pre-R1 pinned the "GET /mcp = 405" contract. Updated to the new contract: GET without `mcp-session-id` → 400; GET with unknown session → 404. Old assertion would have been a false-positive on the new transport.

- **`tests/v2-5-mcp-subscriptions-http.test.ts`** *(new)* — REAL-HTTP integration test. Per R1's MED 3, the load-bearing subscription test must exercise the SHIPPED path:
  - Spawns the actual `dist/index.js` daemon as a subprocess (no in-process Server reuse).
  - Connects via the SDK's `StreamableHTTPClientTransport` (the same transport the VSCode extension uses).
  - Subscribes to `relay://inbox/<X>`, triggers a real `send_message` via the stateless POST (different transport — proves the cross-session bus actually fans out), asserts a real `notifications/resources/updated` frame arrives.
  - Q-HTTP-1: subscribe → real frame end-to-end. Q-HTTP-2: cross-session isolation — subscriber for X is NOT woken by message to Y. Q-HTTP-3: DELETE `/mcp` closes the session and subsequent traffic on the id 404s.

#### Mechanical fixes

- **R1 #2 — extension compile failure**. `extensions/vscode/src/extension.ts:88` accessed `result.contents[0]?.text`, which fails strict-mode (TS2339) because the SDK's contents array union is `{text} | {blob}`. Fix: `if (!first || !("text" in first) || typeof first.text !== "string") return null;` type-guard. R0 didn't catch this because the pre-push gate didn't compile the extension subdir; that gap is closed in R1 (see "Pre-push gate" below).
- **R1 #3 — endpoint precedence bug**. The R0 inline `(cfg.get(...) || "").trim() || env.X ? <env-string> : <fallback>` parsed as `((cfg||"") || env) ? <env-string> : <fallback>` because `?:` binds after `||`. Result: VSCode-configured endpoint was silently ignored when env was set; partial env (RELAY_HTTP_PORT without HOST) produced `http://undefined:3777`. Extracted to `extensions/vscode/src/config.ts:resolveTetherConfig()` — a pure function unit-tested via a 10-case matrix in `tests/v2-5-tether-format.test.ts` covering VSCode setting > env > default precedence for endpoint, agentName, and agentToken.
- **R1 #4 — inbox preview bypassed decrypt-on-read**. R0's `buildInboxSnapshot()` SELECTed `messages.content` directly and surfaced the raw column. With `RELAY_ENCRYPTION_KEY` active, MCP subscribers + the VSCode webview saw `enc1:<ciphertext>` instead of plaintext. Fix: route through `decryptContent()` (the same helper `getMessages` line 2714 + `getMessagesSummary` line 2772 use). Q9 in `tests/v2-5-mcp-subscriptions.test.ts` pins the contract — encryption-active case asserts plaintext preview + `not.toContain("enc1:")`.

#### Pre-push gate update

- **`scripts/pre-publish-check.sh`** — added "extension TS compile (extensions/vscode)" step. Runs `npx tsc -p . --noEmit` in the subdir; one-time `npm install` if `node_modules/` absent. Catches TS2339-class breakage that escaped R0's gate. The order is `tsc → vitest → npm audit → build → extension-compile → drift/sanctioned-helper guards → 25-tool smoke → CI green-gate → split-brain warn` — fail-fast on the cheap checks first.

**Total after R1: 1187 tests pass** across 103 files (1172 from R0 + 10 endpoint-precedence + 1 encryption decrypt + 3 real-HTTP integration + 1 http.test.ts net delta — the existing 405 case was rewritten to assert 400 + a new 404-on-unknown-session case was added; old assertion would have been a false-positive on the new transport).

#### Discipline notes

- **Citation discipline** — re-read `src/transport/http.ts:1225-1267` verbatim before designing the dual-path router; verified SDK `StreamableHTTPServerTransport` semantics from `node_modules/@modelcontextprotocol/sdk/dist/cjs/server/streamableHttp.d.ts:30-52` (stateful vs stateless mode rules) rather than spec memory.
- **Test path matches shipped path** — `tests/v2-5-mcp-subscriptions-http.test.ts` IS the R1 contract test. The R0 InMemoryTransport file stays as a fast unit-level check; the HTTP integration test is now the load-bearing one.
- **STOP-AND-SURFACE on architectural forks** — surfaced the Option A/B/C choice with explicit recommendation + reasoning before coding. Option A was picked. R0 violated this once (got lucky on subscription state — MCP spec mandated per-session); R1 didn't repeat the violation.

### Hall of Fame

- **Design rationale** — bundle for velocity, then split free/paid cleanly (the R1 Option A call). The strict free/paid line is a deliberate architecture choice.
- **Codex** — diagnosed the v2.5 R0 HTTP push gap from line numbers in `src/transport/http.ts` rather than running the test suite. The InMemoryTransport-vs-shipped-path failure mode is exactly the test-path-must-match-shipped-path discipline; Codex caught it on the first pass.
- **The MCP resource subscription primitive** — shipped as a deferred capability in v2.4.0; finally getting used end-to-end in v2.5.0 R1.
- **The TTY guard regression in v2.2.1** — the empirical evidence that operator-side notification was a real product gap, not a hypothetical one. That's what made Tether load-bearing instead of nice-to-have.

## v2.4.5 — 2026-04-28 — stdio/hook split-brain DB hotfix

Hotfix for a multi-agent coordination bug Codex caught during the v2.4.4 R2 audit. v2.4.0 shipped per-instance local isolation; the HTTP daemon's startup path correctly routed through `resolveInstanceDbPath()`, but several auxiliary entry points kept hardcoding `~/.bot-relay/relay.db`. An operator running with an active per-instance setup ended up with silent split-brain: the daemon wrote per-instance, the SessionStart hook read legacy, and an agent registered through one path was invisible from the other. The symptom that exposed it: Codex couldn't authenticate via MCP because its hook was delivering from legacy while its agent row lived per-instance.

### Fixed

- **`hooks/check-relay.sh`** — SessionStart hook's `DB_PATH` resolution now mirrors `src/instance.ts:resolveInstanceDbPath()` exactly. New bash helper `resolve_relay_db_path()` checks `RELAY_DB_PATH` → `RELAY_INSTANCE_ID` → `~/.bot-relay/active-instance` (symlink AND regular-file forms) → legacy fallback. Path-traversal defended via the same `[A-Za-z0-9._-]+` allowlist `instance.ts:instanceDir()` enforces.
- **`hooks/post-tool-use-check.sh`** — same fix to the PostToolUse mailbox-check fallback path. The HTTP path was already correct (the daemon resolves per-instance); only the sqlite-direct fallback was hitting legacy.
- **`src/cli/doctor.ts`** — `getDbPath()` and `getConfigPath()` helpers were hardcoding `os.homedir() + '.bot-relay/relay.db'` (and `config.json`). Now route through `resolveInstanceDbPath()` / `resolveInstanceConfigPath()`. Pre-v2.4.5 `relay doctor` printed PASS against the wrong file, hiding the very split-brain it exists to surface.

### Added

- **`scripts/pre-publish-check.sh`** — non-blocking WARN step that compares agent counts between the legacy DB and the active per-instance DB. Fires when both have agents (the fingerprint of a stale npx-cached bot-relay-mcp under `~/.npm/_npx/<hash>/node_modules/bot-relay-mcp/` writing to legacy in parallel with a current daemon serving per-instance — exactly the deployment state we found on the maintainer's machine while diagnosing this bug). Pure WARN: never blocks publish, since the fix is operator-side (kill the stale process, clear the npx cache, unset stray `RELAY_DB_PATH`).
- **`tests/v2-4-5-stdio-per-instance-db.test.ts`** — 7 cases pinning the contract:
  - Q1: bash hook resolver falls back to legacy when no env + no symlink.
  - Q2: bash hook resolver follows the active-instance symlink target.
  - Q3: bash hook resolver honours `RELAY_INSTANCE_ID` env.
  - Q4: bash hook resolver gives `RELAY_DB_PATH` highest priority.
  - Q5: bash hook resolver rejects malformed active-instance content (path-traversal defense-in-depth).
  - Q6: TS-side `resolveInstanceDbPath()` returns the per-instance path under multi-instance mode (doctor surrogate — doctor calls this exact function).
  - Q7: two TS DB handles under the same `RELAY_INSTANCE_ID` see the same rows (parity surrogate proving the resolver IS the single resolution gate).

The hook tests extract `resolve_relay_db_path()` from `check-relay.sh` via regex + invoke it under a controlled `$HOME`, so the contract test runs the actual bash code shipped to operators rather than a duplicate.

### Out of scope (deliberately deferred)

- Default-transport flip from stdio → HTTP. Bigger UX change, not a hotfix.
- Migration command to copy legacy data into a per-instance DB. Operators who want the migration can `relay backup` → `relay init --instance-id=<x>` → `relay restore` today; a one-shot `relay migrate-legacy-to-instance` is a v2.7+ candidate.
- TS-side `src/db.ts:getDbPath()` already routed through `resolveInstanceDbPath` in v2.4.0; v2.4.5 confirmed it via Q7 but did not need to touch it.

### Cross-platform parity

The bash helper uses `readlink` (POSIX, available on macOS + Linux) for the symlink branch and `head -n 1 / tr -d` for the regular-file branch. Windows operators run the TS surface (which already uses `fs.readlink` + `path.basename` — see `src/instance.ts:resolveActiveInstanceId()`); the bash hooks are macOS/Linux-only by design.

**Total after v2.4.5: 1144 tests pass** (1137 from v2.4.4 + 7 new).

### R1 patch round (dual-5.5 audit, 2026-04-28)

Both 5.5 auditor instances ran independent passes on the R0 branch. Each found different problems — useful confirmation that even same-model auditors diverge by sampling. Consolidated R1 patches:

- **HIGH 1 — `hooks/stop-check.sh` had the same legacy hardcode.** R0 only patched `check-relay.sh` + `post-tool-use-check.sh`; the Stop hook's sqlite-fallback path kept reading `~/.bot-relay/relay.db`. Since `src/cli/generate-hooks.ts` installs Stop by default, every Claude Code session with hooks would have hit the same bug v2.4.5 was supposed to fix. Stop hook now carries the byte-identical `resolve_relay_db_path()` helper.
- **MED 2 — bash mirror now honors `RELAY_HOME` + raises on malformed instance_id.** TS's `botRelayRoot()` reads `RELAY_HOME` (test seam) before falling back to `$HOME/.bot-relay` (`src/instance.ts:70`); R0 hooks hardcoded `$HOME/.bot-relay`. TS's `instanceDir()` throws on `instance_id` failing the `[A-Za-z0-9._-]+` allowlist; R0 hooks silently fell back to legacy. Both gaps closed: `${RELAY_HOME:-$HOME/.bot-relay}` for the root, and stderr error + `return 1` on a malformed id.
- **MED 3 — Q6 / Q7 are now real-subprocess contract tests.** R0 Q6 unit-asserted `resolveInstanceDbPath()` directly with a "doctor surrogate" comment; R1 spawns the real `node bin/relay doctor` subprocess and asserts the per-instance path appears in its rendered output (and the legacy path doesn't). R0 Q7 used `closeDb() + initializeDb()` in the same process to simulate "two processes"; R1 now spawns a separate `node` child that reads the DB written by the parent — actual cross-process parity.
- **MED 4 — Q1–Q5 parameterized over all three hooks.** R0 only extracted the resolver from `check-relay.sh`. The HIGH 1 Stop bug literally existed because the test never invoked the resolver from `stop-check.sh`. R1 loops the test cases over `[check-relay.sh, post-tool-use-check.sh, stop-check.sh]`, plus a new Q-identity case asserts the three function bodies are byte-identical via SHA-256 — drift is now physically impossible.
- **MED 5 — Windows hook story documented + warned.** The three `.sh` hooks are bash-only (POSIX `readlink`, `sqlite3`, `python3`). v2.4.5 ships an explicit non-support note: Windows operators run inside WSL (recommended) or skip hook install + use the HTTP transport. `relay generate-hooks` emits a stderr WARNING on `process.platform === 'win32'`. PowerShell mirrors deferred until inbound demand surfaces — `docs/multi-instance.md §'Windows hook story'` carries the rationale.

**Total after R1: 1158 tests pass** (1137 from v2.4.4 + 21 new — the R0 cut had 7; R1 grew to 21 because Q1–Q5 fan out across 3 hooks plus Q-identity, Q5b RELAY_HOME, real-subprocess Q6, real-subprocess Q7).

### Hall of Fame

- **Codex** — diagnosed the split-brain via direct DB inspection during its first session. Sharp catch.
- **Codex (×2 in R1)** — both auditor instances independently caught different gaps; the dual-audit pattern works. HIGH 1 (Stop hook) would have shipped uncaught from a single-auditor pass.
- **Maintainer feedback** — pushed back that a non-user-friendly path is a bug to fix, which elevated this from a one-off workaround to a proper architectural fix.

## v2.4.4 — 2026-04-27 — tool description quality (Glama A-tier push)

Pure docs release. Zero behavior change. Every one of the 30 MCP tool descriptions audited and rewritten against Glama's Tool Definition Quality Score (TDQS) criteria so the public score is gated by the surface, not by three thin one-liners. Glama scores TDQS as 60% mean + 40% MIN, so a single 30-character `delete_webhook: "Delete a webhook subscription by ID."` was actively dragging the score; that's the exact pattern this release closes.

### Rewrote

All 30 tool descriptions in `src/server.ts` now follow the same structured shape, ~600–1500 chars each:

- **Purpose** — single-sentence what.
- **When to use** — disambiguation against the related tools (`send_message` vs `broadcast` vs `post_to_channel`, `post_task` vs `post_task_auto`, `get_messages` vs `get_messages_summary` vs `peek_inbox_version`, etc.).
- **Behavior** — side effects, state machine, auth requirements, version notes.
- **Returns** — the success-shape object literal so callers don't have to read `src/types.ts` to know what comes back.
- **Errors** — the specific `error_code` strings the dispatcher emits, each with its trigger condition.

Tools that needed the most lift: `delete_webhook`, `list_webhooks`, `leave_channel` (one-liners pre-v2.4.4), plus `send_message`, `broadcast`, `create_channel`, `join_channel`, `post_to_channel`, `post_task`, `get_tasks`, `get_task`, `register_agent`, `discover_agents`. Tools already in good shape (`rotate_token`, `get_standup`, `peek_inbox_version`, `expand_capabilities`) got a format-consistency pass so the whole surface reads uniformly.

### Filled parameter-description gaps

Two parameter schemas in `src/types.ts` lacked `.describe()` calls — caught by the new quality-gate test:

- `register_agent.force` — escape hatch for re-registering an actively-held name. Description now documents the duplicate-name collision check it bypasses + when it's safe to use.
- `get_standup.filter` — narrowing object on the standup snapshot. Description covers the AND-shape across `agents`/`roles` and the `include_offline` flip.

### Tests

- `tests/v2-4-4-tool-description-quality.test.ts` — 7 quality-gate cases that hit `tools/list` against a freshly-spawned `dist/index.js` (the same surface a Glama scanner sees) and assert: every description ≥300 chars (Q1), mentions "When to use" (Q2), mentions Returns or Errors (Q3), mentions Behavior (Q4), tool name follows `verb_noun` convention (Q5), every input parameter has a description (Q6), and the locked tool surface is exactly 30 named tools with stable description hashes (Q7) so future PRs that touch any description surface the diff in review.

**Total after v2.4.4: 1137 tests pass** (1130 pre-v2.4.4 + 7 new).

### Out of scope (deliberately)

No new tools (Glama Completeness anti-pattern: adding tools just to score). No behavior changes. No schema changes. No tool reorganization. The `protocol_version` stays at `2.4.0` because the MCP tool surface contract (names, parameters, return types) is byte-identical — only descriptions changed.

### Cross-platform parity

Pure string edits in TypeScript source. No runtime behavior, no syscalls, no platform branches.

## v2.4.3 — 2026-04-27 — pre-publish `npm audit` resilience

The CI green badge on main has to track code health, not npm registry weather. Pre-v2.4.3 it tracked both — when the v2.4.0 main merge ran post-tests, the legacy `/-/npm/v1/security/audits/quick` endpoint returned 400 ("This endpoint is being retired. Use the bulk advisory endpoint instead.") and the public CI badge swung red even though all 1099 tests passed. Same commit's branch CI was green a few minutes earlier — pure registry-side flake. Goal: registry-side availability must not turn the public CI badge red on every push.

### Added

- **`scripts/audit-with-retry.sh`** — resilient wrapper around `npm audit --json --audit-level=$LEVEL` (modern bulk-advisory endpoint on npm 10+). Classifies outcomes into three buckets: clean (exit 0), real high+ vuln finding (exit 1, no retry), or transient registry-side error. Transient errors retry up to 3 times with 5/15/30s backoff; if all 3 attempts hit transient errors, the wrapper soft-fails to exit 0 with a loud WARN. Real high+ findings still exit 1 immediately. Unknown / malformed responses with no transient marker also exit 1 (don't silently skip a new failure mode). Test-injection seam (`RELAY_TEST_AUDIT_CMD`) lets tests mock the registry without touching the network.
- **`scripts/pre-publish-check.sh`** — `npm audit (high+)` step now invokes the wrapper instead of calling `npm audit` directly. Behavior on a clean registry is unchanged (still gates on real high+ vulns).
- **`.github/workflows/ci.yml`** *(R1)* — same rewire applied to the CI workflow. R0 only patched the pre-publish gate; the CI step itself stayed raw, so the public main badge was still exposed to the same registry endpoint flakes the wrapper exists to absorb. Both invocations now go through the wrapper.
- **`.github/dependabot.yml`** *(R1)* — Dependabot configuration covering both the `npm` and `github-actions` ecosystems on a weekly Monday schedule. Repo-side `vulnerability-alerts` and `automated-security-fixes` settings flipped on at the same time so Dependabot's defense-in-depth claim in this CHANGELOG is verifiable (`gh api repos/<owner>/<repo>/vulnerability-alerts` returns 204; `automated-security-fixes` returns `{"enabled":true,"paused":false}`).

### Transient classifier (narrowed in R1)

Codex caught that R0's classifier blanket-matched HTTP 4xx, which soft-failed through 401 / 403 / 404 — those are deterministic auth/permission/config problems, not registry flakes. The R1 classifier is narrow:

- 5xx — server errors → transient
- 408 Request Timeout → transient
- 429 Too Many Requests → transient (backoff is the correct response)
- `ENETWORK` / `EAI_AGAIN` / `ECONNRESET` / `ECONNREFUSED` / `ETIMEDOUT` / `ENOTFOUND` → transient transport
- `endpoint is being retired` (the explicit npm sunset signal — the v2.4.0 main repro) → transient
- `fetch failed` / `socket hang up` (node-fetch transport messages) → transient
- 4xx other than 408/429 (incl. 401, 403, 404, 410) → **not** transient → exit 1

### Why soft-fail is safe

`npm audit` is one input among many for security gating. **Dependabot is enabled on the repo** (verified via the GitHub API after R1 — see `.github/dependabot.yml` plus the repo-side `vulnerability-alerts` + `automated-security-fixes` toggles), giving an independent network path with no shared failure mode with the audit endpoint. The wrapper exits 1 the moment a real high+ vuln finding parses out of the JSON metadata. The soft-fail only triggers when **three consecutive attempts** all hit narrowly-classified transient transport errors — at that point the registry itself is unreachable, blocking the CI badge would still not surface a new advisory, and we'd rather see the warning in CI logs than a red badge on a public repo with passing tests.

### Tests

- `tests/v2-4-3-pre-publish-audit-resilience.test.ts` — 11 cases (R0: 6, R1: +5) mocking `npm audit` via the test-injection seam. R0 cases: clean first try (exit 0), real high+ finding (exit 1, no retry), the exact 400/"endpoint is being retired" repro from the v2.4.0 main red CI (3 retries → soft-fail to 0), 503 with backoff (3 retries → soft-fail to 0), malformed JSON with no transient marker (exit 1), transient-then-clean (1 retry → success at attempt 2, no soft-fail). R1 cases: 401 Unauthorized → exit 1 immediately, 403 Forbidden → exit 1 immediately, 404 on the audit endpoint → exit 1 immediately, 408 Request Timeout 3x → soft-fail, 429 Too Many Requests 3x → soft-fail.
- `tests/v2-4-3-ci-audit-bypass-guard.test.ts` *(R1)* — sweeps `.github/workflows/*.yml` for raw `npm audit` invocations outside the wrapper. Same drift-grep shape as the pre-publish gate's existing guards. Catches the exact regression Codex flagged: the wrapper exists but the CI step bypasses it.

**Total after v2.4.3 R1: 1137 tests pass** (1124 pre-v2.4.3 + 6 R0 + 5 R1 audit cases + 2 R1 bypass-guard cases).

### Cross-platform parity

Pure bash + `npm` + `python3` (Python is preinstalled on every CI runner the repo targets). The audit step runs on Ubuntu CI runners only; nothing platform-specific in the wrapper itself.

## v2.4.2 — 2026-04-24 — stdio TTY guard refinement (closes v2.2.1 oversight)

Fixes a plug-and-play regression introduced in v2.2.1: the stdio TTY guard (`src/index.ts`) exited immediately when stdin was not a TTY, which is the default configuration for every legitimate MCP client (Claude Code, Cursor, Cline, …). Every new MCP spawn after v2.2.1 silently failed until the operator added `RELAY_SKIP_TTY_CHECK=1` to their `~/.claude.json` `mcpServers.bot-relay.env` block. The regression was masked for six releases by the long-running pre-v2.2.1 daemon, which kept serving existing operators without re-spawning.

### Fixed

- **`src/index.ts` TTY guard** — replaced the immediate-exit branch with a 1500ms grace window. If any bytes arrive on stdin during the window, treat as a legitimate MCP client (MCP clients send their `initialize` frame within the first few hundred ms), cancel the guard, `unshift()` the received chunk so the MCP transport downstream reads the frame unchanged, and proceed. If the timer expires with zero bytes, exit with the existing code 3 + helpful message (background-daemon-attempt detection preserved). Added `RELAY_TTY_GRACE_MS` as an override so tests can drive the grace tight without changing production behavior.
- **`~/.claude.json` workaround is now redundant.** Operators can remove `"RELAY_SKIP_TTY_CHECK": "1"` from their `mcpServers.bot-relay.env` block after upgrading. Leaving it in place also works — the env var stays as an explicit bypass for test harnesses whose first write comes after the grace window.

### Added

- **`docs/deployment.md`** — first version of the deployment doc the guard error message has been pointing to since v2.2.1. Covers transport selection, HTTP-daemon launch pattern, and the TTY-guard heuristic + overrides.

### Tests

- `tests/v2-4-2-tty-guard.test.ts` +5 cases spawning `dist/index.js` with varying stdin configurations: background daemon (stdin=/dev/null, exits code 3 within grace), MCP-style launch (pipe + `initialize` frame within grace, proceeds), `RELAY_SKIP_TTY_CHECK=1` bypass, pipe-but-no-writes (exits code 3), non-stdio transport (guard never fires).

**Total after v2.4.2: 1124 tests pass** (1119 pre-v2.4.2 + 5 new).

### Cross-platform parity

Pure Node `process.stdin` + timers — no platform-specific syscalls. Tests use `spawn` with `/dev/null` stdin on darwin + linux; Windows CI relies on the heuristic being platform-agnostic (`stdin.isTTY` + `stdin.on('data', …)` are portable surface).

## v2.4.1 — 2026-04-24 — dashboard inbox visibility

Operator-facing friction → product feature. Pre-v2.4.1 an operator had to drop to `sqlite3 ~/.bot-relay/relay.db` to answer "which agent has mail piling up?" — the dashboard showed the 20 most-recent messages flat but no per-agent rollup. v2.4.1 closes that gap.

### Added

- **`getInboxSummary()` in `src/db.ts`.** Single GROUP BY query across `agents LEFT JOIN messages` returning `{ agent_name, pending_count, unread_count, last_message_at }` for every registered agent — including agents with zero mail (LEFT JOIN + `m.id IS NOT NULL` guards inside the CASE expressions so no-match rows don't inflate counts). Semantics: `pending_count` = `status='pending'` (un-drained), `unread_count` = `seq IS NULL` (never observed — mirrors v2.3 `peek_inbox_version`), `last_message_at` = `MAX(created_at)` across any status.
- **Snapshot endpoint enrichment in `src/dashboard.ts`.** `snapshotApi` now merges `getInboxSummary()` onto each `agents[]` row as additive fields (`pending_count`, `unread_count`, `last_message_at`). Existing `AgentWithStatus` shape is untouched — the dashboard's AI consumers see a strict superset.
- **Dashboard HTML rendering.** New inbox-badge element on every agent card (yellow pill for `pending_count > 0`, gray for 0) with a hover `title` tooltip carrying `"Inbox: N pending, M unread · last message <relative>"`. New "inbox" option in the sort dropdown sorts agents descending by `pending_count` so backlog bubbles to the top of the grid; tie-breaks on `unread_count`, then name.

### Tests

- `tests/v2-4-1-inbox-visibility.test.ts` +7 cases: empty relay / zero-mail agents still appear (LEFT JOIN) / pending + unread piling up / drain flips `pending_count` to 0 but `last_message_at` survives / snapshot enrichment per agent / snapshot shape stable (existing fields untouched) / HTML contains inbox-badge markup + new sort option.

**Total after v2.4.1: 1119 tests pass** (1112 pre-v2.4.1 + 7 new).

### Cross-platform parity

Pure SQL + TypeScript + HTML. No filesystem, no syscalls beyond existing patterns. Wasm SQLite driver (`RELAY_SQLITE_DRIVER=wasm`) covered via the shared `CompatDatabase` interface — the new query uses only `prepare` + `all`, both supported.

### Out of scope (deferred to v2.5)

Per-agent inbox drilldown page, filtering by sender/priority/age, real-time push (snapshot is poll-only — fine for this scope), mobile-responsive table.

## v2.4.0 — 2026-04-23 — traffic replay + per-instance isolation + MCP prompts/resources split

### ⚠ Codex pre-ship audit patches (applied 2026-04-23)

Codex returned PATCH-THEN-SHIP with 2 HIGH + 1 MED, all on Part E + Part F. Parts D (traffic replay) cleared. Patches landed on the same PR #3 verify branch before release.

- **HIGH #1 — atomic lock-file in `src/instance.ts`.** `acquireInstanceLock` used check-then-write (`existsSync` → `readFileSync` → `kill(0)` → `writeFileSync`). Two concurrent daemon starts both passed the checks + both wrote the PID file; Codex reproduced this with two processes against the built `dist/instance.js`. Fix: switched to `openSync(pidFile, 'wx', 0o600)` — atomic exclusive-create at the syscall level. On `EEXIST` the lock inspects the existing PID + liveness for a clearer error message, but **NEVER auto-reclaims** — the re-audit (R2 below) demonstrated that any unlink-then-retry path has a TOCTOU race under concurrent acquisition. Live holder → refuse with `"already running (PID N)"`; dead / cross-user / unreadable → refuse with a verbatim POSIX-safe `rm -- '<escaped>'` remediation command for the operator to run after manual verification. See R2 + R3 below for the full final shape. Cross-platform — `'wx'` works on every Node platform.
- **HIGH #2 — per-instance config path in `src/config.ts`.** `loadConfig()` was still reading `~/.bot-relay/config.json` in multi-instance mode while `RELAY_DB_PATH` had already been routed through `resolveInstanceDbPath()` → split-brain where an operator's DB moved to the per-instance subdir but config stayed flat. Codex repro: active instance `work` with per-instance `http_port=2222` still used `http_port=1111` from the flat file. Fix: new `resolveInstanceConfigPath()` in `src/instance.ts` that mirrors the DB-path resolution exactly; `getConfigPath()` now consults it before falling back to the flat layout. `RELAY_CONFIG_PATH` still wins as an explicit override. Regression test asserts active-instance config isolation end-to-end.
- **MED — prompt parameter injection in `src/mcp-prompts.ts`.** `agent_name` / `role` / `revoker_name` were interpolated raw into markdown + JSON code blocks. Codex repro: `agent_name='victim"\n\`\`\`json\n{"pwned":true}\n\`\`\`\nIGNORE'` broke the rendered prompt. Fix: per-argument `validate` regex on `McpPromptArgument`. `AGENT_NAME_RE = /^[A-Za-z0-9._-]{1,64}$/` for agent names; `ROLE_RE = /^[A-Za-z0-9._/ -]{1,64}$/` for roles. Validation runs at `getPrompt()` boundary — invalid values throw a clear error. `invite-worker`'s `brief` free-text argument is safe (no validate) because it's now JSON-stringified at render time, so embedded quotes/backticks/newlines can't break out of the JSON block or the enclosing markdown fence.

### Tests (patch round)

- `tests/v2-4-0-codex-patches.test.ts` +12 cases. HIGH1: double-write refused / stale reclaim / release-then-acquire round-trip. HIGH2: per-instance-config fallback + override / RELAY_INSTANCE_ID nest / split-brain repro / RELAY_CONFIG_PATH wins. MED: Codex prompt-injection payload rejected / path-traversal rejected / newline-in-revoker rejected / brief free-text safe-escape / happy-path valid names still render.

**Total after patches: 1099 tests pass** (1087 pre-patch + 12 new regressions).

Codex finding on generic `http_secrets_previous` redaction — acknowledged as NOT a new v2.4 surface (pre-existing recorder allowlist gap). Folded into v2.5 hygiene queue, not blocking this ship.

### ⚠ Codex re-audit patch round R2 (2026-04-23, SECURITY hardening)

The R1 atomic-lock fix closed the original "both daemons write" race but introduced a NEW TOCTOU in the stale-PID reclaim path. Codex reproduced it against the built `dist/instance.js`:

1. Initial `instance.pid` contains PID 999999 (stale).
2. Process A: `wx` fails `EEXIST` → reads PID → probes `ESRCH` (dead) → **pauses** just before `unlinkSync`.
3. Process B: same path → unlinks → `wx` wins → writes its live PID.
4. Process A resumes → unlinks B's LIVE pidfile → `wx` wins → writes its own PID.
5. Both A and B believe they hold the lock.

The auto-reclaim path cannot be made safe without an atomic "test-and-replace specific content" primitive, which POSIX `fs` doesn't provide.

**Fix R2 — fail-closed on every EEXIST**, regardless of PID liveness. `acquireInstanceLock` in `src/instance.ts` no longer unlinks anything. On `EEXIST`:

- Live holder → refuse with `"already running (PID ...)"`.
- Dead holder → refuse with a verbatim `rm <path>` command for the operator to run after confirming no daemon is alive.
- Cross-user EPERM or unreadable file → refuse as unknown-liveness.

Auto-reclaim is deferred to **v2.5+** with a proper atomic primitive (fcntl lock on the open fd, or a directory-based lock) + a regression that mirrors the exact Codex schedule.

`docs/multi-instance.md` gained a "Why auto-reclaim was removed" section with the full race description + manual-cleanup workflow.

### R2 regression tests

- `tests/v2-4-0-codex-patches.test.ts` expanded to 16 cases (from 12). New: H1.2b manual-cleanup round-trip, H1.2c live-holder clear error, H1.2d unreadable-pidfile fail-closed, H1.2e TOCTOU scenario (two observers of the same stale pidfile BOTH refuse — neither silently reclaims).
- `tests/v2-4-0-per-instance-isolation.test.ts` E.2.4 flipped from "reclaims stale" to "fails-closed on stale + manual cleanup succeeds."

**Total after R2: 1103 tests pass** (1099 post-R1 + 4 new R2 regressions).

**Not addressed (deferred):** auto-reclaim itself. v2.5+ with proper primitive + Codex-schedule regression.

### ⚠ Codex re-audit patch round R3 (2026-04-23, MED + LOW)

R2 HIGH cleared (both the fail-closed refusal and the EPERM cross-user path verified by Codex). Two smaller items remained:

- **MED — shell-injectable `rm` command in the stale-pidfile error text.** The R2 patch printed `rm "${pidFile}"` using double-quoted interpolation. Double quotes handle spaces but don't neutralize `$()`, backticks, or `$VAR`. Codex repro: `RELAY_HOME=/tmp/bad"$(touch SHOULD_NOT_RUN)"` produces a printed command that, when copy-pasted, executes the embedded command substitution. Local / operator-controlled input so MED not HIGH. Fix: new `shellSingleQuoteEscape(value)` helper in `src/instance.ts` that wraps the value in single quotes + escapes interior `'` as `'\''` (canonical POSIX-safe idiom). The error + log.warn now emit `rm -- '<escaped>'`. Added `(H1.3 R3)` regression that creates a hostile RELAY_HOME path containing `$()`, backticks, and `$VAR`, captures the printed command, feeds it through `spawnSync('/bin/sh', ['-c', cmd])`, and asserts that (a) no sentinel files were created (no side-effect command substitution), and (b) the pidfile was legitimately removed. Plus `(H1.3b R3)` round-trip helper test across 8 nasty input cases (space, single quote, `$()`, backtick, `$HOME`, newlines, mixed).
- **LOW — CHANGELOG top-bullet drift.** The HIGH #1 top bullet still described R1 behavior ("reclaim stale files (one retry bounded)") while the R2 section correctly explained why auto-reclaim was removed. Rewrote the top bullet to reflect the R2 + R3 final shape: atomic `wx` create, never auto-reclaim, clear error + POSIX-safe `rm -- '<escaped>'` remediation command.

**Total after R3: 1105 tests pass** (1103 post-R2 + 2 new R3 regressions).



Three bundled parts in the v2.4.0 release, shipped the moment v2.3.0 shipped:

- **Part D** — A.3 traffic-replay harness (deferred from v2.3.0 at the explicit escape-hatch).
- **Part E** — per-instance local isolation (per the federation design roadmap, re-slotted to v2.4 since v2.3 took profiles + ambient-wake bandwidth).
- **Part F** — MCP prompts + resources split (the federation memo's "tools/resources/prompts split more aggressively" recommendation).

Schema unchanged (v11 stays). Tool count unchanged (30 stays — prompts + resources are separate MCP capabilities, not tools). CLI subcommands 11 → 13 (`relay list-instances` + `relay use-instance`). Protocol `2.3.0 → 2.4.0`.

### Part D — traffic-replay harness

- **D.1 — `src/transport/traffic-recorder.ts`**. Env-gated via `RELAY_RECORD_TRAFFIC=<path>`. Records every MCP tool call as a JSONL line (`{ts, tool, args, response, transport, source_ip}`). `fsync`-per-write for durability. Sensitive fields (`agent_token`, `plaintext_token`, `recovery_token`, `http_secret`, `password`, `secret`) redacted at capture time. 1 GB safety cap — disables capture when log exceeds that + logs a warn. Never throws.
- **D.2 — `scripts/replay-relay-traffic.ts`**. CLI: `npx tsx scripts/replay-relay-traffic.ts <log.jsonl>`. Spins an isolated relay, re-issues every recorded call, compares responses. Volatile fields (UUIDs, timestamps, tokens, seq/epoch) normalized to `<volatile>` sentinels; prose-embedded UUIDs + ISO timestamps normalized to `<uuid>` / `<iso>`. Exit 0 on full parity, 1 on any divergence. Internal `_requestHandlers` accessor on the MCP server routes through the same dispatch path as the live stdio/http transports.
- **D.3 — tests + docs**. 8 cases in `tests/v2-4-0-traffic-replay.test.ts` covering record disable/enable, redaction, 1 GB cap, replay parity, divergence detection, volatile-field normalization, end-to-end recorded-then-replayed round-trip. `docs/traffic-replay.md` explains when to use + when NOT to.

### Part E — per-instance local isolation

Per Codex federation design memo: isolation unit is `instance_id` (UUID), NOT per-$USER. v2.4.0 supports COEXISTENCE of multiple daemons on the same machine; cross-instance messaging is still out of scope (v2.5+ federation territory).

- **E.1 — `src/instance.ts`**. Instance-ID model: UUID per instance, `~/.bot-relay/instances/<id>/` subdir with `instance.json` metadata, `relay.db`, `config.json`, `instance.pid` lock. Path-traversal guard on `instance_id` (`/^[A-Za-z0-9._-]+$/`). `RELAY_HOME` env-override for test harnesses. Lock file pattern: `acquireInstanceLock` with PID liveness check (ESRCH → stale, reclaim + warn). `resolveInstanceDbPath()` returns per-instance path in multi-instance mode, falls back to legacy `~/.bot-relay/relay.db` otherwise. Symlink-or-file active-instance pointer (lstat-aware, handles dangling symlinks). Wired into `src/db.ts getDbPath` — `RELAY_DB_PATH` still wins as explicit override; otherwise per-instance path; otherwise legacy flat layout.
- **E.2 — two-instance coexistence**. 12 tests in `tests/v2-4-0-per-instance-isolation.test.ts` covering legacy-mode default, `RELAY_INSTANCE_ID` flip, metadata round-trip, path-traversal rejection, separate DB paths, messages non-bleeding, lock-file collision, stale-PID reclaim, `listInstances`, `setActiveInstance` + `resolveActiveInstanceId`.
- **E.3 — CLI + docs**. Two new subcommands: `relay list-instances` (with `--json`) and `relay use-instance <id>` (kubectl-style). `relay init` gains `--instance-id=<id>` and `--multi-instance` (auto-generates UUID) flags. CLI subcommand count 11 → 13. `docs/multi-instance.md` explains the model + when to use + backward-compat.
- **E.4 — backward compatibility**. Operators with existing `~/.bot-relay/relay.db` see NO behavior change. Multi-instance is strictly opt-in via env or CLI flag. 8 additional tests in `tests/v2-4-0-instance-cli.test.ts` covering the CLI surface + legacy/multi coexistence.

### Part F — MCP prompts + resources split

Tool count stays 30 — neither prompts nor resources add tools.

- **F.1 — `src/mcp-prompts.ts`**. Three shipped prompts: `recover-lost-token`, `invite-worker`, `rotate-compromised-agent`. Each is a `McpPromptDefinition` with `name`, `description`, `arguments[]`, and a `render(args)` function that returns the user-role message text. Parameter substitution validated at call time (missing required arg throws a clear error).
- **F.2 — `src/mcp-resources.ts`**. Three shipped resources: `relay://current-state` (agents + active tasks + pending counts + schema_version), `relay://recent-activity` (last 50 audit entries with `params_json` stripped), `relay://agent-graph` (nodes + message-edges + task-edges for visualization).
- **F.3 — capabilities**. `createServer` advertises `prompts: {}` + `resources: {}` alongside `tools: {}` in the initial capabilities exchange. Request handlers registered for `prompts/list`, `prompts/get`, `resources/list`, `resources/read`.
- 11 tests in `tests/v2-4-0-mcp-prompts-resources.test.ts` covering prompt enumeration, parameter substitution, missing/unknown-prompt errors, all-prompts-render smoke, resource enumeration, current-state/agent-graph content shapes, unknown-URI errors, server capability declaration. `docs/mcp-prompts.md` operator guide.

### Tests

- `tests/v2-4-0-traffic-replay.test.ts` (D) — 8 cases.
- `tests/v2-4-0-per-instance-isolation.test.ts` (E core) — 12 cases.
- `tests/v2-4-0-instance-cli.test.ts` (E CLI) — 8 cases.
- `tests/v2-4-0-mcp-prompts-resources.test.ts` (F) — 11 cases.

**Total: 1087 tests pass** (1048 v2.3.0 baseline + 39 new v2.4.0).

### Release hygiene

- `package.json` 2.3.0 → 2.4.0.
- `src/protocol.ts` 2.3.0 → 2.4.0.
- New files: `src/transport/traffic-recorder.ts`, `scripts/replay-relay-traffic.ts`, `src/instance.ts`, `src/cli/list-instances.ts`, `src/cli/use-instance.ts`, `src/mcp-prompts.ts`, `src/mcp-resources.ts`, 3 docs, 4 test files.

### Hall of Fame

- **Release cadence** — v2.4.0 bundled related work and shipped shortly after v2.3.0.
- **Codex** — federation design memo (2026-04-19) that locked `instance_id` (not $USER) as the per-instance isolation unit + the MCP tools/resources/prompts split recommendation.
- **The v2.3.0 Part A infrastructure** — property tests + consistency probe — made the A.3 traffic-replay harness possible without re-deriving ground-truth invariants. Traffic replay now stands alongside them as permanent bug-finding infra.

## v2.3.0 — 2026-04-22 — systemic bug-finding + profiles + Phase 4s ambient wake

### ⚠ Codex pre-ship audit patches (applied 2026-04-23)

Codex's dual-model audit of the v2.3.0 diff returned PATCH-THEN-SHIP. Two HIGH findings in Part C (Phase 4s ambient wake); Parts A + B cleared. Patches landed on the same PR #2 verify branch before release.

- **HIGH #1 — atomic seq assignment race.** Pre-patch `getMessages` snapshotted `mailbox.next_seq` outside the transaction, then blindly incremented `next` for every candidate even when the `UPDATE ... WHERE seq IS NULL` no-op'd because another reader had already claimed the row. Two overlapping cross-process drains could stamp the same seq onto different messages. Fix: mailbox row is now read INSIDE the transaction; `next` advances ONLY when `UPDATE.changes === 1`; mailbox.next_seq persists the actual claim count, not the candidate count; transaction runs as BEGIN IMMEDIATE via better-sqlite3's `.immediate()` modifier so cross-process readers serialize at tx start. Rows claimed by another reader during our tx are re-hydrated via a targeted SELECT so the caller sees consistent seqs. (`src/db.ts`.)
- **HIGH #2 — peek_inbox_version didn't surface new unread mail.** Pre-patch the response exposed `last_seq` as the watch field, but seq is assigned at DELIVERY time so `last_seq` only advances when the agent CALLS `get_messages` — defeating the point of the lightweight control-plane peek. Fix: new `total_unread_count` field computed via `SELECT COUNT(*) FROM messages WHERE to_agent = ? AND seq IS NULL`. Advances on every `send_message`. `docs/ambient-wake.md` promoted to watch-signal; `last_seq` demoted to read-cursor-across-reconnects.

### Tests (patch round)

- `tests/v2-3-0-ambient-wake.test.ts` +4 cases (C.3.4, C.3.5, HIGH1.1, HIGH1.2).
- `tests/v2-3-0-property-based-query.test.ts` +2 fast-check properties (P7 seq uniqueness under overlap, P8 send→peek control-plane visibility). P7 runs fresh-DB per iteration for isolation.

**Total after patches: 1048 tests pass** (1042 pre-patch + 6 new regressions).



v2.3.0 is the first MINOR release since v2.2.0. Three bundled parts: (A) systemic bug-finding infrastructure — property-based tests + a live consistency probe; (B) profiles + surface shaping via `relay init --profile={solo,team,ci}`; (C) Phase 4s ambient-wake model — mailbox table + per-recipient monotonic seq + new `peek_inbox_version` MCP tool + filesystem marker fallback + dashboard wake button. Schema v10 → v11. Tool count 29 → 30. Protocol `2.2.3 → 2.3.0`.

### Part A — systemic bug-finding infrastructure

- **A.1 — property-based tests** (`tests/v2-3-0-property-based-query.test.ts`). 6 `fast-check` properties assert invariants that must hold regardless of input shape: (P1) send-then-peek returns exactly once, (P2) peek is non-mutating across repeated calls, (P3) consume-once drains, (P4) status-partition sums correctly, (P5) limit respected, (P6) round-trip content identity. Default gate runs 30 iterations per property (~180 scenarios); `FAST_CHECK_FULL=1` bumps to 200 (~1200) for the `--full` gate. New devDependency: `fast-check`.
- **A.2 — live consistency probe** (`src/transport/consistency-probe.ts`). Sampling observer that runs inside the daemon when `RELAY_CONSISTENCY_PROBE=1`. Every Nth `get_messages` call (configurable via `RELAY_CONSISTENCY_PROBE_RATE`, default 100), a parallel raw-SQL SUPERSET query runs against `messages.to_agent`; if SQL sees pending rows the MCP path dropped, a structured warning lands on stderr with the missing IDs. Off by default. Never throws, never blocks, pure observation. Designed to catch the v2.2.1 drops-pending class of bug automatically in any environment the probe is on. 4 regression tests.
- **A.3 — traffic-replay harness** — **deferred to v2.3.1** per the explicit escape hatch. A.1 + A.2 deliver the bulk of the bug-finding value; the replay harness adds marginal coverage for the token cost of more test surface. Revisit when we have a reproducible bug that A.1 + A.2 can't surface deterministically.

### Part B — profiles + surface shaping

- **B.1 — `relay init --profile={solo,team,ci}`**. Profiles shape the surface, not just defaults. `solo` (default): stdio transport, core bundle only, info logs, 30-day abandon threshold. `team`: http transport, all feature bundles, 7-day abandon. `ci`: stdio, core only, warn logs, dashboard disabled, 1-day abandon. Explicit flags (`--transport`, `--port`) still win over profile defaults.
- **B.2 — surface-shaping filter in `src/server.ts`**. New `TOOL_BUNDLES` map + `isToolVisible` + `resolveSurfaceShape` helpers. `tools/list` now filters by the active profile's `feature_bundles` + `tool_visibility.hidden`. Calls to a hidden tool return `TOOL_NOT_AVAILABLE` with a hint naming the profile that would expose it. `health_check` + `discover_agents` are always visible (diagnostic/routing primitives). New error code `TOOL_NOT_AVAILABLE` in the stable error-code catalog.
- **B.3 — docs + tests**. `docs/profiles.md` with per-profile settings + bundle table + TOOL_NOT_AVAILABLE error-shape reference. README pointer. 10 tests covering init writes, visibility filter, hidden-override, and a drift guard that asserts every registered tool has a bundle mapping.

### Part C — Phase 4s ambient wake

- **C.1 — schema v10 → v11 migration** (`migrateSchemaToV2_9`). Phase 7q's reserved `mailbox` + `agent_cursor` stub tables are wired up: `mailbox` gets `agent_name` + `created_at` columns + unique index on agent_name; `agent_cursor` gets `agent_name` + `updated_at`. `messages` gets `seq INTEGER` + `epoch TEXT` columns. Index on `(to_agent, seq)` for cursor-based drain. Additive + idempotent.
- **C.2 — delivery-time seq assignment**. `getMessages` now atomically looks up the recipient's `mailbox` row and assigns `seq` + `epoch` to every returned message where `seq IS NULL`. Single transaction wraps the increment + UPDATEs. Per Codex Q9 (2026-04-19): seq reflects the order the RECIPIENT saw messages, not send order. `sendMessage` is untouched.
- **C.3 — new MCP tool `peek_inbox_version`** (`src/tools/peek-inbox-version.ts`). Pure observation: returns `{mailbox_id, epoch, last_seq, total_messages_count}`. No mutation, no read-mark side effect. Tool count 29 → 30. `core` feature bundle — visible in every profile.
- **C.4 — filesystem marker fallback** (`src/filesystem-marker.ts`). Opt-in via `RELAY_FILESYSTEM_MARKERS=1`. Daemon touches `~/.bot-relay/marker/<agent_name>.touch` on every delivery; shell clients can `fs.watch()` + call `peek_inbox_version` on change. HINT only, non-authoritative — SQLite remains the truth. Cross-platform (macOS/Linux/Windows). Path sanitized against traversal.
- **C.5 — dashboard wake-agent button**. 🔔 Wake agent button in the focused-agent panel. POST `/api/wake-agent {agent_name}` touches the marker + writes a `wake_agent` audit entry. When markers are disabled on the daemon, the endpoint returns `markers_enabled: false` + a hint — the button shows a disabled state instead of lying.
- **C.6 — tests + docs**. 13 ambient-wake tests covering schema migration, monotonic seq, epoch rotation, marker opt-in/off/sanitization, wake endpoint round-trip, and audit-log payload shape. `docs/ambient-wake.md` with the full model + shell/Claude Code/Python integration sketches + backward-compatibility notes.

### Schema notes

- `CURRENT_SCHEMA_VERSION` bumped 10 → 11.
- `messages.seq` + `messages.epoch` are NULL for pre-v2.3.0 rows; assigned on first read by the v2.3.0 delivery-time path. No backfill required.
- Epoch is TEXT (UUID) per Codex Q9 locked design. Rotates explicitly on `rotateMailboxEpoch` (called from backup/restore in future phases); does NOT rotate on every daemon restart.
- Phase 7q's `mailbox` / `agent_cursor` stub tables from schema v6 are expanded additively — no table rebuild.

### Tests

- `tests/v2-3-0-property-based-query.test.ts` (A.1) — 6 properties × 30 iterations default.
- `tests/v2-3-0-consistency-probe.test.ts` (A.2) — 4 cases.
- `tests/v2-3-0-profiles.test.ts` (B) — 10 cases.
- `tests/v2-3-0-ambient-wake.test.ts` (C) — 13 cases.

Test-fixture bumps for version/tool-count drift:

- `tests/v2-1-schema-info.test.ts` — `applyMigration(10, 11)` no-op + raise pivot to `11 → 12`.
- `tests/v2-1-3-agent-status-enum.test.ts` — schema version pin 10 → 11.
- `tests/http.test.ts` — tools/list count 29 → 30 + `peek_inbox_version` presence assertion.
- `tests/v2-2-0-full-dashboard-smoke.test.ts` — version pins bumped to 2.3.0.

**Total: 1042 tests pass** (1009 v2.2.3 baseline + 33 new v2.3.0).

### Release hygiene

- `package.json` 2.2.3 → 2.3.0.
- `src/protocol.ts` 2.2.3 → 2.3.0.
- New files: `src/cli/init.ts` profile section, `src/transport/consistency-probe.ts`, `src/tools/peek-inbox-version.ts`, `src/filesystem-marker.ts`, `docs/profiles.md`, `docs/ambient-wake.md`.

### Hall of Fame

- **Rationale** — Part A was driven by the v2.2.1 get_messages-drops-pending incident (find bugs at scale without slowing delivery); Parts A + B + C bundle the related work.
- **Codex (2026-04-19 Prompt B + Q9 reviews)** — the mailbox/seq-at-delivery-time/epoch-as-UUID design locked in Phase 4s. Also the event-sourcing-not-CRDT architectural correction that shapes v3+.
- **The 2026-04-22 four-release sprint** (v2.2.0 → v2.2.1 → v2.2.2 → v2.2.3, all in one day) — every bug that surfaced became a seed for Part A's permanent prevention infrastructure.

## v2.2.3 — 2026-04-22 — Node 18 webhook-timeout hotfix + CI green-gate

Hotfix release. CI has been red on Node 18 since v2.2.1 — both `tests/v2-2-1-bug-sweep.test.ts (B5.1)` and `tests/v2-2-1-codex-patches.test.ts (B5n.2)` timed out at 15s. v2.2.1 + v2.2.2 shipped to npm anyway because the local pre-publish gate didn't run the CI matrix. This release (a) patches the underlying webhook-delivery behavior, (b) adds a systemic guard so we can't ship another CI-red commit, and (c) seeds permanent regression coverage for the Node 18 failure mode.

**Protocol bump `2.2.2 → 2.2.3` (PATCH)**. No API surface change — only delivery-layer behavior.

### Root cause

`sendOnce()` in `src/webhook-delivery.ts` used `req.setTimeout(timeoutMs, handler)`, which is a **socket-level** timeout — it only fires once a TCP socket exists. The B5 test fixture pins to `0.0.0.1` (an invalid-route loopback address) to force a connect refusal. On Node 20/22 the kernel rejects this immediately, `req.on("error")` fires, the failover loop advances. **On Node 18 the kernel sits in `EINPROGRESS` indefinitely** — no socket → `req.setTimeout` never fires → `req.on("error")` never fires → the Promise from `sendOnce` never resolves → the test times out at the vitest 15s harness limit. Real webhook delivery to a misrouted load-balanced target on Node 18 would block the whole failover loop until the process died.

### Fix

- **`src/webhook-delivery.ts` `sendOnce`**: adds a **hard JS `setTimeout`** that fires regardless of socket state. Set up right after `let settled = false;`, before `const req = requester(...)`. The hard timer resolves the Promise with a `timeout` error + destroys `req` when it fires. Every existing `settled = true` path (res.on("end"), res.on("error"), req.on("error"), req.setTimeout callback) gains a `clearTimeout(hardTimer)` so we don't double-resolve. The socket-level `req.setTimeout` stays in place for the "socket connected but stalled mid-stream" case — the hard timer is belt + suspenders.

### Systemic fix — GitHub CI green-gate

`scripts/pre-publish-check.sh` gains a new step after the 25-tool smoke: `GitHub CI green-gate`. Probes `gh run list --commit HEAD` for the current HEAD's CI conclusion. Behavior:

- `success` → PASS.
- `failure` / `cancelled` / `timed_out` / `action_required` → **FAIL** (refuses to publish).
- `in_progress` / `queued` / `waiting` / `pending` → WARN (still running, proceed at own risk).
- `no-run` → WARN (HEAD not pushed yet).
- gh CLI absent → SKIP with a console notice (no hard block on tooling absence).

This closes the "local gate green, CI red" loophole that let v2.2.1 + v2.2.2 ship with known-red matrix tests.

### Tests

- `tests/v2-2-2-regression-from-released-bugs.test.ts` — case `(7) sendOnce hard-timer fires even when no socket exists (Node 18 EINPROGRESS)`. Pins to `0.0.0.1` with `timeoutMs=400ms`; asserts the call completes within 8× timeoutMs wall-clock and resolves with `statusCode=null + error`. Guard against re-landing the socket-only timeout.
- `tests/v2-2-1-bug-sweep.test.ts (B5.1)` + `tests/v2-2-1-codex-patches.test.ts (B5n.2)` — previously Node 18 red, now green on all three runtimes.

**Total: 1009 tests pass** (1008 v2.2.2 baseline + 1 new case (7) regression).

### Release hygiene

- `package.json` 2.2.2 → 2.2.3.
- `src/protocol.ts` 2.2.2 → 2.2.3.
- `tests/v2-2-0-full-dashboard-smoke.test.ts` — version pins bumped to 2.2.3.

### Hall of Fame

- **Maintainer (public-facing CI failure, 2026-04-22):** caught the red badge + correctly escalated.
- **Root-cause diagnosis:** socket-level vs JS-level timer distinction + Node 18 EINPROGRESS behavior + prescribed fix + CI green-gate design.

## v2.2.2 — 2026-04-22 — defense-in-depth + dashboard UX polish + CLI ergonomics + 3 bundled bugs

v2.2.2 bundles 11 items across four buckets: two server-side defense-in-depth touches (A1/A2), five dashboard UX polish items (B1-B5), one CLI ergonomics addition (C1), and three bundled bugs surfaced during release validation (BUG1/BUG2/BUG3). One release, one Codex audit. Protocol bumps `2.2.1 → 2.2.2` (MINOR — additive endpoints + `get_messages.peek` parameter + agent_status enum widening with `abandoned` and `closed`). No breaking changes.

### Part A — server-side defense-in-depth

- **A1 — `/api/send-message` optional `from_agent_token`.** The dashboard endpoint previously trusted the dashboard secret alone — any operator could send-as any agent with no per-agent verification. v2.2.2 adds a *defense-in-depth* path: callers may supply `from_agent_token` in the body OR `X-From-Agent-Token` header. When present, the server verifies against the from-agent's stored `token_hash` and records `from_authenticated: true` in the audit log. When absent, behavior is unchanged (v2.2.1 Option (a) audit-only model); the audit entry records `from_authenticated: false` so incident review can distinguish operator-impersonation from token-verified sends. Mismatch → `403 AUTH_FAILED` + audit `success=0`.
- **A2 — per-human operator identity cookie.** New `relay_operator_identity` cookie (SameSite=Lax, 90-day, not HttpOnly so the dashboard JS can read it). Precedence: cookie > `RELAY_DASHBOARD_OPERATOR` env > `"dashboard-user"` default. New endpoints: `GET /api/operator-identity` (reports resolved identity + source) and `POST /api/operator-identity` with `{identity}` to set/renew, or `{identity: ""}` to clear. Dashboard header ships a button that shows the current identity and opens a `prompt()` to change it. Audit log entries from state-changing dashboard endpoints (`send_message`, `kill-agent`, `set-status`, `set_operator_identity`, `set_dashboard_theme`) now carry the per-human identity in `operator_identity`.

### Part B — dashboard UX polish

- **B1 — rich custom-theme `<dialog>`.** Replaces the `prompt()` JSON-paste flow with a native `<dialog>` modal: 14 color pickers (one per CSS token), live preview pane, paste-JSON fallback for operators who already have a theme file. Save applies locally + persists server-side via new `POST /api/dashboard-theme` + broadcasts `dashboard.theme_changed` to open WS clients. Cancel reverts cleanly (snapshot taken on open). Closes on Escape + click-outside (native `<dialog>` semantics).
- **B2 — per-card resize with snap-to-grid-column.** Bottom-right corner drag on any agent card resizes that card by integer col/row spans (max 4×3), snapping to the computed grid-column width. Per-agent-name state persists in localStorage `bot-relay-card-sizes-v1`, LRU-capped at 50 entries so retired agents don't accumulate forever. Top-right × button resets a single card's sizing.
- **B3 — abandoned agent status + hide toggle + `relay purge-agents` CLI.** New `agent_status: "abandoned"` surfaced by `deriveAgentStatus` when `last_seen` > `RELAY_AGENT_ABANDON_DAYS` (default 7). Distinct from `offline` so dashboards can hide retired terminals by default. Dashboard adds a "show abandoned" checkbox (default off) + the existing status filter dropdown gets an `abandoned` option. Data preserved — operators prune via the new `relay purge-agents [--abandoned-since=N] [--apply]` subcommand: dry-run by default, `--apply` commits, one audit-log entry per deleted row (`purge-agents.cli`). Messages + tasks are NOT touched (use `relay purge-history` separately). New sanctioned-helper pair in `db.ts` — `listAgentsOlderThan` + `deleteAgentIfAbandoned` — so the CLI passes the drift-grep guard.
- **B4 — agent-card sort toggle.** New `sort` dropdown in the filter bar: `status` (active-first: working → blocked → waiting_user → idle → stale → offline → abandoned; default), `role`, `last seen` (most-recent first), `name`. Persists in localStorage.
- **B5 — message search bar.** Search input above the messages timeline. Case-insensitive substring match over `content_preview` + `from_agent` + `to_agent`. 200 ms debounce, empty shows all, last query persisted.

### Part C — CLI ergonomics

- **C1 — `relay open [--url <u>]`.** Opens the dashboard URL in the default browser. Auto-detects host + port from config (`$RELAY_HTTP_HOST`, `$RELAY_HTTP_PORT`, or `~/.bot-relay/config.json`). Platform routing: `darwin` → `open`, `win32` → `cmd.exe /c start "" <url>`, `linux` → `$BROWSER` when set else `xdg-open`. Daemon-down is a warning (prints actionable hint), not a hard failure — the browser still opens.

### Part D — bundled bugs (discovered during v2.2.2 release validation)

- **BUG1 — `get_messages` read-mark race on repeated pending polls.** Pre-v2.2.2 `getMessages(agent, 'pending', ...)` SELECTed pending messages + immediately UPDATEd them to `read_by_session = currentSession`. A second pending poll from the same session excluded those rows because they now matched their own session id — orchestrators that surveyed their own inbox on a polling interval lost visibility of real pending mail the moment they looked at it once. Fix: new optional `peek: boolean` field on `GetMessagesSchema` (default false). Threaded into `getMessages(agentName, status, limit, peek)`: when true, the mark-as-read UPDATE is skipped. Default behavior unchanged — single-shot workers still consume-once. Tool description expanded to document the orchestrator-polling use case.
- **BUG2 — intentional-terminal-close `closed` agent_status.** Pre-v2.2.2 SIGINT/SIGTERM from a stdio terminal set `agent_status='offline'`, indistinguishable from a network drop. v2.2.2 adds a new enum value `closed` (relay-computed, not user-settable via `set_status`) with a sanctioned helper `closeAgentSession(name, expectedSessionId)` mirroring `markAgentOffline` but writing `'closed'`. `performAutoUnregister` prefers the new helper + falls back to `markAgentOffline` on helper-level failure; audit entry tool name is `stdio.auto_close` (vs `stdio.auto_offline`). Auto-promotes to `abandoned` via the existing `RELAY_AGENT_ABANDON_DAYS` chain. Dashboard adds `closed` to the status filter dropdown + a `.badge-closed` style (muted with line-through).
- **BUG3 — regression-from-released-bugs suite.** New `tests/v2-2-2-regression-from-released-bugs.test.ts` with one permanent case per bug discovered in the v2.1.x → v2.2.x arc that made it to an operator's hands. Top-of-file pattern doc explains when to add a new entry. 6 seeded cases: (1) CLI flag parsing, (2) daemon non-TTY guard, (3) `since`-filter trap + hint, (4) NAME_COLLISION_ACTIVE + `force=true`, (5) BUG1 read-mark race, (6) BUG2 closed status. File grows with every release cycle.

### Tests

- `tests/v2-2-2-defense-in-depth.test.ts` (A1) — 4 cases.
- `tests/v2-2-2-operator-identity.test.ts` (A2) — 7 cases.
- `tests/v2-2-2-theme-dialog.test.ts` (B1) — 6 cases.
- `tests/v2-2-2-card-resize.test.ts` (B2) — 5 cases.
- `tests/v2-2-2-abandoned-agents.test.ts` (B3) — 8 cases.
- `tests/v2-2-2-sort-and-search.test.ts` (B4+B5) — 2 cases.
- `tests/v2-2-2-cli-open.test.ts` (C1) — 6 cases.
- `tests/v2-2-2-bug1-get-messages-peek.test.ts` (BUG1) — 4 cases.
- `tests/v2-2-2-bug2-closed-status.test.ts` (BUG2) — 3 cases.
- `tests/v2-2-2-regression-from-released-bugs.test.ts` (BUG3) — 6 cases.

Test-fixture adjustments (v2.2.2 BUG2 semantic change — SIGINT path now `closed` not `offline`):

- `tests/v2-0-2-audit-fix.test.ts` — capturedSid-match case expects `agent_status='closed'`.
- `tests/v2-1-3-mark-offline.test.ts` — round-trip test expects `'closed'`; audit entries tagged `stdio.auto_close`.
- `tests/v2-1-3-name-collision.test.ts` — `(d)` offline-row re-register test expects `'closed'` (semantics preserved; only the label changed).

**Total: 1008 tests pass** (995 after the original 8 items + 13 new BUG1/BUG2/BUG3 regressions).

### Release hygiene

- `package.json` 2.2.1 → 2.2.2.
- `src/protocol.ts` 2.2.1 → 2.2.2.

### Hall of Fame

- **Operator (2026-04-22):** surfaced the abandoned-agents UX pain after v2.2.1 ship filled the dashboard with retired spawns (B3), and requested the per-human identity cookie so the shared-daemon audit log could tell humans apart (A2).
- **Codex pre-ship audit (v2.2.1):** flagged the dashboard-secret-impersonation risk as a defense-in-depth gap → A1.

## v2.2.1 — 2026-04-21 — consolidated bug-sweep + dashboard polish

v2.2.1 bundles 6 bug fixes (caught during v2.2.0 release + operator use) with 5 dashboard polish items queued from the v2.2 spec. One release, one Codex audit. Protocol bumps `2.2.0 → 2.2.1` (MINOR — additive `set_dashboard_theme` tool + new `/api/send-message`/`/api/kill-agent`/`/api/set-status` endpoints + optional `force` field on `register_agent` + optional `hint` field on `get_messages`). No breaking changes. Schema migrates `v9 → v10`.

### Part A — bug sweep

- **B1 — CLI parser (Option A).** `node dist/index.js --transport=http --port=3777` no longer silently ignores the flags. New `src/cli.ts` with a tight allowlist (`--transport`, `--port`, `--host`, `--config`, `--help`, `--version`). Precedence: CLI > env > config file > default. Unknown flags fast-fail with `exit(2) + clear message`. Startup source-log emits one line showing which layer won for each knob.
- **B2 — duplicate-name register race.** Two Claude Code terminals running under the same `RELAY_AGENT_NAME` + shared token previously silently rotated `session_id` and one terminal's mailbox reads silently dropped mail. `register_agent` now hard-rejects with `NAME_COLLISION_ACTIVE` when the existing row's `session_id` is set, `agent_status` is active, and `last_seen` is within 120s. Escape hatch: `force: true` on the register call. Exempt: `recovery_pending` + `legacy_bootstrap` auth states (admin-approved flows). Tests that legitimately exercise re-register flows now pass `force: true` to reach the re-register branch.
- **B3 — daemon non-TTY fallback.** `node dist/index.js` in a non-TTY context (Claude Code bash sandbox, systemd service with no pty) now `exit(3)` with an actionable message pointing at `RELAY_TRANSPORT=http`. Was previously a silent exit-on-stdin-close. Escape hatch: `RELAY_SKIP_TTY_CHECK=1` for test harnesses piping MCP deliberately.
- **B4 — `get_messages` since UX hint.** `status='pending' + count=0 + since < 24h` now surfaces a `hint` field: `"Narrow since window may hide older pending messages. Try since='24h' or since='all'..."`. Caught during v2.2.0 release-validation debugging when a 25-min-old pending message was hidden by `since='15m'` and triggered a false ghost-session diagnosis.
- **B5 — multi-IP DNS round-robin.** `deliverPinnedPost` now accepts `pinnedIps: string[]` and tries each IP in order on connect-refused / timeout / ECONNRESET (max 3 attempts). Load-balanced webhook targets regain the failover-across-replicas semantic that native `fetch()`'s internal round-robin used to provide. Non-retryable errors (TLS, server 5xx) don't loop — they return immediately.
- **B6 — Stop hook payload docs.** New `docs/hook-payload-format.md` consolidates Claude Code 2.1.x hook payload shapes (SessionStart / Stop / PostToolUse / PreToolUse / UserPromptSubmit) with minimal reader templates in Node + bash. Cross-linked from `docs/hooks.md`. Captured during Phase F build when the Stop-hook stdin-JSON format wasn't obvious from existing SessionStart patterns.

### Part B — dashboard polish

- **P1 — themes + `set_dashboard_theme` MCP tool.** New MCP tool `set_dashboard_theme({mode, custom_json?})`. Modes: `catppuccin` (default Mocha palette) / `dark` (tool-neutral) / `light` (tool-neutral) / `custom` (14-token JSON paste). Server-side default stored in new `dashboard_prefs` single-row table (schema v10 via `migrateSchemaToV2_8`). Dashboard client reads default on first connect; localStorage beats server default for repeat visits (**path-1 client-only design**). No WebSocket push of theme changes — operators reload to surface server-side updates on already-connected dashboards. Tool count 28 → 29.
- **P2 — three inline `/api/*` endpoints.** `POST /api/send-message` (body `{from, to, content, priority?}`), `POST /api/kill-agent` (body `{name}` + `X-Relay-Confirm: yes` header), `POST /api/set-status` (body `{agent_name, agent_status}`). All three gated by the v2.1.7 CSRF + rate-limit + host-check infra. Trust model: dashboard access = operator-level trust; no additional agent-token check (same pattern as admin panels). 
- **P3 — CSS extraction.** `src/dashboard-styles.ts` new file exports `DASHBOARD_BASE_STYLES` + `DASHBOARD_THEMES`. `src/dashboard.ts` dropped from 749 → 575 LOC. Themes land as `[data-theme="…"]` selectors in the same module so color-surface changes stay in one place. No new route (single-HTML-response model preserved).
- **P4 — WS test helper extraction.** `tests/_helpers/ws.ts` new shared helper exports `connectWs(port, urlPath, subprotocols?)` with the eager-queue hello-frame-race handling. `tests/v2-2-0-phase-2-websocket.test.ts` delegates to it; new v2.2.1 WS tests import from there.
- **P5 — SECURITY.md CSRF loopback-dev callout.** Added a "Known residual behavior" paragraph documenting that CSRF is skipped in loopback-dev mode (no secret + loopback peer) and what that implies now that state-changing `/api/*` endpoints actually exist. Operators on multi-user machines or machines running unrelated local webservers should set `RELAY_DASHBOARD_SECRET` to activate the full CSRF + auth gate.

### Schema migration (v9 → v10)

`migrateSchemaToV2_8` adds the `dashboard_prefs` table:

```
CREATE TABLE dashboard_prefs (
  id INTEGER PRIMARY KEY CHECK (id = 1),
  theme TEXT NOT NULL DEFAULT 'catppuccin'
    CHECK (theme IN ('catppuccin', 'dark', 'light', 'custom')),
  custom_json TEXT,
  updated_at TEXT NOT NULL
);
```

Single-row (CHECK id=1), seeded with `{theme:'catppuccin', custom_json:null}` on first migration. Additive + idempotent — no data backfill.

### ⚠ Security patches from Codex pre-ship audit

Codex's dual-model audit of the v2.2.1 diff returned PATCH-THEN-SHIP. Patches landed inline (not deferred) before release:

- **M1 (MEDIUM) — inline `/api/*` endpoints bypassed audit/attribution.** `/api/send-message`, `/api/kill-agent`, `/api/set-status` previously went directly to the DB functions, skipping the MCP dispatcher's `logAudit` + verified-caller attribution. A dashboard-secret holder could spoof a `send_message` as any registered `from` agent with no relay record pointing at the operator. Fix: new `logDashboardAudit()` helper in `src/transport/http.ts` wraps every call path with `source='dashboard'` + `via_dashboard: true` + `operator_identity` (sourced from `RELAY_DASHBOARD_OPERATOR` env var, fallback `"dashboard-user"`). Audit entries now record BOTH the operator AND the on-behalf-of agent for forensic replay. Failure paths (validation rejects, SENDER_NOT_REGISTERED, internal errors) also audit with `success=0` + error string.
- **L1 (LOW) — `/api/set-status` missing WS broadcast.** MCP `set_status` fan-outs to connected dashboards via `broadcastDashboardEvent`; the HTTP endpoint skipped that step, leaving connected UIs up to 10s stale until the safety-net `/api/snapshot` poll. Fix: added the same `agent.state_changed` broadcast to the HTTP path.
- **L2 (LOW) — CLI parser drift.** Three sub-fixes: (a) `--help` / `--version` now win over unknown-flag errors so `node dist/index.js --bogus --help` still prints usage instead of exiting 2 silently; (b) `applyCliToEnv()` widens source tracking from `"cli" | "env" | "default"` to `"cli" | "env" | "config" | "default"` — pre-L2 config-file-won values were mislabeled `default` in the startup log; (c) `src/config.ts` gains a small `readConfigFileKeys()` helper so `src/index.ts` can pass file-layer keys into the source classifier. +6 regression tests.
- **B2 doc drift — stale `force` flag comments.** `src/db.ts` had comments claiming `force` was NOT on the MCP surface AND implying DB-layer hard-reject semantics. Neither was accurate post-v2.2.1 (the `force` flag IS a Zod field on RegisterAgentSchema; enforcement moved to the handler layer). Comments replaced with an accurate one-liner pointing at `handleRegisterAgent`.

### Behavior change (post-Codex): ECONNRESET no longer retries

The v2.2.1 B5 multi-IP round-robin helper previously retried across sibling replicas on ECONNRESET. Codex flagged this as an at-least-once delivery hazard: ECONNRESET is ambiguous for POST (the peer may have accepted the request body before resetting), so retrying risks **duplicate webhooks**. Duplicate webhooks are worse than rare one-shot reset losses for operator-run self-hosted tooling (the whole threat model assumes fire-and-forget best effort).

Pre-patch retryable set: ECONNREFUSED + ETIMEDOUT + EHOSTUNREACH + ENETUNREACH + ECONNRESET + EAI_AGAIN. Post-patch: ECONNREFUSED + EHOSTUNREACH + ENETUNREACH + ETIMEDOUT + EAI_AGAIN (all clearly pre-connect failures; peer never saw the body). ECONNRESET now returns immediately with the error surfaced to the caller — `scheduleWebhookRetry` still applies at the outer layer per its own retry policy, so legit retries happen with full `delivery_id`-based dedup semantics.

### Tests

- `tests/v2-2-1-bug-sweep.test.ts` +24 cases (B1 × 12, B2 × 3, B3 × 2, B4 × 4, B5 × 3).
- `tests/v2-2-1-polish.test.ts` +22 cases (P1 × 7, P2 × 10, P3/P4/P5 × 5 artifact checks).
- Existing test updates (schema v10, tool count 29, v2.2.1 version pin): `tests/v2-1-3-agent-status-enum.test.ts`, `tests/v2-1-schema-info.test.ts`, `tests/http.test.ts`, `tests/v2-2-0-full-dashboard-smoke.test.ts`.
- Test fixtures that legitimately re-register an active agent (valid token but no scope-change intent) updated to pass `force: true`: `tests/auth-dispatcher.test.ts`, `tests/phase-4b-1-v2.test.ts`, `tests/phase-4b-2-inherited-token.test.ts`, `tests/regression-plug-and-play.test.ts`, `tests/v2-1-3-name-collision.test.ts`, `tests/v2-1-legacy-migration.test.ts`, `tests/v2-1-sid-recapture.test.ts`.

- `tests/v2-2-1-codex-patches.test.ts` +14 cases (M1 × 5, L1 × 1, L2 × 6, B5n × 2).

**Total: 957 tests pass** (897 v2.2.0 baseline + 46 new v2.2.1 bug-sweep+polish + 14 new Codex-patch regressions). Pre-publish `--full` gate: **11/11 PASS** both pre-audit and post-audit.

### Release hygiene

- `package.json` 2.2.0 → 2.2.1.
- `src/protocol.ts` 2.2.0 → 2.2.1.

### Hall of Fame

- **Operator surfacing (2026-04-21 evening):** CLI-args-silently-ignored (B1), `since` filter hiding pending (B4), daemon non-TTY silent exit (B3) — all caught during v2.2.0 release validation.
- **Codex (v2.2.0 audit):** duplicate-name race pattern context via the scoped-agent-names guidance (B2), multi-IP round-robin note (B5).
- **Phase F discovery (2026-04-21):** Claude Code 2.1.x Stop-hook stdin JSON payload format (B6).

## v2.2.0 — 2026-04-21 — core dashboard observability (+ bundled webhook TOCTOU fix + IP-classifier consolidation)

v2.2.0 is the first MINOR release in the 2.x line. Focus: operator-facing dashboard. Ships core observability (Phases 1-3 of the dashboard spec) plus two bundled fixes from the v2.1.7 audit — the webhook DNS TOCTOU (Item 7, deferred) and the IP-classifier consolidation (Item 9, v2.2 candidate). v2.2.1 polish items (themes / custom-paste / inline send / kill / set-status) ship in a separate cycle.

Protocol bumps `2.1.3 → 2.2.0` (MINOR — new HTTP endpoint + new WebSocket endpoint + additive `terminal_title_ref` field on `register_agent` / `discover_agents` + new `*_preview` fields on `/api/snapshot`). No breaking changes. Schema migrates `v8 → v9`.

### ⚠ Policy change for operators — `/api/snapshot` now returns decrypted content previews

Pre-v2.2.0, `/api/snapshot` returned the raw at-rest-encrypted `content` / `description` / `result` fields verbatim so a dashboard-auth failure could not leak plaintext (see v2.1 Phase 4d design note in `src/dashboard.ts`). v2.2.0 **adds** `content_preview` / `description_preview` / `result_preview` fields alongside the raw encrypted ones — 100-char decrypted previews rendered by the new reactive dashboard.

Narrow scope of the expansion:
- Raw `content` / `description` / `result` stay as ciphertext in the response for clients that want the on-disk form.
- Preview fields are new sibling keys; they do NOT replace the raw fields.
- The dashboard remains gated by `dashboardAuthCheck` + `originCheck` + `httpHostCheck` (plus CSRF on state-changing endpoints). Any caller who can reach the previews can already call `get_messages` with the same decrypted result — the preview is equivalent surface.
- If your deployment is loopback-only with no dashboard secret (dev-mode default), the previews are only reachable from the same machine's local processes — the same trust boundary `get_messages` already assumes.

**If your threat model relies on `/api/snapshot` NEVER decrypting at rest** (e.g. you shipped a custom audit pipeline that pointed at it), set `RELAY_DASHBOARD_SECRET` to gate the endpoint, or stop scraping it and use `get_messages` directly.

### Platform validation status (read before deploying to Linux or Windows)

v2.2.0 introduces three platform-specific click-to-focus drivers. Validation coverage at ship time is **not symmetric across platforms**:

- **macOS** — end-to-end validated. Full unit + integration test coverage (~90 tests across the v2.2.0 phase suites + 9 Codex-patch regressions). iTerm2 raise verified manually during release validation.
- **Linux** — code paths covered by command-construction tests (mocked spawn for `wmctrl -a`). **Not E2E-verified on a real Linux box at ship.** Architecture mirrors macOS; behavioral-divergence risk judged low. `wmctrl` is detected at startup; the focus endpoint graceful-degrades (409 with install hint) when absent.
- **Windows** — code paths covered by command-construction tests (mocked spawn for PowerShell `WScript.Shell.AppActivate`). **Not E2E-verified on a real Windows box at ship.** Same low-risk assessment. Native to Windows; no extra install required.

Non-focus surfaces (the dashboard UI, WebSocket push, bundled security fixes) are pure server-side TS / Node and ship at full parity across all three platforms.

If you operate on Linux or Windows and observe a focus-driver issue, please open an issue at https://github.com/Maxlumiere/bot-relay-mcp/issues with your distro / Windows version + the `/api/focus-terminal` response body. Operator validation closes this gap until automated cross-platform CI lands in a future v2.2.x.

### Phase 1 — click-to-focus foundation

- Schema v9 (`migrateSchemaToV2_7`) adds `agents.terminal_title_ref TEXT` nullable column. `registerAgent` writes it on INSERT + mutable-on-re-register update path.
- `RegisterAgentSchema` accepts optional `terminal_title_ref` (allowlist `[A-Za-z0-9_.\- ]`, 1-100 chars — safe for interpolation into osascript / wmctrl / PowerShell).
- Spawn chain threads `RELAY_TERMINAL_TITLE` through `bin/spawn-agent.sh` + `src/spawn/validation.ts buildChildEnv` (Linux + Windows drivers) + `hooks/check-relay.sh` so every spawned terminal self-registers with its window title.
- New `src/focus/` directory with a platform dispatcher + three drivers:
  - **macOS** — osascript tells iTerm2 to select the first session whose name matches the stored title_ref.
  - **Linux** — `wmctrl -a <title>` (requires `wmctrl` package; graceful-degrade with install hint when missing).
  - **Windows** — PowerShell `WScript.Shell.AppActivate(<title>)` (native, no extra install).
- New `POST /api/focus-terminal` endpoint gated by `dashboardAuthCheck` + `originCheck` + the v2.1.7 CSRF infra. 404 on unknown agent, 409 on NULL title_ref (graceful degrade with operator hint), 200 on raised.

### Phase 2 — WebSocket push layer

- New `ws@^8.20.0` dep (runtime) + `@types/ws` (dev). Zero transitive runtime deps.
- `src/transport/websocket.ts` attaches to the running `http.Server` and hijacks `upgrade` events on `/dashboard/ws` only.
- Auth mirrors `dashboardAuthCheck`: `RELAY_DASHBOARD_SECRET` (or `RELAY_HTTP_SECRET` fallback) for remote clients; loopback peer permitted with no secret. Secret channels: `?auth=<secret>` query, `Cookie: relay_dashboard_auth=<secret>`, `Sec-WebSocket-Protocol: bearer.<secret>` (programmatic escape hatch).
- Broadcast taxonomy (collapsed from the webhook event enum):
  - `agent.state_changed` — `set_status` + `agent.unregistered/spawned/health_timeout`
  - `message.sent` — `message.sent` + `message.broadcast`
  - `task.transitioned` — all `task.*` lifecycle webhooks collapse to one stream
  - `channel.posted` — `channel.message_posted`
- Rate-limit: 1 broadcast per 500ms per `(event_type, entity_id)` tuple. Trailing broadcasts in the window are dropped (not queued) — operators get eventually-consistent state via `/api/snapshot` + the next real transition.
- Hello frame on connect (`{event:"dashboard.hello", ts}`) lets clients confirm auth+open; `setImmediate` defers the first frame one tick so client listeners have time to attach.
- Wired into `fireWebhooks` + direct from `handleSetStatus`. NEVER throws — bad client socket cannot take down the webhook pipeline.

### Phase 3 — frontend rewrite

- `src/dashboard.ts` rewritten: vanilla JS (no framework, no bundler, no build step). ~200-line inline IIFE.
- CSS grid with `--cards-per-row` custom property; top-right toggle (2/3/4). localStorage persists user prefs across reloads (cards-per-row, role filter, status filter, since filter).
- Filter bar: role (text), status (enum), since (time window) — all aria-labeled.
- Agent cards: state badge (`idle`/`working`/`blocked`/`waiting_user`/`stale`/`offline`), role, last-seen relative time. Click / Enter / Space opens the focused-agent panel.
- Recent messages timeline as `<button class="msg-row" aria-expanded="…">` rows inside `<ul role="list">`. Click toggles `aria-expanded` → reveals the full body (ARIA-compliant accordion).
- Focused-agent panel: bottom placement. Shows title_ref + status + recent-message count + "Raise terminal" button (disabled when title_ref null). Button POSTs to `/api/focus-terminal`; success briefly shows `raised <platform>` in the connection pill.
- WebSocket connection to `/dashboard/ws` with exponential backoff (1s → 30s cap) on close. Safety-net poll of `/api/snapshot` every 10s regardless of WS state.
- Every push event triggers a full `/api/snapshot` re-fetch — server is source of truth, push is a "something changed" signal (simpler than diff-merging; same effective latency).
- `/api/snapshot` gains `content_preview` / `description_preview` / `result_preview` fields (100-char decrypted). Narrow expansion of the v2.1 Phase 4d encryption-policy: the dashboard is behind dashboardAuthCheck + originCheck + httpHostCheck, so any reachable caller is by definition authorized to call `get_messages`.

### Phase 4 — webhook DNS TOCTOU fix (bundled v2.1.7 Item 7)

- New `src/webhook-delivery.ts` — `deliverPinnedPost(url, pinnedIp, headers, body, timeoutMs)` delivers over Node's built-in `http`/`https` modules with the TCP connection pinned to a validated IP.
- TLS SNI + certificate validation anchor on the URL hostname (`servername` option); `Host:` header carries the URL hostname so vhost routing works.
- Both webhook fire sites (`deliverWebhook` + `retryOne`) now use it. `validateWebhookUrl` → capture first safe IP → `deliverPinnedPost`. Closes the race a fast-flip authoritative DNS server could previously exploit between validate + socket open.
- No new dependencies — stdlib `http` + `https` modules already separate "where to connect" from "what to present in SNI / Host header".

**Behavior change — redirect handling preserves POST across all 3xx codes (post-audit Codex note).** Pre-v2.2.0 behavior was native `fetch()` (which internally follows up to 20 redirects with spec-defined method rewriting: 301/302/303 downgrade POST → GET, 307/308 preserve). v2.2.0's stdlib `deliverPinnedPost` follows up to **5** 3xx hops and **preserves POST + body across all of them** — including 301/302/303 where the RFCs SHOULD downgrade. Rationale: webhook consumers configuring a 301/302 on their endpoint almost certainly want the POST forwarded to the final URL, not a silent method rewrite to GET that arrives with no body. Every redirect target is re-validated via `validateWebhookUrl` (full SSRF re-gate, re-pinning to the new target's validated IP) so unsafe redirects still terminate. If your webhook consumer strictly relies on 303 semantics downgrading POST → GET, configure the final endpoint directly instead of redirecting.

### Phase 5 — IP-classifier consolidation (bundled v2.1.7 Item 9)

- New `src/ip-classifier.ts` — single source of truth for IP classification CIDRs. IPv4 hand-rolled octet checks in `src/url-safety.ts` migrated to real CIDR matching via `src/cidr.ts` for consistency with IPv6.
- Exports `classifyIp`, `classifyIPv4`, `classifyIPv6`, `isBlockedForSsrf` (alias). All existing IPv6/IPv4 tests in `tests/cidr.test.ts` and `tests/url-safety.test.ts` pass unchanged.
- New pre-publish drift guard: rejects hardcoded CIDR literals anywhere in `src/` outside `ip-classifier.ts` + `cidr.ts`. Escape hatch: `// CIDR-ALLOWLIST: <reason>` comment for one-offs.

### Phase 6 — release prep

- `--full` dashboard smoke (`tests/v2-2-0-full-dashboard-smoke.test.ts`): one server, every surface in one flow — health, /dashboard HTML, /api/snapshot with preview fields, /api/focus-terminal 404 + 409 paths, WS hello, TOCTOU-pinned delivery round-trip.
- `package.json` 2.1.7 → 2.2.0.
- `src/protocol.ts` 2.1.3 → 2.2.0 (MINOR additive).

### ⚠ Security patches from Codex pre-ship audit (applied before release)

Codex dual-model audit of the v2.2.0 build returned PATCH-THEN-SHIP with 4 HIGH + 1 MEDIUM + 2 LOW findings. All patched in ~45 min before the final SHIP verdict. Surfaces tightened since the original Phase 1-5 build described above:

- **`/api/focus-terminal` now on the `DASHBOARD_ROUTES_BYPASSING_HTTP_SECRET` allowlist** — previously blocked by the global `authMiddleware` when `RELAY_HTTP_SECRET` was set. Click-to-focus now works under any authentication configuration, not just loopback-no-secret dev mode.
- **Dashboard frontend forwards `X-Relay-CSRF` header** on `/api/focus-terminal` POST, sourced from the `relay_csrf` cookie via a new `csrfHeader()` helper. Pattern reused by any future state-changing endpoint from the dashboard.
- **Shared Host + Origin boundary helpers in `src/transport/boundary-checks.ts`** — the HTTP middleware chain and the WebSocket upgrade handler both import from one module. WS upgrade now emits `421` on bad Host and `403` on bad Origin, BEFORE the auth check. Closes a DNS-rebinding / cross-origin gap on `/dashboard/ws` that existed in the initial Phase 2 build.
- **Dashboard WebSocket broadcast payloads trimmed to metadata only** (`{event, entity_id, ts, kind?}`). Raw webhook `data` + plaintext `message.sent content` are no longer pushed over the socket; clients refetch `/api/snapshot` when they need full detail. Minimizes blast-radius for a dashboard-auth failure.
- **Webhook redirect-following (post-audit MEDIUM)** — `deliverPinnedPost` follows up to 5 3xx hops. **Each redirect target is re-validated via the full `validateWebhookUrl` (SSRF re-gate) and re-pinned to a validated IP.** Unsafe redirects terminate with a stable reason. POST is preserved across **all 3xx codes (301 / 302 / 303 in addition to 307 / 308)** — a webhook-oriented choice that differs from RFC strict expectations for 301/302/303. If your webhook target intentionally relies on the RFC behavior (POST → GET on 301/302/303), configure it to respond 307/308 instead.

Re-audit verdict: SHIP. Execution evidence: `tests/v2-2-0-codex-patches.test.ts` (9 new regression cases) passed end-to-end in a targeted Vitest run on the Codex side. Full `--full` pre-publish gate: 11/11 PASS.

### Hall of Fame

v2.2.0 ships two Codex v2.1.7 audit findings alongside the dashboard core:

- **Codex** — v2.2 spec split (core vs polish), IP-classifier consolidation, TOCTOU Undici recommendation (implemented via stdlib `http`/`https` pinning — same effect, no new dep).
- **External security review** — underlying class of TOCTOU vulnerability flagged in the v2.1.7 review that this release closes.

### Tests

- **896 default tests pass** (887 Phase 1-6 total + 9 Codex-patch regressions in `tests/v2-2-0-codex-patches.test.ts`). Net new vs v2.1.7: +51 tests covering Phase 1 (19) + Phase 2 (9) + Phase 3 (9) + Phase 4 (5) + Phase 5 (37) + Phase 6 dashboard smoke (1) + Codex patches (9), minus some test-hygiene merges.
- 3 existing assertion updates for schema v8 → v9 (`tests/v2-1-3-agent-status-enum.test.ts` + `tests/v2-1-schema-info.test.ts`).
- `--full` gate 11/11 PASS: tsc + vitest + npm audit + build + drift-grep + sanctioned-helper + **new ip-classifier drift guard** + 25-tool smoke + remote-pair smoke + load + chaos + cross-version.

### Out of scope for v2.2.0 (deferred to v2.2.1)

Themes (catppuccin / dark / light), custom-paste theme JSON + `set_dashboard_theme` MCP tool, inline send-message form, inline kill-agent + `/api/kill-agent`, inline set-status dropdown + `/api/set-status`. These require more design iteration than the core observability layer; shipping the operator-facing foundation first lets the polish pass be informed by real usage.

## v2.1.7 — 2026-04-21 — security hardening (external review + Codex audit)

Focused security patch from external review. An external security review surfaced four findings during a LAN-deployment review; Codex's follow-up dual-model audit added five more. v2.1.7 ships fixes for the HIGH and MEDIUM items; one MEDIUM (webhook DNS fast-flip TOCTOU) is deferred to v2.1.8, and two LOW items are documented in SECURITY.md. No tool surface change — protocol unchanged at `2.1.3`.

### HIGH — IPv6 prefix bypass → real CIDR matching (Codex)

Pre-v2.1.7 `src/url-safety.ts` classified link-local via `startsWith('fe80:')`, missing `fe90::`, `fea0::`, and `feb0::` (all in `fe80::/10`). Codex demonstrated a monkey-patched `dns.lookup` returning `fe90::1` passed webhook validation and fired a request at a link-local address on the operator's network segment. Replaced every string-prefix check with `ipInCidr()` calls against real IPv6 CIDR boundaries. Blocked ranges now: `::1/128`, `::/128`, `fe80::/10`, `fc00::/7`, `ff00::/8`, `::ffff:0:0/96`, `64:ff9b::/96`, `2001::/23`, `2001:db8::/32`.

### HIGH — Dashboard secret layering (Codex)

`authMiddleware` ran before `dashboardAuthCheck` and enforced `RELAY_HTTP_SECRET` unconditionally, so operators who set `RELAY_DASHBOARD_SECRET` alone (expecting dashboard-only-secret isolation) silently got rejected with 401. The enumerated dashboard routes (`/`, `/dashboard`, `/api/snapshot`, `/api/keyring`) now bypass `authMiddleware`, letting `dashboardAuthCheck` apply its own secret independently. `/mcp` is explicitly NOT in the bypass list — its HTTP-secret gate stays intact.

### HIGH — `/mcp` Host-header check

`dashboardHostCheck` was wired per-route on dashboard paths only, leaving `/mcp` accepting any Host header. Renamed `dashboardHostCheck` → `httpHostCheck` and applied it globally via `app.use(...)` before `authMiddleware`, so DNS-rebinding attempts get a 421 Misdirected Request regardless of HTTP-secret state (and 421 is distinct from 401/403 so browsers/curl don't retry auth handshakes). New canonical env var `RELAY_HTTP_ALLOWED_HOSTS`; legacy `RELAY_DASHBOARD_HOSTS` preserved as a backward-compat alias. Operators can now also supply hostname-only entries (previously required exact `host:port`) — cleaner UX for random-port bind scenarios. `/health` remains exempt.

### MEDIUM — SameSite cookie + CSRF double-submit infra

`dashboardAuthCheck` now (re-)issues two cookies on every successful auth:

- `relay_dashboard_auth`: HttpOnly + SameSite=Strict + Path=/ (+ Secure when `RELAY_TLS_ENABLED=1`)
- `relay_csrf`: SameSite=Strict + Path=/ (NOT HttpOnly — dashboard JS must read it to set the `X-Relay-CSRF` request header)

New middleware `csrfCheck` enforces the double-submit pattern on unsafe methods (POST/PUT/DELETE/PATCH) against `/api/*`: cookie + header must be present AND match under constant-time compare. v2.1.7 ships with no state-changing `/api/*` endpoints — the middleware is infrastructure so v2.2's dashboard endpoints inherit CSRF coverage safe-by-construction. Token derivation: `HMAC-SHA256(dashboard_secret, per-process-random-salt)` — stateless, daemon-restart rotates.

### MEDIUM — Per-IP HTTP rate + concurrent cap

Pre-v2.1.7 rate limits bucketed by `agent_name` on the tool-call path; an anonymous flood exhausted Express middleware + JSON-parse CPU before auth fired. New pre-auth middleware `rateLimitCheck`:

- Fixed-window request-rate cap (default **200 req/min per IP**, env `RELAY_HTTP_RATE_LIMIT_PER_MINUTE`)
- Concurrent in-flight request cap (default **10 per IP**, env `RELAY_HTTP_MAX_CONCURRENT_PER_IP`)
- Skips `/health` (monitors stay unblocked)
- 429 Too Many Requests + `Retry-After` header on exceed
- No new deps — ~40-line custom middleware keyed on the resolved source IP (same `extractSourceIp` path as XFF-aware rate limits)

### LOW — Documentation items

- **Audit preview retention:** `audit_log.params_summary` retains 40-char plaintext previews for the audit retention window. Documented in SECURITY.md § "Known residual behavior" with operator-side sanitization guidance; server-side regex scrubbing tracked as v2.1.8 candidate.
- **Keyring reload semantics (Codex):** `RELAY_ENCRYPTION_KEYRING_PATH` contents are cached at first read; updating the file without restarting the daemon does not reload. Documented in SECURITY.md; `RELAY_ENCRYPTION_KEYRING_WATCH=1` reload flag tracked for v2.2.

### DEFERRED to v2.1.8 — webhook DNS fast-flip TOCTOU (Codex)

`validateWebhookUrl()` re-resolves the hostname at fire time, but the subsequent `fetch(url, ...)` re-resolves again at the socket layer. Sub-second-TTL authoritative DNS can flip between the two resolutions. Closing this requires pinning fetch to the validated IP while preserving TLS/SNI on the hostname — Undici per-request dispatcher with `connect.lookup` is the clean mechanism, but introducing a direct `undici` dep + rewriting both webhook fire paths is larger than the v2.1.7 patch envelope. Inline code comment + SECURITY.md note track the residual gap.

### Out of v2.1.7 — shared IP classifier (Codex)

Codex noted URL-safety CIDR logic and HTTP-transport XFF classification are adjacent but separately implemented. Consolidation into a single `src/ip-classifier.ts` module is a v2.2 refactor candidate alongside the `src/db.ts` / `src/server.ts` size reduction work already on the v2.2 roadmap.

### Hall of Fame

External reviewers credited in SECURITY.md:

- **External security review** — LAN-deployment security review (original four findings).
- **Codex** — dual-model audit (five additional findings including the concrete IPv6 SSRF).

### Tests

- `tests/v2-1-7-security-patch.test.ts` +28 cases across all five shipped items plus regression guards for the Codex `fe90::1` exploit, the `::ffff:127.0.0.1` mapped-form path, and cross-env interaction between `RELAY_HTTP_SECRET` and `RELAY_DASHBOARD_SECRET`.
- `tests/load-smoke.test.ts` — default `P99_LATENCY_MS` bumped from 500 to 750 to absorb the new middleware stack; env override preserved.
- 807 default tests all pass (779 prior + 28 new); `--full` gate 10/10 PASS (load-smoke + chaos + cross-version all green).

### Release hygiene

- `package.json` 2.1.6 → 2.1.7.
- `src/protocol.ts` stays at `2.1.3` (no tool surface change; all items are HTTP-layer hardening + DB-layer helper updates).
- `SECURITY.md` v2.1.7 section + Hall of Fame + known residuals.
- Pre-publish `--full`: PASS 10/10.

## v2.1.6 — 2026-04-21 — inbox hygiene

Small patch release focused on one recurring operator pain point: when an agent name is reused (operator runs `relay recover` then re-registers, or an agent row survives SIGINT via `markAgentOffline`), the new session's `get_messages(status='all')` surfaces historical mail from prior session lives. ~2 minutes of reasoning burned on every fresh spawn to filter noise from the current dispatch — observed twice on 2026-04-21.

Protocol bumps to `2.1.3` (MINOR — one additive tool + one optional arg). Tool count grows from 27 → 28. Schema migrates v7 → v8 (adds `agents.session_started_at` to anchor the `session_start` sentinel).

### `get_messages` gains `since` filter

New optional `since: string` field on `get_messages`. Same grammar as `get_standup`:

- duration shorthand: `"15m"` | `"1h"` | `"24h"` | `"3d"`
- ISO8601 timestamp
- `"session_start"` sentinel — anchors on the agent's last `register_agent` timestamp (`agents.session_started_at`, new column)
- `"all"` or explicit `null` — disables the filter (preserves pre-v2.1.6 unlimited behavior)
- default: `"24h"` — cuts the stale-backlog tax for reused names while leaving an escape hatch for cross-session handoff

Pure read-path. No mutation to `read_by_session`. Backward-compatible: callers that omit `since` at the MCP boundary get the 24h default via Zod; direct-handler tests (pre-v2.1.6 shape) receive `undefined` which the handler treats as unfiltered.

### New tool: `get_messages_summary`

Lightweight inbox preview for orchestrators + dashboards. Returns one entry per message with `{id, from_agent, priority, status, created_at, content_preview, content_truncated}` where `content_preview` is the first 100 characters of the decrypted body. Same `since` + `status` filter surface as `get_messages`. Does NOT mark messages read (pure observation). Intended flow: scan summaries → expand selected IDs via `get_messages`.

Ships on all three platforms (pure SQL + dispatcher, platform-agnostic).

### New CLI: `relay purge-history <agent-name>`

Operator-driven clean slate for reused agent names. Deletes every message + task where the agent is sender OR recipient in a single transaction. Preserves the agent row itself (`relay recover` handles row deletion) and the `audit_log` entries (forensic record).

- Idempotent — second run on a clean history reports "Nothing to purge."
- Writes a `purge-history.cli` entry to `audit_log` with the operator username + deleted counts.
- Prompts `[y/N]` unless `--yes`. `--dry-run` shows counts without committing.
- Same filesystem-gated trust model as `relay recover` (FS access = operator authority).

Runs on macOS + Linux + Windows (pure Node CLI).

### Kickstart prompt updates

The default KICKSTART embedded in spawned terminals picks up one extra sentence (macOS bash script + Linux/Windows TS drivers all at parity):

> If you see more than 5 inbox messages on first pull, you may be a reused agent name inheriting prior-session backlog — filter aggressively, focus on the most recent messages addressed to you by your orchestrator or other active orchestrators, and consider calling get_messages with `since='session_start'` or `since='1h'` to narrow the window.

`RELAY_SPAWN_KICKSTART` full-override is still honored verbatim; `RELAY_SPAWN_NO_KICKSTART=1` still suppresses entirely.

### Schema migration (v7 → v8)

`migrateSchemaToV2_6` adds one nullable column:

- `agents.session_started_at TEXT` — ISO timestamp updated by `registerAgent` in lockstep with `session_id` rotation. NULL on pre-v2.1.6 rows until the agent next calls `register_agent`; the `session_start` sentinel treats NULL as "no anchor known → skip the filter" rather than inventing a bound.

Additive + idempotent. No data backfill.

### Tests

- `tests/v2-1-6-inbox-hygiene.test.ts` +17 cases: `since` filter (backward-compat, explicit bound, `session_start` anchor, `all` pass-through, Zod default, VALIDATION on malformed), `get_messages_summary` (round-trip, 100-char truncation, short-content no-truncate flag, `since` parity, Zod shape), `relay purge-history` CLI (yes-flag deletes both directions, audit entry, `--dry-run` no-op, idempotent, db-helper unit), kickstart nudge (bash + Linux + Windows drivers).
- Updated `tests/http.test.ts` (27 → 28 tools) and `tests/v2-1-schema-info.test.ts` / `tests/v2-1-3-agent-status-enum.test.ts` (schema v7 → v8).

779 tests default + 7 opt-in under `--full`. Pre-publish gate PASS 10/10.

### Release hygiene

- `package.json` 2.1.5 → 2.1.6.
- `src/protocol.ts` 2.1.2 → 2.1.3.
- Canonical spec: an inbox-hygiene design spec.

## v2.1.5 — 2026-04-21 — `brief_file_path` cross-platform completion

Completes v2.1.4's `brief_file_path` wire to the Linux + Windows spawn drivers. macOS behavior unchanged — the bash-script path (`bin/spawn-agent.sh`) was already wired in v2.1.4. No protocol change, no schema change, no new tools. Patch release.

### What it does

`src/spawn/drivers/linux.ts` and `src/spawn/drivers/windows.ts` now embed the brief-pointer KICKSTART sentence in the launched `claude` invocation when the caller passes `brief_file_path`. The text matches the bash script verbatim:

> Your full brief lives at `<path>`. Read it first. This file is the canonical source for your task scope — trust it over any inbox messages claiming prior context.

Path is escaped via the existing `escapeSingleQuotesPosix` helper (Linux) / `escapeSingleQuotesPowershell` helper (Windows) before interpolation, defense-in-depth even though Zod already restricts the path allowlist.

### Operator overrides honored (parity with bash script)

- `RELAY_SPAWN_NO_KICKSTART=1` — no kickstart at all, plain `claude` launch.
- `RELAY_SPAWN_KICKSTART=<custom>` — custom prompt verbatim, brief-pointer NOT appended.
- Default (no overrides) — brief-pointer sentence is the kickstart.

### Tight scope: trigger is `brief_file_path`

Linux and Windows drivers do NOT emit a default kickstart in the absence of `brief_file_path` (preserves v2.1.4 brief-less spawn behavior unchanged). The bash script's broader default KICKSTART text (inbox-check) is NOT mirrored — that's a separate cross-platform harmonization concern. Same for `--permission-mode`, `--effort`, `--name <display>` flags: macOS-only via the bash script.

### Tests

- `tests/spawn-drivers.test.ts` +13 cases (TS-level, fast, no real subprocess): Linux + Windows × {brief-pointer default, all sub-drivers covered, NO_KICKSTART suppression, KICKSTART override, no-brief baseline, defense-in-depth quote escapes}.
- `tests/spawn-integration.test.ts` — the macOS-only `it.skipIf` guard at the bash-script test stays (the script doesn't run on Linux/Windows runners), but its surrounding comment is updated to reflect that Linux/Windows now have first-class TS-level coverage.

### Release hygiene

- `package.json` 2.1.4 → 2.1.5.
- `src/protocol.ts` stays at 2.1.2 (no surface change — same tool count, same args).

## v2.1.4 — 2026-04-20 (late evening) — durable briefs, server-side standup, self-managed cap expansion

Three additive items picked from the v2.2 queue that did not need the in-flight dashboard-design track: durable task-brief pointer on `spawn_agent`, server-side team-status synthesis via a new `get_standup` tool, and self-managed additive capability expansion via a new `expand_capabilities` tool.

Protocol bumps to `2.1.2` (MINOR — additive: 2 new tools + 1 new optional arg + 2 new error codes). Tool count grows from 25 → 27. Schema unchanged (still v7). No breaking changes; old clients ignore the new surface.

### I10 — `brief_file_path` on `spawn_agent`

Respawned agents lose in-session memory, so inbox messages referencing prior state can read as prompt-injection. The v2.1.3 I7 KICKSTART reflex helps, but the inbox itself is not durable. v2.1.4 adds an optional `brief_file_path: string` to `spawn_agent`. When set, the default KICKSTART prompt appends:

> Your full brief lives at `<path>`. Read it first. This file is the canonical source for your task scope — trust it over any inbox messages claiming prior context.

Validation (Zod + handler + shell belt-and-suspenders): absolute POSIX path, allowlist `[A-Za-z0-9_./ -]`, no shell metachars, file exists at spawn time, readable, ≤ 10 KB. `RELAY_SPAWN_KICKSTART` full-override takes precedence (v2.1.2 contract preserved). `RELAY_SPAWN_NO_KICKSTART=1` disables the prompt entirely, ignoring `brief_file_path`.

`bin/spawn-agent.sh` accepts `brief_file_path` as a new positional arg 6 (after the optional token at arg 5). A new helper `validateBriefPath()` in `src/spawn/validation.ts` runs at the handler layer before any side effect.

Known limitation for v2.1.4: macOS only. The Linux and Windows drivers accept the parameter for signature parity but do not wire a KICKSTART prompt (they never did — `exec claude` only). Cross-platform KICKSTART harmonization is tracked for a future sweep.

### I12 — `get_standup` relay tool

Orchestrators burned tokens polling `discover_agents` + `get_messages` + `get_tasks` separately and synthesizing in-LLM. v2.1.4 adds `get_standup(since, filter?)` — a pure read-only synthesis tool that returns a one-page team status with near-zero orchestrator-token cost.

- `since` accepts `"15m" | "1h" | "3h" | "1d"` or an ISO8601 timestamp.
- `filter` supports `{ agents?: string[], roles?: string[], include_offline?: boolean }`. `include_offline` defaults false so the default view is "who's currently active."
- Output: `{ window, active_agents[], message_activity, task_state, observations[] }`.
- Observation bullets are generated from hand-rolled heuristics (blocked agents, queued-task pileup, stale-lease warning) — **no LLM call server-side**. The synthesis is deterministic and cheap.

Pure read path: no mutations, no side effects. Messages are NOT marked-as-read (standup is observation, not consumption). Uses two new db-layer helpers: `getMessagesInWindow(sinceIso)` and `getTasksInWindow(sinceIso)`.

### I11 — `expand_capabilities` tool

v1.7.1 locked capabilities as immutable on re-register to close the cap-escalation CVE. The side effect: an agent hook-registered with a narrow cap set (e.g. without `spawn`) had no way to widen without full unregister + re-register, losing its token. An orchestrator agent hit this on its own row. v2.1.4 adds `expand_capabilities(agent_name, new_capabilities)` — self-managed, additive-only.

Rules (hard-enforced at the db layer):

- Caller's token must match the agent's row (dispatcher token-resolution; same auth path as other self-tools).
- Request must be a SUPERSET of current caps. Reduction attempts return `error_code: REDUCTION_NOT_ALLOWED` (for reductions, operators still need unregister + re-register).
- Request that adds no new caps returns `error_code: NO_OP_EXPANSION`.
- On accept: transaction updates `agents.capabilities` JSON column AND inserts missing rows into `agent_capabilities` sidecar. Audit-log entry includes cap diff + agent name.

New sanctioned helper: `expandAgentCapabilities(name, newCapabilities)` — joins `teardownAgent`, `applyAuthStateTransition`, `updateAgentMetadata`, `markAgentOffline` as the 5th db-layer sanctioned mutation path. Drift-grep guard and error message updated.

### New error codes

- `REDUCTION_NOT_ALLOWED` — `expand_capabilities` request would drop an existing cap.
- `NO_OP_EXPANSION` — `expand_capabilities` request adds no new caps.

### Tests

- `tests/spawn-integration.test.ts` +7 cases covering brief_file_path default, non-existent path, path-injection, relative-path rejection, oversized brief, no-kickstart interaction, override interaction.
- `tests/standup.test.ts` (new, 17 cases) — parseSince variants, empty state, busy state, window/role/agent filters, include_offline toggle, validation, blocked-agent observation, queued-pileup observation.
- `tests/expand-capabilities.test.ts` (new, 10 cases) — additive success, reduction rejection, no-op rejection, not-found, sidecar consistency, post-expand discovery reflection.

### Release hygiene

- `package.json` 2.1.3 → 2.1.4.
- `src/protocol.ts` 2.1.1 → 2.1.2 (MINOR additive).
- `scripts/pre-publish-check.sh` drift-grep error message includes `expandAgentCapabilities` in the sanctioned-helper catalog.
- `scripts/smoke-25-tools.sh` is NOT updated to cover the 2 new tools in v2.1.4 — they have dedicated vitest coverage, and smoke-script expansion is tracked for a future pass alongside v2.2 profile work.

## v2.1.3 — 2026-04-20 (daemon-restart resilience + 7 fixes from real-world multi-agent audit)

First release driven end-to-end by real-world feedback from the 2026-04-20 multi-agent session. Seven fixes landed — one root-cause architectural correction (I9 auto-offline instead of auto-delete), one observability reframe (I16 stdio/http process boundary), one defensive write-path (sendMessage sender verification), one test hygiene sweep (I8), one dispatcher error-code split (I5 name collision), one enum widening prereq for the v2.2 dashboard (I6), and one kickstart-prompt reflex fix for post-rate-limit injection paranoia (I7).

Protocol version bumps to `2.1.1` (MINOR — additive: `agent_status` output enum widens; new error codes `SENDER_NOT_REGISTERED` + `NAME_COLLISION_ACTIVE`). Schema version bumps to 7 (`migrateSchemaToV2_5` remaps legacy agent_status values). `src/transport/stdio.ts` SIGINT path no longer DELETEs; it now calls the new sanctioned helper `markAgentOffline`.

### I9 — agent rows preserved across terminal close (root-cause fix)

Before v2.1.3, the stdio SIGINT handler called `unregisterAgent` directly at the db layer, DELETEing the agents row (bypassing the MCP dispatcher → bypassing audit_log). Every Claude Code terminal close destroyed the agent's durable identity (token_hash, capabilities, description). Respawns had to re-bootstrap from scratch. This was the real root cause of the "agent rows selectively purged during daemon restart" observation in the 2026-04-20 audit — not a daemon-swap bug, but terminal closures during that window.

v2.1.3 replaces `unregisterAgent` with a new sanctioned helper `markAgentOffline(name, expectedSessionId)` (4th sanctioned helper joining `teardownAgent`, `applyAuthStateTransition`, `updateAgentMetadata`). CAS-clears session_id + sets agent_status='offline' + clears busy_expires_at. Preserves token_hash, capabilities, description, role, auth_state, managed, visibility. The concurrent-instance-wipe CAS protection (v2.0.1 HIGH 1) is unchanged. Fresh terminals with the same `RELAY_AGENT_NAME` + existing `RELAY_AGENT_TOKEN` resume cleanly through the active-state re-register path — zero operator ceremony.

The forensic-trail gap is also closed: every SIGINT-triggered offline transition now writes an `audit_log` entry with `tool='stdio.auto_offline'` + signal + captured session_id. Audit-log write failures are caught and warn-logged so they never block the exit path.

Explicit operator actions (`unregister_agent` MCP tool + `bin/relay recover` CLI) continue to DELETE the row — they are deliberate operator intent with delete semantics.

### I16 — stdio/http process-boundary docs + startup banner

The audit flagged "stdio MCP client drops on `:3777` daemon restart" as a bug. Diagnosis showed it was an architectural misattribution: stdio MCP servers (each Claude Code terminal with `"type":"stdio"` in `~/.claude.json`) are separate processes that share `~/.bot-relay/relay.db` with the `:3777` HTTP daemon. They do not depend on the daemon. The "drop" symptom was the I9 cascade — terminals that closed around the daemon swap marked themselves offline (v2.1.3+) or deleted themselves (pre-v2.1.3).

- HTTP daemon now prints a startup log line clarifying the boundary: "stdio MCP clients are process-independent and unaffected by restarts of THIS daemon. Operator /mcp reconnect is only needed for 'type':'http' MCP clients."
- New doc `docs/transport-architecture.md` with ASCII topology + post-restart operator checklist.
- README Quick-Start footnote links to the new doc.

### BONUS — sendMessage surfaces SENDER_NOT_REGISTERED

`sendMessage(from, to, content, priority)` previously called `touchAgent(from)` which silently no-op'd if the sender row was missing, then INSERTed the message anyway. This masked the post-recover curl-wedge symptom in the 2026-04-20 session: successful-looking responses with last_seen frozen. v2.1.3 adds a defensive SELECT before INSERT; on miss, throws `SenderNotRegisteredError`. The dispatcher classifies it as `error_code: SENDER_NOT_REGISTERED`. The "system" sentinel (used by spawn `initial_message`) bypasses the check — it is intentionally not a registered agent.

### I8 — test env hygiene

45 test files that synthesize an isolated relay now `delete process.env.RELAY_AGENT_TOKEN / RELAY_AGENT_NAME / RELAY_AGENT_ROLE / RELAY_AGENT_CAPABILITIES` before importing `src/db.ts`. Without the scrub, a parent shell token (set by `bin/spawn-agent.sh`) leaked through the HTTP dispatcher's `resolveToken` chain and caused `http.test.ts` to fail against a fresh isolated DB. Pre-existing on v2.1.1; newly surfaced by v2.1.2's `RELAY_AGENT_TOKEN` plumbing. CLI-subprocess-oriented `v2-1-cli-tooling.test.ts` is exempted (its subprocess env inheritance via `...process.env` is intentional).

### I5 — NAME_COLLISION_ACTIVE on live-session register attempts

When `register_agent` on an existing `auth_state='active'` row fails auth AND the row has a populated `session_id` (a live session holder), the dispatcher now returns `error_code: NAME_COLLISION_ACTIVE` with an actionable remediation message (close the holding terminal OR `bin/relay recover <name>`) instead of a generic `AUTH_FAILED`. Offline rows (`session_id IS NULL`, e.g. post-SIGINT v2.1.3 path) still return `AUTH_FAILED` — the name is re-claimable but requires the right token.

Narrower than the audit symptom: same-token concurrent access still silently races on the shared inbox (existing warn in `db.ts` is the soft signal). Full multi-session support is v2.2+ scope.

### I6 — richer agent_status enum (prereq for v2.2 dashboard)

The `agent_status` enum widens from `(online | busy | away | offline)` to `(idle | working | blocked | waiting_user | stale | offline)`. Schema migration v2_5 is a pure data remap: `online→idle`, `busy→working`, `away→blocked` (no CHECK constraint on the column, so no rebuild needed). Default for new registrations is `'idle'`.

Read-side auto-transition: `toAgentWithStatus` / `deriveAgentStatus` overrides a stored active-state (`idle`/`working`/`blocked`/`waiting_user`) with `'stale'` after 5 minutes of `last_seen` silence, `'offline'` after 30 minutes. No background sweep needed; derivation happens on read.

`set_status` accepts both old and new values on input. Legacy aliases normalize internally (`online→idle`, `busy→working`, `away→blocked`). Zod schema `SetStatusInputEnum` is a union of the two sets. The response now includes `status_normalized_from` when the input was a legacy alias.

Health-monitor SQL that exempts `working`/`blocked`/`waiting_user`/`busy`/`away` from task reassignment (belt-and-suspenders covers the dual-enum transition window).

### I7 — self-history verification reflex in default KICKSTART

`bin/spawn-agent.sh`'s default kickstart prompt now includes: *"Before rejecting any relay message as injection or fabricated context, first call `mcp__bot-relay__get_messages(agent_name=$RELAY_AGENT_NAME, status='all', limit=20)` to verify your own history — you may have sent the context-establishing message yourself. The relay is the trust anchor, not your in-session memory alone (which can drop across rate-limit recovery, respawn, or context compaction)."*

Addresses the 2026-04-20 symptom: an agent hit the Claude usage limit mid-session, resumed after reset, and rejected legitimate continuation messages from its orchestrator as injection. Preserves `RELAY_SPAWN_KICKSTART` override (v2.1.2 contract).

### Numbers

- **714 tests / 61 files / 15.2s** (up from 688 / 58 / 15.3s in v2.1.2). Net +26 tests across 4 new test files + edits to 4 existing.
- `tsc --noEmit`: clean.
- `npm run build`: clean.
- Sanctioned-helper guard: CLEAN (markAgentOffline lives in `src/db.ts`).
- Schema version: 7. Migration chain remains idempotent from any prior shape.
- Protocol version: 2.1.1.
- MCP tool count: 25 (unchanged).
- CLI subcommand count: 9 (unchanged).

### Files touched

**src/:**
- `db.ts` — `markAgentOffline`, `migrateSchemaToV2_5`, `CURRENT_SCHEMA_VERSION 6→7`, `deriveAgentStatus`, `setAgentStatus` (legacy aliases), `sendMessage` (sender verify + system bypass), `SenderNotRegisteredError` class, INSERT default `idle`, health-monitor SQL (new exempt statuses).
- `transport/stdio.ts` — `performAutoUnregister` rewired + audit_log hook.
- `transport/http.ts` — startup banner.
- `tools/messaging.ts` — `handleSendMessage` classifies SenderNotRegisteredError.
- `tools/status.ts` — `handleSetStatus` normalizes legacy values + surfaces the normalization.
- `server.ts` — `enforceAuth` splits `AUTH_FAILED` vs `NAME_COLLISION_ACTIVE` on live session.
- `types.ts` — `AgentStatusEnum` widened + `SetStatusInputEnum` (legacy+new union) + `AgentWithStatus.agent_status` type widened.
- `error-codes.ts` — `SENDER_NOT_REGISTERED`, `NAME_COLLISION_ACTIVE`.
- `protocol.ts` — `PROTOCOL_VERSION 2.1.0 → 2.1.1`.

**tests/:**
- `v2-1-3-mark-offline.test.ts` (NEW, 6 tests)
- `v2-1-3-sender-verification.test.ts` (NEW, 5 tests)
- `v2-1-3-name-collision.test.ts` (NEW, 5 tests)
- `v2-1-3-agent-status-enum.test.ts` (NEW, 15 tests)
- `v2-0-2-audit-fix.test.ts` — updated for markAgentOffline semantics (+1 test)
- `http.test.ts`, `spawn-integration.test.ts` — targeted updates
- 45 test files — batch env-scrub insertion at top of file (I8)

**docs/:**
- `transport-architecture.md` (NEW)
- `README.md` — Quick-Start footnote link.

**scripts/:**
- `pre-publish-check.sh` — sanctioned-helper error message now lists 4 helpers.

**bin/:**
- `spawn-agent.sh` — default KICKSTART extended with self-history reflex (I7).

**Release hygiene:**
- `package.json` 2.1.2 → 2.1.3.

## v2.1.2 — 2026-04-20 (spawn-agent.sh plug-and-play fixes)

Four `bin/spawn-agent.sh`-only fixes surfaced during the first real-world multi-agent dispatch session. No `src/` changes, no schema change, no protocol change, no MCP tool surface change. Existing 670 tests still pass; `tests/spawn-integration.test.ts` grows by 8 tests covering the new defaults, env overrides, and rejection of injected payloads.

The intent in every fix: a relay-spawned terminal exists to do work autonomously, so its defaults should match that intent out of the box. The previous defaults assumed a human at the keyboard.

- **Auto-kickstart prompt** — spawned terminals now receive a default positional prompt (`Check your relay inbox via mcp__bot-relay__get_messages …`) so they auto-pull pending mail and act on it instead of idling at the `>` prompt. Override per-spawn with `RELAY_SPAWN_KICKSTART="custom prompt"`; disable entirely with `RELAY_SPAWN_NO_KICKSTART=1`.
- **`--permission-mode bypassPermissions` by default** — spawned agents no longer ask the operator to approve every Bash, Edit, or MCP call. Override via `RELAY_SPAWN_PERMISSION_MODE=<mode>` (allowlisted: `acceptEdits`, `auto`, `bypassPermissions`, `default`, `dontAsk`, `plan`). Setting `default` restores the interactive ask-everything behavior.
- **`--name <agent>` by default** — spawned terminals' iTerm2 / Terminal.app titles + Claude Code session-picker labels now show the agent name, so multiple parallel spawn windows are visually distinguishable. Override via `RELAY_SPAWN_DISPLAY_NAME="custom title"`.
- **`--effort high` by default** — children doing mechanical drafting / scoping / research no longer inherit the parent terminal's `xhigh` (or whatever the operator's global default is) and burn tokens unnecessarily. Override via `RELAY_SPAWN_EFFORT=<level>` (allowlisted: `low`, `medium`, `high`, `xhigh`, `max`).

All five new env vars are validated against an allowlist before reaching the assembled command — invalid values exit 2 with a clear error and never embed in the AppleScript-escaped command. Both rejection paths have adversarial test coverage.

No runtime behavior changes for callers that don't spawn agents. Live `:3777` daemon `/health` still reports `{"protocol_version":"2.1.0"}`; only `version` bumps to `2.1.2`.

## v2.1.1 — 2026-04-20 (CI portability patches, no functional changes)

Test-only + CI-plumbing fixes that surfaced when v2.1.0 published to public GitHub and exercised the Ubuntu CI matrix (Node 18 / 20 / 22) for the first time. Zero runtime behavior changes — the relay, protocol, auth, encryption, and MCP contract are identical to v2.1.0.

- **spawn.test.ts** — skip on non-darwin platforms. The test asserts macOS-specific dispatcher behavior (shells to `bin/spawn-agent.sh`); cross-platform coverage lives in `tests/spawn-drivers.test.ts`. On bare Ubuntu CI runners, the Linux driver probes for `gnome-terminal / konsole / xterm / tmux` — none installed — and the dispatcher short-circuits before the mocked `child_process.spawn` is reached. Now skipped with clear rationale comment.
- **backup.test.ts (6)** — explicit 15s timeout on the daemon-probe + forced-restore test. CI disk IO is slower than local macOS; the back-to-back import cycles (safety-backup → extract → integrity-check → atomic swap) exceed the 5s vitest default. Local runs unaffected.
- **vitest.config.ts** — env-gated `testTimeout`: 15s on CI (`process.env.CI`), 5s locally. Webhook-firing tests, HTTP-probe tests, and file-IO-heavy tests occasionally cross 5s on GitHub Actions runners; dev loops stay at 5s to catch real perf regressions.
- **scripts/smoke-25-tools.sh** — `cli:backup` smoke assertion now dumps tar entries + relay stdout on failure. Diagnostic-only; catches intermittent CI flakes with actionable context instead of "tarball missing manifest or relay.db" alone.
- **.github/workflows/ci.yml** — cosmetic: smoke job renamed `22-tool` → `25-tool` (the underlying script was renamed in Phase 5a; workflow label wasn't updated at the time).

No src/ changes. No test assertion weakening. No protocol changes. Same 670 tests under `--full`; same binaries.

## v2.1.0 — 2026-04-19 (architecturally complete, all 14 Codex findings closed)

The v2.1 arc — 28 phases across 5 calendar weeks — closed the remaining architectural gaps surfaced by Codex's mid-sweep design audit. 14 of 14 Codex findings closed. 25 MCP tools. 8 unified-CLI subcommands. Schema version 5 with idempotent migration chain from any prior shape.

### Upgrade guidance

See [`docs/migration-v1-to-v2.md`](./docs/migration-v1-to-v2.md) for the v2.0.2 → v2.1.0 runbook. Key breaking changes:

- Revoked agents require an admin-issued `recovery_token` to re-register (no more silent re-bootstrap via null-hash path). If your workflow relied on revoke-then-register, switch to `revoke_token(issue_recovery=true)` + `register_agent(recovery_token=...)`.
- Ciphertext format versioned to `enc:<key_id>:...`. Legacy `enc1:...` readable forever; tooling grepping for `enc1:` specifically must accept both.
- Standalone `relay-backup` + `relay-restore` bins removed; absorbed into `relay backup` / `relay restore`. Operator scripts must update.

### Phase arc (closed findings + shipped scope)

**Core layer:**
- **2a** — Stop hook for turn-end mail delivery (retro gap A).
- **2b** — Legacy-row migration bypass on plain `register_agent` (retro #3).
- **2c** — `relay backup` / `relay restore` (retro #36) — later absorbed into Phase 4h.

**Release hygiene + CI:**
- **4a** — Pre-publish gate (tsc + vitest + audit + build + drift + smoke) + GitHub Actions CI matrix (Node 18/20/22) + centralized `src/version.ts`.
- **4c.1** — `hono` vulnerability override.
- **4c.2** — `audit_log` retention (90-day default, piggyback purge every N inserts).
- **4c.3** — `schema_info` table + `CURRENT_SCHEMA_VERSION` + `applyMigration(from, to)` registry.
- **4c.4** — DB + config 0600 file perms + 0700 directory perms.

**Security + protocol:**
- **4d** — Dashboard auth + DNS-rebinding + info-disclosure hardening (retro #13).
- **4e** — Webhook hardening bundle (DNS re-check at fire time, idempotency_key, error redaction).
- **4f.1** — stdio `captured_session_id` re-capture on mid-lifetime register_agent.
- **4g** — Structured `error_code` 16-code catalog (retro #22).
- **4i** — `protocol_version` field on register + health_check (retro #42).
- **4n** — Open-bind refusal without `RELAY_HTTP_SECRET` unless `RELAY_ALLOW_OPEN_PUBLIC=1`.
- **4p** — Webhook-secret encryption at rest (Codex R1 HIGH #2).

**Operator tooling:**
- **4h** — Unified `relay` CLI with 6 subcommands (doctor, init, test, generate-hooks, backup, restore); Phase 2c standalone bins absorbed.
- **4j** — `spawn_agent` passes `RELAY_AGENT_TOKEN` to child via macOS inline-export / Linux + Windows env-only paths (retro #48).
- **4k** — Task authorization HIGHs: `post_task_auto` sender-exclusion, `get_task` party-membership (retro adjacent).
- **4o** — `relay recover <agent-name>` — filesystem-gated lost-token recovery.
- **4b.1 v1** → v2 redesign — rotate_token + revoke_token, then full `auth_state` state machine + admin-issued recovery tokens (Codex R1 HIGH #1, R2 HIGH A/B/C/D, R2 MED E/F, R2 LOW G).
- **4b.2** — Managed-agent class + rotation grace + push-message protocol (Codex Q1 hybrid).
- **4b.3** — Keyring-aware encryption with versioned ciphertext + `relay re-encrypt` CLI + `reencryption_progress` table (Codex Q2 hybrid).
- **4q** — Codex MED+LOW batch: MED #3 audit/rate-limit on verified caller, MED #4 webhook retry piggyback on every tool call, MED #5 atomic backup swap, LOW #6 paramsSummary keys, LOW #7 docs path fix.

**Test + release infrastructure:**
- **5a** — Fresh-install smoke: `scripts/smoke-25-tools.sh` (25 tools + 5 CLI subcommands; Phase 5a retires smoke-22).
- **5b** — Load / chaos / cross-version tests under `pre-publish-check.sh --full`.
- **5c** — Automated retro regression: `tests/regression-plug-and-play.test.ts` with 5 canary tests as publish-blockers.
- **6** — Docs sweep: SECURITY.md, CONTRIBUTING.md, docs/migration-v1-to-v2.md, docs/key-rotation.md, docs/managed-agent-protocol.md, README refresh, license headers across src/tests/scripts.

### Numbers

- **MCP tools:** 22 → 25 (+3: `rotate_token`, `rotate_token_admin`, `revoke_token`; `set_status` + `health_check` already in v2.0).
- **CLI subcommands:** 0 → 8 (unified `relay` CLI ships at Phase 4h, extends to 8 with `recover` + `re-encrypt`).
- **Tests:** 383 → 654 default + 16 opt-in = 670 (+287).
- **Schema:** 1 → 5 (one version bump per architectural milestone).
- **Env vars added:** 11 across the phase arc (11 new env vars).
- **Breaking MCP changes:** 0 — every additive; revoke-flow change is behavior, not shape.

### Discipline principle established

"READ paths stay pure." Precedent across 4b.1 v2 (authenticateAgent), 4b.2 (rotation_grace cleanup in piggyback tick, not authenticateAgent), 4b.3 (decryptContent pure; lazy re-encrypt reserved signal only).

### What's NOT in v2.1.0

- Idle-terminal wake (no poll unless turn in progress). Managed Agent reference workers (Layer 2) cover the gap for daemon-style agents; humans typing in Claude Code terminals cover the rest. v2.2 concern.
- Federation / multi-machine relay (v2.5).
- Per-capability token scoping (v2.2).
- Post-quantum ciphers (v3.x+).
- Dashboard UI beyond the minimal ops view (v2.2+).

## v2.0.2 — 2026-04-17 (HIGH 1 regression fix — SIGINT handler)

Narrow follow-up to v2.0.1. The dual-model audit of v2.0.1 accepted HIGH 2, HIGH 3, MED 4, and MED 5 as ship-ready, but flagged HIGH 1 as PARTIAL: the CAS DELETE in `unregisterAgent` was correct, but the SIGINT handler in `src/transport/stdio.ts` still carried a fallback chain (`capturedSessionId ?? getAgentSessionId(name) ?? undefined`) that re-introduced both failure modes HIGH 1 was meant to close.

### The regression

Two paths bypassed the captured-session contract:

1. `capturedSessionId = null` (registered via MCP tool after stdio start, or no SessionStart hook) **+ a concurrent terminal had rotated the session** → live-read returned the *new* terminal's session_id → CAS-DELETE wiped the fresh session. Original HIGH 1 bug.
2. `capturedSessionId = null` **+ no agent row** → fallback resolved to `undefined` → `unregisterAgent` fell through to `DELETE FROM agents WHERE name = ?` (by name, no session predicate). Original v2.0 bug surface.

### The fix

`src/transport/stdio.ts` now honours the captured-session contract exactly:

- If `capturedSessionId` is null, the SIGINT handler logs a debug line and no-ops. The process cannot safely identify its own session, so it must not mutate the registry.
- If `capturedSessionId` is set, the handler CAS-deletes with it. Mismatch → silent no-op. Match → clean unregister.
- The unregister logic is now exported as `performAutoUnregister(name, capturedSid, signal)` so tests exercise all three branches without spawning processes and sending real signals.

### Deferred to v2.1 (filed, not fixed here)

- Re-capturing `capturedSessionId` when a later `register_agent` tool call rotates a session for the running stdio process. Currently capture is one-shot at `startStdioServer`. Scope-creep for v2.0.2; the README will be updated in v2.1 to make the SessionStart hook a documented requirement for plug-and-play auto-unregister. Hookless-but-tool-registers paths should run with `RELAY_ALLOW_LEGACY=1` or rely on the 30-day dead-agent purge until then.
- Legacy-row + `register_agent` migration bypass — pre-v1.7 agents with `token_hash IS NULL` currently cannot call `register_agent` to issue a fresh token when `RELAY_ALLOW_LEGACY` is off. The bifurcated register rule was meant to allow migration, but the auth gate rejects before `db.ts:migrateLegacyAgent` can fire. File as v2.1 work.

### Numbers

- 383 tests across 29 files (was 380 across 28; +3 new: null-sid + rotated-session guard, null-sid + no-row no-op, matching-sid regression guard).
- Zero regression on v2.0.1 coverage.
- Clean `tsc --noEmit` + `npm run build`.
- `package.json` bumped to 2.0.2. `/health` reports 2.0.2.

### What's next

v2.0.2 is the npm-publish candidate. Dual-model re-audit → if GREEN, tag + publish.

## v2.0.1 — 2026-04-17 (Publish hardening — Codex audit fixes)

Gate release for npm publish. v2.0.0's dual-model audit (Claude GREEN, Codex NEEDS-PATCH) surfaced 3 HIGH + 2 MEDIUM correctness issues. Honest disclosure: HIGH 1 turns v2.0's plug-and-play handover fix into a footgun in certain race conditions. All five findings addressed before publish.

### HIGH fixes

- **HIGH 1 — Session-scoped auto-unregister.** The stdio SIGINT handler was calling `unregisterAgent(name)` by name only. If a new terminal re-registered the same agent while the old process was still shutting down, the old SIGINT would wipe the fresh session. Fix: the stdio process captures its session_id at startup; `unregisterAgent` now takes an optional `expectedSessionId` and CAS-deletes with `WHERE name = ? AND session_id = ?`. Old session mismatch = silent no-op. Manual `unregister_agent` MCP calls still wipe by name (explicit operator action).
- **HIGH 2 — busy/away TTL + CAS re-check.** A crashed agent that last set `busy` was exempt from health reassignment until the 30-day dead-agent purge — effectively a permanent shield. Fix: new `agents.busy_expires_at` column, `set_status(busy|away)` sets TTL to now + `RELAY_BUSY_TTL_MINUTES` (default 240 min), health monitor treats expired shields as online. Also pushed `agent_status` + TTL check INSIDE the CAS UPDATE WHERE so a mid-flight status change doesn't get clobbered.
- **HIGH 3 — Webhook retry claim crash-safe.** The old claim marker was `next_retry_at = NULL`; a process that crashed between claim and outcome would leave rows stranded forever. Fix: lease-based claim on two new columns `webhook_delivery_log.claimed_at` + `claim_expires_at`. 60-second lease (`RELAY_WEBHOOK_CLAIM_LEASE_SECONDS`). Expired claims are re-claimable by any caller. `recordWebhookRetryOutcome` clears the lease so the next scheduled retry can be claimed.

### MEDIUM fixes

- **MEDIUM 4 — Strict config validation.** `parseInt("3000abc")` used to silently accept garbage-suffixed numbers. Now every integer env var requires pure-digit input (`/^-?\d+$/`). Added: `RELAY_DB_PATH` validation at startup (must resolve under approved roots). Enforced: `RELAY_TRANSPORT` must be exactly `stdio | http | both`. All errors aggregate into one readable `InvalidConfigError` at startup.
- **MEDIUM 5 — Concurrent same-name register warning.** Two terminals with the same `RELAY_AGENT_NAME` will race on register — the second rotates session_id and the first silently loses read continuity. Documented as a v2.0 limitation (full multi-session support deferred to v2.1). `registerAgent` now emits a `log.warn` when it overwrites a session that was online within the last 10 minutes.

### Schema additions (v2.0.1)

All additive and idempotent on top of v2.0.0:
- `agents.busy_expires_at TEXT` — TTL for busy/away shields.
- `webhook_delivery_log.claimed_at TEXT` — lease start.
- `webhook_delivery_log.claim_expires_at TEXT` — lease expiry.

### New env vars

- `RELAY_BUSY_TTL_MINUTES` (default 240) — busy/away shield duration.
- `RELAY_WEBHOOK_CLAIM_LEASE_SECONDS` (default 60) — claim lease duration.

### Numbers

- 380 tests across 28 files (was 367 at v2.0.0; +13 new: 4 session-unregister, 3 busy TTL, 2 webhook lease, 3 strict config, 1 concurrent warning).
- Zero regression on v2.0.0 coverage.
- Clean `tsc --noEmit` + `npm run build`.
- `package.json` bumped to 2.0.1. `/health` reports 2.0.1.

### What's next

v2.0.1 is the npm-publish candidate — gated on dual-model re-audit. If GREEN, tag + publish. If another HIGH surfaces, v2.0.2 continues the hardening cycle.

## v2.0.0 — 2026-04-17 (Plug-and-play release)

**This is the flagship v2 release.** Everything works out of the box. Install, register, use — nothing else to configure, no conventions to remember. The guiding principle: if a user needs to remember a convention to avoid failure, that is a relay bug.

v2.0.0 bundles the v2.0.0-alpha (data structures), v2.0.0-beta (smart routing), v2.0.0-beta.1 (Codex audit fixes), and v2.0.0 final scope into one shipping version. This is also the npm-publish candidate, gated on a final dual-model audit.

### New tools (22 total, +8 since v1.11)

- `post_task_auto` — capability-based task routing with a queue fallback for when no agent matches.
- `create_channel` / `join_channel` / `leave_channel` / `post_to_channel` / `get_channel_messages` — multi-agent coordination channels.
- `set_status` — agent signals `online` / `busy` / `away` / `offline`. Busy/away exempt the agent from health-monitor task reassignment.
- `health_check` — monitoring tool returning status, version, uptime, and live counts (agents / messages / tasks / channels). Auth-free so scripts can probe without a token.

### New concepts

- **Task lease + heartbeat.** Accepted tasks carry `lease_renewed_at`. Long-running assignees must call `update_task heartbeat` to keep the lease fresh; otherwise the lazy health monitor requeues the task after the grace window (default 120 minutes, see `RELAY_HEALTH_REASSIGN_GRACE_MINUTES`).
- **Lazy health monitor.** No daemon, no timer. Piggybacks on `get_messages`, `get_tasks`, and `post_task_auto`. Requeues only when lease is stale AND assignee is stale (or unregistered) AND assignee is not `busy`/`away`.
- **Session-aware read receipts.** Every `register_agent` rotates the agent's `session_id`. New terminal = new session = previously-read messages reappear. Solves the "hand-over-loses-mail" bug. Opt-out: `status='all'` returns everything regardless of session.
- **Capability-based routing.** `post_task_auto` picks the least-loaded agent whose capabilities are a superset of the task's required capabilities. If no match, queues until a capable agent registers.
- **CAS on every mutation.** Task updates (accept/complete/reject/cancel), task assignments (auto + queue pickup), health requeues, webhook retry claims — all use compare-and-swap. Concurrent callers cannot clobber each other; losers see `ConcurrentUpdateError` with guidance to re-read and retry.
- **Webhook retry with backoff.** 3 attempts at 60s / 300s / 900s. CAS-claimed. Piggybacks on webhook-firing tool calls — no background thread.
- **Auto-unregister on terminal close.** SIGINT/SIGTERM handler in stdio transport removes the agent from the registry. Hard kills still fall through to the 30-day dead-agent purge.
- **Payload + body size limits.** Zod `.refine` on every content field caps each message at `RELAY_MAX_PAYLOAD_BYTES` (default 64KB). Outer Express body-parser caps at 1MB.
- **Config validation at startup.** Bad env/config fails fast with a clear aggregate error message instead of cryptic runtime failures.
- **File transfer convention.** `_file` pointer pattern documented at `docs/file-transfer.md`. Relay stays opaque — receivers validate path, size, hash, and never execute without sandboxing.

### Schema additions (v2.0 final migration)

All additive and idempotent. Upgrades from v1.11.x are zero-downtime.

- `agents.session_id TEXT` — UUID, rotates on every register. Powers session-aware reads.
- `agents.agent_status TEXT NOT NULL DEFAULT 'online'` — operational status distinct from presence.
- `agents.description TEXT` — optional human-readable description, shown in discover + dashboard.
- `messages.read_by_session TEXT` — session that read this message.
- `tasks.required_capabilities TEXT` — JSON array for routing + queue-reassignment.
- `tasks.lease_renewed_at TEXT` — task-level liveness signal.
- `tasks.to_agent` — rebuilt nullable (via transactional CREATE + INSERT + DROP + RENAME) so queued tasks can exist without an assignee.
- `webhook_delivery_log.retry_count / next_retry_at / terminal_status` — retry bookkeeping.
- `channels` + `channel_members` + `channel_messages` tables.
- `agent_capabilities` — normalized index for O(1) capability lookup in routing.

### New env vars

- `RELAY_MAX_PAYLOAD_BYTES` (default 65536) — per-field content byte limit.
- `RELAY_HTTP_BODY_LIMIT` (default `1mb`) — Express body-parser outer bound.
- `RELAY_LOG_LEVEL` (default `info`) — `debug` / `info` / `warn` / `error`. Supersedes `RELAY_LOG_DEBUG=1` (still honored for back-compat).
- `RELAY_HEALTH_REASSIGN_GRACE_MINUTES` (default 120) — lease expiry window.
- `RELAY_HEALTH_SCAN_LIMIT` (default 50) — max tasks per lazy scan.
- `RELAY_HEALTH_DISABLED` — emergency off-switch for the health monitor.
- `RELAY_AUTO_ASSIGN_LIMIT` (default 20) — max queued tasks assigned per register sweep.
- `RELAY_WEBHOOK_RETRY_BATCH_SIZE` (default 10) — max retries processed per piggyback.

### Hook script improvements

Both `check-relay.sh` (SessionStart) and `post-tool-use-check.sh` (PostToolUse) now self-check `$0` at entry. If the install path looks truncated (spaces not quoted in `.claude/settings.json`), the hook emits a stderr warning. Never silent-fails on misconfiguration.

### Behavior changes (breaking for long-running assignees)

- `touchAgent` no longer renews task leases as a side effect. Only task-specific updates on that row (accept, heartbeat, complete, reject, cancel) renew `lease_renewed_at`. Long-running work must heartbeat.
- `to_agent` is nullable on `tasks`. Callers that assumed it is always a string must handle null for queued tasks.
- `TaskStatus` union expanded to include `queued` and `cancelled`.
- `TaskAction` union expanded to include `cancel` and `heartbeat`.

### Security

- `ConcurrentUpdateError` on CAS mismatch — no silent overwrites under contention.
- Health-monitor CAS re-checks agent liveness inside the WHERE clause; a heartbeat between scan and requeue wins.
- `processDueWebhookRetries` runs on every webhook-firing tool call but CAS-claims each job, preventing double delivery.
- `RELAY_HTTP_SECRET` must be ≥32 characters when set (enforced at startup).
- `RELAY_ENCRYPTION_KEY` must decode to exactly 32 bytes (enforced at startup).

### Numbers

- 22 MCP tools (was 14 at v1.11.1).
- 367 tests across 27 files (was 306 at v1.11.1; +61 tests).
- Zero regression across v1.x and v2 alpha/beta/beta.1 coverage.
- Clean `tsc --noEmit` + `npm run build`.

### Deferred to v2.1+

Full backlog at `plug-and-play-retro.md`. Highlights deliberately left for later versions:

- Legacy auto-migration (v2.1)
- Token rotation tool (v2.1)
- Message threading / reply chains (v2.1)
- Task dependencies + timeout (v2.1)
- CLI tooling: `relay doctor` / `relay init` / `relay test` / `relay generate-hooks` (v2.2)
- Batch + fan-out operations (v2.2)
- Private channels + per-channel permissions (v2.2)
- E2E message encryption, federation (v2.5)

---

## v2.0.0-beta.1 — 2026-04-17 (Codex audit fixes — rolled into v2.0.0)

A Codex audit ran on beta. Four HIGH + two MEDIUM + one LOW findings — all valid. Beta.1 closes every HIGH before v2.0.0 final work starts (foundation-before-features).

### HIGH fixes

- **HIGH 1 — Health monitor requires both lease expired AND assignee stale.** Previous check looked only at `lease_renewed_at`. An alive-but-observing agent could lose active tasks. Fixed by joining `agents` into the scan + CAS: a task is requeued only when its assignee is offline (`last_seen < grace`) or the agent is no longer registered. Also adds two new health tests: "alive-quiet assignee doesn't requeue" and "unregistered assignee requeues".
- **HIGH 2 — CAS on every `updateTask` mutation.** accept / complete / reject / cancel were pre-read-then-update, allowing concurrent clobber. Now every mutation uses `UPDATE ... WHERE id=? AND status=? AND (to|from)_agent=?` and raises `ConcurrentUpdateError` if 0 rows change. The pre-read still powers authz + transition error messages but is no longer authoritative.
- **HIGH 3 — Lease renewal decoupled from `touchAgent`.** `touchAgent` no longer bumps `tasks.lease_renewed_at`. Only task-specific actions (accept, heartbeat, and any update on that exact row) renew the lease. An agent can no longer keep abandoned tasks alive by doing unrelated work. Side effect: long-running assignees must periodically call `update_task heartbeat` to keep their lease fresh.
- **HIGH 4 — Auto-assign moved into `registerAgent()`.** Was wired only in `handleRegisterAgent`, so non-MCP callers silently lost the sweep. `registerAgent` now returns `{agent, plaintext_token, auto_assigned}` and every caller gets the assignments (handler still fires webhooks off the returned list, circular-import-safe).

### MEDIUM / LOW fixes

- **MEDIUM 5 — `postTaskAuto` pick+insert is transactional.** Wrapped the SELECT candidates + INSERT picked inside a `db.transaction()` (BEGIN IMMEDIATE) so two concurrent callers can't both pick the same least-loaded agent. Strict across-process serialization.
- **MEDIUM 6 — Real OS-process concurrency test.** New test spawns 5 child processes competing for 1 queued task via the CAS UPDATE pattern. Asserts `totalClaimed === 1`. Proves CAS under genuine contention, not just sequential-same-process.
- **LOW 7 — Empty `required_capabilities` throws.** Defense-in-depth guard inside `postTaskAuto` (zod already enforces `.min(1)` at the tool surface, but direct callers bypass that).

### Numbers

- 346 tests across 26 files (was 338; +8: three health-monitor variants, three CAS-on-updateTask variants, empty-caps guard, real-concurrent child-process test).
- Zero regression on pre-beta (all v1.x tests + alpha tests unchanged).
- Clean `tsc --noEmit` + `npm run build`.
- Still no `package.json` bump (same policy as beta — version bump is for v2.0.0 final).

### Files changed since beta

- `src/db.ts` — `touchAgent` (HIGH 3), `updateTask` CAS refactor (HIGH 2), `runHealthMonitorTick` agent-status join (HIGH 1), `postTaskAuto` transaction wrap + empty-caps guard (MEDIUM 5 + LOW 7), `registerAgent` returns `auto_assigned` (HIGH 4), new `ConcurrentUpdateError` class.
- `src/tools/identity.ts` — reads `auto_assigned` from `registerAgent` return instead of calling the helper separately.
- `tests/beta-smart-routing.test.ts` — updated 4 tests to match new semantics, added 8 new tests.

## v2.0.0-beta — 2026-04-17 (Smart routing + lease heartbeat + lazy health monitor — rolled into v2.0.0)

Second v2.0 sub-release. Working tree only — not published, not version-bumped in package.json yet. Checkpointing before the final v2.0.0 sub-release (file transfer conventions + webhook retry with CAS).

### What ships

- **`post_task_auto`** (new MCP tool, 20 total) — picks the least-loaded agent whose `agent_capabilities` rows are a superset of the task's `required_capabilities`. Tie-break: freshest `last_seen`. If no agent matches, task is stored with `status='queued'`, `to_agent=NULL`; it will be auto-assigned when a capable agent registers.
- **Task lease heartbeat** — `tasks.lease_renewed_at` is set on accept, implicitly renewed when the assignee makes any tool call that currently bumps `agents.last_seen` (send_message / broadcast / post_task / update_task / post_task_auto), and NOT renewed by observation tools (get_messages / get_tasks / discover_agents) — same v1.3 presence-integrity split applied to task leases.
- **`heartbeat` action on `update_task`** — explicit lease renewal, no state change. Only the assignee, only when status=`accepted`. No webhook (would be too noisy).
- **`cancel` action on `update_task`** — only the requester (`from_agent`) can cancel. Allowed from `queued`, `posted`, `accepted`; rejected from terminal states. Fires `task.cancelled`.
- **Lazy health monitor** — no background timer. Piggybacks on `get_messages`, `get_tasks`, and `post_task_auto`. Scans for accepted tasks where `lease_renewed_at` is older than the grace window (default 120 min via `RELAY_HEALTH_REASSIGN_GRACE_MINUTES`). CAS-requeues them (`to_agent=NULL`, `status=queued`). Fires `task.health_reassigned`. Bounded by `RELAY_HEALTH_SCAN_LIMIT` (default 50). Emergency off-switch: `RELAY_HEALTH_DISABLED=1`.
- **Auto-assign on register** — when a new or re-registered agent comes online, the server sweeps queued tasks whose `required_capabilities` are a subset of the agent's capabilities and CAS-assigns them. Bounded by `RELAY_AUTO_ASSIGN_LIMIT` (default 20). Fires `task.posted` with `auto_assigned_from_queue=true`.

### Schema

- New column `tasks.required_capabilities TEXT` (JSON array). Null for v1.x tasks.
- Column `tasks.to_agent` is now nullable (rebuilt in-place via a transactional `CREATE new + INSERT + DROP + RENAME` — additive, idempotent, preserves all existing rows).
- Column `tasks.lease_renewed_at` is now wired (was declared in alpha; beta consumes it).

### Adversarial tests (19 new, beta-smart-routing.test.ts)

- Auto-routing (8): no-match → queued, single match routing, least-loaded preference, tie-break on last_seen, strict capability superset filter, queued pickup on register, queue CAS prevents double-assign, requester can cancel queued.
- Health + leases (11): lease stamped on accept, active tools bump lease/observation does not, heartbeat renews/authz/status checks (3), expired lease requeues with full CAS chain, CAS short-circuits on fresh lease, `RELAY_HEALTH_DISABLED` off-switch, cancel by requester (3 variants including terminal-state rejection).

### Bug fix surfaced during beta

- `WasmDatabase.exec()` was unconditionally calling `flush()` even inside an open transaction — `db.export()` mid-transaction silently lost pending DDL. Now gated on `txDepth === 0` (same guard the `prepare→run` path already used). Only observed when beta's in-place column rebuild attempted `CREATE TABLE` + `INSERT SELECT` in a single transaction; existing wasm tests did not exercise intra-transaction `exec()`.

### Numbers

- 338 tests across 26 files (was 319; +19 new beta tests).
- Zero regression on alpha (all 319 still green, including 15 wasm tests that were unaffected once the exec/flush bug was patched).
- Clean `tsc --noEmit`.
- No `package.json` version bump yet — that happens at v2.0.0 final.

### What's deliberately deferred to v2.0.0 final

- File transfer convention docs (`_file` metadata receiver-validation guidelines).
- Webhook retry with CAS (schema already in alpha: `retry_count`, `next_retry_at`, `terminal_status`).
- `package.json` version bump.
- Full CHANGELOG / README polish pass.

## v1.11.1 — 2026-04-17 (Dual-model audit fixes — first Claude + Codex/GPT review)

**Milestone: first dual-model audit.** Claude GREEN'd the v1.11.0 release ("does the code do what it claims?"). A parallel Codex/GPT audit ("what happens when things go wrong?") found 3 HIGH + 1 MEDIUM issue Claude missed. Different model families, different blind spots — pattern proved its value.

### Fixes

- **HIGH 1 — Flush fail-closed:** `flush()` no longer silently swallows disk-write errors. Emscripten `ErrnoError` from `db.export()` is caught as warn-level (in-memory data safe; export glitch non-fatal). Real `fs.writeFileSync` errors (ENOSPC, EACCES) propagate and fail the operation.
- **HIGH 2 — Nested transaction compat:** Inner transactions now use `SAVEPOINT sp_N / RELEASE sp_N` instead of bare `BEGIN TRANSACTION` (which errors on SQLite: "cannot start a transaction within a transaction"). Matches better-sqlite3 semantics.
- **HIGH 3 — Init race condition:** `initializeDb()` caches its promise. Concurrent callers share one in-flight initialization — no more two independent wasm DB instances against the same file.
- **MEDIUM 4 — lastInsertRowid:** Now queries `SELECT last_insert_rowid()` via a prepared statement before flush. Returns real rowid (was hardcoded 0).
- **Bonus fix:** journal_mode pragma now skipped entirely on wasm (was `DELETE`, which threw Emscripten FS error). In-memory databases use memory journaling by default.

### Numbers

- 306 tests across 24 files (was 302; +4 new: reopen persistence, nested transaction, concurrent init, lastInsertRowid).
- Zero regression on native (291 existing tests unchanged).

## v1.11.0 — 2026-04-17 (SQLite WASM driver — zero native compilation)

`better-sqlite3` requires a C++ compiler at `npm install` time. This blocks Windows (Visual Studio Build Tools), Alpine/musl Linux (build-essential), Docker (200MB+ toolchain), CI, and ARM cross-compilation. v1.11 adds `sql.js` (SQLite compiled to WebAssembly) as an opt-in alternative behind `RELAY_SQLITE_DRIVER=wasm`.

### What ships

- **`src/sqlite-compat.ts`** — `CompatDatabase` / `CompatStatement` adapter that wraps sql.js behind a better-sqlite3-compatible API. All 731 lines of SQL queries in `src/db.ts` work identically on both drivers — zero query changes.
  - `WasmStatement`: wraps sql.js's `prepare→bind→step→getAsObject→free` into better-sqlite3's `.run()`, `.get()`, `.all()` API.
  - `WasmDatabase`: wraps sql.js's `Database` with write-back-to-file after every write (`db.export()` + `fs.writeFileSync`), transaction support (depth-tracked `BEGIN→COMMIT/ROLLBACK` with flush suppression inside transactions), `pragma()` interception (WAL gracefully degrades to DELETE, busy_timeout is no-op).
  - `initializeDb()` async factory: native path is sync under the hood; wasm path loads the wasm binary asynchronously at startup then all subsequent queries are sync.

- **`sql.js` in `optionalDependencies`** — `npm install` does not force it. Users who want wasm install it explicitly with `npm install sql.js`.

- **`src/db.ts` changes (minimal):**
  - Import type changed from `Database.Database` to `CompatDatabase`.
  - `getDb()` fallback path preserved for native (backward compat with tests that don't call `initializeDb()`).
  - New `initializeDb()` async export called from `src/index.ts` at startup.
  - All SQL queries, migrations, purge logic: **unchanged**.

- **`src/index.ts`** — `getDb()` call replaced with `await initializeDb()`.

- **`tests/db-wasm.test.ts`** — 11 new tests: agent ops (create, re-register cap immutability, filter by role), messaging (send+get, mark-as-read, broadcast), tasks (post→accept→complete, get_tasks), wasm-specific (getDb works, WAL degrades gracefully, sequential writes don't corrupt).

- **`docs/sqlite-wasm-driver.md`** — when to use, how to switch, performance notes, limitations (single-process only, no WAL, write-back latency, crash durability).

- **README "SQLite Driver Options" section** — quick reference linking to the full docs.

### Limitations (documented, not worked around)

1. **Single-process only.** sql.js operates in-memory with write-back. Two processes sharing the same DB file overwrite each other's changes. Multi-terminal stdio setups (each terminal = its own MCP process) MUST use native. HTTP transport (single daemon) is safe.
2. **No WAL mode.** wasm uses `journal_mode=DELETE`. At our scale this is imperceptible.
3. **Write-back latency.** Full DB export + `fs.writeFileSync` after every write. At < 1MB: sub-millisecond. At larger sizes: could be noticeable.
4. **Crash durability.** Same as native WAL — last write may be lost if process crashes between the write and the flush.

### Numbers

- 302 tests across 24 files (was 291; +11 new wasm tests).
- Clean `tsc` compile.
- Zero regression on native driver (all 291 existing tests pass unchanged).
- sql.js in optionalDependencies (no forced install).

### What was deliberately NOT done

- No refactor of src/db.ts into src/db/ module (adapter pattern kept it surgical).
- No making wasm the default (native is default, explicit opt-in only).
- No multi-process wasm support (documented as single-process-only).
- No custom wasm binary compilation (uses sql.js's pre-compiled binary).
- No changes to spawn / auth / hooks / webhooks / encryption / CIDR / dashboard.
- No v2.0 intelligence layer work. No npm publish.

## v1.10.0 — 2026-04-17 (Layer 2: Managed Agent integration — docs + reference workers)

Layer 2 of the four-layer delivery architecture. Non-Claude-Code agents (Python daemons, Node workers, Ollama/vLLM integrations, custom scripts) can now integrate with the relay using a comprehensive guide and two runnable reference implementations.

### What ships

- **`docs/managed-agent-integration.md`** (~350 lines) — full integration guide:
  - Mental model: how a Managed Agent (non-Claude-Code) fits alongside existing terminals.
  - Three transport options: HTTP (recommended), direct SQLite (same-machine), webhook subscription (event-driven).
  - Auth flow: first-time registration → token persistence → subsequent calls (header / arg / env) → capability declaration (immutable, per v1.7.1) → token rotation.
  - Lifecycle: startup → operating loop (poll and/or webhook) → SIGINT/SIGTERM clean shutdown with `unregister_agent`.
  - Error handling: retry-with-backoff on network, no-retry on 401, rate-limit handling, structured tool-error responses.
  - Security notes: never commit tokens, prefer HTTPS in production, scope capabilities narrowly, enable encryption at rest.
  - FAQ: cross-layer messaging, multiple agents same name, relay URL discovery, client library roadmap, relay restarts, direct agent-to-agent.

- **`examples/managed-agent-reference/python/agent.py`** (~200 LOC, stdlib-only) — single-file Python script using `urllib`. Demonstrates: register, send_message, get_messages, post_task (via check_tasks), update_task, discover_agents, unregister_agent. SIGINT handler. Inline comments teach the protocol. No pip dependencies.

- **`examples/managed-agent-reference/node/agent.js`** (~200 LOC, stdlib-only) — parallel Node implementation using `node:http`. Same coverage and structure. No npm dependencies.

- Both examples have a **`SMOKE.md`** with a 5-step manual verification checklist: start the agent, verify in `discover_agents`, send a test message, post a test task, Ctrl-C and verify clean unregister.

- **README** — new "Layer 2: Managed Agents" section linking to the guide and reference scripts.

- **CLAUDE.md** — `examples/` directory added to file map, status line updated.

### Fold-in from review 10

- **tmux birthday-paradox math precision fix** — `docs/cross-platform-spawn.md` now says "50% at 362, 1% at 36, 0.1% at 11" instead of the directionally-correct-but-imprecise "~256." Per a review note.

### Numbers

- 291 tests across 23 files — **unchanged**. src/ has zero changes; this is a docs + examples release.
- Clean `tsc` compile.
- No new MCP tools, no schema changes, no auth / hook / spawn / server code changes.

### What was deliberately NOT done

- **No src/ code changes.** Managed Agents use the existing 14 MCP tools via the existing HTTP transport. If a future version needs server-side ergonomic improvements, it ships separately.
- **No client library.** The reference scripts (~200 LOC each) are intentionally minimal so integrators see the protocol clearly and can port to any language.
- **No Docker / systemd / launchd templates.** Deployment environments vary; the reference scripts teach the protocol, not the hosting.
- **No webhook-push-via-relay feature.** Agents that cannot accept inbound HTTP should poll.
- **No v1.11 sqlite-wasm work, no v2.0 intelligence layer, no v2.5 federation.**
- **No npm publish.**

## v1.9.1 — 2026-04-16 (Cross-platform spawn hardening)

Closes three blockers + four fold-ins surfaced by a review of v1.9.0. Verdict for v1.9.0 was "foundation solid" with real seams to close — exactly what the v1.9 post-build section predicted. Foundation-before-features: ships before v1.10.

### Blockers closed

**1. Adversarial payload parity on Linux + Windows drivers.** v1.9.0 had zero hostile-input tests against the Linux or Windows drivers — only the zod schema layer protected them. v1.9.1 adds ~30 mock-level adversarial tests covering the same payload classes as the macOS integration suite: name/role injection (`;`, `|`, `&`, `$(cmd)`, backtick, newline, quotes), cwd injection (substitution, CRLF, null byte, relative paths), length limits (name > 64, cwd > 1024), override case-variance + unknown values. Each test asserts either (a) zod throws at the boundary, or (b) the constructed argv is provably safe — payload appears as its own argv element, never concatenated into a shell-interpreted string.

**2. Linux tmux single-quote POSIX escape.** The Linux driver's launch command `cd '<cwd>' && exec claude` single-quotes cwd. Today zod blocks `'` in cwd so this is not exploitable, but it is a single point of defense. v1.9.1 adds the standard POSIX `'\''` escape (close quote, literal quote, reopen) via a new `escapeSingleQuotesPosix` helper in `src/spawn/validation.ts`. A defense-in-depth test fabricates an input bypassing zod and asserts the escape is applied correctly. Mirrors the `printf %q` pattern in `bin/spawn-agent.sh`.

**3. tmux session-name collision.** v1.9.0 used the agent name directly as the tmux session name — two agents with the same relay name silently collided (`tmux new-session` fails on duplicate session names, but nothing surfaces to the caller). v1.9.1 appends a 4-hex random suffix from `crypto.randomBytes(2)` to the tmux session name. The agent's registered relay identity is unchanged (peers discover by relay name; only the tmux binding carries the suffix). Actual session name is logged to stderr at spawn time so operators see what to `tmux attach -t <agent>-<4hex>`. Entropy 16 bits = 65,536 values; collision probability negligible for any realistic workload.

### Fold-ins

**4. Cross-platform cwd rejection.** `normalizeCwd` in `src/spawn/validation.ts` now throws when passed a cwd that is nonsensical for the target platform: drive-letter paths (`C:\`, `D:/`) on POSIX, non-absolute paths on Windows. Three new fold-in tests assert these.

**5. Honest-caveats section updated.** This entry explicitly lists the tmux collision closure. The earlier expected-review-findings note has since been trued up by the post-build section.

**6. Platform-aware `RELAY_TERMINAL_APP` override.** `resolveTerminalOverride` now takes the current platform as its second argument. Cross-platform names (e.g., `gnome-terminal` on macOS) are treated as invalid and fall through to auto-detect with a platform-specific stderr warning listing the valid choices. Previously: silently accepted then ignored by the macOS driver. Now: rejected consistently.

**7. PowerShell single-quote edge.** The Windows PowerShell driver's `Set-Location -LiteralPath '<cwd>'` now routes through a new `escapeSingleQuotesPowershell` helper implementing PowerShell's own `''` doubling rule. A defense-in-depth test fabricates an input bypassing zod and asserts the doubling is applied. Same motivation as blocker 2.

### Numbers

- 291 tests across 23 files (was 260; +31). Spawn-drivers test file: 21 → 53. Zero regression in the 5 `spawn.test.ts` handler tests or the 22 `spawn-integration.test.ts` macOS payload tests — the macOS shell script is still frozen at v1.6.4.
- Clean `tsc` compile.
- No new MCP tools, no schema changes, no auth/hook/server changes. Pure hardening patch inside the `src/spawn/` module.

### What did the adversarial tests actually find?

Honest answer: **they confirmed existing hardening rather than surfacing new bugs.** zod's `SpawnAgentSchema` rejected every hostile input class cleanly at the boundary. The defense-in-depth tests (blocker 2 + fold-in 7) verified the per-driver escape helpers work — both are also no-ops today on legitimate input, so there was nothing to fix operationally. Net value: regression protection for future zod changes, plus explicit coverage that makes the per-platform safety story auditable without reading the whole source tree.

### What was deliberately NOT done

- No changes to `bin/spawn-agent.sh` — frozen at v1.6.4.
- No real-subprocess Linux/Windows CI infrastructure. Manual smoke checklists in `docs/cross-platform-spawn.md` remain authoritative.
- No new drivers. SSH / container / remote spawn still blocked on v2.5.
- No new MCP tools, no schema, no auth, no hook, no server code.
- No v1.10 work. No npm publish.

## v1.9.0 — 2026-04-16 (Cross-platform spawn — Linux + Windows support)

Abstracts the `spawn_agent` backend so the relay works on macOS, Linux, and Windows without modification. macOS keeps the proven `bin/spawn-agent.sh` (untouched from v1.6.4); Linux and Windows get fresh TypeScript drivers.

### Why Node/TypeScript over extending bash

Windows ships without bash by default (no WSL, no Git Bash, no Cygwin). A future `npm install -g bot-relay-mcp` user on stock Windows would hit a wall if the spawn driver were bash-based. Node is already a hard requirement (the MCP server runs on Node), so choosing Node for the driver introduces zero new dependencies while unlocking Windows as a first-class target. `src/types.ts` `SpawnAgentSchema` (zod + allowlist regexes) becomes the single source of validation truth — the bash shell's duplicate allowlist is now one driver among several, not a parallel rulebook to drift out of sync.

### What ships

- **`src/spawn/` module** — new home for the driver abstraction:
  - `types.ts` — `SpawnDriver` interface + `SpawnCommand` shape.
  - `validation.ts` — `resolveTerminalOverride()` (allowlist-gated env-var override), `normalizeCwd()` (POSIX/Windows separator handling), `buildChildEnv()` (principle-of-least-authority env propagation).
  - `dispatcher.ts` — picks a driver via env override > `process.platform` > per-platform fallback. ONLY code path that calls `child_process.spawn`; drivers are pure on the build side (mockable in tests).
  - `drivers/macos.ts` — thin wrapper that shells to `bin/spawn-agent.sh` with the same `[name, role, caps, cwd]` args. Preserves the 3-layer hardening + 19-payload adversarial suite unchanged.
  - `drivers/linux.ts` — fallback chain `gnome-terminal → konsole → xterm → tmux`. tmux fallback creates a detached session (`tmux attach -t <agent>` to enter). Headless servers covered.
  - `drivers/windows.ts` — fallback chain `wt.exe → powershell.exe → cmd.exe`. Forward-slash CWDs normalized to backslashes at the validation boundary. Avoids cmd.exe quoting landmines by keeping args separate (no monolithic command string).

- **`src/tools/spawn.ts` refactored** to delegate to `spawnAgent()` from the dispatcher. Response includes `platform` and `driver` fields so callers can see which sub-driver ran. Error hints are platform-specific.

- **`tests/spawn-drivers.test.ts`** — 21 new mock tests covering:
  - Linux fallback chain (all four sub-drivers picked correctly, error on none available, override honored, override falls through if binary missing).
  - Windows fallback chain (wt/powershell/cmd picked correctly, error on all missing, CWD backslash normalization).
  - macOS driver builds the right bash-script invocation.
  - `resolveTerminalOverride` allowlist gating (accepts every allowed name case-insensitively, rejects everything else).
  - `buildChildEnv` principle-of-least-authority (propagates `RELAY_*` + system essentials; does NOT propagate `AWS_SECRET_ACCESS_KEY` / `GITHUB_TOKEN`).
  - Platform-aware CWD normalization.

- **`docs/cross-platform-spawn.md`** — new docs: driver selection flowchart, per-platform install requirements, `RELAY_TERMINAL_APP` override semantics, env-var propagation policy, manual smoke-test checklists (macOS + Linux + Windows), troubleshooting.

- **README + CLAUDE.md** updated — tool table entry no longer says "macOS only"; new section after the Near-Real-Time Mail Delivery block; CLAUDE.md file map lists the new `src/spawn/` structure.

### Numbers

- 259 tests across 23 files (was 238; +21 new driver tests).
- Zero regression on macOS: `tests/spawn.test.ts` (5) and `tests/spawn-integration.test.ts` (22) still green — `bin/spawn-agent.sh` is unchanged.
- Clean `tsc` compile.
- No new MCP tools. No schema changes. No auth / hook / server code changes.

### Env-var propagation — principle of least authority

Spawned agents receive a minimal env by default:
- System essentials: `PATH`, `HOME`/`USERPROFILE`, `LANG`, `TERM`, `SHELL` (POSIX), Windows-specific (`SYSTEMROOT`, `APPDATA`, etc.).
- Anything prefixed with `RELAY_*`.
- Explicitly set from the spawn call: `RELAY_AGENT_NAME`, `RELAY_AGENT_ROLE`, `RELAY_AGENT_CAPABILITIES`.

Arbitrary parent env vars (`AWS_SECRET_ACCESS_KEY`, `GITHUB_TOKEN`, `OPENAI_API_KEY`, ...) do NOT propagate. Operators who need custom forwarding can prefix their var with `RELAY_`.

### `RELAY_TERMINAL_APP` override

Allowlist-gated string that forces a specific sub-driver: `iterm2`, `terminal`, `gnome-terminal`, `konsole`, `xterm`, `tmux`, `wt`, `powershell`, `cmd`. Unknown values are ignored (fall through to auto-detect) with a stderr warning — never silent. If the forced sub-driver's binary is not on PATH, the driver treats it as unavailable and walks the chain normally.

### Tested on

- **macOS** — full CI (existing 27 spawn tests green).
- **Linux + Windows** — mock-only in CI (21 new driver tests). Real-subprocess testing is **manual smoke** per `docs/cross-platform-spawn.md`. No Linux/Windows CI infrastructure was built — scope creep for v1.9.

### What was deliberately NOT done

- **No pure-TS macOS driver.** The existing shell script has a proven adversarial test suite and is left alone.
- **No rewrite of `bin/spawn-agent.sh`.** Frozen at v1.6.4 state.
- **No Linux / Windows CI.** Manual smoke documented instead.
- **No new MCP tools.**
- **No schema changes.**
- **No auth / hook changes.**
- **No v1.10 / v1.11 work.**
- **No npm publish.**

### Honest caveats

- Linux and Windows drivers have never run a real subprocess in CI — only mock tests. First real user on those platforms may hit seams the mocks did not catch. Expect a v1.9.1 patch cycle.
- Linux tmux fallback uses session-only invocation (no window) — correct for headless servers, but operators used to a GUI window may be momentarily confused. `tmux attach -t <agent>` is documented.
- Windows paths with embedded quotes are not specifically tested. The zod allowlist forbids quotes so this should be unreachable, but flag if found.

## v1.8.1 — 2026-04-16 (Docs correction — quoted-path install guidance)

**Docs-only. No code change.**

v1.8.0 shipped the `PostToolUse` hook correctly, but the install docs did not teach readers that paths containing spaces must be single-quoted inside the JSON string of `.claude/settings.json`. Claude Code passes the `command` field to `/bin/sh`, which splits on whitespace — installations at paths like `/path/to/My Projects/bot-relay-mcp/...` silently fail with `/bin/sh: ... is a directory` and no surfaced diagnostic.

### What ships

- README "Near-Real-Time Mail Delivery" — important callout immediately after the copy-paste JSON block, with a concrete real-world quoted-path example alongside the generic `/path/to/...` template.
- `docs/post-tool-use-hook.md` — fuller callout in the install section explaining the shell-quoting mechanics (outer JSON double-quotes, inner shell single-quotes), plus a diagnostic command (`sh -c "$COMMAND"`). Troubleshooting section gains a named entry: **"Hook silently fails on paths with spaces"** — so readers with a broken install can grep for the symptom and land on the fix.
- Version bump to 1.8.1 across `package.json`, `src/server.ts`, `src/transport/http.ts`.

### Why this earns a patch release, not a fold-in to v1.9

Project rule: v1.8 (the PostToolUse hook foundation) must be solid before v1.9 (cross-platform spawn) starts. A silent-failure install path on any machine whose workspace lives at a space-containing path is a broken first-run experience. Fixing it in a clean patch release preserves the foundation-before-features invariant and the project's clean-patch-release discipline.

### What was deliberately NOT done

- No code changes. The hook script is correct as shipped.
- No test churn. 238 tests pass unchanged; docs are not under test. (The optional embedded-JSON-parse test from the review did not apply — no embedded example exists in source.)
- No SessionStart-hook `docs/hooks.md` parallel patch. Out of scope for this patch — if the same guidance belongs there too, it queues as a separate doc ticket.

## v1.8.0 — 2026-04-15 (Layer 1 PostToolUse hook — near-real-time mail delivery)

Closes the human-bridged latency between agents. A new `PostToolUse` hook checks the mailbox after every tool call and surfaces pending messages as `additionalContext` so the running Claude Code session picks them up immediately — no waiting for the next `SessionStart` or a human-pasted "check mail."

### What ships

- **`hooks/post-tool-use-check.sh`** — per-project hook script.
  - HTTP path preferred when `RELAY_AGENT_TOKEN` is set and the daemon responds on `/health` within 1s (full auth + audit pipeline; server-side atomic mark-as-read).
  - Sqlite-direct fallback for stdio-only deployments (reads pending rows, marks surfaced IDs read via a follow-up UPDATE per-ID).
  - 2s self-imposed budget (1s health probe + 2s `get_messages`).
  - Never re-registers (that is the SessionStart hook's job).
  - Never reads stdin (mail check is tool-agnostic; the tool-call payload is ignored).
  - Silent-fails on any error — empty stdout, exit 0 — never pollutes the conversation with error text or partial JSON.
  - Every env-var input validated against an allowlist before use. Agent name/host/port/token shapes all matched to regex; `RELAY_DB_PATH` resolved under `$HOME`/tmp only.

- **`docs/post-tool-use-hook.md`** — install guide + env-var reference + troubleshooting + what the hook deliberately does NOT do.

- **README "Near-Real-Time Mail Delivery" subsection** — copy-pasteable `.claude/settings.json` block and honest limitation callout.

- **`tests/hooks-post-tool-use.test.ts`** — 8 integration tests covering: HTTP happy path, empty mailbox, idempotency, unreachable-relay graceful fail with timing ceiling, missing-token sqlite fallback, missing `RELAY_AGENT_NAME`, invalid-token-shape fallback, and a behavioral invariant that the hook does NOT mutate the agent's role or capabilities.

### Honest limitations (documented, not worked around)

- **Idle terminals get no delivery.** The hook only fires when the agent is actively running tool calls. If a terminal is sitting idle, it will not see new mail until the next tool call or next `SessionStart`. Continue to rely on SessionStart + human attention for long-idle windows.
- **HTTP path requires python3.** The script uses `python3` to safely construct and parse JSON (no jq dependency; jq is not guaranteed to exist). macOS and most Linux distros ship python3; if absent, HTTP path fails fast and sqlite fallback runs.
- **Tasks are NOT surfaced.** Only pending messages. Task delivery stays in the SessionStart hook for now — focused scope, less context pressure per firing.
- **Per-project install recommended.** Global install would fire in every Claude Code terminal including ones with no relay identity, which is unwanted polling + token exposure. Opt each workspace in deliberately.

### Numbers

- 238 tests across 22 files (was 230; +8 new integration tests for the hook).
- Clean `tsc` compile.
- No new MCP tools. No schema changes. No new dependencies (python3 and sqlite3 are standard OS tooling; already required by the existing SessionStart hook).

### Security posture (unchanged, verified)

- The hook authenticates via `RELAY_AGENT_TOKEN` on the HTTP path, meaning it goes through the full v1.7 auth layer: token bcrypt-verified, capability gate (`get_messages` is always-allowed-for-authenticated, matching the server rule).
- The sqlite fallback path queries only `WHERE to_agent = :name AND status = 'pending'` — agent isolation preserved at the SQL level.
- Env-var inputs validated against allowlists to block URL-/header-injection via crafted env values.
- The hook cannot escalate capabilities — it never calls `register_agent`.

### What was deliberately NOT done (Karpathy rule 2 — surgical scope)

- No cross-platform spawn work (v1.9).
- No Managed Agent reference worker (v1.10).
- No sqlite-wasm migration (v1.11).
- No new MCP tools.
- No schema changes.
- No changes to the MCP server code — this release is additive (shell script + docs + tests).
- No task surfacing in the hook.
- No npm publish.

## v1.7.1 — 2026-04-15 (Auth hardening — security advisory)

**Two blockers from a v1.7 review.** Both are real vulnerabilities in the v1.7 auth layer. No internet-facing deployments exist and nothing has been published to npm, so no external clients are at risk — but the fixes ship before v1.8 per the foundation-before-features rule.

### Security advisory — CVE-equivalent issues in v1.7.0

**CVE-equivalent 1 — Capability escalation via unauthenticated re-register (CRITICAL)**
- In v1.7.0, `register_agent` was in `TOOLS_NO_AUTH` for bootstrap. Re-registration on an EXISTING agent name hit the same no-auth path and silently updated `capabilities`. An unauthenticated attacker could call `register_agent("target-agent", "r", ["spawn", "tasks", "webhooks", "broadcast"])` and grant themselves (or any agent they have a token for) every capability — nullifying the entire v1.7 capability-scoping feature.
- **Fix:** dispatcher now bifurcates `register_agent`:
  - New registration (name does not exist in DB) → no auth required (bootstrap path preserved).
  - Re-registration (name exists, token_hash present) → auth required; the presented token must match that agent's stored hash.
  - Re-registration on a legacy pre-v1.7 agent (token_hash = NULL) → defers to `RELAY_ALLOW_LEGACY` grace.
- **Cap immutability:** even on authenticated re-register, capabilities are PRESERVED unchanged. The `capabilities` argument in re-register calls is ignored; only `role` and `last_seen` update. `registerAgent` in `db.ts` enforces this as defense-in-depth regardless of dispatcher state. Callers that pass a different capability set receive a `capabilities_note` in the response explaining the rule.
- To change an agent's capabilities, operators must `unregister_agent` (with valid token) and then `register_agent` fresh.

**CVE-equivalent 2 — Timing-unsafe HTTP secret comparison (HIGH)**
- `src/transport/http.ts` used `presented === config.http_secret` and `findIndex((s) => s === presented)` — both are byte-by-byte short-circuiting JavaScript string equality. A remote caller could measure response timing to recover the shared secret one character at a time.
- **Fix:** both checks now go through a `timingSafeStringEq` helper that length-checks first (short-circuit on length mismatch; length is operational metadata not a secret) then calls `crypto.timingSafeEqual` on `Buffer.from(s, "utf8")`. Content comparison is now constant-time. Length-mismatched callers get a clean 401 instead of a 500 (timingSafeEqual would otherwise throw).
- A side-channel still exists on WHICH previous secret matched during rotation (loop short-circuits on first match). Judged acceptable since previous secrets are already lower-trust, and will be revisited in a future patch if review rules otherwise.

### Fold-ins from the docs audit

- **README** — new sections: **Per-Agent Tokens**, **Encryption at Rest**, **Rotation Guide** for the HTTP shared secret. CHANGELOG and devlogs already covered these, but README is the public-facing doc.

### Numbers

- 230 tests across 21 files (+12: 6 new adversarial re-register tests, 5 new timing-safety tests, +1 multi-cycle immutability test; the existing `upserts on duplicate name` assertion was rewritten in place to match the new immutable-caps behavior — not weakened).
- Clean `tsc` compile.
- Two existing tests updated: `db.test.ts:55` (upsert caps assertion) and `auth.test.ts:130` (re-register caps assertion). Both previously documented the v1.7 buggy behavior; now they document v1.7.1 correctness.

### What was deliberately NOT done (Karpathy rule 2 — surgical scope)

- **No TOOLS_NO_AUTH membership change.** `register_agent` stays in the set for bootstrap; the re-register gate sits BEFORE the set check.
- **No new `description` agent field.** The spec mentioned "role and description" — there is no `description` column today; adding one would be scope creep. Only `role` is updatable on re-register.
- **No bcrypt-path rework.** `bcrypt.compareSync` is already constant-time by design; only the HTTP shared-secret path was timing-leaky.
- **No constant-time scan across all previous secrets.** Short-circuit on first match retained; documented as an acceptable trade-off.
- **No encryption / CORS / audit-log changes.** Out of scope.
- **No npm publish, no v1.8 work.** Waiting on the v1.7.1 review green-light.

### Upgrade notes

- **Existing agents:** no action needed if you already have a token. The stored `token_hash` is preserved on re-register.
- **SessionStart hook:** the hook's `register_agent` call will still succeed on every terminal open. The `capabilities` argument is now ignored on re-register, so the hook cannot drift an agent's caps — a capability change requires explicit `unregister_agent` + fresh `register_agent`.
- **Shared-secret rotation:** same env vars (`RELAY_HTTP_SECRET`, `RELAY_HTTP_SECRET_PREVIOUS`). Comparison is now timing-safe.

## v1.7.0 — 2026-04-14 (Auth layer — biggest release yet)

Per-agent auth, capability scoping, secret rotation, encryption at rest, structured audit log, CORS. Foundation shipped cleanly from v1.3 → v1.6.4. Now building the secure multi-agent + external integration layer on top of it.

Plus 2 gate items rolled in from v1.6.4 review:
- **G1:** CLAUDE.md status line bumped to v1.7.0 / 218 tests / 14 tools.
- **G2:** `ipInCidr` whitespace asymmetry — now trims both IP and CIDR inputs symmetrically.

### Fix 1 — Per-agent auth tokens
- `register_agent` now generates a random 32-byte token on first registration, stores a bcryptjs hash in `agents.token_hash`, and returns the raw token ONCE in the response + stderr log line.
- Re-registration of an already-tokened agent preserves the existing hash (SessionStart hook can safely upsert without rotating).
- Legacy agents (registered pre-v1.7, `token_hash = NULL`) can be migrated two ways: (a) call `register_agent` again to get a fresh token, or (b) set `RELAY_ALLOW_LEGACY=1` during a grace window.
- Every tool call except `register_agent` validates a presented token against the caller's stored hash. Token source (in precedence order): `agent_token` tool arg → `X-Agent-Token` HTTP header → `RELAY_AGENT_TOKEN` env var.
- Impersonation (claim-to-be-X-with-X's-token → claim-to-be-Y) is rejected.

### Fix 2 — Capability scoping per tool
- Each sensitive tool declares a required capability:
  - `spawn_agent` → `spawn`
  - `post_task`, `update_task` → `tasks`
  - `broadcast` → `broadcast`
  - `register_webhook`, `list_webhooks`, `delete_webhook` → `webhooks`
- Always-allowed (no capability check, token still required): `unregister_agent`, `discover_agents`, `send_message`, `get_messages`, `get_tasks`, `get_task`.
- Capabilities set at register time are immutable. To change, unregister + re-register.

### Fix 3 — Shared-secret rotation with grace period
- New env var `RELAY_HTTP_SECRET_PREVIOUS` — comma-separated list of previously-valid secrets. Accepted during a rotation window.
- Primary secret flows through `RELAY_HTTP_SECRET`. Previous secrets emit `X-Relay-Secret-Deprecated: true` response header so clients can see they should upgrade.
- Audit log entries tag which secret was used (`primary` or `previous[N]`).

### Fix 4 — AES-256-GCM encryption at rest
- New opt-in env var `RELAY_ENCRYPTION_KEY` — 32-byte base64-encoded key. When set, the following fields are encrypted on write and decrypted on read:
  - `messages.content`
  - `tasks.description`, `tasks.result`
  - `audit_log.params_json`
- Not encrypted (queryable metadata): agent names, tool names, from/to, priority, status, timestamps.
- Storage format: `enc1:<base64-iv>:<base64-ciphertext-plus-tag>`. Per-row IV (12 bytes).
- Legacy plaintext rows (predating the key, or rows written while the key was unset) remain readable — decrypt is a safe no-op for non-`enc1:` rows.
- Wrong key → GCM auth tag mismatch → decrypt throws clearly.
- Key rotation DEFERRED to v1.7.1 per the original plan. Single active key in v1.7.

### Fix 5 — Structured JSON audit log format
- New column `audit_log.params_json` (added additively; old `params_summary` column preserved for back-compat readers).
- Every tool call writes a structured `{ tool, agent_name, auth_method, source_ip, result, error_message? }` record, encrypted at rest.
- `getAuditLog()` returns parsed objects (or `{ _parse_error: true }` for malformed rows — never throws).
- Legacy rows (without `params_json`) surface as `{ legacy_summary: "<old text>" }` after migration.

### Fix 6 — CORS / Origin allow-list on dashboard
- New config field `allowed_dashboard_origins: string[]`. Default: `["http://localhost", "http://localhost:*", "http://127.0.0.1", "http://127.0.0.1:*"]`.
- Dashboard (`/`, `/dashboard`, `/api/snapshot`) checks the `Origin` header. Missing origin → allowed (non-browser callers). In allowlist → allowed with `Access-Control-Allow-Origin` echoed back. Outside allowlist → 403.
- `/health` always open.
- Port-glob syntax supported: `"http://localhost:*"` matches any port.

### Gate fixes (rolled in)
- `ipInCidr` now trims whitespace from BOTH the IP and CIDR inputs (previously only CIDR was trimmed, causing false negatives on trailing-space IPs).

### Tests
- 218 tests passing (was 162).
- New test files:
  - `tests/auth.test.ts` — 14 tests, token primitives + authenticateAgent logic + integration with registerAgent
  - `tests/auth-dispatcher.test.ts` — 13 tests, end-to-end via HTTP: token required/wrong/right, impersonation rejected, capability scoping per tool, X-Agent-Token header fallback
  - `tests/secret-rotation.test.ts` — 5 tests, primary + previous secrets accepted, wrong/missing rejected, deprecation header on previous
  - `tests/encryption.test.ts` — 13 tests, primitives + db round-trip + raw SQL verification that plaintext is actually gone from disk
  - `tests/cors-and-audit.test.ts` — 10 tests, Origin allow-list enforcement + structured audit JSON + encryption round-trip + malformed row handling
- All 162 v1.6.4 tests pass without modification (one test env vars updated to `RELAY_ALLOW_LEGACY=1` for legacy compatibility).

### Dependencies
- Added `bcryptjs@^3.0.3` and `@types/bcryptjs` (devDep).

### Migration guide (v1.6.x → v1.7.0)
- Option A (recommended): call `register_agent` for each existing agent to issue a new token. Capture the token from the response and save it in `RELAY_AGENT_TOKEN` env in your shell alias.
- Option B (grace window): set `RELAY_ALLOW_LEGACY=1` on the server. Existing token-less agents keep working; new agents still get tokens. Remove the env var after all agents have been migrated.
- Existing DB files upgrade automatically. `ALTER TABLE agents ADD COLUMN token_hash TEXT` and `ALTER TABLE audit_log ADD COLUMN params_json TEXT` run on startup (idempotent).

### Version bumps
- package.json: 1.6.4 → 1.7.0
- MCP server version: 1.6.4 → 1.7.0
- /health version: 1.7.0
- Webhook User-Agent: bot-relay-mcp/1.7.0

### What was DELIBERATELY not done
- Encryption key rotation (deferred to v1.7.1 per the original plan).
- Token revocation list (for now, revocation = unregister the agent).
- Capability mutation after registration (by design — unregister + re-register).
- OAuth/JWT/SSO — shared-secret + per-agent tokens are sufficient for this threat model.
- IP auth on stdio — stdio is local-user-trust; token is the fence.

## v1.6.4 — 2026-04-14 (IPv6 form coverage + test hygiene — 5 surgical sharpening items)

The v1.6.3 review verdict was GREEN on all functional claims, with 5 sharpening items for v1.6.4. All shipped.

### Fix 1 — Fully-expanded IPv4-mapped IPv6 detection
- Previously `ipv4FromMappedIPv6()` only recognized compressed forms (`::ffff:0102:0304`, `::ffff:1.2.3.4`). Fully-expanded `0:0:0:0:0:ffff:0102:0304` returned null, which would silently fail trust checks if an operator wrote a fully-expanded address in `trusted_proxies` config.
- Rewrote with a structural approach: split the address into 8 hex groups (after :: expansion), check the canonical `[0,0,0,0,0,0xffff,hi,lo]` pattern. Single code path handles all compression forms.
- For mixed-dotted forms (any address containing a `.`), split off the IPv4 tail and verify the IPv6 prefix is structurally `0:0:0:0:0:ffff` via a `padded + expand + check` helper.
- 2 new tests verify fully-expanded form behaves identically to compressed form.

### Fix 2 — IPv4-mapped IPv6 peer in trusted-proxy
- Exported `extractSourceIp` from `src/transport/http.ts` for direct unit testing.
- 4 new unit tests in `tests/trusted-proxy.test.ts` exercise scenarios that are awkward to provoke over real sockets:
  - dual-stack peer `::ffff:127.0.0.1` IS trusted against IPv4 CIDR `127.0.0.0/8`
  - IPv6-mapped CIDR rule `::ffff:127.0.0.0/104` matches IPv4 peer `127.0.0.1`
  - mapped peer NOT in trusted list correctly returns the peer (XFF ignored)
  - empty trusted_proxies always returns peer regardless of XFF

### Fix 3 — Bare-path approved root acceptance
- `bin/spawn-agent.sh` case pattern previously had `/var/folders/*` but not the bare `/var/folders` form. A cwd that resolved to exactly `/var/folders` (no subpath) would be wrongly rejected.
- Added bare-path alternative: `"/var/folders"|"/var/folders/"*`. HOME, /tmp, /private/tmp already had this pattern.
- New test `accepts cwd that resolves EXACTLY to an approved root (no subpath)` confirms the fix.

### Fix 4 — Adversarial IPv6 form documentation
- `::1.2.3.4` (IPv4-compatible IPv6, RFC 4291 §2.5.5.1, deprecated) is intentionally NOT treated as IPv4-mapped. Treating it as such would wrongly grant IPv4 CIDR trust to pure IPv6 callers using a deprecated transition format.
- `64:ff9b::1.2.3.4` (NAT64, RFC 6052) is intentionally NOT treated as IPv4-mapped. Different semantics — represents a translated IPv4 destination, not an incoming IPv4 client.
- Both behaviors verified by 2 new tests.
- Comment block in `src/cidr.ts` explains why these forms are intentionally excluded, citing the relevant RFCs.

### Fix 5 — assertBlocked helper consistency
- Extended `assertBlocked()` in `tests/spawn-integration.test.ts` to take optional `stderrContains` and `stdoutNotContains` opts.
- Refactored all 16 existing attack tests to use the helper. Eliminates inline assertion drift over time and makes adding new attack tests a 1-liner.

### Tests
- 162 tests passing (was 153).
  - +2 CIDR fully-expanded tests
  - +2 CIDR adversarial form tests (IPv4-compat, NAT64)
  - +4 extractSourceIp unit tests for IPv4-mapped peer scenarios
  - +1 bare-path approved root test
  - 0 net change from Fix 5 (pure refactor)
- All 153 prior tests pass without modification.

### Files changed
- `src/cidr.ts` — rewrote `ipv4FromMappedIPv6`, added `expandIPv6ToGroups` + `isAllZeroPlusFfffPrefix` helpers, dropped unused helper, expanded comment block
- `src/transport/http.ts` — exported `extractSourceIp`
- `bin/spawn-agent.sh` — added bare-path approved root, comment update
- `tests/cidr.test.ts` — +4 tests
- `tests/spawn-integration.test.ts` — extended assertBlocked, refactored 16 tests, +1 bare-path test
- `tests/trusted-proxy.test.ts` — +4 unit tests for extractSourceIp

### Version bumps
- package.json: 1.6.3 → 1.6.4
- MCP server version: 1.6.3 → 1.6.4
- /health version: 1.6.4
- Webhook User-Agent: bot-relay-mcp/1.6.4

### Backward compatibility
- Behavior change is only ADDITIVE — fully-expanded mapped form now matches where it returned false before; bare `/var/folders` cwd now allowed where it would have been rejected. No existing passing case turns into a failing case.
- All tool signatures unchanged.

## v1.6.3 — 2026-04-14 (IPv4-mapped IPv6 + deeper attack coverage + doc drift fixes)

The v1.6.2 review found 10/13 items FULL, 2/13 PARTIAL, 1/13 DRIFT. v1.6.3 closes all three.

### Fix 1 — IPv4-mapped IPv6 normalization in CIDR matcher (RFC 7239 §7.4)
- Previously `::ffff:1.2.3.4` compared against `1.2.3.0/24` returned false (cross-family mismatch). Operators writing IPv4 CIDRs would fail to match dual-stack clients arriving in the mapped form.
- Added `ipv4FromMappedIPv6()` helper that detects both textual (`::ffff:1.2.3.4`) and hex (`::ffff:0102:0304`) forms and extracts the embedded IPv4.
- `ipInCidr` now normalizes mapped-IPv6 on both the IP side and the CIDR side, and maps the IPv6 prefix to the corresponding IPv4 prefix (`/120` on `::ffff:a.b.c.d/120` = `/24` on the embedded IPv4).
- Regular IPv6 (non-mapped) still does NOT match IPv4 rules — explicit guard test added.
- 8 new CIDR tests covering: mapped-to-IPv4, IPv4-to-mapped, hex-form mapping, multiple family combos, `ipInAnyCidr` with mixed list.

### Fix 2 — Deeper spawn attack coverage + stderr assertions on every attack
- Added 4 new attack payloads to `tests/spawn-integration.test.ts`:
  - CRLF mixed injection in role (`\r\n`)
  - Unicode NFD normalization bypass (`e` + U+0301 combining acute)
  - Symlink path traversal (creates `/tmp/v163-bad-link-PID -> /etc`, passes as cwd, asserts rejection)
  - Long-payload DoS (cwd > 1024 chars)
- Added stderr-non-empty assertion to every attack test. Catches silent-failure regressions where the script exits 2 but gives no hint why.
- 17 → 21 spawn integration tests.

### Fix 3 — Bash symlink/path-resolution defense in `bin/spawn-agent.sh`
- Added `cd "$CWD" && pwd -P` resolution after text validation. The resolved path must still be under an approved root ($HOME, /tmp, /private/tmp, /var/folders). A symlink pointing outside an approved root is rejected with an explicit "resolves to ... outside approved roots" error.
- Gracefully skips resolution if the path doesn't exist yet (child terminal will fail naturally on `cd`).
- CRLF smuggling in cwd now caught by a dedicated `tr -d` length-comparison block (mirroring the validate_token pattern).

### Fix 4 — Documentation drift
- `CLAUDE.md` was stuck reporting "104 tests, 13 files" (v1.6.1 numbers). Updated to 153 tests / 16 files.
- Added missing entries to the file map: `src/cidr.ts`, `tests/cidr.test.ts`, `tests/spawn-integration.test.ts`, `tests/trusted-proxy.test.ts`.
- Corrected the v1.6.2 CHANGELOG entry: was "validate_token blocks smuggling" (no such function); now points to the real code — `SpawnAgentSchema.SPAWN_CWD_FORBIDDEN` in `src/types.ts` plus the bash `tr -d` check.

### Tests
- 153 tests passing (was 140).
  - +8 CIDR tests (IPv4-mapped IPv6)
  - +4 spawn integration tests (CRLF, Unicode NFD, symlink traversal, DoS)
  - +13 stderr assertions across existing attack tests
- All 140 prior tests pass without modification.

### Version bumps
- package.json: 1.6.2 → 1.6.3
- MCP server version: 1.6.2 → 1.6.3
- /health version: 1.6.3
- Webhook User-Agent: bot-relay-mcp/1.6.3

### Backward compatibility
- `ipInCidr` is stricter only in that it now MATCHES mapped-IPv6 against IPv4 rules where it used to return false — operators who wrote IPv4 CIDRs in `trusted_proxies` now correctly trust dual-stack clients. Non-match cases are unchanged.
- Bash path-resolution defense only triggers when the cwd path exists; non-existent paths pass through as before (child terminal handles the `cd` failure).

## v1.6.2 — 2026-04-14 (Defense-in-depth + trusted-proxy config)

A review of v1.6.1 found 2 items were PARTIAL. v1.6.2 addresses both fully, with real integration tests replacing mocked ones.

### Fix 1 — Spawn shell injection: defense-in-depth
- **TS-layer validation in `src/types.ts` SpawnAgentSchema.** Zod schema now enforces explicit regex patterns matching the bash layer: name/role `[A-Za-z0-9_.-]+`, each capability item `[A-Za-z0-9_.-]+`, cwd absolute path with `/A-Za-z0-9_./-]` allowlist plus a negative-check `.refine` that rejects any shell metacharacter or control character even if the base pattern somehow let it through. This catches attacks at the MCP boundary before the shell ever runs.
- **Real integration tests via `bin/spawn-agent.sh` with `RELAY_SPAWN_DRY_RUN=1`.** 17 tests in `tests/spawn-integration.test.ts` spawn the actual bash script and feed it attack payloads: semicolons, pipes, ampersands, `$()` and backtick command substitution, newlines, quote-mixing, dollar-sign expansion in capabilities, relative cwd, cwd with command substitution, cwd with backtick or semicolon, oversized name, and `RELAY_TERMINAL_APP` env var injection. Every payload is blocked; dry-run stdout is verified NOT to contain the attack.
- **Inline comments in `bin/spawn-agent.sh`** now explicitly document the three layers of defense (TS Zod → bash regex → `printf %q` + AppleScript escape) and warn future maintainers against simplifying.
- Fixed a macOS bash 3.2 bug where `$'\n'` in case patterns didn't match reliably. Replaced with a length-comparison check after `tr -d` strips control chars.

### Fix 2 — X-Forwarded-For trusted-proxy configuration
- **New config field `trusted_proxies: string[]`** (CIDR blocks, default empty).
- **New env var `RELAY_TRUSTED_PROXIES`** (comma-separated CIDRs) that overrides the file config.
- **Behavior change:** when `trusted_proxies` is empty (DEFAULT), the X-Forwarded-For header is COMPLETELY IGNORED. Rate limits key only on the direct socket peer IP. This closes the previous spoofing vector where any caller could send `X-Forwarded-For: 1.2.3.4` and get a fresh quota bucket.
- **When trusted_proxies is configured:** the server only honors X-Forwarded-For if the direct peer IP falls in one of the trusted CIDRs. It then walks the XFF chain right-to-left, skipping trusted hops, and picks the leftmost-untrusted hop as the "real" client IP. This matches RFC 7239 §7.4 and how nginx/Express/Rails normally handle this.
- **New `src/cidr.ts` CIDR matcher** supports both IPv4 and IPv6, with unit tests covering exact matches, /0, /24, /8, /32, /128, /10, IPv4-mapped IPv6 edge cases, malformed input rejection, and cross-family non-matching.

### Tests
- 140 tests passing (was 104).
  - +23 CIDR tests (IPv4, IPv6, /0, /24, /8, /32, /128, cross-family, malformed, ipInAnyCidr)
  - +17 spawn integration tests (real bash invocation, 15+ attack payloads)
  - +2 trusted-proxy HTTP tests (XFF ignored by default, XFF honored from trusted peer)
  - Small tweaks: null-byte/newline/tab/CR smuggling now explicitly blocked by the `SpawnAgentSchema.SPAWN_CWD_FORBIDDEN` regex in `src/types.ts` (TS layer) and the `tr -d` length check in `bin/spawn-agent.sh` (bash layer)

### Version bumps
- package.json: 1.6.1 → 1.6.2
- MCP server version: 1.6.1 → 1.6.2
- /health version: 1.6.2
- Webhook User-Agent: bot-relay-mcp/1.6.2

### Backward compatibility
- `trusted_proxies` defaults to `[]`. Behavior for existing deployments WITHOUT config is unchanged at the behavioral level: we never honored XFF from them meaningfully before (we did in v1.6.1, which a review flagged as a leak). For anyone relying on XFF for rate limiting, they now need to explicitly configure `trusted_proxies`.
- All 104 v1.6.1 tests pass without modification.
- All tool signatures unchanged.

### Files added
- `src/cidr.ts` — IPv4/IPv6 CIDR matching utility (100 lines)
- `tests/cidr.test.ts` — 23 CIDR tests
- `tests/spawn-integration.test.ts` — 17 real-shell integration tests
- `tests/trusted-proxy.test.ts` — 2 HTTP-level XFF tests

## v1.6.1 — 2026-04-14 (review fixes — 3 blockers + 5 fix-this-session items)

A review of v1.6 held npm publish on three blockers + five fix-this-session items. All eight landed.

### Blockers resolved
- **`bin/spawn-agent.sh` shell injection (CRITICAL).** Previously interpolated `$NAME`/`$ROLE`/`$CAPS`/`$CWD` directly into shell + osascript. Rewrote with: input validation regexes (name/role `[A-Za-z0-9_.-]`, caps `[A-Za-z0-9_.,-]`, cwd must be absolute path with no shell metacharacters), `printf %q` quoting for shell interpolation, and an `applescript_escape` helper that handles `\` and `"` for the AppleScript heredoc. 6 injection payloads verified blocked in smoke test.
- **MCP SDK pin mismatch.** Installed was 1.29.0 while package.json pinned `~1.12.1` — lockfile wasn't forced on pin change. Bumped pin to `~1.29.0` (the version all tests were already passing on) and ran `npm install` to lock it in.
- **Concurrent test was sequential.** Replaced the in-process for-loop with a child-process based test that spawns a second Node process writing the same SQLite file. Real OS-level contention; busy_timeout must kick in for both sides to complete. 100 writes per process, all 200 land.

### Fix-this-session items resolved
- **`src/db.ts` path traversal validation.** Mirrors `check-relay.sh` logic: RELAY_DB_PATH must resolve under `$HOME`, `/tmp`, `/private/tmp`, or `/var/folders`. Throws at startup otherwise.
- **HTTP no-auth rate-limit bypass.** Caller could rotate `agent_name` per call to reset their quota. Fix: new `src/request-context.ts` uses AsyncLocalStorage to bind source IP to each HTTP request, and the server dispatcher composes the rate-limit key as `ip:<addr>` when the call is HTTP + unauthenticated. IP-based quotas cannot be bypassed by agent name switching.
- **`CLAUDE.md` stale at v1.4.** Updated to v1.6.1, 104 tests, 14 tools, full file map reflects current src/ layout.
- **Tool-level rate-limit rejection tests.** Added 3 tests in `security.test.ts`: tool rejection after limit hit, IP-keyed bypass prevention, separate IPs get separate quotas.
- **Hardcoded `sleep(200)` replaced with polling.** `tests/unregister.test.ts` "does NOT fire webhook" test now polls up to 500ms, failing early if the webhook wrongly fires.

### Tests
- 104 tests passing (was 100).
  - +1 concurrent OS-level contention test (child process + parent write race)
  - +3 rate-limit rejection tests (tool level, IP-keyed, multi-IP)
  - Fixed: concurrent test had a 5th arg for 4 placeholders — corrected
  - Fixed: concurrent test was sequential — now actually concurrent across processes

### Version bumps
- package.json: 1.6.0 → 1.6.1
- MCP SDK pin: `~1.12.1` → `~1.29.0` (matches installed + lockfile)
- MCP server version: 1.6.0 → 1.6.1
- /health version: 1.6.1
- Webhook User-Agent: bot-relay-mcp/1.6.1

### Backward compatibility
- Zero behavior changes to tool signatures.
- Zero schema changes.
- The HTTP no-auth rate limit now keys by IP when unauth — existing stdio and authenticated HTTP behavior unchanged.

## v1.6.0 — 2026-04-14 (Hardening pass — no new features)

After 3 parallel research agents audited security, architecture, and tech stack, this release fixes the real issues they found. Zero new features. Zero new tools. Just hardening.

### Security fixes
- **SSRF protection on webhooks.** `register_webhook` now resolves DNS at registration time and rejects URLs targeting private IP ranges (10.x, 172.16/12, 192.168.x, 127.x, 169.254.x cloud metadata, fc00::/7, fe80::/10, ::1) and non-HTTP(S) schemes (file://, ftp://, gopher://). Set `RELAY_ALLOW_PRIVATE_WEBHOOKS=1` to opt-in for local n8n at 127.0.0.1.
- **Hook script input validation.** `check-relay.sh` now validates `RELAY_AGENT_NAME`, `RELAY_AGENT_ROLE`, and `RELAY_AGENT_CAPABILITIES` against `[A-Za-z0-9_.-]` before passing them to sqlite3. SQL injection and shell-substitution attacks are blocked at the input boundary.
- **Path traversal protection.** `RELAY_DB_PATH` is now resolved and must live under `$HOME` or `/tmp` (`/private/tmp`, `/var/folders` for macOS test environments). Pointing at `/etc/passwd` is rejected.
- **Dual-key parameter binding in hook script.** Although input validation already prevents SQL injection, the hook now uses sqlite3's `.parameter set` mechanism for defense-in-depth.

### Tech hygiene
- **Stderr-only logger** (`src/logger.ts`) — every log goes to stderr regardless of transport. Replaced all internal `console.error` calls. Stdout in stdio mode is reserved exclusively for the MCP JSON-RPC channel.
- **CI-style test that fails on `console.log` regression** in any source file. Prevents the #1 silent-break failure mode reported on community MCP servers.
- **Pre-log webhook deliveries** before the fetch fires (was after). A process crash mid-delivery still leaves an audit trail.
- **Node 18+ runtime check** at startup with a clear error message (the engines.node field doesn't enforce at runtime).
- **MCP SDK pinned to `~1.12.1`** (patch-only updates) to avoid silent breakage from minor version bumps.

### Tests
- 100 tests passing (was 79).
  - +15 URL safety tests (scheme, IP literal blocking, IPv6, opt-in, public destinations)
  - +2 stdout discipline tests (no console.log in src/, no process.stdout.write)
  - +3 concurrent write tests (busy_timeout effective, WAL mode, two-connection contention)
  - +1 schema migration test (v1.0-era DB upgrades cleanly to v1.6)

### Backward compatibility
- All 79 v1.5 tests pass with minor adjustments (4 webhook tests needed `await` for the now-async `handleRegisterWebhook`, plus `RELAY_ALLOW_PRIVATE_WEBHOOKS=1` for receivers on 127.0.0.1).
- No new MCP tools, no schema changes, no breaking config changes.
- HTTP auth unchanged from v1.5.

### Version bumps
- package.json: 1.5.0 → 1.6.0
- MCP server version: 1.5.0 → 1.6.0
- /health version: 1.6.0
- Webhook User-Agent: bot-relay-mcp/1.6.0

## v1.5.0 — 2026-04-14

### Added — Security hardening (responding to user feedback on built-in security)

#### Shared-secret auth on HTTP transport
- New config option `http_secret` (file) / `RELAY_HTTP_SECRET` (env var).
- When set, all HTTP requests except `/health` require `Authorization: Bearer <secret>` or `X-Relay-Secret: <secret>` header.
- Rejects missing or wrong secret with HTTP 401 and a helpful hint.
- `/health` stays open for monitors to ping without credentials — now reports `auth_required` in its response.
- Solo stdio use is unaffected (no auth required for stdio transport).

#### Audit log
- New `audit_log` SQLite table.
- Every tool call is logged with agent name, tool, param summary (first 80 chars of key fields), success/failure, and error if any.
- Auto-purges entries older than 30 days.
- Queryable via `getAuditLog(agentName?, tool?, limit?)` library function. (No MCP tool yet — add in v1.6 if users request it.)

#### Rate limiting (sliding-window, per agent per bucket)
- Three buckets: `messages` (send_message + broadcast), `tasks` (post_task), `spawns` (spawn_agent).
- Defaults: 1000 messages/hour, 200 tasks/hour, 50 spawns/hour. 0 disables.
- Configurable via `rate_limit_messages_per_hour`, `rate_limit_tasks_per_hour`, `rate_limit_spawns_per_hour` in `~/.bot-relay/config.json`.
- Over-limit calls return structured error with current/limit counts and reset hint.
- Every rate-limit rejection is also logged to the audit log.

### Changed
- Server version: 1.4.0 → 1.5.0
- `/health` response now includes `auth_required` boolean.
- Tool dispatcher wrapped with rate-limit check + audit logging. All 14 tools still function identically — this is purely additive.

### Tests
- 79 tests passing (was 63)
  - 9 new security tests (audit log writes/filters, rate limit per agent + bucket)
  - 7 new HTTP auth tests (401 without auth, Bearer token, X-Relay-Secret, health exempt, dashboard protected)
- All 63 v1.4 tests pass without modification.

### Backward compatibility
- Default config has `http_secret: null` — HTTP mode works with no auth if the user doesn't set one. This is identical to v1.4 behavior.
- Default rate limits are generous (1000/hr messages) and can be disabled by setting to 0.
- stdio mode unchanged.

## v1.4.0 — 2026-04-14

### Added — spawn_agent
- New MCP tool: `spawn_agent(name, role, capabilities, cwd?, initial_message?)`
- Opens a new Claude Code terminal window (iTerm2 or Terminal.app) pre-configured with `RELAY_AGENT_NAME`, `RELAY_AGENT_ROLE`, `RELAY_AGENT_CAPABILITIES` env vars.
- The SessionStart hook auto-registers the agent and delivers any queued mail on arrival.
- Optional `initial_message` queues a message before spawning, so the new agent sees instructions on first wake.
- Fires new `agent.spawned` webhook event.
- Shell script at `bin/spawn-agent.sh` can also be called directly from the command line.
- macOS only for now (uses osascript). Linux/Windows support is a v2 candidate.

### Added — Dashboard
- Built-in HTML dashboard served at `GET /` and `GET /dashboard` in HTTP mode.
- JSON snapshot API at `GET /api/snapshot` (agents, messages, active/completed tasks, webhooks).
- Vanilla JS, no build step, auto-refreshes every 3 seconds.
- Color-coded presence status (online/stale/offline), priority badges, task state badges.
- Dark theme matching common terminal aesthetics.

### Added — Role templates
- New `roles/` directory with drop-in CLAUDE.md snippets for common agent roles:
  - `planner.md` — orchestrator/delegator
  - `builder.md` — worker that accepts and completes tasks
  - `reviewer.md` — skeptical reviewer with structured output
  - `researcher.md` — investigates questions, returns findings
- `roles/README.md` explains three ways to apply a role (per-project CLAUDE.md, spawn initial_message, shell alias).

### Added — Hardening
- SQLite `busy_timeout = 5000ms` — waits up to 5s for write locks instead of throwing SQLITE_BUSY. Prevents spurious errors under burst traffic.

### Changed
- MCP server version: 1.3.0 → 1.4.0
- Tool count: 13 → 14
- Webhook events: 8 → 9 (added `agent.spawned`)
- Added `mcp__bot-relay__spawn_agent` to pre-approved tools in `.claude/settings.json`

### Tests
- 63 tests passing (was 56)
  - 5 new spawn tests (mocked child_process.spawn)
  - 2 new HTTP tests (dashboard HTML, snapshot API)
- All existing tests untouched

### Fixed
- Updated tool count test from 13 to 14 to match new `spawn_agent`.

## v1.3.0 — 2026-04-14

### Fixed — Presence integrity
- `getMessages()` no longer bumps `last_seen` on the agent calling it. Reading your mailbox is observation, not liveness.
- `getTasks()` no longer bumps `last_seen` on the agent calling it. Same reason.
- `registerAgent`, `sendMessage(from)`, `broadcastMessage(from)`, `postTask(from)`, `updateTask(agent_name)` still bump `last_seen` — these are real actions.
- Net effect: `discover_agents` now tells the truth about who is actually doing something vs who is just lurking.

### Added — Agent lifecycle
- New tool: `unregister_agent(name)` — removes an agent from the relay. Idempotent (returns `removed: false` if the name was not registered).
- New webhook event: `agent.unregistered` — fires when an agent is successfully removed. Does not fire on idempotent no-op removes.
- `agent.unregistered` payload: `from_agent` and `to_agent` both equal the removed name (self-event).
- Added `mcp__bot-relay__unregister_agent` to the pre-approved tools in `.claude/settings.json`.
- **Deliberately skipped:** auto-unregister on SIGINT/SIGTERM in the stdio transport. The stdio transport has no per-connection state that maps a process to its registered agent name. Adding that requires richer per-connection state (v2+ scope). Exposing the tool is enough: clients can call `unregister_agent` themselves on shutdown, or hooks can clean up stale entries.

### Changed — SessionStart hook
- `hooks/check-relay.sh` now registers the agent (upsert) before checking mail. Registration is a real liveness signal.
- Agent name, role, and capabilities are read from `RELAY_AGENT_NAME`, `RELAY_AGENT_ROLE`, `RELAY_AGENT_CAPABILITIES` env vars (comma-separated for caps). Sensible defaults: `default` / `user` / empty array.
- Pending messages and active tasks are printed to stdout (injected into Claude's context) AND stderr (shown to the human) on session open.
- Uses `sqlite3` CLI directly — no daemon dependency, works regardless of transport mode.

### Tests
- 56 tests passing (was 48).
- +4 presence tests (getMessages/getTasks don't touch, sendMessage/postTask do).
- +4 unregister tests (removal, idempotency, webhook fires, no webhook on no-op).

### Unchanged (backward-compatible)
- All 48 v1.2 tests pass without modification.
- All existing tool signatures identical.
- No SQLite schema changes.
- stdio and HTTP transports unchanged.

### Version
- `package.json`: 1.2.0 → 1.3.0
- MCP server version: 1.2.0 → 1.3.0 (seen in initialize handshake)
- `/health` version: 1.3.0
- Webhook `User-Agent`: `bot-relay-mcp/1.3.0`

## v1.2.0 — 2026-04-14

### Added — HTTP Transport
- `StreamableHTTPServerTransport` support alongside stdio
- New entry point supports three transport modes via `RELAY_TRANSPORT` env var or config file:
  - `stdio` (default) — current behavior, one server per terminal
  - `http` — HTTP daemon mode on `RELAY_HTTP_PORT` (default 3777)
  - `both` — HTTP server plus stdio (for daemon + local Claude Code simultaneously)
- `/health` endpoint for HTTP mode (`GET /health` returns status JSON)
- `/mcp` endpoint handles JSON-RPC over HTTP with SSE streaming
- Stateless mode — each request gets its own transport; all share the same SQLite

### Added — Webhook System
- New SQLite tables: `webhook_subscriptions`, `webhook_delivery_log`
- New tools:
  - `register_webhook(url, event, filter?, secret?)` — subscribe to relay events
  - `list_webhooks()` — list all subscriptions (secrets hidden)
  - `delete_webhook(webhook_id)` — remove a subscription
- Supported events: `message.sent`, `message.broadcast`, `task.posted`, `task.accepted`, `task.completed`, `task.rejected`, `*`
- Fire-and-forget delivery with 5s timeout (does not block tool responses)
- HMAC-SHA256 signatures in `X-Relay-Signature` header when `secret` is set
- Optional agent name filter (fires only when `from_agent` or `to_agent` matches)
- Delivery attempts logged to `webhook_delivery_log` with status code or error
- Auto-purge: delivery logs older than 7 days

### Added — Config File
- `~/.bot-relay/config.json` for transport mode, HTTP port, webhook timeout, API allowlist
- Environment variables (`RELAY_TRANSPORT`, `RELAY_HTTP_PORT`, `RELAY_HTTP_HOST`) override file config
- Invalid or missing config falls back to safe defaults

### Changed
- Refactored `src/index.ts` into `src/server.ts` (reusable factory) + `src/transport/{stdio,http}.ts`
- Server version bumped to 1.2.0 in MCP handshake
- 12 MCP tools now registered (was 9 in v1.1)
- Version: 1.1.0 → 1.2.0

### Dependencies
- Added `express@^5.2.1` and `@types/express` (devDep)

### Tests
- 48 tests passing (was 28 in v1.1)
  - 21 database layer (unchanged)
  - 7 tool integration (unchanged)
  - 11 new webhook tests (registration, firing on all events, HMAC, filters, failure handling)
  - 4 new HTTP transport tests (health, tools/list, round-trip, method restrictions)
  - 5 new config loader tests (defaults, file, env overrides, malformed input)

### Unchanged (backward-compatible)
- All existing tool signatures and behaviors preserved
- stdio mode identical to v1.1
- Existing SQLite tables unchanged — only new tables added
- All v1.1 tests still pass without modification

## v1.1.0 — 2026-04-13

### Added
- `get_tasks` tool — query your task queue by role (assigned/posted) and status
- `get_task` tool — look up a single task by ID
- 28 tests (vitest) covering database layer and tool handlers
- README.md with Quick Start, tool reference, examples, roadmap
- SessionStart hook (`hooks/check-relay.sh`) for auto-checking relay at session start
- `.claude/settings.json` pre-approving all relay tools (zero friction)
- `docs/hooks.md` — hook setup guide
- `docs/claude-md-snippet.md` — CLAUDE.md instructions for users
- `.gitignore`, MIT LICENSE

### Changed
- Moved project from `side-projects/bot-relay-mcp/` to `bot-relay-mcp/` (top-level)
- Updated MCP path in `~/.claude.json`
- Version: 1.0.0 → 1.1.0

## v1.0.0 — 2026-04-06

### Added
- Initial release
- 7 MCP tools: `register_agent`, `discover_agents`, `send_message`, `get_messages`, `broadcast`, `post_task`, `update_task`
- TypeScript, stdio transport, SQLite shared state
- WAL mode for concurrent access
- Auto-purge for old messages (7 days) and completed/rejected tasks (30 days)
