# Changelog

## 3.4.24

- A repeating watch keeps polling while its triggered delivery waits for a busy room, waits to retry after a failed or `/stop`-killed run, or runs. A new hit folds into that one pending delivery instead of queueing another, and the delivered prompt says how many more times the check passed and whether its latest run had cleared. On 24 Sep the event-bus watch went unpolled for 7.5 h behind a busy room, and after a killed delivery attempt its polling never resumed. One-shot watches are unchanged, and the 15-minute "⏳ Watch … waiting" line still goes out once per occurrence.
- A watch whose check times out 3 times in a row now wakes the agent once, marked "Watch check timing out", and keeps polling; any completed check resets the count. A hanging health check used to be logged `idle (check timed out)` and never alarm. The alarm does not use up a one-shot watch.
- Wakeups more than ~24.8 days out no longer fire at once. `setTimeout` overflows past 2^31-1 ms, which fired a 2028 reminder 61 s after it was set (17 Sep); far wakeups now sleep in chunks and re-plan until due, then fire once. `schedule-wakeup` refuses a `<when>` more than 10 years out, in the CLI and in the loopback route.
- A queued delivery is no longer stranded when a jobs-store write throws (a full disk, a torn journal). Those paths returned without a retry timer, so the alarm was never tried again; they now keep the retry armed and say so in watch.log. This matches the watch alarms that triggered during long turns on 21 and 23 Sep and were never delivered. A regression test also pins that a delivery queued behind a long turn survives the turn's compaction and runs once at the next idle moment, on the compacted session; that path already worked.
- `cron-list` and `watch-list` say where a fired job stands (delivery waiting, running, retrying or dead) instead of listing it as still to come. A triggered watch whose delivery was queued read "EXPIRED (fires on next poll)" (21 Sep), and a fired 2028 reminder still read "fires 2028-08-01" (17 Sep).
- `open-claudia say` and the send-file/voice/photo routes now say which step failed when the channel adapter refuses the upload, instead of a bare `loopback 500: {"ok":false}` (10 Sep). The adapter's own error stays in bot.log.
- Release gate: `test-provider-pack-review.js` gives each probe child 120 s instead of 30 s and reports a timeout as a timeout rather than "failed (null)" (v3.4.19, job 43703). The parity suite's cap for it rises to 8 minutes so it cannot undercut its children.

## 3.4.23

- Watch, wakeup and cron turns now stream their progress (interim narration, and thinking under `/verbose`) like chat turns. The live preview was gated to chat turns only, so a watch-woken turn showed the bell notice and "thinking…" and then nothing until its final answer, however long it worked.
- `/stop` no longer cancels the messages you send after it. It released the room the moment it killed the run, before that run had unwound, so a message sent in those seconds started, met the stop flag at its pre-spawn checkpoint and was cancelled "before the provider started" — three messages lost after one `/stop` on 24 Sep. The room now stays busy until the stopped run's own unwind releases it and drains the queue (a 30 s backstop covers a child that never reports its exit), and the stop flag names the run it stops, so it can never cancel another.
- "Queued." now says what the message is waiting on once the wait passes two minutes, e.g. `Queued behind the watch "Phone alarm" (running 42 min). /stop cancels it.`
- A watch that has triggered but waited 15 minutes behind a busy room posts one plain line per occurrence, so an alarm is never hidden for hours (24 Sep: 2 h and 7.5 h, with nothing in the chat).

## 3.4.22

Supersedes 3.4.21, which was tagged but never published: its build failed because `test-upgrade-path.js` was wired into the test chain without being committed. Nothing shipped from that tag.


- Fix `/upgrade` reporting success while upgrading a copy of the bot that nothing runs. It ran `npm install -g`, then read the resulting version back from `npm root -g` — the one directory guaranteed to agree with itself. Found on a host where the bot executes from `/usr/local/lib/node_modules` while `npm config get prefix` is `~/.local`: three releases (3.4.18, 3.4.19, 3.4.20) were each announced as installed while the running process stayed on 3.4.17.
- The layout check no longer string-matches the docker `/app` path. It compares the directory the running process executes from against the directory npm would write to, so any prefix drift routes to the in-place overlay, not just the container case. An unqueryable `npm root -g` is not treated as a mismatch — the post-install check is the backstop there.
- `/upgrade` and `/downgrade` now read the resulting version back from the directory the process actually executes from. A mismatch reports "Upgrade did NOT take", names both directories, and restarts nothing; the previous behaviour was to restart and claim success.
- Drop the superseded `claude-opus-4-8`, `claude-opus-4-7` and `claude-opus-4-6` buttons from the Claude model picker, which had grown long enough to bury the current models.
- Removal is cosmetic by design. `/model` takes a typed name verbatim without consulting the list, so `/model claude-opus-4-8` still works as a rollback; only the buttons go. The list is still checked on tapped callbacks, whose job is to reject a stale button from an old message.
- `claude-sonnet-4-6` and `claude-haiku-4-5-20251001` deliberately stay: they are the tier models for medium and low utility work, not picker garnish, and dropping them would push recall, pack review and the enforcer onto opus.

## 3.4.20

- Make `claude-opus-5-5` the default Claude model. 3.4.19 only added it to the picker; the default stayed on `claude-opus-4-8`, so a room that had never run `/model` kept the old model.
- The default governs main conversation turns only. Utility work remains tier-pinned — `low` to `claude-haiku-4-5-20251001` and `medium` to `claude-sonnet-4-6` — so recall, pack review, the dedup adjudicator and the enforcer do not reprice with it, and a room's own `/model` choice still wins over the default.
- Note for operators: `CLAUDE_MODEL` in `~/.open-claudia/.env` (or the process env) overrides this. An install that pinned the old default there keeps it until that line is changed. Pod images carry no `CLAUDE_MODEL` and the entrypoint never writes one, so pods resolve the shipped constant.
- `claude-opus-5-5` requires Claude Code >= 2.1.280; an older CLI refuses it with an API 400 rather than degrading. The pod image installs the CLI unpinned, so an image built before that release must not run this default.

## 3.4.19

- Offer `claude-opus-5-5` in the Claude model picker. Verified against the live API rather than assumed: on Claude Code 2.1.280 the API bills the request to `claude-opus-5-5` (1M context, 128k max output), while 2.1.258 refuses it with "does not support this model; version 2.1.280 or newer is required". An older CLI returns that API 400 instead of degrading, so a pod on a stale image errors on the model rather than falling back — the same constraint already recorded for `claude-fable-5-1` (>= 2.1.251).

## 3.4.18

- Pass images and files from recent Kazee group messages and quoted replies into the model, retaining sender and message attribution. Follow-up questions can reuse attachments after the text catch-up watermark advances.
- Keep group participation and capability checks ahead of referenced downloads. Exclude deleted and cross-room attachments; prioritize quoted media, deduplicate files, and bound downloads by count, size, time and free disk space.
- Cache completed downloads under generated filenames and explicitly report unavailable or omitted attachments. Cover real image bytes, repeated follow-ups, reply resolution, runner attachment handoff, held turns, room isolation and failed downloads.

## 3.4.17

- Include resolved caller, channel and room identifiers in external guard judgments so caller-specific permissions can be evaluated for contacts without a saved profile. Identity never grants permission by itself; existing policy and refusal behavior remain in force.
- Bind cached guard decisions to caller, channel, room and displayed person context, preventing one contact from inheriting another contact's verdict when the proposed reply and policy text match.
- Reproduce the missing-identity failure and cover same-caller cache reuse, caller/room/transport separation, missing IDs, claimed IDs in proposed text and display-label changes.

## 3.4.16

- Start live-call answer generation without waiting for an auxiliary memory-relevance model. Calls keep local matches from the current utterance, bounded spoken history, and full policy; context-only keyword matches remain excluded without a relevance judgment. Ordinary chat filtering and explicit on-demand recall are unchanged.
- Log live-call provider preparation, prompt readiness, process start and completion with elapsed time, so recognition delay can be distinguished from reply startup and generation.
- Reproduce the call-path regression through the actual prompt builder and classic matcher, covering both configured recall engines, context-only exclusions, ordinary chat filtering and on-demand recall.

## 3.4.15

- Preserve complete Kazee speech turns with protocol-2 segment matching, bounded pending finals, and cancellation when speech resumes during an endpoint decision. Keep compatibility with older speech servers, including a delayed protocol-ready message.
- Stop obsolete speech and remove superseded queued call turns on interruption or hangup. Use a fresh provider conversation with bounded spoken history while preserving identity, policy and recall. In-progress tools continue under their existing lifecycle.
- Add `open-claudia call-say` for explicit complete answer sentences before the provider finishes. Each request is bound to the active call and run, redacted, and checked by the outbound reply guard. Retries are idempotent; final delivery removes the exact prefix already spoken. Provider previews and reasoning remain unspoken; this does not add token streaming to the provider transport.
- Replace repeated thinking sounds with optional delayed fixed cues from the separate speech controller. Endpoint and cue models require the corresponding speech-service deployment; model quality, live latency and Android acceptance remain separate verification gates.
- Isolate permission-test fixtures from inherited test-launcher approval without changing production permission gates.

## 3.4.14

- Subscribe to managed call audio only when gateway-authenticated publisher metadata identifies a User. Exclude bots, the current bot's other feeds, and missing or malformed classifications before creating subscriber connections or speech streams.
- Preserve metadata through the initial room response and live publisher events. Managed calls use that authenticated classification; legacy calls retain their existing display-prefix exclusion because client-supplied metadata is untrusted.
- Exercise the actual session subscription and offer/answer path for allowed users and rejected bot, self and unknown feeds.

## 3.4.13

- **Codex is now measured the way Claude is.** After each Codex turn the bot reads Codex's own per-request token log (`$CODEX_HOME/sessions/…/rollout-*-<thread>.jsonl`) and uses the last request's size as the live context, instead of the turn's input summed over every tool call. That sum (33M on one turn) tripped auto-compact after almost every substantial Codex turn — 8 needless compactions on 2026-09-05, each a summary run plus a cold restart — while Codex's real window occupancy never exceeded 245k of its 258k window. Codex sessions now compact at 80% of the window Codex reports.
- **Codex usage was mis-recorded three ways; all fixed.** (1) `turn.completed` usage is per turn by Codex's own definition, but the ledger treated it as a running total and subtracted the previous turn — 9 of 36 Codex turns that day were logged as zero. (2) OpenAI counts cached tokens inside `input_tokens`; the ledger and `/usage` now report uncached input for both providers (the raw comparison read 89.7M vs 4.5k for what was really 2.4M vs 0.8M). (3) Sub-agent threads Codex spawns with `spawn_agent` never appeared anywhere: the day's ten hidden threads read ~240M tokens and took a weekly quota from 75% to 100% in two hours. They are now attributed to the parent turn (`subagentThreads`, `subagentInputTokens`, …) and shown by `/usage`.
- **Quota and window visibility.** Ledger rows carry `contextWindowTokens`, `requests`, `providerCompactions` (Codex's own compactions) and the plan quota snapshot (`quotaUsedPercent`, `quotaResetsAt`); `/usage` prints window occupancy, sub-agent totals and quota used.

## 3.4.12

- Keep incoming calls on their per-attempt admission policy when the organization canary setting changes; managed access supplies the assigned gateway and TURN configuration. Owner-only audio calling remains enforced.
- Flatten ICE server URLs for Werift, whose installed parser silently omits TURN when given browser-style URL arrays. Verify actual peer-connection TURN configuration in the tests.
- Refresh Janus grants and relay credentials for future peer connections. Existing Werift TURN allocations keep their original credentials and expiry, so the bot closes before the earliest active allocation deadline. It does not hot-swap existing allocation authentication; the default one-hour TURN lifetime exceeds the bot's thirty-minute call limit.
- Exclude the bot's own publisher by its Janus-assigned feed ID, including when the gateway replaces display names with canonical email. Keep the legacy bot display-prefix exclusion.

## 3.4.11

- Verify timed-out guard processes from their parent, including forced termination and released concurrency, without racing a child-written receipt.
- Retain the calling credential-renewal and scheduler fixes from the preceding release candidates.

## 3.4.10

- Include managed device ownership generations on bot call signals, reject stale refresh generations, and report local termination once.
- Renew transport credentials before expiry on the existing Janus session.
- Give the durable migration regression a bounded two-minute subprocess budget and report spawn errors; the CI probe exceeded its previous thirty-second limit.

## 3.4.9

- Keep concurrent Opus codecs valid when WebAssembly memory grows; test 100 simultaneous codecs and legacy packet interoperability.
- Use the assigned Janus endpoint for managed calls and reject placement changes during credential refresh.
- Route release builds through the dedicated VM runner after cluster runner startup failures.

## v3.4.8 — Authenticated calling rooms and attempt-safe bot setup

- Kazee calling: stage server-authorized room allocation and expiring room credentials, authenticate Janus frames and subscriber joins, preserve call attempt identity across signals, and cancel stale setup. Legacy calls remain supported while room enforcement is disabled. Owner-only audio calling policy is preserved.

## v3.4.7 — Keep scheduled follow-ups queued while the chat is busy

- Scheduler: busy chats keep wakeups queued without consuming execution retries. Watch triggers and expiries are saved once, survive restart, and stop polling while delivery is pending. Concurrent callbacks cannot start overlapping runs; advisory notifications are attempted once per occurrence. Real execution failures retain bounded retries.

## v3.4.6 — Codex runs again: the gate stopped asking for the impossible

- **Every Codex turn since Codex CLI 0.153.x failed with `OpenAI Codex is incompatible: Codex rejected the isolated PreToolUse configuration or could not disable uncovered tool transports`.** The mandatory tool-hook gate ran `codex features list` with our overrides and required `unified_exec` to read `false`. Codex 0.153.2, 0.153.3 and 0.153.4 all ignore `-c features.unified_exec=false` (the feature is locked on and `codex features list` reports `unified_exec stable true` regardless), so the gate refused the provider on every turn — the failure Sumeet hit right after switching to gpt-6-astra on 3.4.5.
- **Why it is safe to stop requiring it off.** Probed live on 0.153.4 (2026-09-05): a shell command run through unified exec still arrives at the PreToolUse hook as tool_name `Bash` with `tool_input.command`, and the real deny gate caught the probe command (deny audited, marker file never written). Unified exec is a covered transport, not an uncovered one. Plugins stay forced off and hooks stay forced on; project trust stays untrusted.
- **What changed.** `core/providers/codex-hook.js` no longer emits the `features.unified_exec=false` override nor checks its state. Fixtures and the tool-hook tests now model Codex as it is (`unified_exec stable true`) and assert no unified_exec override is generated. Fleet pods are affected identically (the 3.4.4 and 3.4.5 images carry Codex 0.153.x), so Codex on pods works only from this release onward.

## v3.4.5 — Astra on the picker, two shortcuts fewer

- **gpt-6-astra is offered by the codex provider.** OpenAI rolled GPT-6 (Astra) out to ChatGPT-plan Codex accounts on 2026-09-05 — the earlier 400 (`not supported when using Codex with a ChatGPT account`) was a staged rollout, not client staleness. It sits first in `CODEX_MODELS`, so it appears as a `/model` button; the typed form `/model gpt-6-astra` already worked because the typed path never validated against the list. The default model is unchanged (`gpt-5.6-sol` unless `CODEX_MODEL` is set).
- **`/claude` and `/codex` are gone.** They were one-line aliases of `/backend claude|codex`, and the `/model` picker switches the backend by itself, so three ways to do one thing became one. `/backend` stays.
- **Codex CLI: the Dockerfile installs `@openai/codex` unpinned, so every image build bakes the then-latest CLI; a running pod keeps whatever its image carried until the next build + roll.** This release's image picks up 0.153.4.

## v3.4.4 — The room's model was saved, shown, and never once used

- **A model pinned in a room was silently ignored by every turn in it.** `/model` wrote the choice to the room, `/status` read it back from the room, and the picker highlighted it — all three already room-aware. Only the code that actually launches the provider was not: `captureRunContext` read `state.providerSettings[provider]`, the *speaker's* own bucket, so the room's pinned model was never consulted. The setting looked applied everywhere a human could look while every turn ran the personal default.
- **This was a v3.4.0 regression, not a long-standing bug.** `run-context.js` dates from 10 Jul, when the speaker's bucket was the only scope that existed and reading it directly was correct. v3.4.0 introduced rooms and the room-aware resolver but never revisited that read. Channels without a room are unaffected — the resolver returns the same personal bucket there.
- **Three more call sites had the same shape.** `applyProviderSetting` wrote `/effort`, `/budget`, `/plan` and `/worktree` to the speaker instead of the room; the provider-switch payload reported the speaker's bucket; and usage records attributed the wrong model, so cost accounting named a model the turn never ran on.
- **`getProviderSettings` was not identity-stable.** It rebuilt the bucket (`{ ...defaults, ...bucket }`) and assigned it back on every call, so any caller holding the previous reference wrote into a detached copy that was then discarded. Nothing depended on the rebuild; it now backfills missing defaults in place.
- **Reads must not mutate.** Routing the run-context read through `getProviderSettings` made merely capturing a turn seed room buckets into `savedState` — side-chat execution began mutating main state. Reads now go through a new pure `resolveProviderSettings`, which creates nothing; the mutating accessor is reserved for the setter that genuinely needs the live object.
- `test-room-provider-settings.js` pins the behaviour and was verified to FAIL against unmodified code with exactly the reported symptom — `got "personal-default-model", expected "room-pinned-model"` — and also pins that an unpinned room does not inherit another room's choice.

- **Kazee replied to everything twice.** chat-central fans `message:new` out to both `chatId` and each participant's `user:<id>` room so delivery survives a failed join. The bot only ever sat in its own user room until v3.4.2 joined chat rooms for the typing indicator — which put it in both, so every inbound message was handled twice and `/model` answered with two identical pickers. The join is correct; the missing half was the id-dedupe every other client already does. Inbound ids are now remembered (bounded, oldest evicted); an id-less message is still let through rather than silently swallowed.

## v3.4.3 — The chip knew it had started typing; nobody asked it in time

- **v3.4.2 restored the indicator and immediately exposed a second bug underneath it.** With delivery fixed, the chip reached the client — and then sat on "thinking" for the whole turn and stopped, never showing "typing". Same symptom, different layer: 3.4.2 was DELIVERY, this is CONTENT.
- **The phase was pushed by events but only read by a timer.** `onEvent` flips `state.thinkingPhase` correctly (`text_delta` → typing, `tool_start` → thinking), but `startTypingHeartbeat` re-read that flag on a 4s interval. An agentic turn spends nearly all of its time in tool calls, then streams its final answer and finishes — routinely inside one interval. So no tick ever sampled the flag while it was false, and the one state the user was waiting to see was the one state never transmitted. A slow, chatty turn would show it; the common case never did.
- **Transitions now push.** `state.typingTick` fires on an actual phase change, guarded so a long stream emits one frame rather than one per delta, and the stopper clears the pusher so a stale tick cannot outlive its turn (the same ownership discipline the heartbeat's stopper already used).
- `test-typing-phase-push.js` pins the sequence and was verified to FAIL against unmodified code with precisely the reported symptom — `["thinking"]` where `["thinking", "typing"]` was expected — rather than merely passing against the fix.

## v3.4.2 — The bot was typing into a room it had never joined

- **The typing and thinking indicators had been dead since 2026-08-23, and nothing said so.** chat-central `7aa9276` ("authorize typing") made `handleTypingStart` return early unless the sending socket had joined that chat's Socket.IO room — a correct fix, since `socket.to(chatId)` addresses a room by name and would otherwise let any authenticated user inject an indicator into any chat whose id they knew. But the Kazee adapter only ever joined the bot's own user room, which is how `message:new` reaches it. So messages kept flowing while every `typing:start` was discarded server-side, which is exactly why this read as a rendering bug in the web client.
- **The adapter now joins the chat room, and waits for the ack before signalling.** The ordering is the fix, not an optimisation: `chat:join` awaits a participant lookup while `typing:start` is handled synchronously, so emitting both back-to-back loses the race and the indicator is dropped anyway. Membership is cached so a per-second heartbeat does not re-join on every tick, concurrent ticks share one in-flight join, and the cache is cleared on connect *and* disconnect because server-side room membership dies with the socket. A refused join now emits nothing rather than something silently discarded, and `typingStop` never joins a room merely to leave it.
- Regression coverage in `test-kazee-typing-join.js`, wired into `pretest`: join-before-emit ordering, the `thinking` state surviving, one join per chat rather than per tick, shared in-flight joins, no indicator after a refusal, re-join after reconnect, and stop-only-for-joined-rooms. The reconnect case earned its place immediately by catching a real defect in the first cut of the fix — a synchronous ack settled the join promise before it was registered, stranding a resolved entry that would have suppressed the re-join.

## v3.4.1 — Fable 5.1, and the CLI floor it quietly needs

- **`claude-fable-5-1` joins the Claude model picker.** The identifier is dash-form; `claude-fable-5.1` does not exist and fails as an unknown model.
- **It carries a CLI floor: Claude Code >= 2.1.251.** An older CLI does not degrade gracefully — it returns an API 400 naming the required version. The local install was taken 2.1.150 -> 2.1.258 before the model was listed, and the picker offering a button the local binary cannot serve is the failure mode to avoid. Pod images install Claude Code unpinned, so a rebuild clears the floor by itself, but pods still on older images will 400 on this model until they are rolled.
- Verified by building the real main invocation through the provider and spawning it, rather than by a hand-written CLI probe: the API reported `claude-fable-5-1` as the serving model, and `claude-opus-4-8` was re-run through the same harness to prove the incumbent path was untouched. Full suite 135/135.

## v3.4.0 — Who you are, where you are, and what you may do there are three different questions

- **The bug that started it: a wakeup scheduled in a Codex room kept running on Claude.** `getActiveProvider` resolves the room's overlay from `state.roomKey`, but that key was stamped from `currentConversationKey()` — and a scheduled job resumes its own saved thread, whose conversation key is the job's *project name*, not a room. The overlay never resolved, so every job silently fell back to the creator's personal default. The thread and the room are now separate values (`roomKey`, defaulting to the conversation key so every inbound path is unchanged); the scheduler passes the room the job was created in. The same fusion is why a group answered as Codex to one person and Claude to another. Relatedly, a run that reverts to the provider it was scheduled under because the room's current choice is not runnable now **says so** — in the prompt and in a ⚠️ line on the advisory — instead of leaving a Codex room reporting a Claude quota error with nothing to explain where Claude came from.
- **A room is not a person.** Authority was resolved from `channelId`, which on every transport where a room holds many people is not an identity at all. Principal is now always the transport USER id, carried separately from the room all the way to the enforcer. A group carrying `isOwner` in `auth.json` can no longer promote every member of that group to owner.
- **Capability is a lookup, not a judgment.** A live `(principal × room)` grant short-circuits the enforcer *before* any model call, so an owner's own decision is never re-litigated by a judge. The grant unit is the **pair**: trusting someone in one group never trusts them in another, nor trusts the rest of that group — "🌍 Always — everywhere" is a separate, explicit tap next to the default "✅ Always — here". Grants carry a capability ladder (`read` < `act` < `destructive`, so an `act` grant can never authorize a destructive run), an optional subject scope, and optional expiry or one-turn burn-down. Re-granting replaces rather than stacks, so downgrading actually downgrades. Deliberately additive: *no grant* is not a denial — the existing mandate and channel-rule paths still run underneath. New `/grants` lists and revokes; every issue is audited.
- **Group admins carry zero authority in the pod.** Presence is surface-owned (whoever administers a room decides whether the agent is *in* it — an agent is a user like any other); capability is pod-owned and set by owners alone. Otherwise sharing a room with someone would make them an admin of your account.
- **A second owner could never approve anything.** Escalation resolved `owners()[0]` and posted every card to that one person, so a card went unanswered whenever they were away. Cards now fan out to every reachable owner, deduped by send target; any owner may decide whichever copy they see, and the decision is idempotent.
- **A grant that expired looked exactly like one you never gave.** The bot simply started asking again with nothing to explain why. Escalation cards now name the lapsed line and its date, at the moment the lapse bites.
- **A torn write could silently demote the owner.** `people.json` was written without fsync or a backup and read as `[]` on any failure — no person found means "external", so one interrupted write guarded the owner in their own pod. Worse, a corrupt `auth.json` parsed as *no owner*, and no owner is what lets the next `/auth` from a stranger seize the pod. Both stores now write atomically with a `.bak` snapshot, recover from it transparently, and serialize read-modify-writes through a cross-process lock (the bot and the CLI are separate writers). `hasOwner()` now distinguishes *absent* (fresh pod, still claimable) from *unreadable* (fails closed). The dashboard and installer carried their own `loadAuth` that would have overwritten the recoverable copy with an empty roster; both fixed.
- **Surface-agnostic by construction.** A room is `<surface>:<id>` — a Spaces task thread resolves identically to a Telegram group, so a Spaces-only pod needs no chat surface at all.
- Regression coverage: `test-grants.js` (the pair is the unit, the capability ladder, reach vs strength, scope, revocation/expiry/one-turn, replace-not-stack, default-deny) and `test-store-durability.js` (`.bak` recovery, fail-closed vs still-claimable, lock exclusivity and stale reclaim) join the chain.

## v3.3.2 — Known people are recognized by transport id, approved owners stay owners, and a mandate can no longer land on the wrong note

- **Every known person got interrogated on a fresh Kazee DM.** The intro flow matched inbound chats only by `(adapter, channelId)` — but Kazee DMs key the envelope's channelId by the conversation *document* id while people handles store the sender's *user* id, so nothing ever matched: known teammates, and the owner on pods without `KAZEE_OWNER_USER_ID`, all got "who are you" (the Coders Africa incident, 2026-08-19/20). Now a sender whose transport user id already maps to a person is silently **adopted**: the DM channel is linked to their record, an auth entry is written, any stale pending intro is resolved, and `handleInbound` returns false so the router falls through and handles the turn normally. Adoption is gated on adapter-attested `surface: "dm"` — the Kazee adapter sets it only on a positive chat-meta read (`isGroupChat: false`); a failed meta read attests nothing (fail closed) — because auth.json standing is per-CHANNEL and adopting a room would hand every member the sender's standing. Display names are never matched; only the transport id, the same trust level as the env owner match.
- **Approving an owner onto a second channel demoted them there.** `applyApproval` hardcoded `isOwner: false` into the auth entry while `isChatOwner` reads exactly that file. The entry now mirrors `person.isOwner`, and a stale `false` entry from the bug era is upgraded in place on the next approval.
- **The mandate writer can no longer be steered by a display name.** `appendMandateLine`'s legacy fallback upserted by `personName` when the record carried no slug — a spoofable transport display name could land owner-authored security text on a same-named OTHER person's note (reachable from telegram-group externals, where there is no person record). Writes are now strictly slug-keyed: `entitySlug`, else the person REGISTRY (`personId` → `resolveEntitySlug`), else **refused** — never the name. The router's rule-capture path passes `personId` so per-person rules keep landing after the fallback's removal.
- **Stranded grants are surfaced the moment they bite.** The dual-file trap — a mandate granted on the name-obvious note (`rahil-malde.md`) while the enforcer reads the person's hashed record slug — made grants silently do nothing. Escalation cards now carry a ⚠️ hint naming the file holding the dead grant and pointing at Verify / Make-a-rule, whenever the speaker's live mandate is empty but a same-named sibling note (that no other person owns) holds one. Detection only: the stranded text is never applied and never quoted.
- **Router plumbing for the fall-through:** the cross-channel dedup check is memoized per envelope object, so the adopted first message isn't swallowed by `handleText`'s re-check while true redeliveries (fresh envelope, same message id) still dedup globally.
- Regression coverage: `test-intro-adoption.js` joins the chain (adoption unit + through-the-router integration, fail-closed attestation, approval owner-mirror and in-place repair); `test-mandate-hygiene.js` pins the name-fallback refusal and personId resolution; `test-kazee-group-gate.js` pins the dm attestation and its fail-closed absence; `test-enforcer.js` pins the stranded-mandate hint and that the stranded text never reaches the live mandate.

## v3.3.1 — Message length can no longer kill the voice: chunked pocket TTS (re-cut carrying unpublished v3.3.0)

- **Long voice notes died at 60 seconds of synthesis.** `pocketToVoice` sent the ENTIRE cleaned text to the speech gateway as one request under `AbortSignal.timeout(60000)`. A long reply blew that window, the pocket route returned null, and on pods the fallback chain has nowhere to go (no ElevenLabs key, no macOS `say`) — so `say` bounced with "no TTS route available" and agents improvised manual per-paragraph chunking (sales-bot, 2026-08-19). Long text is now sentence-packed into ≤450-char chunks (`chunkForTTS`; run-on sentences hard-split at word boundaries so nothing can recreate the timeout), each chunk synthesized in its own 120s window with one retry (a transient gateway blip no longer wastes the chunks already done), and the wavs concatenated via ffmpeg's concat demuxer into ONE ogg/opus. Short text keeps the exact single-request behavior. Benefits every consumer of `textToVoice`/`textToVoiceFile`: `say`, `tts`, and whole voice-note replies; the voice channel's streamed per-sentence path is untouched. Regression pinned in `test-speech-loopback.js`: multiple gateway calls each inside the pack limit, same voice across chunks, lossless reassembly, temp cleanup, and chunker unit coverage.
- **Why this is 3.3.1 and not 3.3.0:** the v3.3.0 release commit bumped `package.json` but not the release-guard pin (and skipped this changelog), so the tag's build job failed its own `npm test` — npm still served 3.2.3 and no fleet image was adopted. Never retag a failed version — re-cut as +1, pin repaired here.
- **Also ships (committed since 3.2.3):** the v3.3.0 verify-contact batch (below) and the Kazee calls batch — trickle ICE consumption (calls died with both legs stuck ICE-checking), streaming TTS playback (~3s faster to first byte) with per-voice greeting/goodbye PCM cache, hybrid word-guarded turn-end VAD, `call:ringing` on invite with publish deferred until the caller's feed is up, and werift ICE unhandled-rejection suppression.

## v3.3.0 — Never published (build failed on a stale release-guard pin); content ships in v3.3.1

- **An unverified contact stays unverified until the owner says otherwise.** `ensureContact` keeps minting fresh unverified stubs (never link-by-name — names are claimable); the owner gets the explicit act that closes the loop: a "Verify contact…" button on escalation cards and `/people verify [link <slug>]`, both living ONLY on chat-envelope surfaces the agent cannot forge. Verification settles identity, never trust — mandates and relationship upgrades stay separate owner grants; retired stubs rename to `.superseded`, not deleted.
- **Three holes from the same audit, closed and pinned by tests:** `callerIsOwner` trusted the CLI's claimed canonicalUserId (forgeable env) → owner-gated loopback kinds now require a per-boot HMAC attest minted by the runner at spawn; `recent`-fetch let a guarded external turn read other channels' transcripts → default-deny; `relay-send` let a guarded turn message third parties as the bot → a guarded requester may relay only to the owner.

## v3.2.3 — No watch can end silently: every watch carries a timeout that fires as a wake-up

- **The silent-failure hole in watches.** A watch whose check never exits 0 polled forever without a sound — if the watched thing hung, was misconfigured (the unauthenticated-`glab` watch from 2026-08-12 would never have fired), or simply never happened, the agent slept through it and the user had to notice the silence themselves. The working pattern was to pair every watch with a manual backup wakeup; that safety net is now built in.
- **Every new watch gets a hard expiry.** `watch-add` grows `--timeout <duration|ISO>` (90m, 6h, 1d, or an ISO datetime), default **2h**, stamped as `expiresAt` on the job. On the first poll tick at or past the deadline where the check *still* fails, the watch FIRES anyway: the woken prompt opens with `[Watch expired ...]` plus an explicit "the watched condition did NOT happen — this is a deadline wake-up, not a trigger" brief, and the chat advisory reads `⏰ Watch "label" timed out without triggering — waking anyway`. A check that passes exactly on the deadline tick still reports as a normal trigger.
- **Expiry ends the watch's window.** An expiry fire is one-shot even for `--repeat` watches (the woken agent re-arms if still wanted), and the expiry mark is persisted so a failed expiry fire retries with the same semantics across restarts — while a retry whose condition came true in the meantime reports as a real trigger, not an expiry. `watch-list` now shows each watch's expiry (or `EXPIRED (fires on next poll)`); the global `/watch off` master switch still pauses everything, expiry included, until re-enabled.
- **Backward compatible by choice.** Watches persisted before this release carry no `expiresAt` and keep their old poll-forever semantics rather than being retro-expired on upgrade (a pre-upgrade watch older than the default would otherwise fire the instant the bot restarts). `watch-list` flags these as `no expiry (pre-timeout watch)`.
- Regression coverage in `test-provider-scheduler.js`: future-expiry idle stays silent, lapsed-expiry fire (header, brief, advisory, removal despite `--repeat`), deadline-tick pass is a trigger, default and explicit expiry stamping, and static end-to-end wiring (scheduler default → loopback passthrough → CLI flag).

## v3.2.2 — The chip thinks from first contact, keeps thinking through the tools, and the dedup gate actually runs

- **The dead minute before the chip appeared.** The typing heartbeat only started inside `executeAdmittedRun` — but a turn-bound message first crosses the enforcer's inbound triage (an LLM call), and voice notes add a media download plus STT before that. All of it ran with zero presence: on a slow triage the chat sat silent looking ignored, then "thinking" popped up long after the work had begun. The router now starts the heartbeat the moment a turn-bound message arrives (`startPreTurnTyping` in all five inbound handlers — text, voice, audio, photo, document), covering triage, download, and transcription. Non-turn traffic (wizards, vault prompts, credential pastes, commands) never shows a false chip because the start sits after those gates.
- **The chip decayed to a stale "typing" during tool work.** Any `tool_start` permanently cleared `thinkingPhase`, so from the first tool call onward every 4s tick sent a stateless `typing:start` — clients rendered "typing…" through minutes of tool execution, and "thinking" never came back within a turn. Terminal and text events (`text_delta`, `text_final`, `result`, `error`) still end the thinking phase, but `tool_start` now re-enters it: streaming text reads as "typing", tool stretches read as "thinking" — which is what they are. (Kazee mobile renders a stateless tick as "typing" and auto-clears 5s after the last tick, so the 4s cadence keeps the chip alive and the state accurate on old and new clients alike.)
- **Handoff without a scramble.** Two starts can now coexist safely: `startTypingHeartbeat` is a singleton per chat state — a new start supersedes the old interval, and every stopper is ownership-aware, acting only while its own interval is current. The router's pre-turn heartbeat hands off to the run's own heartbeat at `executeAdmittedRun` with no gap and no double interval; a stale stopper (a router `finally` after handoff, a superseded run) is a structural no-op. `/stop`, which clears the interval directly and thereby orphans the run's stopper, now emits the explicit `typing:stop` itself in both branches. `test-typing-heartbeat.js` joins the release chain: immediate thinking tick, supersede/ownership semantics, no orphaned timers across a full tick window, and the event-steering contract pinned.
- **The journal dedup gate never actually ran.** v3.2.0's write-time dedup adjudicator (A1) calls its utility with `purpose: "dedup"` — a purpose `utility-policy` never registered, so every invocation died at validation with `INVALID_UTILITY_PURPOSE` and failed open, silently, by design. Log-proven on the local install. "dedup" is now a first-class purpose riding the review pipeline: `MEMORY_PROVIDER` steers it and `PACK_DEDUP_MODEL` is honored as its model override (both the prefix and the legacy Claude alias, matching what the adjudicator already passes). Pinned in `test-utility-provider-policy.js`.
- **Not changed, on purpose:** the 5× duplicate voice notes from 2026-08-09 were client multi-tap creating *distinct* messages (fixed client-side in kazee-chat-mobile v3.25.23's optimistic-upload guard); the router's existing `isDuplicate` LRU already covers true socket redeliveries by message id, so no bot-side dedup was added.

## v3.2.1 — Speech as a general capability, sockets that survive rejection, and calls that greet back

- **A rejected handshake no longer strands a bot forever.** socket.io-client permanently stops reconnecting when the *server* rejects a handshake (auth middleware `next(new Error())`): `connect_error` fires with `socket.active === false` and the manager never dials again. During the 2026-08-08 outage chat-central booted before Central's MongoDB and rejected every handshake for a few minutes — terminally stranding the whole bot fleet, pods "Running" but deaf until manually restarted. New `core/socket-reconnect.js` `wireTerminalReconnect()` watches for exactly that terminal state and drives a manual `connect()` with 2s→30s capped backoff; transport-level failures (where the manager retries on its own, `socket.active === true`) are left alone. A deliberate client-side `disconnect()` is intercepted — a wedged socket emits no event for it — and cancels retries until `connect()` re-arms them. Wired into both the kazee and spaces sockets; `test-socket-reconnect.js` joins the pretest chain.
- **The bot can now place calls, both directions get a greeting, and the media leg works behind NAT.** Call-platform batch in the kazee channel, live-tested end to end on 2026-08-08:
  - ICE servers now come from chat-central's `GET /webrtc/config` instead of a hardcoded STUN list, filtered to `stun:`/`turn:` URLs without `transport=tcp` (werift's ICE is UDP-only), falling back to defaults if the fetch fails. This fixes the zero-RTP calls where STUN alone couldn't punch through (diagnosed 2026-07-30).
  - `outgoing({ chatId, toUserId })` places a bot-initiated call: the caller side owns room lifecycle, so the session creates the Janus room itself (mirroring the web client's params, tolerating 427 already-exists) and destroys it on close; `call:invite` goes out with a 60s ring timer — FCM's TTL is 60s and chat-central keeps ringing calls fresh for 90s, while the mobile app auto-ends its ring UI at 30s.
  - That 30s device auto-end fires `call:end`, which previously tore down the room while a human was still walking to the phone — a late answer (44s in the live test) then joined a destroyed room and hung on "Connecting". `ended()` now ignores end events while the call is still ringing; the ring timer and `call:rejected` own that phase.
  - Callers heard silence until they spoke first and read the call as dead. The session now speaks a greeting exactly once, when the first remote feed comes up — both directions ("Hey, it's Claudia…").
  - Double-subscribe race fixed: the join response and the async publishers event both list the same feed, and while `attach()` was in flight the second frame passed the `has()` check and subscribed the feed twice, orphaning a PeerConnection (diagnosed 2026-07-30). The slot is now reserved before the first await.
  - `call:accepted`/`call:rejected` are handled, and both call-test harnesses grew coverage (+45/+47 lines) for the outgoing flow, ring timeout, ignore-while-ringing, and greeting.
- **Received videos and mobile uploads no longer save as `.bin`.** Telegram messages without a fileName (videos especially) and Kazee Media docs from mobile uploads (which often carry no fileName at all) fell through to the `.bin` fallback, so an mp4 landed on disk unplayable. Extension derivation is now tiered: fileName ext → transport hint (Telegram's `file_path`, which carries the real extension; Kazee's blob-URL pathname) → mimeType map (16 common types) → per-media-type default (video→.mp4, audio→.m4a/.ogg) → `.bin` only as the true last resort. Kazee documents that keep their original name get the derived extension appended when the name has none.
- **TTS and STT become first-class agent capabilities, not just the voice-note reply path.** Fleet agents asked to produce a voice-over or transcribe a sent recording had nothing to reach for: they probed `open-claudia say --help` (which would have *spoken* "--help"), then tried direct gateway/vendor API calls and concluded "the key is missing" — correctly, because the provider subprocess is credential-stripped by design. Two new loopback verbs run in the credentialed bot process: `tts` (`open-claudia tts "<text>" --out <path> [--voice <id>]` — synthesize to a FILE without any chat send; extension picks the container, .ogg native / .mp3 / .wav transcoded; voice override reaches the gateway for cloned voices) and `transcribe` (`open-claudia transcribe <file>` — prints the transcript of any audio OR video file; `transcribeMedia` ffmpeg-normalizes to 16kHz mono wav first, which extracts video audio tracks and shrinks the gateway upload, then falls through the existing gateway → local-Whisper chain). No guardrail on `tts` because the artifact stays on local disk — the send verbs that could ship it carry their own R8 guard. A new "Speech services (TTS / STT)" system-prompt block — gated on real capability (`canSpeak()` / new `canTranscribe()`) — documents both plus `say`, and tells agents explicitly to never call the gateway or vendor APIs directly. `say`/`tts`/`transcribe` all answer `--help` now instead of synthesizing the flag. Regression chain: `test-speech-loopback.js` pins the loopback contract (empty text 400, no route 503, synth failure 500, relative/missing transcribe paths 400, empty transcript is a valid 200, auth 401) and the media mechanics (copy vs transcode by extension, mkdir -p, voice override in the gateway body, temp-wav cleanup, prompt-block presence). Verified live: a real gateway round-trip synthesized a sentence to mp3 and transcribed it back byte-identical.

## v3.2.0 — The employee batch: memory hygiene at write time, prospective memory, earned trust, and learning from the owner's judgment calls

Three approved plans land together — memory adoption A1-A4 (from the TencentDB-Agent-Memory review), codegraph Phase 1, and employee capabilities C1-C6 — plus the Kazee group speaker-identity fix. Two rules hold throughout: **report-only where trust is involved** (nothing here grants, acts, or loosens an approval rule by itself), and **no new standing context cost** (everything injects only when already-matched content fires, or lands in the nightly dream summary).

- **A1 — write-time dedup adjudicator, gated (the only hot-path item).** The post-turn reviewer's journal facts used to append unconditionally; duplicates accumulated all day for the one nightly pass that can lose things (merge) to clean up. Now cheap extractive recall checks each proposed fact against existing journal lines first. The gate keeps the hot path free: no candidates → zero LLM cost, store as proposed; exact-normalized duplicates auto-skip mechanically; only the fuzzy remainder goes to ONE batched low-tier call returning per-fact `store | update | merge | skip`. Fail-open everywhere — any error, timeout, or malformed judgement stores facts as proposed, because dedup failure must never make memory worse than append-only.
- **A2 — merges keep the absorbed packs' history.** Dream merges used to delete the absorbed pack's entire Journal (surviving only in the backup dir); a merged pack read younger and thinner than the history it represents. Absorbed Journal lines now carry into the survivor — normalized-exact dedup only (never fuzzy: similar-but-distinct events must both survive), date-ascending, the cap keeps newest and drops undated lines first.
- **A3 — substitutability stamps.** The reviewer self-rates each journal line 0-10 for how fully its wording substitutes for the turn's original words — a checkable claim about the text, not a guess about future importance. The score rides the stored line as a trailing `⟨s:N⟩` marker (announcements stay clean, near-dup similarity ignores it) and protects low-scored, verbatim-precious lines from the dream's journal collapse.
- **A4 — section budgets are enforced, not requested.** "Keep State under ~150 words" was prose; States drifted. LLM-edited sections now have mechanical ceilings: pack State trims at the cap with a visible cut marker + flag, entity Notes and persona sections trim at theirs, and the review prompt measures State against the target so over-budget turns the reviewer's job into an edit, not an append. Mandate (owner-set policy) and user-origin writes are never trimmed.
- **codegraph — code-structure queries as a reusable read-only tool (Phase 1).** Machine-side `codegraph` tool wrapping `@colbymchenry/codegraph`, pinned exactly 1.5.0 (native binaries from a solo author — upgrades are deliberate, reviewed events, never `^` drift), `CODEGRAPH_TELEMETRY=0` baked into every invocation, 12 read-only verbs with query verbs auto-syncing the index first. TOOL.md carries the graded trust model: `explore`/`node`/`search`/`files` are high-trust retrieval; `callers`/`impact`/`affected` are **advisory, positive-signal only** (verified misses: CJS namespace calls, some TS class-method sites) — grep remains the completeness check for destructive refactors. Two-week usage evaluation follows; MCP exposure was dropped against the token-efficiency goal.
- **C1 — prospective memory.** "Next time X comes up, mention Y" — reminders bound to a TOPIC (pack or entity), not a time; crons and watches already own time and condition. They fire as a `📌 Pending:` line when the target surfaces in recall injection or is opened with `pack show`/`entity show`, and are consumed ONLY at commit points — a dropped render can never eat a reminder. One-shot by default with a fired ledger; `--sticky` survives; the dream GCs expired and orphaned entries. New `open-claudia remind` CLI.
- **C2 — recall misses become permanent tests.** `open-claudia recall-case add "<query>" --expect <slug>` saves a case whenever a miss is corrected; the dream re-grades the suite nightly against the deterministic recall floor (matchPacks + matchEntities — the candidate stage every engine builds on) so the suite runs, and means the same thing, on nights the providers are down. Capped with oldest-first rotation; orphaned cases drop; pass-rate + delta land in the dream's memory-health section. "It feels degraded" is now a number.
- **C3 — skill certification: trust is earned per verb, and granting stays human.** Tools already log per-verb runs in state.json; a certification tier is now derived from that telemetry — `certified` at ≥5 clean runs with no recent failure, `probation` for 7 days after a failure (only clean runs SINCE the failure re-certify, read from history), `learning` otherwise. `tool show` renders the tier per verb, labeled informational — nothing in the run gate consults it. The dream proposes at most 2 certified gated-tier verbs per night as Always-allow candidates; read-only verbs need no grant, and no auto-grant exists, ever.
- **C4 — apprenticeship radar: it asks to be taught.** A deterministic 7-day transcript scan spots repeated manual work — the same ad-hoc command re-derived by the agent (inline code + `$`-lines; its own CLI excluded) or the owner asking for the same thing again — at ≥3 occurrences across ≥2 distinct days. The dream's introspection report offers at most 2 as `/learn` candidates with the counts as evidence. Report-only: `/learn` stays owner-triggered, and no evidence means no section, no standing cost.
- **C5 — denied-approval mining: trained by the owner's judgment calls.** Every Deny tap was training data we dropped. Recently DENIED approval records (never approvals — an approve confirms, it corrects nothing) now feed the post-turn reviewer as fenced DATA with explicit never-instructions framing, on owner-turn reviews only — deny content can never surface via an external speaker's turn. Each deny is fed exactly once (stamped after a parseable decision; parse failure re-feeds). Forge-proof by structure: a deny-sourced lesson's trigger must cite a real record with status denied — the model cannot invent a refusal, and an approved id confers nothing.
- **C6 — weekly cost self-report.** The dream summary gains a cost section: rolling 7-day totals over the existing per-turn usage log, top burns by project, week-over-week delta, and ONE concrete suggested cut picked by deterministic heuristics. Pure arithmetic — no model call, no new logging surface.
- **Kazee groups: the Speaker block now knows the env-configured owner.** The prompt's Speaker block resolved identity ONLY through `people.findByHandle`, so an env-configured owner (`KAZEE_OWNER_USER_ID` — no people record, because auto-contact skips owners by design) and any not-yet-recorded group member spoke anonymously, and in group chats the model refused to attribute their messages (the "Ok go" incident, 2026-08-03). The block now falls back to the canonical owner check and the transport-reported display name — without a name ever conferring owner standing (R1 intact).
- **Eleven new release-chain tests** (chain now 67): journal union, substitutability stamps, section budgets, dedup adjudicator, prospective memory, recall regressions, cost report, speaker context, approval mining, tool certification, apprenticeship radar. The 360° dangling-pathway pass found and closed one real gap pre-release: the C5 deny-feed originally fired on ALL reviewed turns, including external speakers' — now owner-gated with tests proving external turns see nothing and records stay un-mined.

## v3.1.28 — The Telegram reboot loop was self-inflicted: the heal kept tearing down its own recovery

- **Symptom.** v3.1.27 shipped to all four Telegram pods and the loop continued (euvy: 6 restarts in 3h). Its fix was real but addressed a different failure — the dominant one never enters `_enterOutagePause` at all (zero "Telegram unreachable" lines in the crash logs). The decisive observation was that the bot reproduced it identically on a laptop — 155 self-exits in a local `bot.log` — which retired the cluster-network and conntrack theories outright. The shape is bursts: healthy for 2.7–4.2h, then ~10 minutes of failed heals, then `process.exit(1)`, repeat.
- **How it was found.** Reasoning from prod logs had already produced one wrong fix, so this time the local install was instrumented directly (unthrottled lifecycle events to a side file) and a real burst was captured. Two facts invisible to the console log settled it: every polling error after the first carried `_polling._abort === true` — meaning *we* aborted it, not Telegram — and `_lastUpdate` read 0 for 200+ seconds straight across 9 successful restarts. It also retired a number the previous investigation had reported as fact: the "~62s failure cadence" was an artifact of the 1/min log throttle. The real rate was one every ~15s.
- **Root cause — three stacked deadlocks, each independently fatal to recovery.** (1) `_heal` → `_stopPollingClean` → `_agent.destroy()` kills the in-flight `getUpdates`, which surfaces in `_onPollingHiccup` ~1ms later as `EFATAL: socket hang up`. `_healTimer` had just been nulled at the top of `_heal`, so the "a heal is already scheduled" guard didn't hold and the teardown scheduled *another* heal — which tore the socket down again. Self-sustaining. (2) The backoff capped at 15s while `polling.params.timeout` is 30s, so every restarted poll was aborted at the halfway mark: `_lastUpdate`, the only evidence of recovery, could never advance. (3) `_armHealthyCheck` cleared its own pending timer on every re-arm, so a heal every 15s starved a check that needs 35s of quiet — **zero healthy checks fired in the entire captured burst** — and that callback owns the only reset of `_wedgedSince` and `_healBackoff`. With nothing able to clear the wedge clock, it ran monotonically to the 10-minute backstop. One ordinary `ECONNRESET` was enough to cost ten minutes and a reboot.
- **Fix.** (1) A time-boxed self-abort window opened in `_stopPollingClean` makes `_onPollingHiccup` ignore the error our own teardown raises; it is never explicitly cleared, so a path that forgets to close it cannot permanently deafen the handler — the worst case is a dropped error that `_armHealthyCheck` re-arms 35s later anyway. (2) The heal ladder (now `_scheduleHeal`) escalates to 45s, clearing the 40s request timeout that bounds one `getUpdates`, so the late rungs give a poll genuine room to land; early rungs stay short because they run while polling is down. (3) Re-arming the healthy check no longer pushes its deadline back — the first arm always fires — while the baseline tracks the newest restart, so a poll that completed *before* the latest restart can't be mistaken for proof the latest restart worked. A failed `startPolling` now schedules its own retry rather than waiting for a hiccup that the suppression window means will never come.
- **Proof, not assertion.** `testOneNetworkBlipDoesNotSpiralIntoTheWedgeExitLoop` models the poller instead of stubbing it: `destroy()` raises the teardown error for real, a healthy poll takes the full 30s hold to complete, the library re-dials ~300ms after a fault (which is why a heal always had something to destroy), and a virtual clock runs a single `ECONNRESET` well past the 10-minute budget. Against unmodified code it reproduces the production failure exactly — a storm of `socket hang up`, then "Polling wedged >10min and Telegram is reachable — exiting for a clean restart" — and against the fix the bot recovers with the wedge cleared and the ladder reset. Three narrower tests pin each deadlock individually. The previous release's test passed against a bug it did not fix because it drove `_enterOutagePause()` directly, a path production never takes; the lesson — model the interaction, don't stub past it — is why this harness raises the error from `destroy()`.
- **Correction to v3.1.27's closing note.** That entry wondered whether the wedge-exit backstop earned its keep, observing that "210 restarts never helped". They did: a fresh process is the only thing that escapes the loop, which is exactly why each reboot bought 2.7–4.2h of health. The backstop stays until this fix is proven in the fleet.

## v3.1.27 — The outage pause stops feeding the wedge-exit clock (the fleet-wide Telegram reboot loop)

- **Symptom.** Every Telegram-channel bot was rebooting itself in a loop while no other bot was: rashmi 214 restarts in 7d15h, euvy 210 in 7d17h, ruk-ai 62 in 44h, rahil-coders 11 — 4 of 4 pods running the `telegram` channel. All 12 pods running only `kazee, spaces, web` sat at 0. All 16 share one node and one egress IP, so node, address and NetworkPolicy preset were all controlled for; the *unrestricted* bot (euvy, allow-all egress) restarted most, which retired the egress theory rather than confirming it. The channel was the only discriminator.
- **Root cause.** `_wedgedSince` is stamped on the first polling hiccup and cleared in exactly one place — `_armHealthyCheck`, on a completed `getUpdates`. `_enterOutagePause()` cleared `_healthyTimer` (the only thing that can reach that reset) and never touched `_wedgedSince`. So every second spent inside the outage pause — a state whose entire documented purpose is *"a reboot can't fix a dead network... never exits"* — still counted toward the 10-minute reboot budget. A link that flapped a few times inside ten minutes rebooted the pod anyway, moments after this very method had twice decided a reboot could not help. Both bots exited almost exactly 10 minutes after their FIRST hiccup regardless of what happened in between: ruk-ai 07:41:02 → 07:51:11 = 10m09s across three outage pauses; euvy 07:33:37 → 07:43:42 = 10m05s.
- **Fix.** `_enterOutagePause()` now retires the wedge clock (`_wedgedSince = 0`, `_healBackoff = 0`). A failed *fresh* probe is positive evidence that the streak so far was the network, not an in-loop poll wedge — and the exit backstop exists only for the in-loop case. Clearing is safe because `_onPollingHiccup` no-ops while paused: a real in-loop wedge that survives the resume starts an honest new streak.
- **Proof, not assertion.** New regression test `testOutagePauseDoesNotFeedTheWedgeExitClock` in `test-telegram-poll-recovery.js` replays the production sequence line for line (hiccup → three genuine outage pauses with successful re-probes → ruk-ai's exact 10m09s mark on a reachable link). It **fails against unmodified code** with the reboot firing, and passes with the fix. `testWedgeExitFiresWhenIdle` still passes, so the genuine in-loop backstop is intact, not disabled.
- **Still open, deliberately not fixed here.** Why the path to `api.telegram.org` flaps at all is unexplained — the probe genuinely fails, three confirmed ~30s outages inside ten minutes on ruk-ai. This change stops the pointless reboots; it does not restore connectivity during a flap. Separately worth asking whether the wedge-exit backstop earns its keep at all: `_heal()` already destroys the agent, nulls `bot._polling` and dials a fresh socket — everything a new process does for the poll path — and 210 restarts never helped.

## v3.1.26 — Claude Opus 5 in the model picker

- **`claude-opus-5` added to the Claude provider model list** (`CLAUDE_MODELS`), slotted between `claude-fable-5` and `claude-opus-4-8`. Opus 5 launched 2026-07-24 (near-Fable-5 capability at roughly half the price; Anthropic's new default on Max). Verified live before listing: a real CLI round-trip with `--model claude-opus-5` succeeded. Free-text `/model claude-opus-5` already worked (the `/model` handler doesn't validate against the list); this puts it on the picker buttons for `/model` and `/login`. Default model unchanged (`claude-opus-4-8`).

## v3.1.25 — Permissions hardening: audit fixes, the 360° dangling-pathway sweep, and capability-scoped approvals

- **Six audit fixes (full ledger: `docs/permissions-audit-2026-07.md`).** (F1) the enforcer judge gets a 60s timeout (`ENFORCER_TIMEOUT_MS`) + one retry on transient transport errors — cold-start overruns no longer fail legitimate replies closed into spurious owner cards. (F2, critical) only Bash was hooked: an external speaker could have the agent write files via provider Write/Edit tools, bypassing the approval path entirely — a second PreToolUse matcher now routes `Write|Edit|MultiEdit|NotebookEdit` to the same speaker-gated deny. (F3) approval cards are person-first: the owner authorizes a PERSON's ask, not a wall of bot output. (F4) up-front triage: inbound is judged BEFORE any work happens; held turns generate nothing. (F5) turn grant: an approval id rides scheduler job → AsyncLocalStorage (the agent subprocess can never forge it), honored in reply vets, validated fail-closed — kills the double-approval fatigue that trains blanket-allow habits. (F6) mandate writes are slug-keyed, whitespace-deduped, and capped at 4000 chars by REFUSAL, never truncation — it's security text.
- **The 360° sweep: guarded externals are policy INPUTS, never policy writers.** (D1, critical) pack-review and dream applied model-proposed entity edits wholesale — a poisoned inbound could rewrite a person's Mandate or relationship through the background reviewer; policy fields are now stripped from every model-output applier, with attack fixtures pinning the strip. (D2) media-only turns skipped triage — every entry lane now gates. (D3) dream entity merges skip any slug carrying a Mandate or people-record pointer. (D4) an external speaker's shell may not reference the config dir, package root, or any `.open-claudia` path — reads included, because lessons/entities/crons ARE the policy. (D5) live-probed codex-cli 0.144.0: file edits arrive as `apply_patch`, so the codex hook matcher widens to cover it. (D6) every mandate writer funnels through one `appendMandateLine`. (D7) the keyring CLI is owner-only on external turns.
- **Capability-scoped approvals (P3): an Always-allow is a bounded grant, not a blanket.** New `core/mandates.js` gives mandate lines the same lifecycle global/channel rules already had: stable 8-char slug-scoped ids (the same capability on two people revokes independently); a trailing ` [until YYYY-MM-DD]` expiry — tap-grants default to `MANDATE_GRANT_TTL_DAYS` (30d, 0 = permanent), owner-authored rule cards stay permanent unless the text ends "for Nd" / "until YYYY-MM-DD". Expiry filtering is deterministic and read-side at the single `speakerFor` choke point: the judge never sees a lapsed capability and never does date math, and the verdict cache busts naturally. Re-granting the same capability REFRESHES its expiry instead of stacking; successful writes prune expired bullets; refused/duplicate writes mutate nothing. `/guardrails` lists per-person grants alongside global + channel rules and `revoke <id>` resolves across all three layers; approval-card acks quote the grant id + the revoke command.
- **The `apr:` button handler finally has a harness (closes audit gap G3).** New release-chain test `test-approval-cards.js` pins the owner's authorization surface end-to-end with the real approvals + entities + mandates modules and stubbed transport: Always-allow delivery + scoped grant + revoke-command ack, double-tap idempotence, deny holds with nothing sent and nothing written, non-owner taps bounce, triage approve wakes the agent in the origin channel with the approval id as a one-off turn grant (request framed strictly as data), triage deny declines politely, make-a-rule offers person/channel/global scopes, stale cards ack cleanly.
- **Also ships:** wakeup/cron model pinning + guard-timeout plumbing in the scheduler and schedule CLI; two stale tests caught and fixed (the release-guard version pin, and `test-persona-pipeline.js` still asserting the pre-D1 behavior that reviewer/dream may write relationship/mandate — rewritten to pin the strip, plus an owner-grant-survives-model-overwrite case).

## v3.1.24 — Catalog-first recall: the walker sees the whole index, nothing is cut before a model judges it

- **Root cause (proven, `diag-recall-ranking.js`).** Recall misses were 4-layer candidate starvation, not bad judging: (1) lexical pack scores count distinct matched terms, so vocab-dense packs outbid the true pack; (2) the graph fire-threshold scales off the MAX seed score, so junk seeds suppress true edges; (3) per-node top-K fan-out on hub nodes cuts true edges before they transmit; (4) killer — the dream tuned `walkerMaxCandidates` down to 10 on 2026-07-09 while seeds are 6 pack + 4 entity = exactly 10 and assembly orders all seeds before all graph-activated nodes, so **graph-activated packs/entities got zero walker slots on every full-seeded turn since**. The walker judged well; it just never saw the right nodes.
- **Catalog-first routing.** The walker's SYSTEM prompt now carries the ENTIRE memory index — every pack, entity, and tool as a one-liner (~50KB for 276 nodes, `buildCatalog()`, 60s TTL cache). Set once per warm-walker session (exempt from the warm per-message char budget) and identical on the cold path. Lexical seeds + graph activation are demoted to per-turn RANKING HINTS: the same top-K candidate assembly, but now it only decides which few nodes carry live excerpts — never membership. The walker may keep ANY catalog id, hinted or not; zero-hint turns still walk. Kept catalog picks inject exactly like hinted ones (packs render extractive lines, un-hinted tools resolve through the registry).
- **Simpler under catalog mode.** The history cascade is bypassed (an auto-drop would cut on history alone — never-cut principle) and tool pack-riding is removed (every tool is already in view). Both return when the `catalog` knob is off — new dream-tunable enum (`RECALL_CATALOG` env pin) that restores the legacy capped-menu contract as a one-flip rollback.
- **Honest observability.** `recall query` now reports `catalog` counts + `hintedActivated`, and the CLI banner says "N kept of catalog of 276 nodes (K hinted)" with "graph-activated in hints: X of Y" — the old "+N graph-activated" counted raw expansion and overstated what actually survived into the walker's view (measured: it said +9 while 0 survived). Unjudged fail-open output no longer claims catalog scope. Per-turn metrics log catalog counts, never the full id list.
- **Validated on the two proven misses.** Sim A (update-app button) now keeps the in-app-updater pack at 85 `[via catalog]` — a node the old pipeline structurally could not surface; sim B (group permission system) keeps the governing pack at 88 as the sole result. Corollary: catalog one-liners make pack descriptions load-bearing for recall — the nightly dream already tightens them.
- New release-chain test `test-recall-catalog.js` pins the contract: full-index build + TTL cache + stub-safe null, catalog in system prompt (never the user prompt), un-hinted pack/entity/tool keeps injecting end-to-end, score-gate on catalog picks, knob-off legacy filtering, and query provenance (`via: "catalog"`).

## v3.1.23 — Scoped external rules (person / channel / everyone), auto-contacts, meet-greet lane

- **Make-a-rule now offers three scopes.** The owner's guardrail card gained a middle option: alongside "just this person" and "everyone", rules can now bind to **this channel** — stored in `channel-rules.md` keyed generically on `adapterType:channelId`, so every surface (kazee groups, Spaces tasks, telegram chats, future public domains) participates simply by existing. No adapter-specific policy logic anywhere; caps at 20 rules × 300 chars per channel.
- **The enforcer judges the union of all layers.** Verdicts now weigh the hardcoded security floor → global guardrails → channel rules → per-person mandate: an action is permitted if ANY owner-authored layer clearly allows it, a prohibition in ANY layer wins, and anything uncovered stays default-deny. The verdict cache keys over every layer (plus the unverified flag), so editing any one of them re-judges previously cached payloads.
- **Auto-contacts.** The first time an authorized speaker with no people record talks to the bot, it creates a fresh contact (tagged `autoCreated`, relationship `external`) so the meet-greet lane and per-person rules have somewhere to live. Security invariants: never linked by display name (spoofable — a name match must not inherit someone's mandate); entity slugs are made unique per (adapter, userId) so two same-named people can never share a mandate; room-keyed envelopes (telegram groups: `userId` = the room, `from.id` = the sender) are refused, so a group can never mint one contact that misattributes everyone; owner channels resolve through canonical identity and are never duplicated.
- **Meet-greet lane.** A brand-new unverified contact gets warm introductions without needing explicit rule coverage — greetings, who they are, what they do, how they know the owner, what they need — while information disclosure, commitments, and actions stay held for escalation. The system prompt steers the same lane, and the background reviewer folds what's learned into the contact's entity as usual.
- **Cards say where it happened.** Escalation cards now read "From Lucy Osweta in Dev Team": adapters may implement an optional `channelLabel(channelId)` (kazee: chat name via meta cache; spaces: task title via notification/thread cache with API fallback), carried on the approval record (clipped at 120 chars) and shown on owner buttons.
- **Slug-keyed entity writes.** `upsertEntity` accepts an explicit slug so record-held slugs (auto-contacts, person-scope rule writes) land on exactly their entity — name-keyed upserts could hit the alphabetically-first same-named entity and leak a mandate across people.
- `/guardrails` lists global and per-channel rules and removes either by id; new release-chain test `test-external-rule-scopes.js` pins the store, the auto-contact invariants (impostor isolation, slug-keyed mandate writes, room-keyed refusal), the three-layer enforcer prompt with UNVERIFIED marker, cache busting on channel-rule edits, and no cross-channel rule leakage.

## v3.1.22 — Group silence is no longer context-blind: engaged turns catch up from the cursor

- Follow-through on the v3.1.21 group gate, per the owner's direction: the bot must not reply in a group unless tagged (owner included) — "but should get messages since its last cursor."
- The kazee adapter now implements the router's `buildThreadContext` hook (the same contract Spaces uses): an engaged group turn opens with a context block of every message since the bot's last engaged turn — exactly the messages the participation gate stayed silent through — with sender attribution ("You" for its own), `[type] caption` rendering for media, deleted messages dropped, bodies clipped at 280 chars, at most 40 messages.
- The cursor is a persisted per-chat watermark (`kazee-state.json`, atomic writes) holding the newest `createdAt` already folded into a turn. The delta is server-derived (`GET /messages/:chatId/:botUserId?since=`, oldest-first, server-capped at 500), so missed socket events and restarts self-heal. A chat's first-ever engage pulls a recent 15-message page instead of the whole backlog.
- Best-effort: a fetch failure returns no block, leaves the watermark alone, and never delays the reply.
- New release-chain test `test-kazee-group-context.js`: cold-start window, warm since-delta, You-attribution, media/deletion/clipping rendering, trigger exclusion, trigger-only watermark advance, failure semantics, DM no-op, on-disk persistence.

## v3.1.21 — Kazee public participation: owner seeding, group tag-gating, guarded external turns

- **Owner seeding for kazee-only pods.** A pod whose only owner signal is `KAZEE_OWNER_USER_ID` (empty `auth.json`) never got a people-store owner record, so every unknown inbound — group or DM — was refused with a false "no owner configured" (observed on network-mon, 2026-07-23). `seedOwnerFromLegacy` now falls back to the kazee env owner; legacy `auth.json` owners keep precedence and seeding stays idempotent.
- **Group chats engage only when addressed.** The kazee adapter fetches chat meta (5-min cache) and, in group chats, engages only on an @-mention of the bot or a reply to the bot's own message — matching chat-central's own `assistantAction` default `assistant_tagged_reply`. Chats set to `assistant_free_reply` engage on everything. Untagged group chatter is dropped silently (no roster prompts, no refusals — owner included). 1:1 chats are unchanged; an unresolved metadata fetch is not cached and fails closed because the room cannot safely be classified as a DM.
- **Tagged externals get the guarded turn.** An engaged group message authorizes the turn via adapter-attested `envelope.engaged` (never derived from message content), but the speaker still classifies external: replies are vetted by the enforcer and escalate to the owner's approve-once / deny / make-a-rule card — the same pattern as Spaces. Slash commands stay closed to externals.
- **Escalation cards reach the kazee owner.** `ownerTarget` used to return env-seeded kazee owner handles as the raw user id, but kazee sends need a chat-document id, so approval cards would 404. User-id handles now resolve through `openDirectChat`, falling back to the configured owner target.
- New release-chain tests: `test-people-owner-seed.js` (env seed, legacy precedence, no-signal refusal) and `test-kazee-group-gate.js` (mention/reply engagement, free-reply, 1:1, meta failure, meta cache, engaged access pass-through).

## v3.1.17 — Call bridge joins through the caller's room-setup race (fixes "Call declined")

- Answering a Kazee web call could fail deterministically fast: the caller creates the Janus videoroom during its own dial-time setup, and the bot answers quicker than any human — its `join` hit Janus before the room existed, got 426 "No such room", and the whole call was rejected as `connection_failed` (surfacing as "Call declined" in the web UI).
- `session.start()` now polls the publisher join through transient 426 (10 attempts × 500ms), mirroring the web receiver's own retry loop in `CallContext.startVideoRoom`. Persistent 426 still fails the call.
- Janus error frames now carry `error.code` alongside the message for future discrimination.
- New assertions: join retries through transient no-such-room and still gives up when the room never appears.

## v3.1.16 — CI-deflake: no live network in unit tests, CI-safe enforcer deadlines

- Identical bot behaviour to v3.1.15 (which never shipped: its pipeline died twice on loaded-runner test flakes — the tag is retired unretagged per release hygiene).
- `test-telegram-conflict-heal-guard.js` stubbed `_probeTelegramReachable`/`_anyTurnActive`: the wedge-exit assertion used to probe the real api.telegram.org, so a slow/blocked CI egress sent `_heal` down the outage-pause path instead of the exit under test. It also now isolates `OPEN_CLAUDIA_CONFIG_DIR` — the exercised exit path was writing a bogus "polling wedged" `last-exit.json` into the real `~/.open-claudia` on every local test run.
- `test-provider-enforcer.js` deadlines made CI-safe: the judged-probe service backstop 2000→8000ms (children answer in ms; 2000 was missed by child boot under runner load), the deliberate-timeout probe 2500→5000ms boot headroom.

## v3.1.15 — Voice-note replies speak a conversational summary, not the chat text

- Replying to a voice note used to TTS the full formatted chat answer verbatim — bullets, emoji, pipeline IDs and all ("this isn't human hearable"). On voice-reply turns the provider is now asked to end its reply with a dedicated spoken script in `<voice-script>` markers: the chat message keeps the full text answer (script stripped), and only the script — two to four conversational sentences written for the ear — is synthesized into the voice note.
- If the provider omits the script, a sanitizing fallback (`speechFallback`) strips code blocks, HTML, URLs, emoji, and list markers, caps the speech at the opening sentences, and closes with "The rest is in the text message." — a verbatim readout can no longer reach TTS on text channels.
- The Kazee call bridge keeps its existing in-call guidance and streams `voiceScript || finalText`, so stray markers are never spoken there either.
- New extraction/fallback/wiring assertions in `test-media-speech.js`.

## v3.1.14 — Call bridge production hardening: serialized speech, barge-in, group attribution, limits

- **TTS is serialized through a queue with cancellation.** Concurrent `speak()` calls used to interleave two 20ms-frame RTP loops over the same sequence/timestamp counters — garbled audio whenever the provider emitted more than one message. Utterances now play strictly one at a time; `cancelSpeech()` (generation-based) stops the current utterance mid-frame and resolves everything queued, and `close()` cancels so no pending `speak()` can hang the adapter.
- **Barge-in + echo gate.** While the bot is speaking, the VAD threshold rises from −38 to −26 dBFS so the caller's device echoing the bot's own TTS can't trigger STT; genuine speech over the gate cancels playback immediately — you can interrupt her mid-answer like a real call.
- **Group calls get per-feed pipelines and speaker attribution.** Each subscribed feed now has its own Opus decoder (they're stateful — one shared decoder produced garbage on 2+ feeds), its own VAD, and its own STT stream. Transcripts carry `{feedId, display, participants}`; with more than one participant in the room the provider sees `alice@…: text`, so group calls stop reading as one merged anonymous voice. STT send/commit failures now log and drop the frame instead of killing the whole call (the stream self-heals on the next send).
- **Calls end on their own when they should.** 30-minute duration cap (`KAZEE_CALL_MAX_MS`) and 2-minute idle timeout (`KAZEE_CALL_IDLE_MS`), both announced with a spoken goodbye; if the caller barges in during the idle goodbye the call stays up. A leaked session can no longer hold Janus/STT resources forever.
- **First syllables reach STT.** A 240ms pre-onset ring buffer flushes into the STT stream when a turn starts, so VAD onset latency stops clipping the start of every utterance.
- The five existing call tests plus new `test-kazee-call-session.js` (serialization, cancellation, echo gate, barge-in, prebuffer, attribution, idle/max limits, close semantics) are wired into the main `npm test` chain — call regressions now block every release, not just `test:kazee-call` runs.

## v3.1.13 — Owner-only Kazee calls reach Open Claudia

- Accepts incoming audio calls from the configured Kazee owner only; external, video, malformed, and concurrent calls are rejected before any media session is created.
- Joins the existing Janus VideoRoom as a headless publisher/subscriber using pure-TypeScript WebRTC, decodes inbound Opus, commits VAD-delimited 16 kHz PCM to hosted streaming STT, and routes final utterances through the owner's normal Kazee identity and conversation.
- Speaks only the settled, guardrail-vetted answer through hosted Pocket TTS PCM and Opus RTP. Provider narration and transient progress previews are suppressed in calls; audio is never persisted.
- Adds deterministic audio, Janus transaction, WebRTC RTP, speech framing, owner gate, lifecycle, and AI-routing tests. Updates `ws` to the patched 8.21 line across direct and Socket.IO use.

## v3.1.10 — `say` synthesizes in the bot process, not the stripped subprocess

- v3.1.9 carried this exact fix but its build never published: the version pin in `test-provider-language.js` still read `3.1.8`, so the guard failed and the docker stage was skipped. Re-cut as v3.1.10 per "never retag a failed version."
- **Root cause of the persistent "no voice" reports.** `open-claudia say` synthesized the audio *inside the CLI*, which runs in the credential-stripped provider subprocess — with no `AGENTSPACE_POD_TOKEN`. `pocketToVoice` had no speech bearer and returned null, so hosted pods told users they couldn't send voice (the rahil-coders and kazee-people reports), even though the parent bot process — which *does* hold the token — evaluated `canSpeak() === true` and advertised "Voice notes: enabled." The capability check and the actual send were running on opposite sides of the credential boundary.
- **Fix.** Synthesis moves into the credentialed bot process behind a new loopback `say` JSON kind. The CLI now POSTs only text; `handleJson` runs `textToVoice` in-process, delivers via `adapter.sendVoice`, and cleans up the temp ogg afterward. A restart alone won't adopt this — it ships in the image and every pod is rolled onto it.
- **Egress guardrail preserved.** The old `say` inherited the R8 external-contact voice-egress default-deny from `handleSend`; the new in-process handler re-applies the same `relationship.guardActive` check, so `say` cannot become a guard bypass to non-owner contacts. Returns 400 (empty text) / 403 (guarded target) / 503 (no TTS route) / 500 (synthesis or delivery failure).
- Regression coverage in `test-say-loopback.js`: success (synthesis + delivery + temp cleanup), empty text, no-route, delivery failure, the R8 guard, and loopback auth.

## v3.1.8 — Hosted pods can actually send voice notes

- **FFMPEG is discovered on PATH.** Provisioned pods ship `/usr/bin/ffmpeg` but never set `FFMPEG=` in config, and `config.js` read only the explicit setting — so `FFMPEG` resolved empty and every TTS route (`pocketToVoice`, ElevenLabs, `say`) silently returned null: voice input got text-only replies. `FFMPEG` now falls back to PATH discovery via the existing `executableCandidate` helper; an explicit `FFMPEG=` still wins.
- **The system prompt's voice-notes flag reflects real TTS capability.** It gated on `WHISPER_CLI && FFMPEG` — the pre-gateway local-Whisper world — so whisper-less hosted pods told the model "Voice notes: disabled" and it refused user requests (the smart-euvy report), even with working gateway speech. New `media.canSpeak()` mirrors the actual pocket → ElevenLabs → macOS `say` fallback chain and now drives the flag; local Whisper is irrelevant to speaking.
- **New `open-claudia say "<text>"` CLI verb.** `send-voice` needs a ready-made ogg and the model had no way to synthesize one on request. `say` runs `textToVoice` and sends the result through the loopback voice path (same egress guardrail), and is documented in the Delivery prompt block.
- Regression coverage in `test-media-speech.js`: a whisper-less pod shape must report `canSpeak() === true`, the prompt gate must key on it, and FFMPEG resolution precedence (explicit > PATH discovery > empty) is asserted via subprocess config loads.

## v3.1.7 — CI spawn-race hardening (first published build of the v3.1.5 changes)

- v3.1.6 also never published. The loaded shared CI runner (node cold-boot 100-370ms vs ~50ms locally) keeps picking off the next-tightest spawn-timing test: this time `test-utility-provider-policy`'s shutdown-drain case, which flat-slept 100ms then cancelled — killing the fake agent before it booted, so its SIGTERM handler never wrote the expected signal capture (pipeline 10383; the v3.1.6 run-lifecycle fix itself passed there).
- Instead of fixing one test per failed pipeline, audited all 13 fake-agent-spawning test files for the pattern and hardened the three racy ones: `test-utility-provider-policy` (abort and drain cases now wait for the child's boot capture before cancelling; timeout case widened to 2500ms deadline / 400ms kill grace; helper wait deadlines to 15s), `test-provider-enforcer` (timeout verdict widened to 2500ms so the hung child's boot capture and SIGTERM record exist before assertions; kill grace 300ms), `test-provider-subagent` (helper wait deadline to 15s). Semantics unchanged everywhere — the same hangs, kills, aborts, and drains are asserted, just on deterministic readiness signals instead of guessed sleeps.
- Ships everything listed under v3.1.5 below; no product code changes.

## v3.1.6 — CI test-timing fix (also never published — see v3.1.7)

- The v3.1.5 tag never published: its CI build failed twice on `test-provider-run-lifecycle`'s forced-termination case, whose 120ms timeout lost the node-boot race on the loaded shared runner (the fake agent was SIGTERMed before it could write the descendant capture). Widened to 2000ms timeout / 500ms kill grace — semantics unchanged (agent still hangs, is force-killed, and the detached TERM-resistant descendant must survive).
- Ships everything listed under v3.1.5 below; no other code changes.

## v3.1.5 — Pod-identity hosted speech + dream memory access

- **Hosted speech now authenticates with the pod's own identity.** The speech-gateway bearer becomes `AGENTSPACE_POD_TOKEN` (minted per pod at provisioning, injected via the `openclaw-secrets` secret) with `SPEECH_API_KEY` kept as fallback — so provisioned bot pods get Pocket TTS/STT without any shared API key. The gateway introspects the token against the Agent Space backend (`/pods/self/speech-authorization`), which scopes it to `speech.tts`/`speech.stt` for bot pods only. Wired through config, media, router, runner, voice-policy, and `/doctor` (which now reports which credential the speech path is using). New hermetic `test-media-speech.js` covers bearer selection and fallback order.
- **Dream introspection can read memory again.** The nightly dream pass's introspection step was denied read access to the memory substrate; it now runs with explicit read grants (provider allow-list + utility-agent pass-through), with regressions in the claude-provider and dream test suites.

## v3.1.4 — Preserve Kazee inbound image captions

- Maps Kazee media message content onto the transport-neutral inbound `caption` field, so image instructions reach the agent prompt instead of being replaced by the generic image-description fallback.
- Leaves plain Kazee text and Telegram normalization unchanged.
- Adds a hermetic adapter regression covering captioned images, captionless images, and ordinary text.

## v3.1.3 — Pocket TTS voice-note provider

- Adds Agent Space Pocket TTS as the preferred voice synthesizer for Telegram and streamed voice replies.
- Requests WAV from the authenticated speech gateway and transcodes it to Ogg/Opus with ffmpeg.
- Keeps local Whisper transcription unchanged and falls back to ElevenLabs, then macOS `say`, on synthesis failure.
- Adds configurable `SPEECH_SERVICE_URL`, `SPEECH_API_KEY`, `POCKET_TTS_VOICE`, and `TTS_PROVIDER` settings while stripping the speech credential from provider subprocess environments.

## v3.1.2 — On-demand recall query + cold-walker fallback fix

Two things fold in: a new `recall query` CLI so the working agent can ask long-term memory a question mid-task, and a latent reliability fix in the recall walker's fallback path.

- **New `open-claudia recall query "<question>"`** — the agent's side door into the seed → graph → walker pipeline. Where per-turn recall seeds on the *user's* message, this seeds the graph with the agent's OWN working context (via `--context`), walks the typed-edge graph with the **active-provider** judge, and returns the kept pack sections with why-bullets and pointers. It closes the documented cold-start hole: mid-build, when the vocabulary you're using shares no words with the pack that governs the work, ordinary recall never resurfaces it — an explicit query carrying your working context does. Read-only (no edge reinforcement), no pre-gate (an explicit ask always runs full), and fail-open (a failed judge returns the whole candidate list unjudged rather than nothing). Flags: `--context`, `--json`, `--no-walker`, `--no-episodes`, `--limit`, `--timeout`.
- **Cold-walker outputSchema 400 fix** — the recall walker prefers a warm persistent judge (Claude only) and falls back to a cold spawn otherwise. The cold path was passing an array-typed `outputSchema`, which the Anthropic API turns into a custom-tool `input_schema` and rejects with `400 input_schema.type: Input should be 'object'`. So any time the warm walker errored and fell back, recall silently 400'd and failed-open — judgement quietly disabled on the fallback path. Removed the schema entirely (the warm path never sent one, and `walk()` already extracts the JSON array from freeform text, filters to known ids, and clamps scores), so freeform is the proven contract on both paths.
- **Agent-space / OpenClaw:** new `core/recall/query.js` (`query()` + `mergeMatches`), `excerptFor` exported from the discoverer, `recall query` wired into `bin/recall.js`. New hermetic `test-recall-query.js` in the pretest chain (unjudged fail-open, judged keep-only, judged-empty, garbage fail-open, context-seed merge, blank-question reject).

## v3.1.1 — /logout and switch-account: re-auth is finally reachable from chat

Reported live: with Claude or Codex already connected, there was **no way to sign into a different account** — `/login` showed "connected" and tapping a model just switched models, because the auth-aware picker (correctly, since v3.0.36) only starts a login flow when a provider is *not* authenticated. And `/clear_oauth_token` alone didn't help: it clears the bot-held token but the CLI credential store keeps `authStatus` reading "authenticated", so the picker still never offered a sign-in. Re-auth needs a *full* sign-out first — this release adds it.

- **New `/logout [claude|codex]`** (owner-only) — a full sign-out: runs `claude auth logout` / `codex logout` against the CLI credential store **and** (for Claude) clears the bot-held OAuth token from `.env`, process env, and vault in the same step. Partial clears were exactly the trap. Reports each step honestly (including "CLI not found — store untouched"), then offers a **Sign in** button so the fresh login is one tap away. Without an argument it shows provider buttons.
- **`/login` picker grows a 🔁 Switch account button** under each *connected* provider. Tapping it asks first — showing who is currently signed in (email where the CLI exposes it, else the auth method: OAuth token source / ChatGPT account / API key) and warning that the bot loses that provider until the new login completes — then signs out and chains straight into the fresh login flow. A stray tap can't kill a working auth: the destructive step is behind the explicit confirm.
- **`/clear_oauth_token` is now documented as the narrow tool** (bot token only); `/logout claude` is the full story.
- **Agent-space / OpenClaw:** new auth-flow primitives `logoutClaude`/`logoutCodex`/`providerLogout`/`canLogout`/`providerAccountSummary` (core/auth-flow.js), shared executor `performProviderLogout` + `startProviderLogin` (core/handlers.js), new callback family `sw:`/`lo:`/`losw:`/`li:` (core/actions.js, owner-gated). New test `test-logout.js` in the pretest chain.

## v3.1.0 — Memory becomes required: safe-mode boot, crash-loop rescue, and /downgrade

Three waves folded into one release, graduated to a minor because boot semantics change. Wave 0: the memory substrate is now **required** — a missing `node:sqlite` boots into safe mode instead of silently degrading, a boot crash loop lands there too, and a new `/downgrade` completes the rescue-from-chat story. Wave 1: the memory-substrate revival, root-caused from a live incident — a pod ran 16 days repeating corrected branding mistakes because its entire memory substrate silently never initialized. Wave 2: the first fix batch from the 2026-07-20 blind code review (18 verified findings; deeper architectural items ship in later batches).

### Wave 0 — boot safety (why this is a minor)

- **The memory substrate is now required, not a silent fallback.** A boot without `node:sqlite` (Node < 24) no longer quietly runs with keyword-only memory — it enters **safe mode**: adapters and diagnostic slash commands stay up, but no agent turns, crons, watches, reviewer, or memory writes run. The owner gets one loud alert explaining exactly what's missing and every way out. Running degraded is still possible, but only as an explicit choice: `MEMORY_ALLOW_DEGRADED=1` in the environment, or a one-shot `/safemode continue` from chat.
- **A boot crash loop can no longer strand you with a dead pod.** If 3 consecutive boots die before reaching 5 minutes of stable uptime, the next boot comes up in safe mode instead of crashing a fourth time — the chat channel stays alive precisely so you can rescue it *from chat*. The alert carries the last recorded exit reason (or names it a hard kill — OOM/SIGKILL — when there is none). The counter is honest about edge cases: a clean deliberate exit (restart/upgrade) never counts as a crash, a hard kill *after* proven-stable uptime starts a fresh count, and safe mode's own stability doesn't launder the streak away — only a stable full boot, `/safemode continue`, or a deliberate version change resets it.
- **New `/downgrade` — roll back to what ran before.** With no argument it targets the previously-running version from the new boot version history; `/downgrade <version>` picks any published release (strict semver-validated, existence-checked against the registry before anything executes). Works from safe mode, which turns a bad release into a two-command recovery: crash loop → safe mode → `/downgrade` → back online. On source-overlay pods the change is flagged as ephemeral (a pod recreate returns to the image version — roll the image for permanence). Both `/upgrade` and `/downgrade` clear the crash counter, since a deliberate version change is an operator fix attempt.
- **`/safemode`** shows the active reasons, last exit, and ways out at any time (and reports "running normally" when not in safe mode). `/status` gains a ⛔ banner line while safe mode is active. Non-allowlisted messages in safe mode get a one-line banner pointing at `/safemode` instead of silence.
- **`package.json` now declares `engines.node >= 24`.** npm warns at install on older runtimes; the bot itself still *boots* on them — deliberately, straight into safe mode, so an old pod becomes a diagnosable chat instead of a crash loop.
- **Agent-space / OpenClaw (wave 0):** new module `core/safe-mode.js` (boot gate consulted at the top of `bot.js`; router gates every message/action through it); new state files under the config dir: `boot-marker.json`, `safemode-continue.json`, `version-history.json` (last 20 versions seen at boot). New commands `/downgrade`, `/safemode`; safe-mode command allowlist: help, start, version, status, doctor, requirements, restart, upgrade, downgrade, safemode, whoami, cluster. New test `test-safe-mode.js` in the pretest chain (which also now runs the previously-unwired `test-boot-health.js` and `test-lexical-floor.js`).

### Wave 1 — memory-substrate revival

- **Why a pod forgot everything for 16 days — and why that can never be silent again.** The pod base image ran Node 20, which has no `node:sqlite` — so pack FTS matching, the recall graph, and episodic memory all bailed out quietly at require-time. 542 turns produced zero recall seeds; three separate user corrections were reviewed and dropped; nothing anywhere said a word. `Dockerfile.base` now builds `FROM node:24-slim` (node:sqlite stable), and the CI build job tests on `node:24` to match — with the `sim-no-sqlite` shim tests keeping the degraded path covered.
- **/doctor grew a full memory-health section.** Substrate presence (hard fail with the exact fix when node:sqlite is missing), recall-graph size, the last 24h of the recall pipeline (turns, seeded, injected, walker failures — 10+ live turns with zero seeds is a hard fail: "matching is returning NOTHING"), reviewer success rate (<50% fail, <90% warn, deferred counts, last error code), mutation-lock state (free / stale / long-held / active), deferred review batches on disk, corpus size vs the lesson cap, and disk headroom.
- **Substrate degradation announces itself at boot.** A degraded boot sends the owner one loud alert ("⚠️ MEMORY SUBSTRATE DEGRADED … I will remember poorly until this is fixed") — deduped per capability fingerprint, re-fired every 24h while broken, cleared the moment a healthy boot is seen so a later regression alerts fresh.
- **Every upgrade is now followed by a doctor run on the NEW process.** All three `/upgrade` paths (AgentSpace rollout, source-overlay, npm) record a marker before restarting; the next boot consumes it and sends a full doctor report back to the channel that requested the upgrade. The pre-restart report can only ever vouch for the *old* process — this closes the gap where an upgrade lands on a broken substrate and nobody notices.
- **Recall now has a floor: no sqlite means keyword matching, never nothing.** `matchPacks`/`matchEntities` fall back to a pure-JS lexical scorer that mirrors the FTS scoring exactly (strong-field hit 2, body hit 1, same thresholds and limits, porter stemming approximated by substring + suffix-strip variants). Both recall engines get seeds through the same two matchers, so the floor fixes classic *and* discoverer with zero engine changes — proven equivalent by a parity test that runs the same assertions with and without the no-sqlite shim.
- **The memory reviewer can no longer be starved by a stale lock.** In a container the bot is PID 1 — and after a crash, the *restarted* bot is also PID 1, so the old liveness check saw a dead incarnation's mutation lock as alive forever (one pod: 87% of reviewer runs starved with `MEMORY_MUTATION_LOCK_TIMEOUT`, feeding a 43-restart loop). Locks now carry a per-process boot id: same pid + different boot id breaks immediately, ownerless locks break after a 5s grace, and a genuinely live holder is still never broken. The reviewer also gained a seatbelt: on lock timeout it retries once after 60s, and if still starved it persists the decision batch to `pending-review-mutations.jsonl` (capped, surfaced by /doctor) and records the run as "deferred" instead of losing the memory write.
- **LOCKED stance lines now carry rule authority.** A Stance line containing `LOCKED` (uppercase — prose "locked" stays put) is hoisted under a "BINDING CONSTRAINTS" header in every injection mode, so a frozen choice reads as a rule, not background prose. The reviewer prompt gained the matching contract: never weaken, reword, or drop a LOCKED line unless the user explicitly unlocks it, and when a reply violates one, journal the violation ("VIOLATION: …") so drift becomes visible instead of silently absorbed.
- **Compaction stops matching memory against its own boilerplate.** The summarize meta-turn now skips recall entirely (its prompt is instructions about summarizing — matching packs against it was noise). The seed turn — which opens a fresh session and previously lost all pack context — now matches on what the work *is*: the condensed brief plus open task titles, so the packs active before compaction re-attract immediately. Lessons stay always-on, and re-attracted packs bring their LOCKED constraints back with authority.
- **Agent-space / OpenClaw (wave 1):** base-image change `FROM node:20-slim` → `node:24-slim` (rebuild `:base` required — the CI base job triggers on the Dockerfile.base change), CI build image `node:20` → `node:24`. New modules `core/doctor-memory.js`, `core/boot-health.js`; boot wiring in `bot.js`; marker writes in the three `/upgrade` paths. New state files under the config dir: `substrate-alert.json`, `pending-doctor.json`, `pending-review-mutations.jsonl`; mutation-lock `owner.json` gains a `bootId`. New tests `test-boot-health.js`, `test-lexical-floor.js`; expanded `test-memory-mutation-queue.js`, `test-extractive-injection.js`, `test-provider-compaction.js`, `test-provider-recall.js`.

### Wave 2 — blind-review fix batch

- **The recall graph no longer eats its own memory.** The nightly decay pass rewrote every edge weight by reapplying the *full* age since last reinforcement — every night. That compounds: an edge on the intended 60-day half-life actually emptied in about ten nights, so long-lived Hebbian links quietly bled away and recall got worse the longer you used it. Decay is now computed **on read** (the stored weight is the value as of its `last_reinforced` stamp; every read maps through `effectiveWeight()`), reinforce/weaken settle the accrued decay before applying their delta, and the nightly rewrite path is gone entirely — the bug class can no longer exist. A one-time, idempotent migration (`recall-graph.db` PRAGMA `user_version` 0→1) treats existing weights as current and restarts the clock. The half-life is now a dream-tunable knob (`decayHalfLifeDays`, env-pinnable) read lazily, so a tuning change applies on the very next read.
- **Keyring and tool state can no longer be destroyed by a crash mid-write.** The operational keyring, tool `state.json`/usage books, and recall tuning were bare `writeFileSync` calls — an OOM kill or power loss mid-write left a truncated, unparseable file, and for the keyring that meant every stored credential gone. All of them now go through the crash-safe writer (temp file + fsync + atomic rename, 0600). The keyring additionally snapshots the last good version to `keyring.json.bak` before each overwrite and *reads through* it transparently: a corrupt primary logs one warning per process, serves the backup, and the next successful write repairs the primary in place.
- **When memory recall breaks, you now hear about it instead of the bot quietly getting dumber.** Recall has always been fail-open (a crashed builder must not block the turn) — but it was also fail-*silent*: the error vanished and every turn just ran memoryless. The failure now logs to the bot log, shows immediately as a "🧠 Recall FAILED this turn" banner when `/recall` is on, and after **3 consecutive** failed turns the owner gets one alert per streak ("running without pack/entity/tool context", with the last error) — quiet again until it recovers or breaks anew.
- **"ok deploy kazee now" no longer skips memory recall.** The recall pre-gate skipped short turns that *start* with a pleasantry — an ack prefix hid a real command, so deploy/restart requests wearing an "ok"/"thanks" ran without any pack context. The gate now strips leading pleasantry tokens (stacked ones too) and tiers what remains; pure ack chains ("ok thanks sounds good") still skip at any length.
- **A chat approval is no longer burned by a run that could never start.** `tool run --approval <id>` consumed the one-shot token *before* checking keyring credentials — a missing key meant the human's Approve was spent on nothing and they had to be asked again. The creds preflight now runs before any approval machinery (including owner escalation, which is no longer raised for unrunnable commands): the failure names the missing keys, the token survives, and the byte-identical command redeems it after `keyring set`.
- **Tool-sequence learning now sees piped compositions.** The graph that learns "tool A is usually followed by tool B" split command lines on `&&`, `||`, `;` and newlines — but not on a single `|` or background `&`, so `tool run a … | tool run b …` credited only the first tool. It now splits on all shell composition operators, with quoted arguments masked first so a pipe inside a string doesn't split the command.
- **Two source files no longer masquerade as binaries.** `core/recall/discoverer.js` and `core/state.js` contained raw NUL bytes inside string literals (composite-key separators), which made grep and most tooling treat the files as binary and skip them. The bytes are now written as `\u0000` escapes — byte-for-byte identical strings at runtime, plain text on disk.
- **README caught up with reality on recall engines.** It claimed **classic** was the default and **discoverer** opt-in; the shipped default has been `discoverer` (per-chat `/engine` → `RECALL_ENGINE` env → default). The engine docs, command table, and env table now say so, and the discoverer section documents its cost profile: one small utility-model call per non-trivial turn, skipped entirely by the pre-gate on trivial turns, model pinnable via `RECALL_DISCOVERER_MODEL`.
- **Agent-space / OpenClaw:** no new deps, no new env (all knobs pre-existed). One schema touch: the automatic, idempotent `recall-graph.db` `user_version` 0→1 restamp above; keyring writes now leave a `.bak` sibling. Changes span `core/recall/graph.js`/`tuning.js`/`discoverer.js`, `core/keyring.js`, `core/tools.js`, `core/tool-graph.js`, `core/state.js`, `core/system-prompt.js`, `core/turn-observer.js`, `core/dream.js`, `bin/tool.js`, README. Covered by new `test-turn-observer.js` (in the pretest chain) plus expanded cases in `test-recall-graph.js`, `test-recall-discoverer.js`, `test-tools.js`, `test-tool-graph.js`, `test-tool-manifest.js`, `test-provider-core-prompt.js`; full suite green. Existing pods pick it up on their next approved upgrade.

## v3.0.44 — Safety refusals explain themselves, and one refusal no longer poisons every turn after it

- **A Claude safety refusal now tells you why instead of dumping a raw API error.** When the model declines a request (`stop_reason: "refusal"`), the API attaches a category and sometimes an explanation — but all you saw was "API Error … appears to violate our Usage Policy" with a request id buried inside. The refusal is now detected structurally from the stream (`stop_details` category/explanation, with a text-parse fallback when the CLI flattens it into the error string) and surfaced as a clear ⛔ diagnostic: the reason, the request ID (for support), and your options (/new, rephrase, /model or /backend).
- **One refusal no longer bricks the whole conversation.** The safety classifier scores the *entire* conversation, so resuming a session whose context tripped it re-sent the flagged material and refused again — every turn, even for innocent new messages; the only escape was knowing to type /new. Now the poisoned session is **quarantined** the moment a resumed turn is refused: its id is recorded in user state (bounded to 20, survives restarts), the stale active pointer is healed, and nothing resumes it again — the natural next turn and even an explicit `/continue <id>` are forced fresh. Session ids the CLI rotates mid-run are quarantined too.
- **Your refused message is retried for you, once, in a fresh session.** After quarantining, the same prompt is automatically replayed in a brand-new session — you see "♻️ That refusal came from flagged context in the old session — quarantined it, retrying your message in a fresh one" and then the actual answer, so you keep working without re-typing anything. Strictly one-shot: if the fresh run is also refused, the new message itself is the problem and you get the full diagnostic with no retry loop; capture/background runs never auto-retry.
- **Agent-space / OpenClaw:** no new deps, env, or schema change (a `quarantinedSessions` map is added to per-user state; old state files load cleanly); Docker image builds FROM the existing `:base`. Changes span `core/providers/events.js`/`claude-events.js` (detection), `core/state.js` (quarantine bookkeeping), `core/run-context.js` (resume guard), `core/runner.js` (diagnostic + one-shot retry); covered by new `test-provider-refusal-recovery.js`, a fake-agent `refusal` scenario, and runner e2e cases. Existing pods pick it up on their next approved upgrade.

## v3.0.43 — Telegram stops rebooting into a dead network: outages now pause-and-wait instead of restart-storming

- **A host-wide network blip no longer triggers a reboot storm.** v3.0.42 fixed the *in-loop* poll wedge (a poisoned reused poller) so it heals in place — but it still kept the 10-minute `process.exit(1)` backstop, which fired even when the *entire host* had lost the network (a laptop sleep/wake, or pod egress starvation where Telegram poll+send, Kazee WS and Spaces WS all drop together). Exiting into a dead network just relaunched straight back into the same outage, so you'd see "Back online and ready! Running v3.0.42" every ~14 minutes, busting the prompt cache and spamming the chat. Now the self-exit is **reachability-gated**: before it ever restarts the process it runs a fresh, connection-fresh HEAD probe to `api.telegram.org`, and only exits if that probe genuinely **succeeds** — proving the API is reachable and it's a real in-process wedge worth a clean restart, not the host being offline.
- **When the host is actually offline, the bot now pauses and waits instead of hammering.** On a confirmed-unreachable probe the adapter enters an *outage pause*: it **stops polling entirely** (so `node-telegram-bot-api` no longer respins `getUpdates` every ~300ms into a dead socket — the "not waiting between reconnects" hammering) and re-probes on a calm **30-second** cadence. The instant connectivity returns it resumes polling cleanly from a fresh poller (the update offset is preserved, so nothing is refetched or skipped) — no process restart, no "Back online" spam. The pause also suspends the health watchdog and hiccup handler so they don't fight the slow re-probe. Poll-health transitions are now timestamped for easier post-mortem.
- **Agent-space / OpenClaw:** no new deps, env, or schema change; Docker image builds FROM the existing `:base`. Change is confined to `channels/telegram/adapter.js` (the `_heal` backstop plus the new `_enterOutagePause`/`_armHealthyCheck` paths); covered by new cases in `test-telegram-poll-recovery.js` (the gated exit defers on a live host, pauses instead of rebooting on a dead one, and resumes when the network returns). Existing pods pick it up on their next approved upgrade.

## v3.0.42 — Telegram stops rebooting itself: the long-poll wedge now heals in place

- **The bot no longer restarts the whole process to recover from a stuck Telegram long-poll.** When the network stalls a poll mid-flight (a half-open socket after a laptop sleep/wake, or a NAT/IPv6 connection eviction), `node-telegram-bot-api` was left holding one reused internal poller whose abort flag got stuck "on". Restarting polling *in place* revived that same poisoned object, which fired a single `getUpdates` and then silently refused to schedule the next one — a dead loop that threw no error, so the heal looked like it "worked" while nothing was actually polling. The only thing that ever truly recovered was the 10-minute backstop calling `process.exit(1)` and letting the supervisor relaunch — i.e. the reboot storm you'd see. The heal now **discards the poisoned poller so a fresh one is built** (the update offset is preserved, so no messages are refetched or skipped), and it only declares polling "healthy again" once a `getUpdates` has genuinely **completed since the restart** — otherwise it re-arms another heal instead of blindly declaring victory. Net effect: transient poll wedges recover in seconds without a process restart; the 10-min self-exit becomes the rare true-last-resort it was meant to be.
- **Agent-space / OpenClaw:** no new deps, env, or schema change; Docker image builds FROM the existing `:base`. Change is confined to `channels/telegram/adapter.js` (the `_heal` path); covered by three new cases in `test-telegram-poll-recovery.js` (proven to fail before the fix and pass after — the poisoned-object revival and the false "healthy" signal). Existing pods pick it up on their next approved upgrade.

## v3.0.41 — Verbose off now means a clean chat: the "Let me…" narration no longer lingers

- **With `/verbose` off, the interim progress message is now removed once a turn settles, so the chat is left with only the final answer.** During a multi-step turn the bot posts a throttled "working…" preview that shows the short "Let me check…/Let me look at…" narration emitted before each tool call. That preview used to *stay* — settling into a permanent message — regardless of `/verbose`, which only ever gated the extended-thinking (🧠) stream. So even with verbose off you'd be left with the narration trail as a leftover message that reads like thinking-out-loud. `finalizeStreamPreview` now deletes that preview when verbose is off; `/verbose on` still keeps the narration trail. Live progress during long turns is unchanged either way — the cleanup only affects what remains after the turn ends.
- **Agent-space / OpenClaw:** no new deps, env, or schema change; Docker image builds FROM the existing `:base`. Pure delivery-path change in `core/runner.js`; existing pods pick it up on their next approved upgrade.

## v3.0.40 — Sharper recall: surfaces the actual matched facts, drops weakly-relevant ones (third cut of the Pass-4 release)

- **On-open pack/entity injection is now extractive, not a headline.** Opening a pack used to inject its one-line description, hiding the facts that actually matched in State/Procedure/Journal. It now injects the Stance verbatim plus the specific fuzzy-matched lines (section-tagged), so the model sees the real matched content rather than a summary of it. Toggle with `RECALL_EXTRACTIVE` (default on).
- **A walker score-gate drops weakly-relevant nodes.** The cheap recall walker is a pure router now — it emits `{id, why, score 0-100}` and never paraphrases. Nodes it keeps below the relevance floor (`RECALL_MIN_RELEVANCE_SCORE`, default 55) are gated out on full-tier walks, and the cascade is tightened so sticky auto-keeps are routed back to the judge instead of leaking in. Fails open (a garbled walk bypasses the gate).
- **A KPI ruler stamps each version.** Per-turn recall records now carry kept-node scores, gate activity, and warmings; the KPI table/HTML report aggregate them per version so each release freezes as a comparable baseline (`relevance`, `gate`, `warm` columns).
- **Co-kept Hebbian warming ships dark (default OFF).** A weaker-than-co-open reinforcement for pairs the walker co-injected, guarded by score floor + a message-grounded endpoint + a top-N cap. It writes persistent graph state so it can't be validated in one session — left off (`RECALL_COKEPT_WARMING`) pending real-traffic evaluation.
- **Escape hatch + measurement caveat.** `/engine classic` reverts to the prior behaviour if anything regresses. Two reviewer-side measurement refinements (used-the-right-fact line-level adjudication, and shadow-delta scoring) were deferred to the separate in-flight shadow-delta task — they change how recall is *measured*, not what ships here.
- **Agent-space / OpenClaw:** no new deps, no new env required (all `RECALL_*` knobs have defaults), no schema change; Docker image builds FROM the existing `:base`. Covered by new `test-pack-search.js`, `test-walker-score.js`, `test-extractive-injection.js`, `test-score-gate.js`, and `test-cokept-warming.js`. Existing pods pick this up on their next approved upgrade. (Two earlier cuts failed CI before publishing anything: v3.0.38 bumped the version but not the release-pin guard in `test-provider-language.js`; v3.0.39 fixed that but `test-cokept-warming.js` asserted a sqlite-graph write that no-ops on CI's `node:20` — it now skips that path when `node:sqlite` is absent, like `test-recall-graph.js`. Both were verified locally by running the suite with `node:sqlite` simulated absent, reproducing the CI environment.)

## v3.0.37 — Silent background watches, and pasted secrets are accepted on any provider

- **You can now arm a watch: a background poll that stays out of the chat and only wakes me when its condition is actually met.** A watch is just a cron carrying a shell `check` command — each tick runs the check with output discarded, and *only* an exit 0 wakes the agent (with a `🔔 Watch "<label>" triggered` notice and a normal reply). A non-zero exit is the idle case: nothing hits the chat, the poll's activity goes to an append-only `watch.log`, and the job stays pending for the next tick. Watches default to **one-shot** (fire once, then remove) so a matched condition can't spam you; add `--repeat` to keep polling. Create/manage from the CLI (`open-claudia watch-add "<check>" "<wake-prompt>" [--every "<cron>"] [--repeat]`, `watch-list`, `watch-remove`), and the owner can flip the global master switch with **`/watch on|off`** (watches are owner-gated because the check is raw shell). Rides the migration-verified jobs store untouched — `check`/`repeat` survive as extra fields, zero schema change.
- **Pasting a credential into the chat now stores it out-of-band regardless of which provider is active.** Codex refuses secrets in its prompt, so a pasted key under the Codex backend used to be rejected — and being told to re-paste it through keyring commands *in the same chat* was no safer and plain annoying. A loose paste is now classified (`core/credential-detect.js`) and captured **before** the turn reaches the model: a Claude OAuth token → `.env`+vault, an OpenAI key → handed to the Codex CLI, a recognised service token → the operational keyring, an ambiguous long token → scrubbed from chat. Every captured secret is registered for redaction and the original message deleted. The capture is provider-agnostic and never forwards the secret to any model.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds FROM the existing `:base`. Covered by new `test-credential-capture.js` and watch probes in `test-provider-scheduler.js`. Existing pods pick this up on their next approved upgrade.

## v3.0.36 — Picking a Claude model no longer forces a re-login

- **Selecting any Claude model claimed Claude wasn't connected and started a login — even while Claude was working fine.** The model-button handler (and the `/login` picker) decided "connected?" from a shallow token probe (`getClaudeOAuthToken().value`) that only sees an OAuth token in `.env`, the process env, or an unlocked vault. Anyone whose Claude login lives in the Claude CLI's own credential store (the normal `claude` login, no bot token) therefore read as *disconnected*, so every model tap answered "Claude isn't connected yet. Starting login…". The re-login it launched (`claude auth login --claudeai`) never stores a bot token, so the shallow check stayed false and the prompt returned on the very next tap — an unbreakable loop.
- **Both surfaces now resolve Claude's real auth state via `claudeProviderAuthStatus()`** — the same probe agent runs rely on, which counts a bot OAuth token *and* a `claude auth status` CLI-store login as authenticated. A model tap only starts a login when Claude is genuinely unauthenticated, mirroring the Codex branch beside it. Codex was never affected (it already checked `codex login status`). Cost: one `claude auth status` (~1s) on a model tap when no bot token exists — fine for a settings screen, and skipped entirely when a bot token is present.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds FROM the existing `:base`. Existing pods pick this up on their next approved upgrade.

## v3.0.35 — Spaces comments are readable rich text, not a run-on wall

- **The bot's Spaces replies rendered as unreadable walls.** `buildCommentContent` posted the entire reply as one text node inside one paragraph. Spaces comments are TipTap/ProseMirror docs: inside a single text node newlines don't render, so paragraphs, numbered lists and `---` separators collapsed into one run-on blob — and the model's markdown (`**bold**`, `*italic*`, backticks) printed its asterisks literally.
- **Outgoing comments now convert markdown → a real ProseMirror doc.** Blank lines become paragraphs, single newlines hardBreaks; `**`/`*`/`` ` `` become bold/italic/code marks (word-bounded, so snake_case identifiers stay plain); `-`/`1.` runs become real bullet/ordered lists; `#`–`###` become headings; ``` fences become code blocks; `---` becomes a divider. Every emitted node type was verified against the Spaces frontend editor schema (StarterKit + HorizontalRule + CodeBlockLowlight), so nothing renders broken.
- **Mention chips unchanged.** The proven tag-chip shape still leads the first paragraph, so notification fan-out and the "@Name text…" lead-in behave exactly as before. `flattenContent` also learned `codeBlock`/`horizontalRule` so the bot reads its own comments back cleanly.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change. Covered by new `test-spaces-comment-content.js` in the pretest chain. Bots pick this up on their next upgrade, together with v3.0.34's /upgrade redelivery guard.

## v3.0.34 — /upgrade no longer triggers a rollout storm

- **One `/upgrade` used to mean up to five rollouts.** On AgentSpace pods the control plane answers `/upgrade` by rolling the deployment, which SIGTERMs the bot mid-turn. Telegram long-polling keeps no persisted offset, so the still-unconfirmed `/upgrade` update was redelivered to the freshly booted pod — which triggered *another* rollout, and so on. Observed on rahilcoder 2026-07-17: deployment revisions 47–51 (five rollouts in three minutes) from a single command; every earlier `/upgrade` had silently double-fired the same way (revs 42/43, 45/46).
- **Fix: an upgrade-in-flight marker (`upgrade-in-flight.json` in the config dir, 10-minute freshness) guards the handler.** A redelivered or impatiently re-run `/upgrade` inside the window is acknowledged but not re-triggered: if the pod has already restarted since the request it replies "that upgrade already went through", otherwise "already in flight". The marker is cleared on a definitive trigger failure so an immediate retry stays possible, and simply ages out otherwise.
- **The ack now goes out BEFORE the rollout is requested.** Previously the "Upgrade requested" confirmation was sent after the control-plane call, so the SIGTERM routinely killed it — Rahil saw reboot greetings *before* any acknowledgement. Ack-first means the user always sees the confirmation, then the reboot.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds FROM the existing `:base`. Existing pods pick this up on their next upgrade — that upgrade itself will still multi-fire (old code is doing the triggering), all later ones are single-shot.

## v3.0.33 — Full sudo inside the pod (re-cut of v3.0.32; the pin guard bit us again)

- **`claudia` now has `NOPASSWD:ALL` sudo in the pod image.** The rebuilt `:base` widens the baked sudoers from the old command allow-list to full passwordless sudo, so the agent can install packages and operate inside its own container like a normal machine. This pairs with agent-space's IaaS-style security-policy system: the blast radius is fenced at the pod's NetworkPolicy (egress/ingress presets, DB-rendered), not by crippling the user inside the container.
- **Re-cut of v3.0.32.** That tag bumped `package.json` but not the version pin in `test-provider-language.js`, so `pretest` asserted 3.0.31, the **build** job failed, and the **docker** job was skipped — nothing published, `:latest` stayed on 3.0.31. Identical payload here with the pin bumped. (Same failure mode as v3.0.30 → v3.0.31.)
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds FROM the rebuilt full-sudo `:base`. Existing pods pick this up on their next approved upgrade.

## v3.0.31 — Re-cut of the approval-card release (the version-pin guard bit us)

- **Same feature payload as v3.0.30, this time it builds.** The v3.0.30 tag's build failed before publishing anything. `test-provider-language.js` pins the release version and asserts it "must remain unchanged without explicit approval" — a deliberate gate so every bump is consciously acknowledged in the same commit. `package.json` was bumped to 3.0.30 but that pin wasn't, so `pretest` asserted 3.0.29, the **build** job failed, and the **docker** job was skipped (no npm publish, no image — `:base` never entered it). The pin now tracks the release, so the build goes green.
- **Nothing else changed.** All of v3.0.30's approval-card UX — the "From `<name>`: `<ask>`" card, in-place Approve/Deny, "Make a rule" → per-person mandate or global guardrails, and the three-layer enforcer — ships unchanged under this version. See the v3.0.30 entry below for the detail.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.30 — The approval card shows who asked, acks your tap, and can become a rule

- **The card now shows the context of the decision.** An external-guardrail approval used to show only the proposed reply, with the sender rendered as "an external contact" (an unmapped speaker's people record has no name). The card now leads with **From `<name>`: `<what they actually asked>`** and labels the held text **Proposed reply** (or **Proposed action**). The inbound message and the sender's transport display-name are threaded through the async request context → `escalate()` → onto the approval record, and the display-name is the fallback when there's no people record.
- **Approve / Deny now visibly change the card.** Tapping a button edited nothing on Kazee, so it was unclear whether the press registered. The handler now **edits the card in place** to `✅ Sent to <name>` / `⛔ Held` / `✅ Approved …`, using the same `adapter.edit(channel, sourceMessageId, …)` path the auth-request buttons already use, and **falls back to a fresh message** if there's no source id or the adapter refuses the edit (Kazee has 403'd edits before). Telegram still gets its `answerCallbackQuery` ack.
- **"Always allow" → "📝 Make a rule…".** The blanket button is replaced by a two-step, intuitive flow: it approves-and-delivers this one, then asks **👤 Just `<name>`** or **🌍 Everyone**; you type the rule in plain language and it's saved to that person's entity **Mandate** (per-person) or to the new **global guardrails** (every external). No people record → only the global option is offered.
- **The guard now vets against three layers.** `enforcer` composes a hardcoded **security floor** (never leak secrets/infra, default-deny — unchanged, absolute) → owner-authored **global guardrails** (new) → the per-person **mandate**. Authority is the union; a prohibition in any layer wins. Global guardrails are managed with **`/guardrails add|remove|list`** (owner-only, capped, deduped) and stored in `external-guardrails.md`. They're folded into the guard's judgment prompt and the verdict cache key.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.29 — The owner actually gets the approval request

- **The external-speaker guardrail no longer swallows approval requests.** When a non-owner talks to a bot and the guard holds the reply, the bot escalates to the owner with Approve/Deny buttons and tells the external "let me check with the owner." That escalation resolved the owner's channel through `relationship.ownerTarget()`, which read **only** `people.owners()`. On a bot seeded from legacy `auth.json` the people store is empty, so it returned no owner, escalation failed with "no reachable owner," and the owner got **nothing** — while the external was still told a follow-up was coming. (Seen live: `kazee-people` deferred to the owner in a Kazee Spaces thread and no approval ever arrived.)
- **`ownerTarget()` now falls back to the configured-owner env identity** — the same operator-declared signal the guardrail already trusts *inbound* (`identity.isConfiguredOwnerChannel`: Telegram `TELEGRAM_CHAT_ID`, Kazee/Spaces `KAZEE_OWNER_USER_ID`). When no people-record handle exists it derives a deliverable target from env, preferring Telegram (native buttons, and the chat id is the send target) and otherwise resolving the owner's Kazee 1:1 DM via the adapter's idempotent `openDirectChat` (a Kazee send target is a chat-document id, not a user id). Only channels actually loaded on the bot are considered; `ownerTarget()` is now async.
- **Fail-closed contract preserved.** If no configured owner channel is reachable the guard still blocks (returns no target) — the fix only *restores* the owner's approval prompt, it never widens what an external can do. Loopback's Spaces-approval binding gets the same fallback. No autonomy change; B2+ and Track A stay gated.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.28 — /upgrade stops crying wolf on a slow roll

- **A client-side `/upgrade` timeout is no longer reported as a failure.** When the bot runs inside an AgentSpace pod, `/upgrade` calls the control plane, which triggers a fire-and-forget rollout and returns 202. If the pod is mid-roll its own in-flight request gets cut off, so the 15s HTTP wait would expire and the bot announced "Upgrade request failed … timeout" — a false negative that invited repeated re-runs (and a storm of redundant rollouts, seen as 11 ReplicaSets in 7 minutes). A timed-out request now surfaces as **pending** ("sent but the reply timed out — likely because I'm already restarting; give it a minute"), and the request timeout is raised 15s → 30s.
- **Pairs with a control-plane probe fix (the actual root cause).** The pod's `/health` probes carried no `timeoutSeconds`, so k8s defaulted it to 1s; a cold-start event-loop stall past 1s flapped the liveness probe and SIGTERM'd the pod in a restart loop. That fix — explicit `timeoutSeconds: 5` on every probe and liveness `failureThreshold` 3 → 5 — ships in the agent-space control plane and reaches existing pods on their next reconcile (`/cluster sync`).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.27 — Agency ledger, advisory "what's next", owner-guardrail capture, and a reboot that waits for the turn to end

- **`/agenda` — one read-only work list across everything (Track B0).** A new aggregator normalises your persistent tasks, `spaces mine`, and connector inbox/notifications into a single list (source, title, status, due, staleness, blocked-on, last-touched). Pure visibility — it reads, it never acts. `/agency` is an alias.
- **`/next` — the bot proposes what it would pick, and does nothing (Track B1, advisory).** A triage→rank→judge pipeline (deterministic prefilter → algorithmic rank → a small judge on the ambiguous top-N) says "here's what I'd do next and why." It acts on nothing without you — bounded autonomy (B2+) stays gated behind a shadow-mode review, exactly as planned.
- **Owner-guardrail capture — say a rule mid-chat and it sticks (Track C0/C1).** When *you* say "from now on, always X" / "never Y" / "stop doing Z", a post-turn detector turns it into an always-loaded lesson immediately — no code change, no deploy. High-confidence phrasings are auto-captured with an announcement and a one-tap **Undo**; softer ones ask first (Save / No). Hard owner gate: a standing rule is **never** captured from an external speaker, and captured rules layer *above* lessons but can never override the code-fixed hard rules. Default-on; set `GUARDRAIL_CAPTURE=off` to disable.
- **Always-on prompt budget — instrumented, not yet trimmed (Track X0).** Every real turn now measures the fixed always-on floor (soul, lessons, runtime state, slash list) versus the variable recall payload, recorded to telemetry so we can baseline cost before capping it. Pure measurement — the prompt itself is unchanged, so caching is unaffected. Exposed via `agency ledger` / `prompt-budget` CLIs and the recall stats.
- **A reboot that never interrupts a turn.** The Telegram long-poll has a 10-minute wedge backstop that self-exits for a clean restart when polling is truly stuck. It now checks first whether any turn is running, queued, or compacting — and if so **defers** the exit, keeps attempting in-place recovery, and only restarts once the process is idle (all replies delivered). If polling self-heals meanwhile, no restart happens at all. Recovery only touches the poll socket, so an in-flight turn's subprocess and its outbound sends are untouched.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade. `GUARDRAIL_CAPTURE` is the one new (optional, default-on) env knob.

## v3.0.26 — Spaces replies actually post (and stop saying "Queued.")

- **An engaged Spaces reply no longer fails with 400 "Parent comment not found."** Core threads every agent reply under the triggering message id — a Telegram-ism. On Spaces that id is a *notification* id, not a comment id, so the backend rejected the whole post and the bot silently failed to answer any mention. `send()` now only threads under an explicit `comment:`-prefixed id and otherwise posts a top-level task comment; if a stale/deleted parent still 400s, it retries once at top level so the reply always lands.
- **The bot threads its reply under the exact comment it's answering.** When a notification names a specific comment (a reply's `replyId`, or a mention/comment that targets a comment), the reply is threaded under it instead of floating at the task root. Task-level mentions and assignments stay top-level, as before.
- **No more "Queued." noise on a task thread.** The transient queue/compaction acknowledgement is a chat-UX affordance for Telegram/Kazee. On a Spaces task it posted a throwaway "Queued." comment when a second turn stacked (and, pre-fix, that post itself 400'd). Spaces now stays silent while queued and simply posts the real reply when its turn runs. Per-identity run serialization is unchanged and intended — one being, one hand.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.25 — The bot types back on a Spaces task

- **"Open Claudia is typing…" now shows on a Spaces task thread while the bot composes.** Core already runs a typing heartbeat around every agent turn (calling `adapter.typing()` every ~4s and `adapter.typingStop()` at turn-end); on Spaces those were no-op stubs. They now re-emit the exact `comment:typing` event a human client sends over the adapter's existing authenticated socket, so a teammate watching a task sees the bot compose like any other member — the same three-dot indicator the web and mobile apps already render.
- **Part 4 of the cross-repo typing indicator.** The backend (`comment:typing` re-broadcast to the `space:<id>` room, identity stamped from the authenticated socket, no persistence), web, and mobile surfaces shipped in prior deploys — human↔human typing was already live end-to-end. This closes the loop so the bot participates too.
- **Minimal by construction.** The adapter caches each task's `spaceId` (from the notification payload and the thread-context fetch) so the heartbeat — which only knows the `task:<id>` channel — can name the room. The emit sends just `targetId + spaceId + isTyping`; the server derives who is typing. Fully best-effort: no socket, no cached space, or a transport blip is a silent no-op, so a typing dot can never affect or delay the actual reply.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.24 — Spaces threads read like a group chat; memory notes go quiet unless you're watching

- **The "🧠 Jotted this down" memory notes are now gated behind `/recall`.** After each turn the pack reviewer still writes to long-term memory exactly as before — but the chat announcement of what it filed away only surfaces when `/recall` is on (`state.settings.showRecall`), the same toggle that already gates the recall/tool banners. Off by default, so ordinary turns stay clean; flip `/recall on` to watch memory (and recall) work. That announcement was the one learning-surface bypassing the gate; it no longer does.
- **Re-ship of v3.0.23, which never published.** Its tag build failed the version-approval CI gate — `package.json` was bumped without mirroring the assertion in `test-provider-language.js`. Same gate that caught v3.0.21; corrected here as v3.0.24, exactly as v3.0.22 re-shipped v3.0.21.
- **An engaged Spaces turn now arrives knowing the whole thread, not just one bare comment.** When a tag/reply/assignment engages the bot on a task, the router pulls a **thread-context block** from the adapter and prepends it to the prompt: the task's metadata (title, status/priority/due, assignees, a clipped description) **plus every comment posted since the bot last replied**. So the bot answers like a member who's been reading along, instead of replying blind. The full task + entire comment history stays one `spaces read` tool call away.
- **"Group chat the bot is a member of."** Untagged comments still never trigger a turn (the engage gate is unchanged — mentions/replies/assignments only). But the next time the bot *is* addressed, everything said in between is caught up in one batched injection. The delta is derived server-side from a per-task **context watermark**, so a missed poll or a restart self-heals — the next engage still pulls the true set. The bot's own prior reply is re-surfaced too, which keeps continuity across auto-compaction.
- **Bounded by construction.** The catch-up is capped (newest 12 comments, each body clipped) so a busy thread can't blow the prompt; the very first engage on a long-idle task shows recent history, not the entire backlog. Auto-compaction absorbs the rest — no new wiring, it's already per-thread and token-triggered.
- **Rendered cleanly, not stitched into the message.** Context is delivered via a new channel-agnostic `buildThreadContext()` adapter method that the router renders — `envelope.text` stays pristine, so command detection and the credential scrub are unaffected (the old approach mutated the message text). The seam is generic: Kazee group rooms can implement the same method later, and the since-last-reply delta is exactly the input a future lightweight activation agent would judge to decide whether to chip in unprompted — provisioned, not built.
- **The bot now leaves a read receipt, like a teammate.** When an engaged turn catches up on a thread, it marks every comment it just read as *seen by* the bot (`comments/mark-seen` → server `$addToSet`s the bot into each comment's `seenBy` and broadcasts `comment:seen`). So a human watching a task sees the bot appear on "seen by" exactly like another member the moment it reads their message — not only when it replies. Fire-and-forget and fully swallowed: it rides the same member/admin `comments.update` scope the reply already uses, and a scope or network failure degrades to silent (never breaks or delays a turn). Only ever fires on an engaged turn — the awareness poll still never mutates Spaces state.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.20 — Spaces answers tags, and knows the owner

- **A tag now actually gets a reply.** The Spaces server labels notifications with **plural** categories (`mentions`, `comments`, `reactions`, `task_updated`), but the conversational plane's engage set only matched the singular `mention` — so every real `@`-tag was silently downgraded to awareness-only and the bot never spoke. `ENGAGE_TYPES` now matches the plural forms (`mentions`/`replies`, singular aliases kept for safety). Plain comments, reactions and task updates stay awareness-only, so the "untagged comments hit the radar without auto-answering" rule is unchanged. Confirmed live: a real `mentions` notification was arriving and being dropped.
- **The owner is recognised on Spaces.** Spaces principals are central/Kazee user ids, but `isConfiguredOwnerChannel` had no `spaces` branch, so the owner's own comments resolved to a fresh `spaces:<id>` canonical instead of their owner identity — classing the owner as **external** (default-deny) on every Spaces task. Added a `spaces` branch mapping to `KAZEE_OWNER_USER_ID` (the same id Spaces/Kazee/central share), so the owner is recognised, their Spaces turns unify into their one canonical brain, and the relationship guardrail no longer misfires on them.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.19 — flattened to the top level (the folder-project model is gone)

- **The bot always runs at the top-level workspace now — no more "Pick a project first."** The per-folder session model (pick a subfolder of `WORKSPACE`, scope the conversation to it) has been removed. `currentSession` defaults to the workspace root (`{ name: "Workspace", dir: WORKSPACE }`) at state creation and on every migration/load, so it is never null. This directly fixes Spaces (and any non-Telegram channel that never picked a folder) replying with a dead "Pick a project first:" project-picker instead of engaging — the picker gate was the only thing blocking it.
- **Removed the folder-picker surface.** Gone: `/projects`, `/session`, the project-selection inline keyboard, and the `show:projects` / `s:<name>` button callbacks. Also removed the four media "Pick a project first." gates (voice/audio/photo/document) and the text-handler gate in the router — all dead once a session always exists. `core/projects.js` is deleted.
- **Conversation-level controls are unchanged.** `/sessions` (list past conversations), `/new` (fresh conversation), `/continue` (resume), `/end` (end current conversation — now resets to the root instead of clearing to null) all work as before, just always scoped to the one top-level workspace.
- **`/cron add` no longer takes a `<project>` argument** — it schedules at the workspace root: `/cron add "<schedule>" "<prompt>"`. Existing crons with a stored project still resolve via the scheduler's descriptor.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.18 — Spaces reconciles its task list (assigned tasks no longer silently missed)

- **Assigning a task to the bot now actually engages it.** An assignment creates the task but emits **no notification**, and the conversational plane only woke on a socket "notification" push — so a task assigned to the bot could sit there while the bot never noticed (confirmed live: bot could `listTasks` its assigned task, but had 0 notifications and never engaged). The adapter now runs a self-contained **reconciliation poll**: on its own timer it calls `listTasks()` (bot-scoped to its assigned tasks) and synthesises a `task_assigned` engage for any active task it hasn't engaged yet.
- **Robust to socket churn.** The Spaces socket transport reconnects constantly; a push emitted during a gap is lost with no replay. The timer also re-runs the existing notification catch-up `poll()` and, crucially, doesn't depend on the connected-apps poller (which can be disabled) — the Spaces plane drives itself. Cadence is `OC_SPACES_POLL_MS` (default 60s).
- **Per-task engaged-dedup.** Engagement is now tracked per task (`markTaskEngaged`/`taskEngaged`), not per notification-id, so a task the socket already handled is skipped by reconcile and vice-versa — no double replies. Terminal (completed/archived/cancelled) tasks are skipped.
- **A successful `listTasks()` is a connection signal** (`markRestConnected`) — the third leg of the hybrid auto-enable, independent of the flaky socket. `/spaces status` now shows a "Last poll" line. Still owner-gated: `/spaces pause` stops replies; reconcile records connection but engages nothing while paused/off.
- **Agent-space / OpenClaw:** no new deps, no new required env (one optional `OC_SPACES_POLL_MS`), no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.17 — Spaces auto-enables on connection; control moves into chat

- **The Spaces conversational plane no longer hides behind an env flag.** `OC_SPACES_CONVERSATIONAL` is gone. Auto-replies now enable themselves the moment the bot is **connected** to Spaces, on a hybrid signal: central authoritatively lists Spaces among the bot's connected apps, **or** the socket has actually received a Spaces notification (which only happens for a task the bot is a member of). Either one flips it on — so it works today over the socket even before the central `/api/principal/applications` endpoint is deployed. It stays self-gating: no Spaces membership ⇒ no notifications ⇒ no replies.
- **All control is in chat now — nothing to set on the pod.** New owner command `/spaces`: `status` (connection + whether replies are live + task mode), `pause` / `resume` (the off-switch for auto-replies; mentions still hit the radar, I just don't answer), and `mode advisory|autonomous|off` (task-working posture, overriding the env default live). State persists in `spaces-state.json`; the owner never needs filesystem or env access.
- **The Spaces channel auto-includes whenever `KAZEE_BOT_TOKEN` is set**, even if `CHANNELS` doesn't list it — Spaces is fixed infra on the same token as Kazee chat, so "has token ⇒ Spaces available" is what makes auto-enable reachable without editing env.
- **Agent-space / OpenClaw:** no new deps, no new env required, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.16 — Spaces task loop + approval thread-mirror (dormant)

- **Kazee Spaces connected-apps gains its two remaining pieces, both shipped OFF.** The task-working loop runs in **advisory** mode by default: when the bot engages a Spaces task thread, its turns for that task are serialized (one at a time, while different tasks proceed in parallel), a persistent cursor + cross-restart dedup lives in `spaces-state.json`, and a new `OC_SPACES_TASK_LOOP` mode (`advisory` | `autonomous` | `off`) selects the framing. The **approval thread-mirror** routes any write/destructive action from a Spaces-origin turn to the owner's Telegram Approve/Deny button (fail-closed if no reachable owner channel — the run stays blocked rather than posting a dead button into the thread) and mirrors a non-binding status comment into the task thread on request and on outcome.
- **No behaviour change anywhere until you opt in.** All of it stays dormant behind the existing kill-switches (`OC_CONNECTED_APPS`, `OC_SPACES_CONVERSATIONAL`, plus the new `OC_SPACES_TASK_LOOP` which is advisory-but-inert while the others are off) and the bot still needs Spaces app-membership to receive anything. Telegram and Kazee flows are byte-identical.
- **Agent-space / OpenClaw:** no new deps, no new env required, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.15 — the owner keeps their memory on Kazee

- **Episodic recall no longer goes dark on Kazee / group chats.** Cross-conversation episodic memory (relevant snippets from the owner's past transcripts) is deliberately owner-only. The gate that decides "is this speaker the owner?" resolved the person by `findByHandle(adapter, channelId)` — but on Kazee the `channelId` is the *chat/room id*, while the owner's stored handle is their Kazee *user id*, so the lookup missed and fail-closed suppressed episodes. The owner's canonical brain, settings, and session were shared via identity unification, but this one recall layer silently went quiet — which read as "less context on Kazee." The gate now resolves the speaker by the envelope's **canonical user id** first (derived from the authenticated sender and unified across a person's linked channels), falling back to the channelId handle lookup when no canonical is in context. Fail-closed semantics for non-owners are preserved: a non-owner's canonical never matches an owner record. Regression pinned in `test-recall-relationship-gate.js` (owner recognised via canonical on a room-id channel; external still blocked).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.14 — the Telegram poll stops wedging, and every reboot says why

- **Root-cause fix for the restart flap.** In the cluster, DNS hands back a dead IPv6 address for `api.telegram.org` first and IPv6 egress is unreachable; under Node's default `verbatim` resolution the long-poll dials the dead address first (a latency penalty always, and a hard wedge on any AAAA-only window). `bot.js` now sets `dns.setDefaultResultOrder("ipv4first")` before any socket opens, eliminating that path for every request the process makes.
- **Half-open long-polls now fail fast instead of wedging for 10 minutes.** The Telegram `request` block gained `timeout: 40000` — with keep-alive off, `request` arms a socket-idle timeout, so a poll whose socket silently dies is killed (`ESOCKETTIMEDOUT` → `EFATAL` → hiccup → heal on a fresh socket) in ~40s rather than sitting dead until the 10-minute backstop calls `process.exit(1)` and the pod restarts. 40s sits safely above the 30s long-poll hold so healthy idle polls are never chopped; `ESOCKETTIMEDOUT` was also added explicitly to the transient-error classifier.
- **Every code-initiated reboot now reports its reason.** New `core/restart-reason.js` records *why* right before any deliberate exit (polling-wedge, uncaught exception, boot failure, graceful SIGTERM/SIGINT, `/restart`, `/upgrade`, runtime-mode switch); the "Back online" greeting reads and clears it and appends `↳ Reboot reason: …`. No record on boot ⇒ an external hard restart (SIGKILL, OOM, host or k8s force-restart) that nothing could log — the greeting says so explicitly. This turns a silent "Back online" into an actionable one.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up on their next approved upgrade.

## v3.0.13 — chat with the bot from its own dashboard

- **The web dashboard gains a Chat tab — a real channel, not a side door.** A new in-process `WebAdapter` (`channels/web/`) registers as a first-class channel whenever the web UI is up: messages typed in the dashboard flow through the same router, identity, packs, transcripts, and approval pipeline as Telegram/Kazee/voice, and the bot's replies (including inline approval buttons, files, voice notes, live typing) stream back over `/api/chat/stream` (SSE with history replay on connect). New session-gated routes: `GET /api/chat/stream`, `POST /api/chat/send`, `POST /api/chat/action`, `GET /api/chat/media/<id>`. The channel is fixed single-owner (`web-owner`): the dashboard session already gates every call, so `access.js`/`identity.js` authorize it as the owner and — under unified identity — it resolves to the owner's canonical brain like every other owner channel.
- **Platform SSO: `POST /api/otp/mint` (control-token only).** AgentSpace can mint a one-time magic-link token (`/otp/<token>`, 10-min TTL, single-use — the existing `core/web-otp` machinery behind `/dashboard`) over `BOT_CONTROL_TOKEN`. Dashboard sessions get a 403: they are already logged in and gain nothing, so the surface stays strictly machine-to-machine.
- **SSE unsubscribe bug fixed before it shipped.** Cleanup was hooked on `req.on("close")`, which in modern Node fires as soon as the GET body completes — every stream subscriber was torn down instantly after the history frame. Cleanup now hooks `res.on("close")` (actual connection teardown), verified by an end-to-end live-frame test.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Existing pods pick this up only on their next approved image upgrade; new pods get it immediately. `channels/` already ships in `files`, so the adapter rides the existing package layout.

## v3.0.10 — Kazee-only pods are first-class

- **/doctor no longer demands Telegram env on a Kazee-only install.** `health.js` required `TELEGRAM_BOT_TOKEN`/`TELEGRAM_CHAT_ID` unconditionally — a pod provisioned with `CHANNELS=kazee` failed its env check despite being perfectly configured. Required keys are now derived from `CHANNELS` (default `telegram`, preserving legacy behaviour): telegram installs require the Telegram pair, kazee installs require `KAZEE_BOT_TOKEN`, and `WORKSPACE` stays universal.
- **Kazee owner bootstrap stores the right identity in the right key.** `/auth` bootstrap on a kazee transport wrote the chat-document id into `TELEGRAM_CHAT_ID` — the wrong id (owner checks compare the Kazee *user* id) in the wrong key (polluting Telegram's authorized list). It now writes `KAZEE_OWNER_USER_ID` with the sender's user id and updates the in-memory config so ownership is live without a restart.
- **An env-provisioned Kazee owner counts as an owner.** `hasOwner()` only looked at `TELEGRAM_CHAT_ID` and auth.json — on a Kazee-only pod provisioned with `KAZEE_OWNER_USER_ID` (the agent-space wizard path), the first stranger to send `/auth` would have been crowned bootstrap owner. `KAZEE_OWNER_USER_ID` now counts.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). This is the bot-side half of Kazee pod provisioning via the agent-space wizard (wizard + provisioner land in agent-space itself). Coverage: new hermetic `test-kazee-channel-health.js` (required-key matrix; kazee bootstrap env writes; provisioned-owner /auth guard), wired into `npm test` and shipped in `files`.

## v3.0.9 — one fat record no longer kills the turn

- **Decoder failures are warnings, not run failures.** The v3 JSONL stream decoder capped provider records at 1MB and surfaced any oversized or malformed line as a terminal `error` event — so a single base64-image tool result (routine in image workflows) killed an otherwise healthy live turn with `Claude Code run failed: Provider JSONL record exceeded 1048576 bytes` (rahil, v3.0.8, 2026-07-12). `LINE_TOO_LARGE` / `MALFORMED_JSON` now normalize to a new non-terminal `warning` event: the record is skipped, the stream continues, and the warning is logged. If the dropped line happened to be the terminal record, the existing `PROVIDER_MISSING_TERMINAL` net still fails the run cleanly.
- **Line cap raised 1MB → 16MB.** Image-bearing records are legitimate; the cap now exists to bound memory, not to police normal payloads.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically (FROM the pre-baked `:base`). Coverage: warning-shape assertions in `test-provider-events.js`, default-cap pin in `test-provider-stream-decoder.js`.

## v3.0.6 — cache-bust burn is now a one-line query

- **Usage records carry a `coldStart` flag.** The restart→cache-bust burn (each bot restart re-bills the whole accumulated window as cache *writes* instead of ~13×-cheaper cache *reads*) was invisible to every token-count analysis — same token counts, different price class — and took a forensic reconstruction of the 22-day usage history to quantify (~8% of all spend, ~$12/day at its worst). Now the first usage record a process writes for each session is tagged `coldStart: true`: a cold start mid-conversation is the restart fingerprint, so the burn is one jq filter away instead of a modelling exercise. Records already carried `cacheReadTokens`/`cacheCreationTokens`; this adds the missing "was the cache necessarily cold?" dimension.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change (additive JSONL field); Docker image builds identically. `usageSessionColdStart` covered by new assertions in `test-usage-accounting.js`.

## v3.0.5 — the conflict handler owns the poll

- **Hiccup heals no longer fight the 409 backoff.** During a 409 conflict streak the conflict handler pauses polling and retries on its own schedule — but a network hiccup (`ECONNRESET`) arriving mid-streak would schedule its own heal, which called `startPolling` straight back into the conflict and reset the streak clock. This was the live v2.15 burn mechanism: its own ghost long-poll after an ECONNRESET heal raised a 409, the old code exited on any 409, launchd respawned it (~1/hr, 208 lifetime startups), and every restart busted the prompt cache so the next turn re-billed the whole accumulated window uncached. Hiccup handling is now a no-op while a conflict streak is active — the conflict handler owns the poll lifecycle.
- **The 10-minute wedge backstop stands down during a conflict.** A conflict streak looks exactly like a wedge (no updates flowing), so the backstop could `exit(1)` mid-streak — the one remaining path where a token fight still killed the lock-holding bot (2 wedge-exits observed in the field). The backstop now skips while a conflict is active; a genuine >10-min wedge with no conflict still exits for a clean restart.
- **Investigated and cleared:** the v3 "token doubling" report was fully root-caused. Per-turn recall injection measured at parity (~4k tokens, 3 packs) across v2.15-classic, v2.15-discoverer, and v3-discoverer; the 160k→347k warning was the v3.0.0 cumulative-vs-peak display bug already fixed in v3.0.2; real spend was flat across the upgrade. The discoverer walker is identical in v2 and v3 (same haiku model via tier "low", same ~$0.015/turn) and stays on by default.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. New hermetic test `test-telegram-conflict-heal-guard.js` (wired into `npm test`) proves all four paths: hiccup no-op during conflict, hiccup still heals without conflict, backstop skipped during conflict, genuine wedge still exits.

## v3.0.4 — surviving the second mount, keeping the crons

- **A Kubernetes PVC remount no longer crashloops the bot.** The migration snapshot validator demanded exact owner-only modes (`0700` dirs, `0600` files). But kubelet's fsGroup ownership pass ORs group `rwX` + setgid onto every file and directory each time the volume is mounted — so the first v3 boot passed (snapshot created after the mount), and the **second pod recreation** hit `INVALID_MIGRATION_SNAPSHOT: migration snapshot has unsafe permissions` at boot and went CrashLoopBackOff. Every v3 pod on an fsGroup-mounted PVC was one recreation away from this (euvy hit it live on 2026-07-12 and needed a manual chmod rescue). Validation now rejects only world-access bits — group access is the pod's own supplemental group, not an exposure — while snapshots are still *created* `0700`/`0600`.
- **Legacy crons adopt your saved defaults instead of dying in the archive.** Pre-provider-aware jobs carried no pinned provider or project — they fired with whatever the bot had active — so the v3 migration classified every one of them "no unambiguous project" and archive-disabled it, silently killing real production crons (euvy's 6-hourly Airbyte data-sync among them). The jobs migration now looks up the owning user's saved state (`settings.backend` and `currentSession` name/dir, in both the v3 `users` shape and the pre-v3 global shape), adopts those as the job's provider/project, and marks the record `legacyDefaultsApplied` — preserving the old fire-time behaviour instead of erasing it. Jobs pinned to a removed provider (Cursor) keep their pin, stay archived, and are never reassigned; a job whose owner has no saved defaults still archives safely.
- **Re-reading the jobs store no longer rewrites stored nulls.** `Number(null)` is `0`, so the re-migration pass that runs on every read flipped `nextAttemptAt`/`lastFireAt` from `null` to epoch-`0` — harmless at runtime but a silent mutation on every load. Migration output is now byte-stable across passes.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. New hermetic probes: fsGroup-mode tolerance (group bits pass, world bits still refused) in `test-provider-migration-backup.js`; legacy-defaults adoption, v2/v3 state shapes, Cursor guard, and migration idempotency in `test-provider-scheduler.js`.

## v3.0.3 — the lock-holder stands its ground

- **A persistent 409 no longer restarts the bot.** v3.0.1 taught the poller to ride out short 409 conflicts but still `exit(1)` once a conflict outlived 2 minutes, on the theory that the survivor had to be a real second instance. In practice the rogue poller is lock-blind — a debugger-run dev checkout, a copy on another config dir, or another host on the same token — so the exit killed the *legitimate*, lock-holding service and launchd respawned it straight back into the same fight (five boot→409→exit cycles on 2026-07-12). The poller now never exits over a 409: it pauses 15s per retry, slows to 60s once the streak passes 2 minutes, sends the owner a throttled Telegram alert (max one per hour) naming the likely culprits, and posts an all-clear once the rogue goes away.
- **The single-instance lock is atomic.** Acquisition was read-check-then-write, so two processes booting together (a launchd respawn racing a manual start) could both judge the lock stale and both claim it — each believing it was the only copy. The lock is now claimed with an exclusive-create (`O_EXCL`) in a bounded retry loop: stale, corrupt, or own-pid locks are unlinked and re-claimed atomically, a live holder still turns the newcomer away, and a process that loses every claim attempt refuses to boot rather than assuming ownership.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. `test-telegram-409-grace.js` now proves the never-exit + owner-alert + all-clear path; `test-single-instance.js` passes unchanged against the atomic rewrite.

## v3.0.2 — honest context numbers, dream summaries that arrive

- **The context warning counts the window again, not the bill.** The v3.0.0 provider rewrite computed the live context number from the provider's *cumulative* turn usage — every tool round-trip re-counted the cached prompt prefix, so a normal turn read as 347k+ "context" (crossing 1M on busy turns), spooked the usage alert, and could prematurely force-compact sessions whose real window was fine. The number now tracks the **peak single API call** in the turn (`peakContextTokens` over per-call usage events) — true window occupancy, the same semantics v2 reported (~160k on a heavy turn). Billing is untouched: ledger records keep the cumulative figure in `billedContextTokens`, and API-reported spend was never affected (v3 bills the same or less than v2 — the display was wrong, not the meter). The usage alert says "Context this turn" whenever the peak is available.
- **Nightly dream summaries reach the chat again.** Dreams ran every night at 04:00 — and for over a week every summary died with Telegram's `400: message is too long` (the report outgrew the 4,096-char hard cap and the adapter had no chunking), so consolidation happened invisibly. Two fixes: the Telegram adapter now splits any oversized message on paragraph → line → word boundaries and sends the chunks in order (reply anchor on the first, buttons on the last, per-chunk HTML→plain fallback preserved), and the dream's tool-health note caps long lists (unused tools, drifted docs) to a short digest in chat while the full lists still land in the on-disk dream log (`~/.open-claudia/dreams/<date>.md`).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. New hermetic assertions: peak-vs-cumulative derivation in `test-usage-accounting.js`, chunk-size and content-preservation in `test-delivery-contract.js`.

## v3.0.1 — streaming restored, /stop queue survival, boot resilience

- **Live streaming is back, and the final message stays clean.** The v3.0.0 provider rewrite dropped the in-chat streaming feel: no live progress while a run worked, and the settled reply concatenated every text block into one blob. Now an active run edits a throttled preview message in place with the agent's narration (and its thinking, when `/verbose` is on — new command + button toggle), and when the run settles only the **last** text segment is delivered as the final answer; earlier pre-tool text is preserved as narration instead of being glued onto the reply. Providers emit a normalized `thinking` event (Claude thinking blocks, Codex reasoning items) that feeds the preview but can never leak into settled output. On success the preview becomes the narration record (or is deleted when there was none); on failure or cancel the partial preview is kept so you can see how far the run got. Preview edits are channel-guarded, so a bubble minted on one channel is never edited from another.
- **`/stop` no longer throws away your queued messages.** Cancelling killed the current turn *and* silently emptied the message queue — every turn you'd typed behind the running one vanished. `/stop` now cancels only the current turn (including a pre-spawn abort checkpoint for runs still building their prompt), keeps the queue intact, and tells you how many queued messages will run next. Cancelled runs unwind quietly instead of delivering an error-shaped reply.
- **The Codex model menu now offers the GPT-5.6 generation.** The retired `gpt-5.1-codex` trio is replaced by `gpt-5.6-sol` / `gpt-5.6-terra` / `gpt-5.6-luna`; effort tiers map low→luna and medium→terra so utility work stays cheap, with sol as the default.
- **One 409 doesn't kill the bot anymore.** A single Telegram 409 Conflict (another poller on the same token — usually a ghost poll from a restarting predecessor) used to `exit(1)` immediately, which is how one stray poll became a restart storm. The poller now rides out a conflict streak: pause 15s and retry, clear the streak after 45s of quiet, and only exit if the conflict persists beyond 2 minutes (a real second instance). The watchdog defers restarts during a streak.
- **A single-instance lock refuses double-boots.** Boot now claims `bot.lock` in the config dir (pid, start time, entrypoint); a second bot pointed at the same config refuses to start while the holder is alive, claims a stale or corrupt lock safely, and release never deletes a successor's lock. Combined with the 409 grace this makes accidental duplicate launches self-identifying and harmless.
- **Startup polling stalls recover on their own.** A Telegram long-poll that went silent right after boot (dead socket, no error) is now detected and restarted instead of leaving a deaf bot.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. New hermetic tests (`test-streaming-split.js`, `test-telegram-409-grace.js`, `test-single-instance.js`, `test-telegram-poll-recovery.js`) wired into `npm test`; CI test commands now carry the fake-agent env explicitly so the suite is hermetic on runners with no provider CLIs.

## v3.0.0 — provider parity and Cursor removal

- **The runtime is now a provider-agnostic coding-agent harness.** Claude Code and OpenAI Codex implement one registry contract for prompts, immutable run admission, normalized events, native sessions, capability reporting, utility work, and pre-tool safety policy. Setup, web configuration, status, and doctor work with either, both, or neither provider installed; only model turns require a compatible authenticated provider.
- **Active Cursor Agent runtime support has been removed.** Open Claudia now runs model turns through the registered Claude Code and OpenAI Codex providers only. The command, model/backend buttons, executable discovery, auth/doctor branches, configuration key, active session pointer, and runtime parser/invocation path are gone. Buttons from an older message return a removal notice and direct the user to `/backend`; they never switch to another provider silently.
- **Cursor data is archived, not rewritten.** The verified provider-migration snapshot retains the byte-for-byte original state, session, job, and legacy cron files. Cursor selections, settings, session pointers/history, and pinned jobs move only into non-selectable/disabled archive records; no native session ID or model is reassigned to Claude or Codex.
- **Migration activation is proven, not asserted.** Startup snapshots the allowlisted originals, migrates state, sessions, and jobs before the scheduler/router/adapters load, and records each live output's schema, size, path, and SHA-256. A partial, stale, or legacy output cannot satisfy provider-removal readiness; rerunning startup safely completes an interrupted component.
- **Rollback requires the matching older runtime.** Stop every bot/scheduler process, validate and restore the snapshot described in `docs/PROVIDER_MIGRATION.md`, then restart a pre-removal Open Claudia release that still understands the restored schema. Do not run old and new schedulers against the same restored files.
- **Authenticated release validation passed locally on 2026-07-10.** Claude Code 2.1.158 and Codex CLI 0.144.0 each passed a fresh turn, resumed turn, read-only plan, harmless tool command, utility JSON output, and prompt-context sentinel through the Open Claudia adapters. Claude-only and Codex-only startup/status probes also passed with the opposite executable absent. These were isolated existing-auth checks only: no login, credential change, publish, deploy, or production action was performed.
- **Real-CLI validation tightened three compatibility edges.** Claude `structured_output` is now the authoritative terminal value for schema-constrained utility work even when the CLI also streams presentation text. Codex hook trust uses a strict inline `projects` table accepted by the current CLI while user configuration remains ignored, and the supported model tiers no longer select deprecated `o4-mini` for utility work.
- **Breaking-release recommendation accepted — released as `v3.0.0` (2026-07-10).** Removing an active provider and changing configuration/session schemas are breaking changes; package metadata and the git tag are v3.0.0.
- **Post-review hardening (2026-07-10).** stdin EPIPE from a fast-exiting provider child no longer crashes the bot — it surfaces as a stderr diagnostic and the exit code decides the run. Crons/wakeups archived by the provider migration are announced once at boot and listed in `open-claudia cron-list` and `/cron`. Healthy provider status is cached for 30 seconds so per-message admission stops respawning CLI probes; disappearance detection stays instant and unhealthy states always re-probe. `/doctor` now runs the real-binary Codex tool-hook enforcement probe behind an explicit opt-in, so boot/health/setup never spawn a model turn. Merged main's Telegram polling recovery fix.

## v2.15.0
- **Production-hardening pass (P0). Two root causes, closed at the source.** A deep-dive audit traced the reliability and safety gaps to two patterns: (1) **overloaded return contracts** where a falsy value meant both "no id" and "failed", so callers guessed — the exact shape behind v2.14.8's Kazee double-send; and (2) **unsafe shared state under concurrency**, amplified because the owner's Telegram + Kazee turns now share ONE in-memory state object *and* load-modify-write the same JSON files under unified identity. This release fixes both, plus the credential-exposure and crash-durability gaps found alongside them. No new deps, no new required env, no schema change.
- **One explicit delivery contract, so no caller guesses.** Adapters historically overloaded `send()`'s return: a message id on success, a bare `true` for an idless success (server fast-ack), or a falsy value on failure. `core/io.send()` now collapses every adapter's raw return through `normalizeSendResult` into `{ ok, messageId, editable }` — `messageId` is always a string or null, so `typeof`/edit-target checks are uniform across transports. The runner's final-delivery guard is now `if (!sent.ok)` (an idless success no longer reads as failure → no phantom re-send), and the streaming-preview path keeps the real string id when editable or a truthy sentinel otherwise, so `canEditStatus()` still declines to edit an unaddressable bubble. The loopback approval-prompt send routes through the same normalizer. New `test-delivery-contract.js`.
- **Crash-safe atomic writes for every durable JSON store.** `state.json`, `sessions.json`, and `identities.json` were each written with a bare `fs.writeFileSync(JSON.stringify(...))` — a crash or a full disk mid-write left a truncated file that fails to parse, and the loaders swallowed the parse error and returned `{}`, i.e. **silent total data loss** (every session pointer, per-user setting, and identity link gone). New `core/fsutil.js` provides `atomicWriteFileSync` (write to a temp file, `fsync`, then `rename` — atomic on POSIX — keeping a `.bak` sibling) and `readJsonWithFallback` (try the file, then its `.bak`, then the default). Every loader/writer in `core/state.js` and `core/identity.js` now uses them, and a failed write logs instead of vanishing. An interrupted flush now falls back to the last good file rather than wiping state.
- **A synchronous single-flight run-lock closes the double-spawn race.** The turn guard was `if (state.runningProcess)`, but `runningProcess` isn't set until *after* the pre-spawn awaits (recall, auto-compaction, auth preflight). Two turns for one canonical user — the owner's Telegram and Kazee messages sharing one state object — could both pass the check during that window and both spawn, racing the shared state and JSON files. New `state.acquireRunLock()` arms `state.preparingRun` **synchronously** (no `await` between test and set), so the second turn queues cleanly; the lock is released if a run bails before spawning (auth preflight) so a never-spawned turn can't wedge the next one. New `test-run-lock.js` proves the single-flight property on a shared state object.
- **Dashboard credential exposure closed on three fronts.** (1) `GET /api/config` returned the **raw `.env`** — every bot token, model API key, and the web control token in plaintext — to anyone holding a dashboard session; it now passes through `maskEnv`, which masks any secret-looking key to a 4-char tail (the config UI only ever edits the non-secret `SAFE_KEYS`, and masked values can't round-trip back in). (2) The session cookie was a **static `sha256(password)`** — identical for every login, unexpiring, and impossible to revoke short of a password change; replaced by `core/web-sessions.js`, which mints a random 256-bit token per login, tracks it server-side with a 24h expiry, validates the cookie against that store, and supports individual `revoke` + `revokeAll` (a password change now drops every outstanding session, then re-issues one for the caller). Sessions persist via the atomic writer, so a login survives a restart/upgrade. (3) The bot's own control-plane secrets (`TELEGRAM_BOT_TOKEN`, `KAZEE_BOT_TOKEN`, `BOT_CONTROL_TOKEN`) are now **stripped from the ambient subprocess env** (`botSubprocessEnv`) — the model CLI is authed by its own token/key and `open-claudia` CLIs re-read `.env` from disk, so nothing legitimately needs them, and a prompt-injected agent dumping its environment can no longer read them. New `test-web-sessions.js` (token lifecycle, targeted + mass revoke, expiry sweep, prototype-pollution-key rejection, disk persistence).
- **Crashes and shutdowns no longer lose state or die silently.** In-memory per-user state (settings, session pointers, usage, compaction timestamps) was only ever flushed opportunistically — a `SIGTERM` (every deploy/restart) or an uncaught exception dropped whatever hadn't been saved. A new best-effort synchronous `persist()` (→ `saveState()`, made crash-safe by the atomic writer above) now runs in `gracefulShutdown` and in the `uncaughtException` handler before exit. And `notifyError` — the last-ditch crash notifier — now **redacts secrets** before sending (`core/redact`), hardens the raw Telegram POST (`req.on("error")` guard so the notifier can't itself throw; correct `Content-Length` via `Buffer.byteLength`), and adds a second path that relays to the owner's **primary channel** via a live adapter, so a Kazee-only owner is reached too (skipped when the raw Telegram ping already covers a Telegram-primary owner — no double-notify).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. All changes are correctness/safety hardening — inert-by-default and back-compatible, with one **visible behaviour change on upgrade**: because the dashboard session model changed, any existing dashboard cookie is invalidated once, so a logged-in browser is asked to log in again after this upgrade (expected, one-time). Three new hermetic tests (`test-run-lock.js`, `test-web-sessions.js`, `test-delivery-contract.js`) wired into `npm test`; full suite green.

## v2.14.8
- **Kazee replies no longer post twice.** On a successful send, the Kazee adapter's `send()` returned the created message's id so the runner can edit its streaming preview in place — but it read only `res.message._id`, then fell back to the ActionHero envelope's numeric `messageId`, then `res._id`. chat-central serves the new message's id under `message.id` on some routes (the adapter's own reply-to fetch already handles `m._id || m.id`), so when only `.id` was present `send()` returned `null`. The final-delivery path in `core/runner.js` reads that return as "did the send fail?": `const sent = await send(firstChunk, { replyTo }); if (!sent) await send(firstChunk);`. A successful-but-idless send therefore looked like a failure and the reply was posted a second time — once as a reply-quote, once as a plain bubble (exactly the pair seen in-chat). `send()` now returns `message._id || message.id` (or the top-level equivalents) and **never** the numeric envelope `messageId` — that counter is a per-request id, not a chat message id, and using it also made later `PUT /message/<counter>` edits 500.
- **A streaming preview minted on one channel is never edited on another.** Under unified identity the owner's Telegram and Kazee turns share ONE per-canonical state object, and `statusMessageId` (the id of the live "typing…/preview" bubble) sat on it untagged. A Telegram message id (e.g. `24991`) could then be edited under the Kazee adapter → `PUT /message/24991` → *"Cast to ObjectId failed"* 500, a frozen preview, and a duplicate final send. `statusMessageId` now carries the channel it was minted on (`statusMessageChannel`, `core/state.js`); every edit-in-place site in `core/runner.js` (streaming tick + final delivery + the external-hold notice) only edits when the channel matches, otherwise it discards the stale id and sends fresh on the correct channel. Telegram-only deployments are unaffected (single channel → always matches).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Both are correctness fixes to multi-channel message delivery — the double-send fix helps every Kazee deployment; the channel-scope guard is inert unless a canonical spans two channels. New hermetic test `test-kazee-message-id.js` (id extraction across `_id`/`id`/top-level shapes + the numeric-envelope-counter guard) wired into `npm test`; full suite green.

## v2.14.7
- **New `/link remove <key>` — prune a single identity link through the code path, not by hand-editing `identities.json`.** Until now the only identity commands were `/links` (list) and `/link` (create); there was no way to drop a specific channel→canonical mapping. `/channel remove <transport>` clears a whole transport's links, but only when that adapter is currently live (`removeAdapter` bails with *"No adapter …"* otherwise), and it never touches `telegram` entries by design — so a stray or stale link (e.g. a leftover test fixture, or a link whose transport is no longer added) had no clean removal path, and hand-editing the file doesn't stick because `identity.js` loads the map once into a shared in-memory object and re-serializes the whole thing on every write (a disk-only edit is resurrected by the cache). New `identity.removeIdentityMapping(key)` removes one channel key from the in-memory map **then** persists (memory + disk consistent), and prunes the canonical's `preferred`-channel pointer only when that was its last remaining channel. Owner-only; refuses to remove your own configured owner channel (`identity.isConfiguredOwnerChannel` — a live Telegram chat id or Kazee owner id) so you can't accidentally drop yourself out of unified identity — re-point with `/link` instead. `/links` now shows a *"Remove one with /link remove <key>"* hint.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Purely an owner-driven identity-maintenance command — inert for any deployment that never runs it. New hermetic test `test-identity-prune.js` (last-link pruning + memory/disk consistency + unknown-key no-op) wired into `npm test`; full suite green.

## v2.14.6
- **`/channel add kazee` started from the inline "Add Kazee" button now completes — owner-capture confirm lands on the operator's channel, not on Kazee.** The wizard's owner-confirm prompt (*"Is this you — the owner?"*) is meant to go back to the channel the operator ran `/channel` from, via `channel-wizard.sayToOrigin`, which needs `session.origin` — captured from the envelope passed to `start()`. The typed `/channel add kazee` path (`handlers.js`) passed the envelope; the button path (`actions.js` `chn:add:kazee`) did **not**, so `origin` was null and the prompt fell back to ambient `io.send` scope — which, while processing the operator's inbound Kazee trigger message, is the Kazee room. The Yes/No/Skip buttons then rendered on Kazee, where a tap can never resolve the session: it's keyed by the operator's canonical id, but a Kazee tap resolves to the owner's *unlinked* Kazee user id (the very thing owner-capture is establishing) → `handleAction` finds no session → *"That setup step expired."* `actions.js` now passes the envelope to `wizard.start`, exactly like the typed path.
- **`/channel remove` and the inline "Remove <id>" button now share one teardown, so they can't drift.** v2.14.5 made the typed `/channel remove` a complete teardown (adapter + `.env` + `people.unlinkAdapterHandles` + `identity.removeTransportIdentities`), but the button (`actions.js` `chn:rm:`) still ran the pre-v2.14.5 version that cleared only `.env` — so removing via the button left exactly the handle/identity residue v2.14.5 was built to clear. The full teardown is now a single `channel-wizard.removeChannel(id)` called by both entry points; callers only guard `telegram` and format `r.summary` / `r.error`. Net −10 lines.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Both fixes close divergences between the button and typed entry points of the owner-driven `/channel` flow — inert for deployments that never add/remove a channel. Full `npm test` green.

## v2.14.5
- **`/channel remove <id>` is now a complete teardown — a re-add starts from a genuinely clean slate.** Removal used to clear only the `.env` channel config (`CHANNELS` + the `KAZEE_*` keys), but the wizard's v2.14.4 owner-capture also links the owner's Kazee handle onto their people record (`people.linkHandle`) and mirrors an identity mapping into `identities.json` (`setIdentityMapping`). Those were left behind, so a subsequent `/channel add kazee` stacked on stale handle + identity state. `/channel remove` now also drops every handle for that transport across all people records (new `people.unlinkAdapterHandles`) and removes its identity mappings — channel→canonical links and preferred-channel pointers (new `identity.removeTransportIdentities`) — and clears the previously-orphaned `KAZEE_BOT_USER_ID`.
- **The cleanup goes through the in-memory stores, so disk and the running process stay consistent.** `identities.json` is loaded once at boot into a shared in-memory object (`identity.js`), and `saveIdentities()` re-serializes that whole object on every write. Because `router.js` calls `autoLinkOwnerChannel` → `setIdentityMapping` → `saveIdentities` on each owner message, any entry lingering in memory is rewritten to disk continuously — so an out-of-band file edit can't stick, and a partial teardown leaves ghosts that reappear. Removing through the in-memory maps (then persisting) means the teardown holds. The `/channel remove` reply now reports what it cleared (e.g. `Removed channel: kazee — cleared 1 handle + identity links`).
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Scoped entirely to the owner-driven `/channel remove` path — inert for deployments that never remove a channel. Full `npm test` green.

## v2.14.4
- **The owner no longer gets asked to double-approve their own Kazee reply — split-brain owner resolution is fixed.** On Kazee a person is identified by `userId` while the room is `channelId` (the chat-document); the two differ. The auth layer already recognised the owner correctly (via `access.matchesTransportOwner` → `currentUserId()`), but the independent relationship guardrail re-derived ownership from `channelId` — which never equals the configured `KAZEE_OWNER_USER_ID` — and so classed the owner as *external*, escalating the owner's own approval a second time. `relationship.speakerFor` now takes an optional third `userId` and keys identity/owner classification off `speakerId = userId || channelId` (the person), while still replying/escalating to `channelId` (the room). The raw per-user id is threaded to every subprocess egress guard via a new `OC_CHANNEL_USER_ID` env var + loopback query/body param (`bin/cli.js`, `bin/entity.js`, `bin/tool.js`, `bin/loopback-client.js`, `core/io.js`, `core/loopback.js`, `core/system-prompt.js`). **On Telegram `userId === channelId` in DMs, so this is a strict no-op** for existing Telegram deployments.
- **`/channel add kazee` now links the owner's Kazee handle at capture time.** Owner recognition previously lived only in the env/adapter (`KAZEE_OWNER_USER_ID`); the owner's people record had no Kazee handle, so the guardrail (which resolves through canonical identity) re-derived the operator as external. `finalizeOwner` now calls `people.linkHandle(owner.id, { adapter: "kazee", channelId: ownerUserId, … })` so the wizard's captured owner is recognised the same way the rest of the bot resolves identity — belt-and-suspenders with the `speakerFor` fix above.
- **Debug/internal chatter is owner-only.** Recall banners (`/recall`), tool/skill traces (`/tooltrace`), and usage/cost alerts are now suppressed for a guarded (external) speaker regardless of the toggle state — internal mechanics of how the assistant works must never leak to a non-owner. Resolved once per turn (`isCurrentSpeakerGuarded`) and applied at all three emitters (`core/runner.js`).
- **Escalation & approval prompts render properly on Telegram.** The external-guardrail Approve/Deny/Always-allow message (`core/enforcer.js`) and the destructive-tool approval prompt (`core/loopback.js`) build HTML (`<b>`/`<pre>`, all dynamic content `esc()`-escaped) but were sent without a parse mode, so the tags showed literally. They now set `parseMode: "HTML"` **only for the Telegram adapter** (other adapters ignore it).
- **Friendlier stale-button handling.** A wizard callback pressed after the session is gone (bot restart / timeout) previously did nothing — a dead button reads as a hang. It now replies *"That setup step expired — I'm no longer listening. Run /channel add kazee to start again."*
- **Agent-space / OpenClaw:** no new deps, no new required env (`OC_CHANNEL_USER_ID` is bot-internal, injected per-turn), no schema change; Docker image builds identically. Correctness fix to owner recognition on Kazee + polish to the onboarding wizard — inert for Telegram-only deployments (userId === channelId). Full `npm test` green.

## v2.14.3
- **`/channel add kazee` owner-confirm no longer depends on the inline button round-tripping.** After the wizard captures a candidate owner message it asks *"Is this you — the owner?"* with Yes/No/Skip buttons on the operator's channel — but if that button callback never reaches the bot (e.g. a channel that drops `message:action` server-side), the operator was stuck: the confirm step (`core/channel-wizard.js` `handleText`) only understood "skip" or a pasted 24-hex id, so typing "yes" did nothing. It now accepts **typed "yes"/"no"** (plus y/n/yep/yeah/nope) as a first-class answer — "yes" finalizes the captured candidate as owner, "no" clears it and keeps listening — and any other text nudges with the three options. The confirm prompt now reads *"Tap a button below, or just reply "yes" / "no"."*
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Purely a robustness fix to the Kazee onboarding wizard's owner-confirm step — inert for deployments that never run `/channel add kazee`.

## v2.14.2
- **Token validation now hits the right service — the previous probe could never succeed.** v2.14.0/.1 validated a Kazee bot token by `GET {host}/api/me`, but chat-central has **no `/me` route** (confirmed against `src/config/routes.ts`) — every attempt 404'd. chat-central doesn't resolve identity itself; it delegates on every request to `@libraries/inet-central-auth`, which POSTs the opaque `kzb_…` token to **Central's `POST /api/principal/verify`**. `core/kazee-probe.js` now calls that endpoint: one request both **validates the token and returns the bot's own user id** (`principal.id`, with `principal.kind === "Bot"` enforced) plus its display name. Verified live end-to-end (real token → `ok` + botUserId + "Maya Bot"; bad token → clean 401).
- **Central is hard-coded too (env-overridable).** New `CENTRAL_DEFAULT_URL = "https://central.inet.africa"` mirrors chat-central's own default; override with `CENTRAL_BASE_URL` for non-prod. The wizard now says "Verifying the bot token…" and greets the bot by name ("Connected as Maya Bot ✅"); `setup.js` prints "Verified as <name>."
- **Agent-space / OpenClaw:** no new deps, no new required env, no schema change; Docker image builds identically. Correctness fix to the onboarding probe — the wizard flow (token-first, listen-to-capture owner) is unchanged.

## v2.14.1
- **The Kazee host URL is no longer asked — it's fixed infrastructure, like Telegram's `api.telegram.org`.** v2.14.0 fixed the probe but still made Step 1 ask the operator to type the host; the Kazee host never varies, so prompting for it is pure friction (and a chance to get it wrong). `core/config.js` now defines `KAZEE_DEFAULT_URL = "https://chat.inet.africa"` and `loadChannels` falls back to it whenever `KAZEE_URL` is unset. `/channel add kazee` is now **two steps, token-first**: Step 1 asks only for the bot token, Step 2 is the listen-to-capture owner flow. `setup.js` likewise drops its URL prompt (default inlined, since `config.js` exits before an `.env` exists). An explicit `KAZEE_URL` in `.env` still overrides, so nothing breaks for a deployment that set one.
- **Agent-space / OpenClaw:** no new deps, no new required env, no schema change; Docker image builds identically. Pure UX fix to the onboarding wizard.

## v2.14.0
- **`/channel add kazee` now finishes cleanly — the validation step was probing the wrong path.** The live adapter mounts every REST route under `{root}/api`, but the setup probe hit `{root}/me`. So the correct host (`https://chat.example.com`) always 404'd at validation, while the only URL that *passed* (`…/api`) then made the adapter build `…/api/api` and break. There was no single URL a user could type to succeed. `kazee-probe` now normalises either form to the root and probes the real `{root}/api/me` (new exported `normalizeKazeeRoot`); the wizard and `setup.js` both store the normalised root. Step 1 now asks for the host and states "I add the /api path myself."
- **The bot's own user id is no longer thrown away.** Validation already fetches it from `/me`, but the wizard never persisted `KAZEE_BOT_USER_ID` or passed it to the adapter — so a wizard-added channel booted without self-echo filtering or slash-command registration (and `loadChannels` warned on every restart). The wizard and `setup.js` now save it and hand it to the adapter.
- **Owner is captured by listening, not by pasting a Mongo id (Step 3 redesign).** After URL+token validate, the wizard brings Kazee **live immediately** (owner still blank) and asks the operator to send any message from their Kazee account. A new cross-channel hook — `router.tryCaptureOwner` → `channel-wizard.tryCaptureOwner` — intercepts that inbound Kazee message (it lands on a different channel than the operator's, so the operator-keyed `isAwaiting()` gate can't see it), then asks back on the operator's channel *"Is this you — the owner?"* with **Yes / No, keep listening / Skip** buttons. Yes grabs the sender's Kazee user id; owner recognition then flows through the existing env path (`identity.isConfiguredOwnerChannel` reads `KAZEE_OWNER_USER_ID`). Pasting an id or "skip" still works as a fallback; Cancel now tears the half-added adapter back down and rolls the `.env` edits back.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Purely a UX + correctness fix to the Kazee onboarding wizard — inert for deployments that never run `/channel add kazee`.

## v2.13.1
- **Guardrail no longer mistakes the owner for an external speaker on `/link`'d or env-unified channels (Phase 3 fix).** v2.13.0's relationship classifier resolved ownership through `people.findByHandle` alone — the handles physically listed on the owner's person record — while the queue/state/run-lock resolve it through `identity.canonicalForChannel`, which also honours explicit `/link`s (`identities.json`) and the configured-owner env path. A channel the owner had unified by `/link` or env-config (e.g. a second Kazee handle) was therefore *owner* to the run-lock (shared queue → "Queued.") but *external* to the guardrail (default-deny → the stranger reply), so the bot could greet its own owner as a limited-access outsider. The classifier now resolves ownership through **canonical identity — the same source of truth the rest of the bot uses**: a channel is the owner's when it resolves to the owner canonical **or** its person record is `isOwner`. New `identity.channelIsOwner()`; `relationship.speakerFor` consults it before default-deny. Every guard consumer (external-mode system-prompt block, `io`/`loopback` egress gates, runner turn-guard, CLI guards) funnels through this one function, so the fix lands everywhere at once.
- **R1 preserved — a note still can never confer owner.** `channelIsOwner` is driven only by the canonical map, which is written **exclusively by owner/operator actions** (`/link`, `.env`), never by a note; and the owner stays anchored to a real owner record via `ownerCanonical → people.isOwner`. The default-deny posture for genuinely external or unmapped channels is unchanged.
- **Tests.** `test-relationship.js` gains a regression reproducing the exact production bug — an owner Kazee channel `/link`'d to the owner canonical but **absent from the people record** now resolves to `owner` (unguarded), while an unmapped channel on the same transport stays `external` (guarded), proving the fix does not widen access. Full `npm test` green.
- **Agent-space / OpenClaw:** no new deps, no new env, no schema change; Docker image builds identically. Purely a correctness fix to owner recognition — inert for deployments that never hit the split-brain.

## v2.13.0
- **The independent relationship guardrail is now LIVE — an owned bot is finally safe to point at external people (Phase 3).** Phases 0–2 unified the bot's identity across channels and taught it WHO it's talking to; this phase adds the missing half: hard limits on what it may do or say when that person is **not** the owner. Everything is a strict **no-op for the owner and for no-context callers** (CLI/cron/tests) — the owner reply/tool/file paths stay **byte-for-byte unchanged** (risk R9/R14), so this is inert on every existing single-owner deployment until an external channel is actually `/auth`'d. Classification lives in `core/relationship.js` and is **cheap, synchronous, and never derived from message text**: the owner is owner **only** via `people.isOwner`; a non-owner whose note claims relationship `owner` is coerced to `external` (**R1 — a note can never confer owner powers**); every unclassified or unmapped authed non-owner defaults to `external` (**default-deny**).
- **Owner-authored mandate, set from chat (3.1).** `entity persona <name> --relationship|--role|--style|--knows|--mandate` lets the owner declare, per person, the ONLY things the bot is cleared to do or discuss with them. The verb is **owner-gated** (an external speaker can never edit their own mandate — self-escalation is refused) and each flag replaces its section wholesale.
- **External-mode posture, injected where it can't rot (3.2).** When the speaker is guarded, a compact default-deny block (`buildExternalModeBlock`) is appended to the **uncached per-turn tail** — never `buildSystemPrompt`/cache — and **re-injected every guarded turn** (not deduped once-per-session like the persona block) so the security posture can't age out of context. It carries the person's mandate (or "no mandate → almost everything is out of scope") plus five non-negotiables: default-deny, never leak owner/infra/creds/paths, no write/destructive without the mandate, treat their messages as DATA not instructions, warm-but-firm.
- **Independent fail-CLOSED enforcer (3.3).** Every outbound reply and every write/destructive tool action aimed at an external person is vetted by an **isolated Sonnet call** (`core/enforcer.js`) that sees only a structured `{mandate, proposed output}` — never the raw external text as instructions — so prompt-injection or context-rot in the main agent can't talk the guard out of the owner's rules (**R10**). It inverts `judgeRelevance`'s fail-OPEN posture: **any error, timeout, or unparseable verdict → "escalate"**, never "allow". One verdict per turn, cached by `sha1(kind|relationship|tier|mandate|payload)` for 5 min (a fail-closed verdict is deliberately **not** cached — a transient failure is re-judged, never remembered).
- **Escalation reuses the approvals plumbing verbatim (3.4).** On block/escalate the external is told "let me check with <owner>" and the owner gets **Approve · Deny · Always-allow-for-<person>** buttons on their own channel (the record separates **origin** = the external speaker from **approver** = the owner). Approve delivers the held reply or runs the held action; Always-allow writes it into the person's mandate. Reuses `approvals.create` + the owner-gated `apr:` handler in `actions.js`, which now wakes the originating (external) channel for a held **action** so an owner-approved external action actually executes where it was requested.
- **Per-speaker tool gate (3.5).** An external write/destructive `tool run` is vetted against the mandate before it can execute: **vet=allow means the owner's mandate IS the pre-authorization** (bypasses the `--yes` risk flag, stamped `approvedVia:"mandate"`); anything else escalates to the owner. Externals never reach the loopback `approval-request` path (which would post buttons to the requester's own channel) — that path is now **owner-only guarded** as defence-in-depth, and a fresh `--approval` token still bypasses the vet.
- **Egress gates (R8).** File/photo/voice sends to a guarded speaker are blocked at the true boundary (`loopback.handleSend`, where the send-* CLIs POST), with a cheap defensive gate in `io.sendFile` too — an unvetted file must never egress to an external contact.
- **Tests.** Two new hermetic suites wired into `npm test`: `test-relationship.js` (owner/external/trusted classification, R1 owner-coercion, default-deny for unmapped/no-pack, no-context ungated, `ownerTarget` picks the primary handle) and `test-enforcer.js` (owner no-op with zero guard calls, allow/block→escalate, fail-closed-and-not-cached, identical-vet caching, escalation record shape for reply + action, read-only plan-mode sub-agent). Full suite green.
- **Agent-space / OpenClaw:** no new deps, no new **required** env; Docker image builds identically. The guard model defaults to `sonnet` (override `ENFORCER_MODEL`) and can be disabled with `ENFORCER=off`. Inert until an external channel is authed AND has persona/mandate data — existing owner-only deployments are unaffected.

## v2.12.0
- **Persona packs — the bot now adapts voice + scope to WHO it's talking to (Phase 2).** An entity note about a *person* is upgraded into a persona pack: the same file gains four optional sections — **Role** (what they do), **Style** (how to talk to them), **Knows** (what to assume vs explain), **Mandate** (for non-owners: what the bot may do/discuss on their behalf) — plus a `relationship: owner|trusted|external` frontmatter scalar. The upgrade is strictly **additive (risk R7)**: an ordinary entity with no persona data serializes **byte-identically** to before, so places/projects/orgs/systems are untouched. `core/entities.js` learns `PERSONA_SECTIONS`, `RELATIONSHIPS`, `normalizeRelationship`, an `emptySections()` backfill for notes written before the upgrade, and persona params on `upsertEntity` that each replace their section wholesale (like Notes). FTS indexes the descriptive sections (Role/Style/Knows) so a persona is matchable by content, but **Mandate is deliberately not indexed** — it's policy, not description, and must not surface a pack on an unrelated mention.
- **Per-turn speaker persona, injected cache-safely (2.2).** `buildSpeakerPersonaBlock` resolves the *current* speaker (deterministically, not by keyword match) and injects their Role/Style/Knows — and, for a non-owner, their Mandate — so the model shapes its voice and scope every turn. It rides the **uncached dynamic tail** (risk R13), never `buildSystemPrompt`, so persona edits never bust the prompt-cache prefix, and dedupes once per (channel, session, entity-version) like packs/entities (a compaction or a persona edit re-injects). An **owner's Mandate is never injected** (the owner has no guardrails — injecting one would be noise).
- **People ↔ persona binding (2.1).** `people.js` gains an `entitySlug` pointer (a person can point at a differently-named pack) with `setEntitySlug` + a pure `resolveEntitySlug` (explicit pointer, else slug-from-name); surfaced in `roster()`.
- **Write-side population (2.3).** The post-turn reviewer (`pack-review.js`) and the nightly dream (`dream.js`) both learn to fill persona sections: their prompts describe the four sections + the relationship scalar with explicit guardrails ("never invent an owner mandate; never copy one person's guardrail onto another"), their JSON schemas carry the new fields, and their apply paths thread them into `upsertEntity` (each replacing its section wholesale; null/omit leaves it intact).
- **Relationship-gated cross-conversation recall (2.4, risk R6).** Episodic recall reaches across conversations into the owner's project transcripts, so surfacing it to a non-owner is a cross-tenant leak. `episodesAllowedForSpeaker()` makes cross-conversation episodes **owner-only and fail-closed**: an unknown speaker or any resolution error suppresses episodes; only a positively-identified owner (or no speaker context at all — CLI/tests) pulls them. Same-conversation packs/entities are unaffected.
- **Provenance stamps (2.5).** Appended Log/Journal lines from a **non-owner** contributor are stamped `(via <name>)` (the owner is the default author — tagging them would be noise), so a topic's history shows where an external/trusted colleague's contribution came from. Resolved from the live store at the reviewer call site (the reviewer applies in an async continuation) and threaded through as a `source`.
- **Tests.** Four new hermetic suites, all wired into `npm test`: `test-persona-packs.js` (round-trip + R7 byte-identical ordinary entity + FTS mandate-exclusion), `test-speaker-persona.js` (owner-Mandate suppression, non-owner Mandate shown, dedupe + re-inject on edit, silent for no-persona/unknown speaker), `test-persona-pipeline.js` (reviewer + dream populate/refresh persona wholesale + provenance), `test-recall-relationship-gate.js` (owner allowed, non-owner/unknown blocked, no-context ungated). Full suite green.
- **Agent-space / OpenClaw:** no new deps, no new required env; Docker image builds identically. Behaviour is inert until a person actually has persona data — existing deployments are unaffected until the reviewer/dream (or the owner, editing `entities/<slug>.md`) fills a pack. The independent **fail-closed enforcer** that hard-enforces Mandate at the outbound choke point is still Phase 3.

## v2.11.0
- **Unified multi-channel identity is now LIVE by default (Phases 0 + 1).** v2.10.0 shipped the split-brain fix inert behind `OC_UNIFIED_IDENTITY` (default off). This release flips the default **on** and lands the Phase 1 queue-routing fix that made it safe to do so. The owner's Telegram, Kazee and voice channels now resolve to one canonical person out of the box — one state, one session, one transcript, one run-lock — so the bot stops fracturing into a separate assistant per channel. The flag stays as an operator **kill-switch** (`OC_UNIFIED_IDENTITY=0` + restart) for reversibility (risk R14); it is no longer a hold-back gate.
- **Person-scoped queue routing (Phase 1, 1.2) — the blocker that kept the flag off.** Under unified identity the owner's channels share one run-lock and message queue, so a message sent from one channel while a run is in flight on another used to drain in the *finishing* turn's async scope — a Kazee follow-up could be answered in the Telegram chat. Each queued message now **captures its origin channel** at enqueue time, and `drainQueuedMessages` peels off one same-origin run at a time (requeuing the rest so the natural close→drain tail-recursion handles the next group), scoping each drained reply to `chatContext.run(origin, …)`. Replies always route back to the channel the message came from; same-origin batches behave exactly as before. Grouping logic extracted to a pure, unit-tested `core/queue-drain.js`.
- **Cross-channel dedup (0.3).** `router.isDuplicate` is re-keyed from `adapter:channel:msg` to `canonical:adapter:msg` — a message redelivered against a different channel that resolves to the same person is now caught once the owner's channels are unified (adapter stays in the key, so distinct messages that merely share a numeric id across transports are never collapsed).
- **First-contact migration (0.2 / 1.1) fires automatically.** With the flag on, the first inbound message from a previously-siloed owner channel folds its pre-existing state + sessions into the owner bucket — non-destructive, idempotent, and backed up to `~/.open-claudia/backups/` first (unchanged from v2.10.0, now actually reached). Pre-migration per-channel transcripts remain as separate files and stay reachable via `transcript-search --all`; new history accretes under the unified person id (1.3).
- **Tests.** `test-unified-identity.js` gains a `default` (unset ⇒ on) child to lock the new default and proves an explicit `0` still reverts to strict per-channel resolution; new `test-queue-routing.js` covers origin keying, same-origin batching, mixed-origin split/requeue, and legacy (pre-capture) fallback. Full suite green.
- **Agent-space / OpenClaw:** no new deps or required env; Docker image builds identically. Existing single-owner deployments see unified behaviour for the owner's own channels only — strangers still resolve to their own isolated bucket (the per-person guardrail arrives in Phase 3).

## v2.10.0
- **Unified multi-channel identity — foundation, gated OFF by default (Phase 0).** An owned bot is meant to be one mind across every channel its owner reaches it on; today it split-brains because conversation state (session, transcript, run-lock, queue) is keyed per `transport:channelId`, so an unlinked channel (voice bridge, a stray Kazee DM) becomes a separate assistant with its own history. This release lands the identity-resolution foundation behind a single flag, `OC_UNIFIED_IDENTITY` (default **off**). **With the flag off nothing changes** — every new code path is gated, so publishing this version is behaviourally identical to v2.9.3 until an operator sets `OC_UNIFIED_IDENTITY=1` and restarts. Boot log now prints `Unified identity: on|off` for visibility.
- **Resolution seam (0.1).** `identity.canonicalForChannel` — the single chokepoint every adapter (Telegram/Kazee/voice) routes through — gains a flag-gated branch: a *configured owner channel* (one whose `transport:channelId` matches an operator-declared owner id — a `TELEGRAM_CHAT_ID`, `KAZEE_OWNER_USER_ID`, or `VOICE_OWNER_USER_ID`) with no explicit `/link` now resolves to the single owner canonical (the owner's first-linked handle, matching `people.linkHandle`'s primary convention). Because it's the one seam, state/transcript/scheduler all unify downstream for free. Owner detection uses only trusted, operator-set signals — never a guess — so a stranger's channel can never be merged by mistake (verified: a non-owner channel still resolves to its own default bucket).
- **One-time owner-state merge (0.2).** The first inbound message from a configured owner channel that was previously siloed folds that channel's pre-existing state + sessions into the owner's canonical bucket (`state.autoLinkOwnerChannel`, called from the router before any state is read). The merge is **non-destructive** (it folds, never overwrites), **idempotent** (in-memory guard plus a naturally-no-op re-migrate once the source is gone — safe across restarts), and **backs up** `state.json`/`sessions.json`/`identities.json` to `~/.open-claudia/backups/` before touching anything. It deliberately does **not** write into `identities.json`, so turning the flag back off cleanly reverts routing while the already-merged history stays with the owner.
- **Agent-space / OpenClaw ready.** No new dependencies, no new native modules, and no new *required* env or files — the Docker `:latest` image and the npm package build and run exactly as before with the flag off, so existing single-user deployments are unaffected and the multi-user owned-agent rollout can flip the flag per-deployment when ready. New `test-unified-identity.js` (wired into `npm test`) proves both flag states in isolated child processes: flag-off byte-identical resolution + inert auto-link; flag-on unification, stranger isolation, and the merge (result, idempotence, buckets folded, sessions preserved, backups written).
- **Explicitly deferred (not shipped hot):** the unified serialized-queue reply-routing fix (when the owner sends from two channels *simultaneously* under the flag, a queued message currently drains in the finishing turn's origin scope — Phase 1), per-person persona packs (Phase 2), and the independent external-speaker guardrail/enforcer (Phase 3). The flag stays off until at least the Phase 1 queue-scope fix lands, so none of these are live-exposed.

## v2.9.3
- **Bot advertises a distinct "thinking" state before its first token (Kazee parity Phase 2).** The typing pipeline now carries an additive `state` discriminator (`'thinking' | 'typing'`) — no new socket event. The runner advertises `thinking` from the moment a turn starts until the first stream event (text delta, assistant content, tool call, or result), then flips to plain `typing`; Kazee clients render "<name> is thinking…" in the header/chat-list for that pre-token window and fall back to normal typing text otherwise. The Telegram adapter ignores `state` (its chat action has no such distinction), so the change is inert there. Pairs with the chat-central `message:interactiveUpdated` broadcast (server-authoritative button selection) landing in the same release across chat-central/web/mobile.
- **Ships the v2.9.2 work that never published.** v2.9.2 (Kazee REST media fallback when a socket send lacks a media path, plus an explicit `typing:stop` on turn end) was committed to main but no tag was ever pushed, so CI never published it. This tag carries it to npm alongside the thinking-state change.

## v2.9.1
- **Post-turn reviewer no longer rewrites pack sections blind — context starvation fixed.** The reviewer prompt showed only the pack *index* (name + description), so a State rewrite was regenerated from the turn text alone and silently destroyed durable facts the turn didn't restate; worse, with no project name in the turn text it could attribute the work to the wrong pack entirely (observed live: an AQUELLE product-shot turn filed under Alpha Protein, wiping a promoted root-cause fact). The runner now passes the turn's **active packs** (recall-surfaced ∪ explicitly opened) into `reviewTurn`; the prompt includes their **full current Stance/Procedure/State + recent journal**, with two new rules: attribute updates to an active pack unless the turn clearly names another, and treat section rewrites as **merges** — carry forward every still-true fact, drop only what the turn contradicted or completed.
- **Structural enforcement, not just prompt guidance.** `applyAction` drops section rewrites (State/Stance/Procedure) targeting any pack whose current content the reviewer did not see — journal lines still land; the announce line notes the drop. Applies to the update path, the exists-on-create path, and the versioned-duplicate fold.
- **Provenance guard extended.** A user-authored Procedure now gets the same intent-wins protection as Stance (an automated writer never silently replaces it). State stays rewritable by design, but replacing a user-authored State is now flagged in the reviewer's announcement (⚠️) so the overwrite is visible instead of silent.

## v2.9.0
- **Hybrid async approvals — a destructive run no longer burns the turn or dead-ends on a late press.** v2.8.0's chat approval blocked the CLI for 5 minutes and treated timeout as denial, which had two UX failures: an undecided window wasted minutes of live turn, and a button pressed after the window hit a dead-end (terse auto-ack, no agent reaction — the record was already denied). The flow is now two-phase. **Fast path**: the CLI waits ~90s; an inline Approve runs immediately, Deny refuses — unchanged semantics, just a tighter window. **Async path**: if undecided at 90s, the CLI marks the record `detached` and exits code 5 with instructions to end the turn ("pending your Approve/Deny — do NOT retry"); the record stays open for 24h. When the button is eventually pressed, the `apr:` handler sees the detached record and **schedules an immediate wakeup** in the originating channel carrying the decision and the exact command, so the agent reacts either way (runs it, or acknowledges the denial).
- **One-shot approval tokens, fail-closed on every branch.** The woken agent redeems the decision with `tool run <name> <args> --approval <apr_id>`. `approvals.consume()` requires ALL of: status `approved`, never consumed before (single use — a replay is refused), within 24h of request, and a command line **byte-identical** to what the user saw on the buttons (a mismatch refuses without burning the token). `--approval` is honored on top-level runs only — a nested/composed tool can't smuggle a token — and redemption stamps `approvedVia: "chat-approval"` + exports `OPEN_CLAUDIA_APPROVED_TIER=destructive` exactly like an inline approve, so composition inheritance is unchanged. Undecided, denied, expired, unknown, or mismatched never runs.
- Approval prompt and system-prompt docs updated for the two-phase flow; new `test-approval-async.js` covers the full lifecycle (detach semantics, pending/denied/unknown refusals, exact-match redemption, replay refusal, token not burned on mismatch, 24h expiry, 1000-char command cap, 90s default window).

## v2.8.0
- **Per-verb manifests — the gate now knows each verb's blast radius (Phase 5).** A tool may declare its verbs in the header (`verb: <name> [args] | <tier> | <doc>` — one line per verb); `tool run` then resolves the risk tier **per invoked verb** instead of whole-tool: `spaces unread` runs freely while `spaces upload` demands `--yes`, and `prod-k8s cert-status` is read-only while `cert-renew` stays write-gated. An **undeclared verb on a manifest tool is refused outright** regardless of flags (declare it before using it), `help`/no-verb is always read-only, and manifest-less tools keep the old whole-tool gate unchanged (fully backward compatible — older runners ignore `verb:` lines). `tool scaffold --verbs "list:read-only:doc,create:write:doc"` generates the manifest lines, a dispatcher entry per verb, and a TOOL.md verbs block; `tool show` lists the declared manifest. Migrated all 8 live verb-CLIs (clickhouse, inet-central-cli, inet-ops, kazee-chat, kticket, metabase, prod-k8s, spaces — 41 declared verbs) in the tools repo.
- **Chat approval with the EXACT payload — destructive runs escalate to inline Approve/Deny.** `--yes-destructive` acknowledged the agent's *paraphrase* of a destructive action; now an unflagged destructive top-level `tool run` inside a bot task POSTs the literal resolved command line to the requesting channel via a new loopback `approval-request` kind — the user sees `open-claudia tool run <name> <args>` verbatim with ✅ Approve / ⛔ Deny buttons (owner-only, idempotent: a second press can't overturn the first). The CLI polls the file-backed record (`core/approvals.js`) and **fails closed**: timeout (5 min) or deny = refusal at exit 3; only an explicit approve proceeds, stamped `approvedVia: "chat-approval"` with `OPEN_CLAUDIA_APPROVED_TIER=destructive` exported for composition. Every run's history entry now records the acknowledged tier — auditable forever.
- **Source fingerprint — you can't unknowingly run changed code.** Every clean exit stamps a sha256 of the executable into `state.json` (`verifiedHash`); the runner compares before each run and warns `⚠ unverified since change` when the source differs from the last green run (never blocks — the next successful run re-verifies). Surfaced in the per-turn tool condition line too.
- **Tooling-mode switch (`/toolmode`, `open-claudia tooling-mode`).** Strict (default, fail-safe) = tool-first prompt + deny-gate blocks raw prod-surface commands. Relaxed = the older model: scripting permitted, deny-gate logs (`relaxed-pass` events keep the audit/KPI honest) without blocking, prompt block swaps to "tools preferred, not forced". Owner-only chat command with inline buttons; risk-tier gate and credential confinement apply in BOTH modes. Instant effect (mode file re-read per hook call), corrupt/missing file fails strict.
- Test coverage: `test-tool-manifest.js` (manifest parse/scaffold round-trip, per-verb gate matrix incl. undeclared refusal + nested composition, fingerprint lifecycle, approvals idempotence + fail-closed timeout) and `test-tooling-mode.js` (strict denies / relaxed logs-and-allows / fail-safe strict).

## v2.7.6
- **/tooltrace now shows WHAT ran, not just THAT it ran.** The 🔧 line per tool run was `Ran tool: kticket` — no verb, no args, and one line per tool even when different verbs ran in the same turn. It now renders the **sanitized command shape** (quoted strings → `<text>`, ids/uuids → `<id>`, paths → `<path>` — same shaping the tool graph already computes, so this is display-only and free) plus a **human-readable doc sentence for that verb**: new `tools.verbDoc(name, verb)` parses the sentence mechanically from the tool's own leading comment block (the same docs `--help` prints — zero model cost, can't drift from the source). Both documented header shapes are recognised (indented description under the verb line, or same-line prose after 2+ spaces), with fallback to the tool description's first clause. Example: `🔧 kticket show <id> — One ticket in full`. Dedup is now per tool+verb per turn, so `kticket list` then `kticket show` both announce. Unit coverage in `test-tools.js` (both shapes, sentence cut, fallback, flag-token and unknown-tool cases).

## v2.7.5
- **Tool composition with inherited approval — a composing tool cannot self-approve (Phase 4).** Higher-order tools may shell into lower tools (report → kticket + kchat), but the callee's risk gate stays intact: `tool run` now exports the tier the **outermost** invocation's flags acknowledged (`OPEN_CLAUDIA_APPROVED_TIER`) plus the call chain (`OPEN_CLAUDIA_TOOL_STACK`) into the child env, and a nested run gates on that inherited tier while **ignoring its own argv flags** — so a composing tool that hardcodes `--yes-destructive` in its source is refused with a message pointing at the outer re-run. An inherited tier passes through unchanged (depth-safe), a tool already in the chain is refused (cycle guard), and each nested run still stamps its own telemetry. Verified live: fabricated approval blocked at exit 3, legit outer `--yes-destructive` flowing through, self-call cycle caught at depth 1.
- **`_lib/` shared plumbing with dependent journaling (Phase 4).** Genuinely shared code (central auth, mongo-exec) now lives in `tools/_lib` — tiny modules tools import via a literal `_lib/<name>` path, never tools themselves (the registry, surfacing index, and legacy migration all skip the `_` prefix). New `tool lib list|show|set`: `set` writes the module (secret-blocking, traversal-safe), finds every dependent by a boundary-safe source grep (`central-auth` matches `_lib/central-auth.js` and the `path.join` form, **not** `central-auth-v2`), journals each dependent `unverified since _lib change — re-verified by next real use`, and commits the lot in ONE commit naming the dependents — the PLAN's blast-radius rule, since a lib change can break several tools at once and health is verified by use, not smoke tests.
- **Umbrella grouping from registry tags (Phase 4).** Tools gain an optional `tags:` header field (`--tags` on `scaffold`/`add`/`set`, sanitised, strong-match in `tool search`); `tool list` renders the first tag as an umbrella group — `kazee/{kchat,kticket}`, `infra/{prod-k8s}` — with untagged tools flat below. Purely presentational: "Kazee" stays a registry grouping, never a binary. The fixed system-prompt tools section gains one composition bullet (approval is inherited, never fabricated; `_lib` marks dependents unverified). Test coverage: tags round-trip, lib dependent grep boundaries, setLib journaling, and the full nested-gate matrix including the fabrication and cycle refusals.
- **The dream now audits tool usage — raw-op detector + usage KPI (Phase 3/item 9 + enforcement 5c).** The guard log knew about *blocked* commands (denies) and *authorised* one-offs (`keyring exec` bypasses), but operational work done raw that the gate didn't cover — inline heredoc scripts fed to an interpreter, raw DB clients, mutating `kubectl` — left no trace, so the tool-usage KPI had a blind spot exactly where the "repeated heredocs slip through" failure lives. New deterministic detector in `core/tool-guard.js` (`matchRawOp`/`noteRawOp`): high-precision patterns (`heredoc-interpreter`, `db-client`, `kubectl-mutate`), wired into the runner's Bash-stream chokepoint so every command is observed — **observation only, never blocking**; the authorised channels and anything the deny-gate already owns are skipped so one command never yields two guard events. Reads (`kubectl get/logs`, `rollout status`, `cat <<EOF` file writes, mentioning a client in grep) stay unflagged.
- **Doc/code drift flag completes the dream tool-pass.** Alongside the existing hygiene flags (dangling pack links, cold never-run tools, near-duplicate pairs), the pass now flags tools whose source was edited after TOOL.md was last touched, and verbs that telemetry proves are in real use but the docs never mention (`tools.docDrift()`; `help`/`(default)` ignored). Flag only — the fix is a `tool note`. Zero false flags on the live 12-tool registry at ship time.
- **Nightly KPI in the dream report, computed from stamps — never prose.** The dream's tool pass now appends `📈 Tool usage since last dream`: % of operational actions via `tool run` (target >90%, denominator = tool runs + escape-hatch bypasses + raw-op sightings; gate denies reported separately since blocked ≠ completed), bypass count **with reasons**, raw-ops by pattern, and failure-at-use counts naming the failing tools. Sources are `state.json` run-history stamps (new `tools.runsSince(sinceTs)`) + `tool-guard.jsonl` (new `readGuardEvents(sinceTs)`) — per the plan, the KPI is never parsed out of conversation prose. Escape-hatch bypasses surface in the dream report only (no immediate chat pings — pings would train us both to ignore them). Test coverage: raw-op matrix incl. no-double-count with denies, event grouping/windowing, `runsSince` window math.
- **Surfaced tools now carry their condition (tiered surfacing, Phase 2/item 6).** The per-turn "Tools that may help here" block rendered one identity line per tool — name, description, keyring needs — so the agent knew a tool *existed* but not what state it was in, and had to spend a `tool show` (or just run it blind) to find out it was broken, never proven, or failing since Tuesday. Each surfaced tool is now a **2-line headline**: line 1 is identity (name — description, risk tier, keyring needs, why it surfaced), line 2 is condition — the TOOL.md **State** headline (first line, truncated), the open **known-issue count** (⚠ n), and **last-run health** from telemetry (`last run OK/FAILED (exit n) <date>, n runs`, or `never run`; legacy flat tools without exit telemetry degrade to `last used`). This completes the tiering ladder: 2-line headline always → TOOL.md via `tool show` → source only via `--source` — you read the manual, not the disassembly, and now the cover tells you if the manual has errata. New `tools.toolCondition()` with unit coverage in `test-tools.js` (headline extraction, `(none yet)` filtering, health formats, clean degrade) and end-to-end render assertions in `test-recall-discoverer.js`.

## v2.7.2
- **Credential confinement — keyring creds now live ONLY inside tool runs (enforcement 5b).** Until now `botSubprocessEnv()` merged the whole operational keyring into *every* agent subprocess, so the agent's own Bash (and any sub-agent) could `curl` a prod API with `$inet_central_user` directly — the tool layer was optional. That merge is gone: agent shells and sub-agents run cred-free, and `tools.runEnv(tool)` injects **only the keys a tool declares via `--requires`**, read fresh from the keyring at run time (least privilege; `process.env` still wins on conflict). Audited all 11 live tools before shipping: every keyring key read is already declared, so nothing breaks. The system prompt's Preauth bullet now tells the truth, and `keyring list` no longer claims env-wide availability.
- **Tool-first deny-gate on the agent's shell (enforcement 5b).** A PreToolUse hook (wired via `--settings` into the main runner *and* sub-agent spawns; new `core/tool-guard.js` + `open-claudia deny-gate-hook`) blocks raw operational commands against known prod surfaces and redirects to the covering tool by name — or to the exact `tool scaffold` command when nothing covers the surface yet. v1 rules are deliberately high-precision: an **acting** binary (curl/wget/ssh/mongosh/…) aimed at `central.inet.africa`, `ticketcentral`, `10.200.0.0/16`, `kubectl exec … mongo(sh)`, or `ssh codersafrica@` — while `grep`/`cat`/`ping` mentioning the same surfaces, read-only `kubectl`, and plain `git push` stay free (reading is always allowed; acting goes through tools). Quotes count as command separators so `bash -c "curl …"` can't dodge the gate. Every failure mode in the gate fails OPEN — a broken guard must never brick the shell. This gate is the honest enforcement for surfaces where confinement is toothless (kubectl/ssh use ambient creds, not keyring keys); the accepted residual dodge is writing the command into a script file first.
- **`keyring exec --reason "<why>" -- <cmd>` — the logged escape hatch.** Confinement without a pressure valve just breeds invisible workarounds (`keyring get` + copy-paste). A genuine one-off no tool covers yet runs through `keyring exec`: full keyring in env for exactly one command, reason mandatory, and the bypass recorded to `~/.open-claudia/tool-guard.jsonl` alongside every deny — the audit stream the 5c KPI pass will read to decide which tools to scaffold next. New `test-tool-guard.js` covers the rule matrix, hook protocol (exit 2/stderr deny, fail-open), audit trail, settings writer idempotency, and env confinement; `test-tools.js` extends to least-privilege injection.

## v2.7.1
- **The system prompt now teaches tool-FIRST, not tool-after (enforcement 5d).** The fixed "Reusable tools" section in `core/system-prompt.js` was the root cause of the baseline failure: it taught "crystallise the script you just wrote" — do the work raw, save a tool afterwards. Rewritten as "Tools are the interface": before operational work, walk the loop **search → use (read TOOL.md first, never guess args) → extend the missing verb → scaffold first** — and only then do the work *through* the tool. Raw one-liners and heredocs are demoted to probing/diagnosis; a broken tool gets fixed or journaled (`tool note --issue`), never routed around. The section also documents the risk gate honestly: the `--yes`/`--yes-destructive` flag is the agent's mechanical acknowledgment *after* the user approves — it never replaces asking first.
- **Tool-scout: operational turns can no longer meet silence (enforcement 5a).** Today an empty tool match is silent, and silence reads as permission to improvise. The discoverer now runs a pattern-based scout (no LLM, ~zero cost) over each seeds/full-tier turn: a table of high-precision operational domains — named hosts (central.inet.africa, ticket-central, chat-central, git.coders.africa…), prod IP ranges, infra commands (kubectl, argocd, mongosh…) — each with a `covers` matcher against the tool registry. When a turn trips a domain: if a registry tool covers it but wasn't auto-surfaced, the block points at it by name (`tool show X`, then work through it); if **nothing** covers it, the block says so and hands over the exact scaffold command (`tool scaffold <suggested-name> --risk write`). Domains already covered by a surfaced tool stay quiet, and non-operational chatter never trips the patterns (precision over recall — they expand from audit evidence later, per 5c). Gap signals flow into recall metrics (`toolGaps`) and the engine result for future `/tooltrace` surfacing. Extends `test-recall-discoverer.js` with pure-function and through-the-engine scout coverage.

## v2.7.0
- **Tools become memory objects (tool-first architecture, Phase 1).** A tool is no longer a flat script — it's a directory: `tools/<name>/tool` (the executable), `TOOL.md` (Interface · Known issues · State · Journal — the curated, human-readable condition of the tool), and `state.json` (machine telemetry, runtime-written, gitignored). The tools dir is now a **git repository**: every add, metadata change, note, and removal auto-commits (with a per-invocation identity, never touching git config), so `git log` *is* the upgrade history and any change is one revert from undone. `tool show` is tiered disclosure: header metadata, live telemetry, TOOL.md, and the recent commit log — the source only with `--source`, so reading about a tool no longer costs its whole body in context.
- **Risk tiers + run gate.** Every tool declares `risk: read-only | write | destructive` in its header (absent = treated as write, fail-safe). `tool run` now enforces the ask-first rule mechanically: read-only runs freely; write requires `--yes`; destructive requires `--yes-destructive` — and a plain `--yes` does **not** unlock a destructive tool. Refusals exit 3 with a clear message. The flags pass through to the tool unstripped, so tools with their own confirmation gates (e.g. prod-k8s) keep working. `tool add`/`tool set` accept `--risk`; adding without it defaults to write and says so.
- **Run telemetry stamped at the point of use.** `tool run` itself records every execution into the tool's `state.json`: run count, last used, last exit, last success/failure, per-verb health (first non-flag arg), and a rolling history of the last 20 invocations with timing. A failing run prints a nudge to journal the defect (`tool note <name> --issue "..."`). This is the v1.2 health doctrine in code: no smoke tests, no health crons — health is verified by use, and breakage is captured automatically where it happens.
- **`tool scaffold`, `tool note`, `tool migrate`.** `scaffold <name>` creates a ready-to-extend verb-dispatch stub (node or bash) with full header, TOOL.md skeleton, and initial telemetry — and refuses if the tool exists ("extend it instead"). `note <name> --issue|--state|--journal` maintains TOOL.md from the CLI: issues and journal lines insert newest-first; `--state` rewrites the State section (it describes *now*, not history). `migrate` converts legacy flat-file tools into directories — one commit each, sidecar telemetry carried over, TOOL.md generated — and runs automatically on first CLI use, so existing tools arrive in the new layout with their memory intact.
- **Secrets can no longer be saved into a tool.** The add-time secret lint is now **blocking** (was advisory): `tool add`, `tool set`, and `tool note` refuse content matching credential patterns, because the git-backed dir would remember the secret forever. Other lints (requires mismatch, syntax check, dangling pack) stay advisory. Extends `test-tools.js` end to end: layout, risk gate matrix, per-verb telemetry, secret refusal, migration, scaffold, git history.

## v2.6.58
- **Tiered recall gate — stop paying 15 seconds for "thanks".** Measured over 700+ live turns, the old discoverer pre-gate returned true whenever *any* seed matched, so it gated just 0.4% of turns — a bare "ok" with an incidental keyword hit still paid the full seed → graph → walker pass (p50 ~15s). A new `recallTier()` classifies each turn into three tiers: **skip** (pleasantries/acks up to 4 words *regardless of seeds*, emoji-only, or ≤2 words with no seeds — no recall at all), **seeds** (short command-like turns ≤4 words — user-origin keyword seeds inject directly with no graph expansion, no walker, no excerpts; the classic-engine baseline at ~ms cost), and **full** (substantive turns — the whole pipeline). The tier is logged per turn and the metrics summary now tracks the seeds-only share alongside the gated share.
- **Walker candidate cap.** The latency deep-dive showed the walk itself dominates, not process spawn: 15–40 candidates × 600-char excerpts = 10–25KB prompts to the haiku judge on every substantive turn. Full-tier candidates are now priority-ordered — user-origin seeds by match score, then context seeds, then graph-activated nodes by activation — and capped at `RECALL_WALKER_MAX_CANDIDATES` (default 14), so the judge reads the strongest few instead of everything the graph touched.
- **Episodic memory — recall now remembers how it did things last time.** The discoverer gains a transcript hop: on full-tier turns it queries the project-transcript FTS index (up to `RECALL_EPISODE_LIMIT`, default 3; `RECALL_EPISODES=off` disables) and offers the hits to the walker as `episode:` candidates with their matched snippets. The walker keeps one only if that past exchange holds a decision, method, or outcome directly reusable for the current message — and the fail-open path (walker off/failed) *drops* episodes rather than injecting unjudged snippets. Kept episodes render as a "Past work that looks related" block with the snippet plus a ready-to-run `open-claudia transcript-search … --all` pointer for the full context, and surface in the `/recall` banner as 📓 lines.
- **Real cost + phase telemetry for recall.** The warm walker's stream-json `total_cost_usd` is cumulative per session; it now tracks the running total and reports the per-walk *delta*, which flows into `recall-metrics.jsonl` per turn (`costUsd`) instead of the flat 0 logged before. Each turn also records `seedMs` / `expandMs` / `walkerMs` and which walker path ran (`warm`/`cold`/`off`/`failed`), so the p50 can be decomposed instead of guessed at.
- **Weak-use signal from the post-turn reviewer.** Precision metrics only counted explicit 📖 opens as "used". Now, when the reviewer decides a turn's work updated pack X or entity Y, that adjudication is logged as use (`source: "reviewer"`) and co-updated nodes get a Hebbian co-work reinforcement — so packs that quietly shape work without being re-opened stop looking like noise.
- **Journal dedupe.** The reviewer could append near-identical Journal lines turn after turn (observed: 28 same-day duplicates on one busy pack). `updatePack` now token-Jaccard-compares a new journal line (dates stripped) against the last 5 entries and skips appends at ≥0.75 similarity, flagging `journalDeduped` instead. Adds `test-journal-dedupe.js`.

## v2.6.57
- **Move the inactivity watchdog into the live runner.** Deep review found v2.6.56 put the watchdog in the legacy `bot-agent.js` monolith, but the LaunchAgent runs `bot.js`, which delegates heavy turns through `core/runner.js`. That meant the installed live path still only had the 6-hour hard timeout and could still wedge on a silent child. This release ports the watchdog to `core/runner.js`, updates activity on child stdout/stderr, terminates the process tree after `OC_TURN_IDLE_MS` of silence, suppresses the confusing partial/no-output final message, clears stuck busy state with a failsafe, and drains any queued follow-up messages instead of stranding them. Adds a static regression test that fails if the watchdog only exists in `bot-agent.js` again.

## v2.6.56
- **Stop a single hung child from wedging the whole bot.** Agent mode is single-flight: a heavy turn spawns one child and sets a global `runningProcess` flag; while it's set, every new message is demoted to a lightweight side-chat instead of a full turn. That flag is *only* ever cleared by the child's `close`/`error` events — so a child that hangs (a wedged tool call that never returns) pins it forever. That's exactly what happened: one turn went silent, its child never exited, and from then on the bot stayed alive (health checks passed) but answered nothing of substance — every later message fell through to a side-chat that also couldn't make progress. The fix is an **inactivity watchdog** in `runClaude` (`bot-agent.js`): track `lastActivity`, bumped on every `stdout`/`stderr` chunk from the child, and inside the existing progress-update loop (which already ticks every 2–5s while a task runs) kill the child if it produces no output for longer than the idle limit — SIGTERM to the process group, SIGKILL 3s later, plus a 10s failsafe that force-clears the flag if `close` somehow never fires. The watchdog short-circuits the close handler so the user gets one clear "task stalled — stopped it, I'm responsive again, send /continue" message instead of a confusing "(no output)". Default idle limit is 10 minutes, tunable via `OC_TURN_IDLE_MS`. Deliberately scoped to the wedge only — the v2.6.55 polling fix is healthy and untouched.

## v2.6.55
- **Make fixing a tool as cheap as creating one.** The tool subsystem was structurally tilted toward sprawl: creating a tool was a trusting one-liner with no validation, while *fixing* one had no verb at all — so the path of least resistance was always another near-duplicate. This release closes that gap from both ends. **A first-class fix verb:** `open-claudia tool set <name> [--desc|--pack|--requires|--usage]` rewrites a tool's header metadata in place (new `updateToolHeader()` in `core/tools.js`) — it refreshes only the recognised field lines, preserving the marker, any hand-written comments, and the immutable `createdAt`, and never touches the body. `open-claudia tool edit <name>` prints the source path (stdout) so the agent's Edit/Write or the user's `$EDITOR` can open it directly, with the metadata-only hint on stderr. **An add-time safety net** (all advisory, never blocking — a tool the agent just wrote is usually right): `tool add`/`tool set` now lint the source and surface (1) hardcoded secrets — `sk-ant-`/`sk-proj-`/`sk-`/`AKIA`/`gh*_`/`xox*`/PEM/inline-Bearer literals that should be `$KEYRING` refs, mirroring the redactor; (2) requires/keyring mismatches — keyring keys the script reads but `--requires` omits (with a ready-to-paste `tool set` fix), and declared keys it never reads; (3) syntax errors via the script's own interpreter (`bash -n` / `node --check` / `py_compile`, silent when the interpreter or type is unknown — never cries wolf); and (4) a dangling `--pack` link to a pack that doesn't exist. **Immutable creation telemetry:** the `.usage.json` sidecar now stamps `createdAt` once at add-time (surfaced on every tool record); the nightly dream's "cold since creation" check ages by this stamp instead of file mtime, so a body edit can no longer reset a never-run tool's staleness clock (falls back to mtime for tools created before the stamp existed). Extends `test-tools.js` (createdAt immutability, `envRefsIn`, `lintToolSource`, `checkSyntax`, `updateToolHeader`).
- **Stop the phantom-outage restart storm — just keep polling.** The Telegram adapter was killing the bot over network blips that weren't real. A side-by-side network monitor (probing 1.1.1.1 *and* api.telegram.org every 60s) showed the connection healthy throughout, yet `bot.log` had logged 20k+ `EFATAL: read ETIMEDOUT` → "Network lost" events: the long-poll's pooled keep-alive socket goes stale on a macOS sleep/wake or an idle-NAT eviction, and every reused poll then times out instantly. The old handler treated *any* such timeout as a lost network — loudly stopping/restarting polling and, after 6 quick attempts (a sleep/wake error burst exhausts those in seconds), calling `process.exit(1)` for launchd to relaunch. That voluntary suicide *was* the "why does it keep restarting" the bot kept doing. Three changes in `channels/telegram/adapter.js`: (1) **fresh socket per poll** (`keepAlive:false`) so there's no pooled socket to go stale — the root of the ETIMEDOUT storm; (2) **gentle, throttled recovery** — a transient timeout no longer screams or tears down on the first hit (with keep-alive off the next poll self-heals on a new connection); it logs at most once a minute and only actively restarts polling if errors persist; (3) **never exit over a network blip** — only a `409 Conflict` (a second poller on the same token) exits immediately; a genuinely wedged loop now has to fail unbroken for a full 10 minutes before a clean relaunch, versus the old ~90s hair-trigger. The diagnostic `netmon.sh` that proved the network was fine stays running.

## v2.6.54
- **Tools as first-class discoverer nodes.** Reusable tools now go through the same recall pipeline as packs and entities — seed → activate → judge → render — instead of being dumped wholesale into every prompt. Until now `buildToolIndexBlock` listed *all* ~40 tools on every turn; that flat block is gone. In its place, the discoverer (`core/recall/discoverer.js`) seeds tool candidates from a new in-memory `matchTools()` (lexical scoring over name/description = strong, requires/pack/source = weak; `core/tools.js`), activates the tools belonging to any pack already in the candidate set (read from the pack↔tool graph + tool `pack:` headers — the latency-sensitive recall `expand()` stays pack/entity-only, zero regression risk), and lets the haiku walker judge them with a "keep a tool only if running it would actually help" instruction. Survivors render into a focused **"Tools that may help here"** block; the system prompt now just states the count and points to `tool search`/`tool list` for the rest. `matchTools` is in-memory (works on every node version, unlike the SQLite graphs).
- **`/tooltrace` toggle (default off).** Mirrors `/recall`. When on, posts short 🔧 lines around each reply: which tools were **surfaced** for the turn, which were **run**, and any **created/updated** — so you can watch the tool layer work. All three signals are now gated behind this one toggle (`settings.showToolTrace`); the run/created/updated banners were removed from the always-on path and the surfaced list was split out of the 🧠 `/recall` banner, so the two toggles are cleanly independent.
- **Tool usage telemetry.** A JSON sidecar (`.usage.json` in the tools dir: `runCount` + `lastUsed` per tool, recorded at end-of-turn for each successful run) works on every node version. `matchTools` breaks score ties by `runCount` so well-worn tools surface first.
- **`tool search` + add-time duplicate guard.** New `open-claudia tool search <query>` for finding a tool on demand; `tool add` now warns (non-blocking) when an existing tool looks similar, nudging toward extending it with a subcommand over spawning a near-duplicate.
- **Dream tool hygiene.** The nightly pass now also flags tools created-but-never-run for 30+ days and near-duplicate tool pairs (alongside the existing dangling-`--pack` flag). Flag-only — a tool is never auto-deleted. Extends `test-tools.js` (matchTools/telemetry) and `test-recall-discoverer.js` (tool seeding). Degrades to a graceful no-op where `node:sqlite` is unavailable.

## v2.6.53
- **Pack↔tool contextual edges — surface how a tool is used in a topic.** Completes the v2.6.49 tool story. The directed tool-graph answers "what runs *next*"; this adds the orthogonal question "how is a tool *used in this context*". A new `pack_tool_edges` table (`core/tool-graph.js`, sharing the existing SQLite store, decay and pruning) records, for each turn, the reusable tools that ran while a context pack was open — keyed by `(pack, tool, command-shape)` so distinct invocations of the same tool coexist and the strongest surface. Capture is automatic and deterministic (no hand-authoring): `runner.js` already tracks the tools run and the packs opened each turn, so at end-of-turn it records a pack↔tool edge per (opened-pack × run-tool), Hebbian-reinforced. Commands are stored as **shapes, not literals** — a new `shapeCommand()` strips ids, paths, numbers and quoted text to placeholders (`spaces docs --task-id <id> <path>`) so the memory is a reusable pattern, not a stale one-off; `parseToolRuns()` extracts `{name, shape}` from a shell command (split across chained invocations) and feeds both graphs. Surfacing: when a pack is injected into context, `formatPackForContext` appends a **"Tools used in this pack"** block with the 1–2 strongest invocations, so the agent sees the concrete command instead of re-deriving it; `tool show` gains the inverse "Used in packs" view; an optional `intent` column lets a one-line *why* be enriched later without a migration. The nightly dream tends these edges alongside the follows-edges (decay + prune orphaned tools/packs). Degrades to a graceful no-op where `node:sqlite` is unavailable, exactly like the existing tool-graph. Extends `test-tool-graph.js`.

## v2.6.49
- **Reusable tools + a directed tool-usage graph.** Two linked features. First, **reusable tools** (`core/tools.js`, `bin/tool.js`): the executable sibling of context packs — when the agent works out an operational procedure (hit an API, drive an interface, transform a file) it crystallises the script into a re-runnable command instead of a throwaway heredoc. Tools are one executable file with a parseable comment header, run pre-authed (the operational keyring is merged into their env so they reference `$NAME` and never hardcode a secret), can link to an owning skill pack, and are surfaced every turn as an always-on index (full docs/source on demand via `open-claudia tool show <name>`). Full CLI: `open-claudia tool list|show|add|run|remove`; `/learn` now crystallises the executable part too. Second, a **directed tool-graph** (`core/tool-graph.js`): a "tool B follows tool A" graph kept deliberately separate from the recall graph — for tools the useful signal is the *order* a chain runs in (auth → list → download), so edges are stored and traversed directionally. Each run-sequence reinforces consecutive pairs (Hebbian, decayed nightly, no structural floor so unused chains prune away). `tool show` and the system-prompt index surface "usually followed by …"; `pack show` gains the reverse view (a pack's linked tools); the nightly dream tends the graph (decay + orphan prune) and flags tools whose `--pack` link dangles. Adds `test-tools.js` + `test-tool-graph.js`.

## v2.6.43
- **Fix the `🧠 Recall this turn` banner rendering raw HTML.** The recall debug banner (from `/recall on`) is built with `<b>` tags and HTML-escaped names, but its two `send()` calls in `runner.js` omitted the Telegram `parseMode: "HTML"` opts that every normal reply already passes via `telegramHtmlOpts()` — so the tags and `&amp;` entities showed up literally in chat instead of rendering. Both recall sends now pass `telegramHtmlOpts()`, matching the rest of the reply path. Display-only; no change to recall behaviour itself.

## v2.6.42
- **Read-time `📖 Recalled my notes on …` now fires for direct file reads, not just the CLI.** The read-side recall banner — and the Hebbian co-use signal that reinforces the recall graph — was wired only to shell commands, watching for `open-claudia pack show <dir>` / `entity show <slug>`. When the agent opened one of its own notes by reading the raw `…/packs/<dir>/PACK.md` or `…/entities/<slug>.md` straight through the Read tool, the detector never saw it: no banner, and the open never counted as co-use. A new `core/recall/read-signal.js` resolves a file path to its memory node (mirroring the existing Write/Edit path-detection), and `runner.js` calls it on Read tool-uses across the Claude and Cursor backends — deduped against the CLI path via the shared notify key, so reading the same node both ways still announces once. Adds `test-read-signal.js`.

## v2.6.37
- **`/recall [on|off]` — watch recall work.** A per-chat debug toggle that, when on, posts a short `🧠 Recall this turn` line just before each reply listing the packs/entities that surfaced — and on the **discoverer** engine, the one-line why-bullet for each. On a gated turn (pre-gate skipped recall) it says so; on a quiet turn with no matches it stays silent. The engines now return a `why` map + `gated` flag, `promptWithDynamicContext` captures a recall summary into the per-turn `consumeLastInjected()` buffer, and `runner.js` renders it when `settings.showRecall` is set. Off by default; flip with `/recall` (buttons) or `/recall on`.

## v2.6.36
- **Docs: document the dual-engine recall feature.** v2.6.34/v2.6.35 shipped the pluggable recall engine and `/engine` switch but the README and `.env.example` were never updated. This release fills the gap: README gains a "Pluggable recall" Features bullet, the `/engine` command row, a "Recall engines" narrative section, the `recall-stats` / `recall graph` CLI commands, and env-table rows for `RECALL_ENGINE` / `RECALL_GRAPH_DB` / `RECALL_METRICS` / `DREAM_TIER` (plus the corrected `DREAM_MODEL` default). `.env.example` gains `RECALL_ENGINE` and `DREAM_TIER`. Docs-only — no code changes.

## v2.6.35
- **Fix `/status` silently dying + show the active recall engine.** The status handler referenced an undefined `activeCrons`, throwing a ReferenceError that `router.js` swallows by design — so `/status` did nothing. It now counts this channel's crons via `jobs.listForChannel(...)` and adds a `Recall engine:` line so you can confirm which engine (`classic`/`discoverer`) the chat is on. Note: `/engine` already worked when typed; it just may not appear in Telegram's slash-command autocomplete until the client refreshes the cached `setMyCommands` menu.

## v2.6.34
- **Dual-engine memory recall: typed-edge graph + discoverer (opt-in).** Recall is now pluggable behind a narrow engine interface (`core/recall/`), selected per chat via `/engine` (callback `eng:`). The default stays **classic** (FTS keyword match + relevance judge + headline injection) — unchanged behaviour. The new **discoverer** engine adds a typed-edge graph (`parent`/`governed-by`/`related` edges with weight + last_reinforced in a SQLite store at `recall-graph.db`) and runs: heuristic pre-gate → FTS seed → spreading activation across the graph (1–2 hops, top-k, firing threshold — auto-pulls cross-cutting concerns the query never named) → a haiku "walker" that reads each candidate's Stance/State excerpt and returns the genuinely-relevant set with one-line why-bullets (fail-open to keyword seeds, so it never recalls worse than classic). Edges form structurally from pack `parent` frontmatter and `[[links]]` (a `[[link]]` to a `shared`-tagged pack becomes `governed-by`) and strengthen via **Hebbian co-use** — when the agent actually opens packs together in a turn (📖), their `related` edge is reinforced; weights decay over time. A metrics layer logs every discoverer turn (seeds, activated, kept+why, pre-gate, latency) — view with `open-claudia recall-stats` and `open-claudia recall graph [--sync]`. The nightly dream now runs on a high tier (opus; `DREAM_TIER`/`DREAM_MODEL` override) and tends the graph (structural sync, weight decay, orphan prune). Switch back any time with `/engine classic`.

## v2.6.30
- **Cluster self-management for AgentSpace pods.** When running as an AgentSpace pod, the bot can inspect and manage its own deployment through the backend broker — `/cluster status|logs [n]|restart|start|stop|scale <0|1>|sync` (owner-gated) for humans, and `open-claudia cluster …` for the agent. Both are thin clients over `core/cluster-client.js`, which posts to the broker's capability-gated `/pods/self/cluster` endpoint with the pod's bearer token; the bot holds no Kubernetes credentials. `brokerConfigured()` requires both `AGENTSPACE_API_URL` and `AGENTSPACE_POD_TOKEN`, so an unconfigured local install safely replies "not available" while a provisioned pod reaches the broker. Also tightens the `core/pack-guard.js` exfil heuristic: the broad "URL co-occurring with a sensitive word" rule (which false-positived on legitimate infra notes) is replaced by two precise carriers — a send/store/upload verb pointed at a URL/address, and a credential attached directly to a URL as a query/fragment param.

## v2.6.29
- **Kazee channel hardening.** `send()` now retries up to 3× with 1s/2s backoff (skipping 4xx), `edit()` normalizes text the same way as `send()`, inbound handling is async and detects ogg/opus voice notes (promoting them from "audio" to "voice" so the router dispatches correctly), reply-to context best-effort fetches the replied-to message before emitting the envelope, and `format.js` converts stray Telegram HTML tags (`<b>`, `<code>`, `<pre>`, `<i>`, `<s>`, `<a>`) to Markdown as a safety net.

## v2.6.28
- **Usage: report true peak context, not the inflated per-turn sum.** The result event re-counted the cached prefix once per tool round-trip, so `contextTokens` ballooned into the millions and the "Context this turn" alert fired every turn. It now tracks the peak single-call prefix (input + cache_read + cache_creation) and reports that as `contextTokens`; the summed value is kept as `billedContextTokens` for cost attribution.

## v2.6.27
- **Tasks: recency-ranked injection, staleness footer, and dream hygiene.** Open-task injection now ranks by activity (max of root and child `updatedAt`) rather than status — the most-recently-worked render in full, colder ones render title-only with a `task get <id>` pointer (new accessor + loopback endpoint + CLI sub), and anything untouched >14d (`OPEN_CLAUDIA_TASK_STALE_DAYS`) collapses into a capped one-line stale footer. The dream pass gains a task-hygiene step that surfaces the stalest tasks in the morning report — flag-and-ask only; it never edits or deletes tasks.

## v2.6.26
- **Memory: threat-scan guard + FTS-miss instrumentation.** `core/pack-guard.js` scans reviewer/dream content before write (override/exfil/base64 heuristics; strict on the auto-injected Stance/Procedure/State/description, lenient on Journal) and is wired into the pack reviewer and dream merge/umbrella/retag — a hit skips the write and announces one line. `logRecall` now records FTS-miss rescues (packs the judge kept that FTS scored ~0) into `recall-stats.json`, so a vector backend becomes a data-driven decision.

## v2.6.25
- **Memory: provenance tracking for pack writes.** A per-pack `.provenance.json` sidecar records which origin (user / reviewer / dream) last authored each section; `updatePack`/`createPack` take an origin, and an automated writer never silently overwrites a user-authored Stance. An absent sidecar defaults to "user" (protect-by-default).

## v2.6.24
- **Memory: pack usage telemetry + lifecycle retirement.** Packs gain `usage_count` telemetry (bumped on each use alongside `last_used`), a dream archive action retires cold packs (idle ≥30d, used ≤3×) into `.archived` with safety guards, and `pack archive`/`restore`/`archived` CLI commands are added; `/skills` surfaces the archived count. Also adds a post-dream chat summary (`DREAM_SUMMARY`), toggled with the new `/dreamsummary on|off` command (default on); when off, the nightly consolidation runs silently.

## v2.6.23
- **Progressive pack disclosure on by default.** Injects Stance + State teasers with on-demand Procedure/Journal and widens the pack match limit to 6. Override with `PACK_PROGRESSIVE=off`.

## v2.6.22
- **Wire `/skills` and `/learn` into the context-pack store.** Both previously read from / wrote to the empty legacy `~/.claude/skills` dir; now `/skills` lists packs with show/remove, and `/learn` folds work into a matching pack's Procedure plus a dated Journal line.

## v2.6.21
- **CI: allow tests to run without bot env.** The test suite no longer requires bot environment variables to be present.

## v2.6.20
- **AgentSpace pod durability.** `/upgrade` and password-change callbacks now preserve the `/api` base path expected by AgentSpace, CLI paths are resolved to concrete executables at startup, and runs repair a stale/missing session workspace before spawning Claude, Cursor, or Codex. This prevents misleading `spawn ... ENOENT` failures when the saved session cwd no longer exists.

## v2.6.15
- **Tokenomics guardrails.** The Usage dashboard now shows the latest context-token count, recent baseline, current rate multiplier, and active alert policy above the existing per-version/token trend charts. Completed turns already logged to `~/.open-claudia/usage-history.jsonl`; the dashboard now uses that history to make regressions visible immediately after `/upgrade`.
- **Hard memory recall budget.** Auto-recalled packs/entities now share a total `MEMORY_RECALL_MAX_CHARS` budget (default `9000`) after relevance filtering. Matched memory beyond the budget is omitted with a short pointer to inspect packs/entities manually, so long-term memory remains available without silently bloating every turn.
- **Token usage alerts.** After each completed real turn, Open Claudia compares the new context-token count against an absolute ceiling (`USAGE_ALERT_CONTEXT_TOKENS`, default `120000`) and the recent baseline rate (`USAGE_ALERT_RATE_MULTIPLIER`, default `1.75x`). If either trips, the bot posts a concise chat alert with the current version/model, measured context, baseline, and reason. `USAGE_ALERT_COOLDOWN_MS` prevents repeated noise. Set either threshold to `off` to disable that side.
- **Dashboard-tunable thresholds.** The web Settings page can edit `MEMORY_RECALL_MAX_CHARS`, `USAGE_ALERT_CONTEXT_TOKENS`, `USAGE_ALERT_RATE_MULTIPLIER`, `USAGE_ALERT_BASELINE_TURNS`, `USAGE_ALERT_MIN_BASELINE_TURNS`, and `USAGE_ALERT_COOLDOWN_MS`; restart the bot after saving so the runner picks up the new policy.

## v2.6.10
- **Sessions: wakeups join the live conversation instead of forking it.** A scheduled wakeup/cron captured the channel's `sessionId` at schedule time and resumed that *frozen* id when it fired — but every run writes its result id back to the shared `state.lastSessionId`, and a home-conversation compaction mints a brand-new id. So a pending wakeup carrying a pre-compaction id would yank the channel's live pointer onto a dead thread, and the user's next typed message landed on the stale session. Multiple concurrent wakeups froze different ids and forked the channel into parallel sessions. Fix (`core/scheduler.js`): a wakeup now resumes the channel's live `state.lastSessionId` (already ownership-guarded in the runner); the frozen `job.sessionId` is used only as a cold-start fallback when there is no live pointer. Wakeups scheduled after upgrade join the current context; already-pending wakeups still carry their old id until they fire.

## v2.6.8
- **Tasks: a plan can't be completed while subtasks are still open.** `task done` (and the loopback task-update path) now refuses to flip a parent to completed if any child is not yet completed — it returns a blocked result listing the open children, and the CLI prints a clean "Can't complete <id>: N subtask(s) still open" message instead of silently corrupting the tree. Standalone tasks and completing the last open child (which removes the whole plan) are unchanged. Prevents the inflated-count drift where parents showed done over still-open children.

## v2.6.6
- **Recall: recent-turns window as judge-gated context.** v2.6.5 fixed quotes, but a bare follow-up like "Push" with no quote still recalled nothing — recall only ever saw the current message. The last ~6 user/assistant turns on this channel (read from the project transcript, current turn excluded) now feed recall the same way the quote does: keyword-matched as origin "context", kept only when the haiku judge confirms relevance, dropped on fail-open. A stale thread can never force-inject packs on keywords alone, but terse follow-ups finally inherit the topic of the conversation.

## v2.6.5
- **Recall: replied-to quote becomes judge-gated context.** v2.6.4 stripped reply-quotes from matching entirely, which threw away real signal — replying "Push" to a release message recalled nothing. The quoted text now drives keyword matching as origin-"context" candidates that only survive when the haiku relevance judge confirms them (the judge sees the quote to resolve references like "it"/"the app"); on judge failure context-derived candidates are dropped while the user's-own-words baseline is kept, so a quote can never force-inject on keywords alone.

## v2.6.4
- **Recall matcher: kill false-positive pack/entity injections.** Three compounding fixes after Sharon/Euvencia/higgsfield got recalled on turns that had nothing to do with them:
  - **Match only the user's own words.** The FTS routers used the full composed prompt — so a Telegram reply-quote fed the entire quoted message (~40 terms) into matching, and any entity named in the quote scored straight past the threshold. The reply-quote block and any 📖 recall-announcement echoes are now stripped before matching (`recallMatchText`); the verbatim prompt still goes to the model untouched.
  - **Haiku relevance gate (`core/recall-filter.js`).** When keywords produce candidates, a cheap model judges which the message is genuinely about — "zoom" no longer drags in the video-gen pack via word collision. Fail-open: error/timeout/garbage keeps the keyword baseline. Costs one haiku call (~3-6s) only on turns that matched something; `RECALL_FILTER=off` to disable, `RECALL_FILTER_MODEL`/`RECALL_FILTER_TIMEOUT_MS` to tune.
  - Quoting a 📖 announcement no longer re-injects every entity it names (the self-reinforcing feedback loop).

## v2.6.3
- Inject local time + off-hours awareness into the runtime state block, so suggestions respect working hours.

## v2.6.2
- Announce freshly recalled packs/entities in chat (📖 line), mirroring the write-side notes.

## v2.6.1
- **Fix: replying to a bot message now carries the quoted context.** Since v2.0.0 the router skipped reply context when the replied-to message came from the bot itself (assuming it was already in conversation context) — which silently dropped the link whenever that message was old or compacted away. Reply context is now always injected, labelled with provenance ("your earlier assistant message" vs "this message").

## v2.6.0
- **Dream: nightly memory consolidation (Phase 3 of the memory system).** A quiet-hours pass (`core/dream.js`, default 4am via `DREAM_CRON`, model `DREAM_MODEL` default sonnet, `DREAM=off` to disable) reviews the entire pack + entity memory on a stronger model and returns a strict-JSON decision that bot code applies:
  - **Merges** packs that drifted into the same topic (sections synthesised, not concatenated) — merged-away packs are backed up to `~/.open-claudia/backup/dream-<stamp>/` before removal.
  - **Umbrella/parent pack trees**: creates umbrella packs over 3+ sibling topics (the umbrella's State is a family map, a router not a duplicate) and assigns parent/sub relationships, with cycle protection. Injected packs now show their parent ("Part of: …") and list sub-packs with dig-deeper hints.
  - **Retags**: tightens pack descriptions and tags — the FTS router matches on those fields, so tighter metadata means fewer false injections.
  - **Entity dedupe + cross-linking**: merges duplicate entities (aliases unioned, source backed up) and rewrites entity Notes with `[[pack-dir]]` cross-links.
  - Reports in chat (owner Telegram channel) whenever it changed something; silent when there was nothing to tidy. Manual trigger: `open-claudia dream`, preview with `--dry-run`.
- **Personality system.** `~/.open-claudia/persona.md` (override with `PERSONA_FILE`) holds Open Claudia's voice — tone, quirks, emoji habits — on top of the user-owned soul. It feeds into the system prompt as a `## Personality` block that explicitly never overrides soul rules, ships with a warm default, and the dream pass may evolve it gently (bounded 80–2400 chars, previous version backed up, change announced).
- **Memory announcements got a face-lift.** Reviewer announcements are now grouped under a friendly header ("🧠 Took some notes while we worked:"), use per-kind emojis (📦 new pack, ✏️ update, 👤/🚀/🏢/📍/🖥️/🔖 entities by type), clip at word boundaries instead of mid-word, carry more of the actual content, and include an inspect hint for new packs. In-turn pack/entity writes announce in the same voice.

## v2.5.0
- **Context packs: skills become living topic documents (skills + memory merged).** Each pack is `~/.open-claudia/packs/<dir>/PACK.md` with four sections — Stance (how to think about the topic: user preferences, hard rules), Procedure (verified how-to, what skills used to be), State (where work stands now; replaced wholesale as things move), Journal (dated one-line log of past sessions, capped at 30 entries).
  - **Pre-turn router** (`core/packs.js` + `system-prompt.js`): the incoming message is FTS5-matched against all packs (porter-stemmed, field-weighted — name/description/tags hits count double, threshold keeps a single incidental body word from dragging a pack in; BM25 was unusable because its IDF collapses on small corpora). Matched packs are injected into the per-turn context block, once per (channel, session, pack-version) so re-injection only happens when a pack actually changed. Rides the user message, not the system prompt, to preserve the prompt-cache prefix.
  - **Post-turn reviewer** (`core/pack-review.js`): after every substantial turn (>400 chars) a haiku subagent reviews the exchange and returns a strict-JSON decision — update a pack (journal line + State rewrite), create one for a new durable topic, or nothing. The model returns JSON only; all file writes are applied by bot code, so the reviewer needs no tools or permissions. Active by default (most working turns produce at least a journal line), `PACK_REVIEW=off` to disable, `PACK_REVIEW_MODEL` to change model. Every applied mutation is announced in chat (one line) — no silent learning.
  - **In-turn writes announced**: `Write`/`Edit` aimed at a `PACK.md` triggers `Updating pack: X` / `New pack: X` chat lines, same as skills did.
  - **`open-claudia pack list|show|match|migrate|remove|reindex`**: inspect packs, debug the router (`match` shows scores), and `pack migrate` folds existing `~/.claude/skills` into packs (body → Procedure section) with originals backed up to `~/.open-claudia/backup/skills-pre-packs/` — run it once after upgrading.
  - `spawnSubagent` learned `model` and `systemPrompt` overrides (used by the reviewer; also available to `open-claudia agent`).
- **Entity memory (phase 2): contextual notes on people, places, projects, orgs and systems.** Each entity is a short file at `~/.open-claudia/entities/<slug>.md` — frontmatter (name, type, aliases, description) plus Notes (current truth, replaced wholesale) and Log (dated one-line observations, capped at 40).
  - **Same router, second store** (`core/entities.js`): incoming messages are FTS5-matched against entities with the pack scoring scheme, except the strong fields are name/aliases — a single mention of "Emmanuel" injects the Emmanuel note, but one stray body word can't. Injected as a `## Known entities` block beside packs, deduped per (channel, session, entity-version).
  - **Same reviewer, one call**: the post-turn haiku reviewer now also returns an `entities` array — it creates or updates entity notes (merge semantics: aliases union, Notes replace, Log appends) in the same pass that maintains packs, at no extra subagent cost. Entity changes are announced in chat like pack changes.
  - **`open-claudia entity list|show|match|note|remove|reindex`**: inspect the store, debug the router, or jot a note manually (`entity note Emmanuel "owns inet-central org cleanup" --type person`).
  - Direct `Write`/`Edit` to an entity file announces `Updating entity: X` / `New entity: X`, same as packs.
  - This completes phase 2 of the packs system; "dream" consolidation (pack merging + skill trees + entity linking) is next.

## v2.4.2
- Fix replies losing everything written before tool calls. The backend's final `result` event carries only the LAST text segment of a turn, but the stream parser assigned it over the accumulated `assistantText` — so a turn shaped "long explanation → tool calls → short closing line" delivered only the closing line, which read as nonsense without its context. `evt.result` is now a fallback used only when nothing was accumulated (some Cursor turns). Fixed in both `runClaude` and the auxiliary runner.

## v2.4.1
- Fix Telegram messages occasionally arriving as raw HTML (`<b>`, `<code>` shown literally). Root cause: the Markdown→HTML pass could inject `<i>` tags *inside* model-authored `<code>` spans (two snake_case identifiers on one line pair their underscores as italics), producing unbalanced HTML that Telegram rejects — and the parse-failure fallback then resent the converted body verbatim, tags and all.
  - `<code>`/`<pre>` spans are now stashed whole (content included) before Markdown conversion, so nothing can be injected inside them.
  - The italic `_..._` rule only fires at word boundaries, so snake_case identifiers anywhere in the text can no longer pair up.
  - Model-authored entities (`&lt;`, `&amp;`, `&#…;`) are preserved instead of double-escaped — `<code>&lt;name&gt;</code>` now renders as `<name>` instead of `&lt;name&gt;`.
  - Last-resort fallback (send + edit) now strips tags and decodes entities (`htmlToPlain`) so a rejected message degrades to clean plain text, never raw markup.

## v2.4.0
- **FTS5 transcript index (cross-session recall).** Project transcripts are now indexed in SQLite FTS5 via Node's built-in `node:sqlite` (no new dependency), turning "did we discuss X last week?" into a ~50ms ranked lookup instead of a linear grep over JSONL.
  - `core/transcript-index.js`: WAL-mode DB at `~/.open-claudia/transcripts/index.db` (0600), FTS5 table with porter tokenizer, per-file byte offsets for idempotent incremental indexing. Partial trailing lines wait for the next pass; a replaced/truncated transcript drops its stale rows and reindexes. Fail-soft: without `node:sqlite` every call no-ops and transcript-window remains the path.
  - **Live indexing**: `appendProjectTranscript` indexes each entry as it lands (text is already redacted at that point), so the index is always current. The CLI also runs a cheap catch-up pass before each query, making it self-healing — no boot-time backfill needed.
  - **`open-claudia transcript-search <query>`** (alias `ts`): bm25-ranked snippets with role/timestamp/project/line pointers; defaults to the current project's transcript (OC_TRANSCRIPT_PATH), `--all` for every project, `--project <name>` filter, `--raw` for full FTS5 syntax, `--rebuild` to reindex from scratch (5.2k entries across 22 transcripts rebuild in ~0.6s). Natural queries are term-quoted so FTS5 operators can't error.
  - Recall guidance updated everywhere it lives: transcript pointer note, compaction seed prompt, and CLI help now lead with transcript-search → transcript-window for context.

## v2.3.0
- **Learned skills (Hermes-style autonomous skill creation).** The agent now captures battle-tested procedures as reusable skills and the bot surfaces skill activity in chat.
  - **Autonomous capture policy** in the appended system prompt: after a complex task (5+ non-trivial tool calls, dead-ends overcome, or a generalising user correction) the agent writes `~/.claude/skills/<name>/SKILL.md` — or patches an existing one rather than duplicating — and must announce it in its reply. The Claude Code harness already auto-loads personal skills with progressive disclosure (cheap name+description listing, full body on use), so a captured skill is available in every future session and project for free; the bot only had to build the capture/manage/notify half. Skills are procedures, not memories: no user facts, no secrets, nothing copied from untrusted output.
  - **`/learn [hint]`**: explicitly capture the most recent piece of work as a skill — same rules, but skips the "worth it?" gate.
  - **`/skills [show|remove <name>]`**: list learned skills with descriptions, show a skill's full content, or delete one. Backed by the new `core/skills.js` (frontmatter parsing without a YAML dep, symlink-aware dir listing, traversal-safe skill-path recognition).
  - **Chat announcements (Hermes-style)**: the stream parser in `runClaude` now recognises the `Skill` tool (`Using skill: X`) and any `Write`/`Edit` aimed at a personal `SKILL.md` (`Learning new skill: X` vs `Updating skill: X`, decided by whether the file existed before the write) and sends a one-line channel message, deduped per turn. Wired for both Claude (tool_use blocks) and Cursor (tool_call events) stream shapes.

## v2.2.24
- Fix one-shot wakeups being lost across restarts when their fire collided with a busy channel. `fireJob` removed the wakeup from `jobs.json` as soon as it returned — but on a busy channel the actual run was deferred to an in-memory 30s retry timer, so the job was already gone from disk and a bot restart during the deferral window (e.g. `/upgrade`) silently dropped it. The wakeup now stays persisted until it actually runs (or is conclusively skipped after retries); a restart mid-deferral re-arms it from `jobs.json` via the existing missed-wakeup grace pass.

## v2.2.23
- **Token economy: tasks (R2).** The per-channel todo list no longer rides every prompt at full size.
  - **Done means gone**: completing a task now deletes it from the store instead of leaving an `[x]` row. Finishing the last subtask of a plan deletes the whole plan. `tasks.complete()` wraps the status flip plus a `prune()` pass; the loopback `task-update` handler routes `status: completed` through it and reports how many entries were removed, so the CLI can tell the agent what disappeared. `prune()` also retires legacy debris — plans whose subtasks are all completed but whose own status was never flipped.
  - **Once-per-session tree injection**: the full pending-task tree is only injected into the first turn of a session (after restart, new conversation, or compaction) — exactly when the agent needs to rediscover where it left off. Later turns get a one-line count (`## Pending tasks: N open (M in progress)`) with a pointer to `open-claudia task list`. Previously a large tree (~60 plans on a busy channel) was re-sent uncached on every single user turn.
  - `formatTree` gained a `hideCompleted` option so any injected tree only shows remaining work; CLI help and the system-prompt task docs updated to describe the new semantics.

## v2.2.22
- **Token economy: stable prompt-cache prefix (R1).** The per-turn churning lines — `Vault: unlocked/locked`, `Session: resuming/new`, and the full pending-tasks tree — moved out of the `--append-system-prompt` block into a "Runtime state (current turn)" header prepended to each user prompt (`promptWithDynamicContext`). The appended system prompt precedes the entire conversation in every API request, so any byte that changed between turns invalidated Anthropic's prompt-cache prefix and re-billed the whole history at 1.25x write price instead of 0.1x read price; on long sessions this churn was the dominant bot cost. `buildSystemPrompt()` is now byte-stable within a session (remaining interpolations — project path, channel, voice, team/speaker — only change on rare events), and the dynamic state rides at the end of the request where it is always uncached anyway. Transcripts still log the bare prompt; Cursor/Codex paths unchanged.
- Compaction is much less lossy. Three changes to `compactActiveSession`:
  - **Two-tier brief**: every compaction brief is now archived verbatim to `~/.open-claudia/briefs/<project-prefix>-<timestamp>.md` before seeding the fresh session, and the seed prompt tells the new session where the archive lives. Repeated compactions previously re-summarized prior summaries, blurring older facts each round — now the new session can reread earlier briefs at full fidelity instead of guessing. The summary prompt also drops length pressure ("length is NOT a constraint") since the brief is archived anyway. The summarizer additionally emits a `=== CONDENSED SEED ===` section (300-500 words: goal, next step, gated actions, load-bearing paths); when present and the full brief made it to disk, only the condensed section is seeded into the live context — cutting the per-turn context cost of the carried summary — while the detailed brief stays on disk for on-demand reads. Missing marker or failed archive falls back to seeding the full text exactly as before.
  - **Machine-generated repo state**: the seed prompt now includes a git-generated block (branch, dirty files, ahead/behind upstream, last commit) for the session cwd — or its direct child repos (capped at 10) when the cwd itself isn't a git repo. Repo state was exactly the class of fact summarizers hallucinated or dropped; this block is produced by `git`, not model recall, and the seed says to trust it over the summary on conflict.
  - **Verbatim recency tail**: the last 6 user/assistant turns (1,500 chars each, compact-summary/seed entries filtered out) are carried into the seed raw from the project transcript. Users most often reference the exchange immediately before compaction ("as I said", "do that thing you showed"), which summaries flatten; now those turns survive as exact quotes. Captured before the summarizer runs so the tail reflects real conversation.
- All three collectors are fail-soft: any error (missing transcript, unreadable repo, git timeout) degrades to the previous behaviour rather than blocking compaction.

## v2.2.21
- Add Claude Fable 5 (`claude-fable-5`) to the Claude model picker and make it the default Claude Code model used when no per-user model override is selected. The default can still be overridden with `CLAUDE_MODEL`.

## v2.2.19
- Add Claude Opus 4.8 (`claude-opus-4-8`) to the Claude model picker and make it the default Claude Code model used when no per-user model override is selected. The default can still be overridden with `CLAUDE_MODEL`, now documented in `.env.example`.

## v2.2.18
- Fix the legacy Telegram send/edit path so replies default to Telegram HTML parse mode and go through the shared formatter, preventing raw `<b>`, `<a>`, and related tags from leaking in chat.
- Preserve safe Telegram HTML tags before Markdown conversion, so underscores inside link URLs such as Higgsfield `wan2_2_video` are not misread as italics and do not break rendered links.

## v2.2.17
- Telegram output now uses `parse_mode: "HTML"` instead of legacy Markdown. The Telegram adapter normalizes both model-authored Telegram HTML and ordinary Markdown/CommonMark into Telegram's safe HTML subset before sending, so `<b>...</b>`, `**bold**`, backticks, links, headings, code blocks, strikethrough, and spoilers render cleanly instead of leaking literal markup or being rejected by Telegram. Messages still fall back to plain text if Telegram rejects the markup.
- Updated the Telegram system prompt to ask agents for short mobile-readable Telegram HTML (`<b>`, `<code>`, `<pre>`, `<a>`) with bullet-style layout. Relay sends and slash-command responses now use the same HTML parse mode as normal assistant replies.

## v2.2.16
- New `/compactwindow` slash command (alias `/autocompact`) lets each user override the auto-compact token threshold from chat instead of editing `AUTO_COMPACT_TOKENS` in `.env` and restarting. Quick-pick buttons for 200k / 300k / 380k / 500k / Off / Default, or free-form `/compactwindow 250k` / `/compactwindow 0.5m` / `/compactwindow off` / `/compactwindow default`. Stored per-user as `settings.compactWindow` (persists across restarts via `state.json`); `null` falls back to the env default, `0` disables auto-compact entirely (manual `/compact` still works). `/status` now reports the effective window, and `/usage`'s "context is large" tip reflects the user's override.

## v2.2.15
- Fix compaction loop: `compactActiveSession` no longer holds `isCompacting=true` for the full duration of its two-step flow. The flag now clears after step 1 (the summarizer) finishes, so step 2 (the seed-the-fresh-session call) runs as a regular long-running task. Previously, the seed step could pick up where the prior conversation left off and do real work — dev servers, package installs — for hours, all while the bot reported "Compacting context, will pick this up next…" to every incoming message. After a `/restart` the in-memory flag reset but `lastSessionId` still pointed to the same huge session, triggering an auto-compact on the next message and looping the same trap. New behaviour: the summarizer-only phase shows the compaction message; once the summary is written the bot returns to its normal "Queued." reply for any messages that arrive while the seed continuation runs.
- Add `COMPACT_SUMMARY_TIMEOUT` (10 minutes) and thread it through `runClaudeCapture` via `opts.timeoutMs`. The summarizer is a single-shot summarisation call — if it hasn't returned in 10 minutes it's hung, not slow. Previously it could sit on the 6-hour `MAX_PROCESS_TIMEOUT` and lock the bot for a quarter of a day. The seed continuation keeps the full 6-hour budget since it can legitimately be a long-running agent task.

## v2.2.14
- Dockerfile: bake `openssh-client` and `rsync` into the image. These were being installed at runtime via `sudo apt-get install` on pods that needed to push code over ssh or rsync to dev servers; baking them in means they survive pod restarts and `/upgrade` overlays. Companion change in the AgentSpace backend flips the bot-pod container `securityContext` to allow privilege escalation + adds `SETUID,SETGID,DAC_OVERRIDE,CHOWN,FOWNER` capabilities so the existing `claudia ALL=(ALL) NOPASSWD: /usr/bin/apt-get` sudoers rule (added in v2.2.7) actually works — without these, kernel `no_new_privs` blocks sudo from elevating. The same backend change also opens 22/TCP egress on the bot's NetworkPolicy so the in-pod ssh actually reaches dev hosts.

## v2.2.13
- **Security fix**: `intro-flow.handleInbound` no longer auto-claims ownership on *any* first inbound message. Previously, if `people.json` had no owner record (fresh pod / fresh install), the very first inbound — including a `/start` tap from a t.me deep-link — would register the sender as the bot owner. In AgentSpace pods this meant anyone who guessed or was sent the bot username could take over a pod simply by clicking Start. The bootstrap path now requires either (a) pod mode (`AGENTSPACE_POD_TOKEN` set) with the inbound chat id pre-seeded in `TELEGRAM_CHAT_ID` env, or (b) local mode with the inbound message literally starting with `/auth`. Anything else gets a "no owner configured" reply and an `intro.bootstrap-refused` audit entry. Existing owner records are unaffected. Follow-up work needed in the AgentSpace backend provisioner to seed `TELEGRAM_CHAT_ID` at pod creation time from the provisioning user's Telegram chat id; until that lands, new pods will refuse all inbound until an operator manually seeds the env.

## v2.2.12
- Cross-channel relay (`open-claudia send-to`, and by extension cron/wakeup-fired messages and any other caller of `relay.send`) now passes `parseMode: "Markdown"` when the resolved adapter is Telegram. Previously `relay.send` called `adapter.send(channelId, text)` with no opts, so the Telegram adapter never set `parse_mode` and every relayed message went out as plain text — `*bold*`, backticks, and `_italic_` rendered as literal characters. The main reply path already injected this via `core/runner.js:25-28`; relay was the missing twin. Non-telegram adapters are unaffected.

## v2.2.11
- Telegram replies now get channel-specific formatting guidance in the system prompt: short mobile-friendly sections, single-asterisk Telegram Markdown labels, no tables/headings/Markdown links, and no noisy quotes/backticks around normal business terms.
- Telegram final/progress messages now request Markdown parse mode, and edited messages support the same parse mode with a plain-text fallback if Telegram rejects the markup.

## v2.2.10
- System prompt: explicitly tell the model not to use the Claude Code harness's built-in `ScheduleWakeup` / `CronCreate` tools. Those register timers inside the per-turn Claude Code subprocess, which exits at end-of-turn, so the wakeup silently never fires. Only the `open-claudia schedule-wakeup` / `cron-add` CLI commands (owned by the long-running bot) actually survive. Mirrors the existing warning we already had for harness `TaskCreate` / `TaskUpdate` vs `open-claudia task`.

## v2.2.9
- Startup self-heal for grandfathered AgentSpace pods: if `.password-changed` exists on the PVC but AgentSpace was never notified (because the pod changed its password before `AGENTSPACE_API_URL` was injected, or before the `/pods/self/password-changed` callback shipped), the web server fires the callback once on startup and writes a `.password-changed-notified` sentinel so subsequent restarts don't retry. `notifyAgentSpacePasswordChanged` now returns a status so `setPassword` can also drop the sentinel after a successful live notification. Pre-v2.2.x grandfathered pods that never wrote the marker file still need manual `webPasswordUserSet=true` in Mongo.

## v2.2.8
- `/upgrade` now works on docker containers that run our baked-in `/app` source (the common self-host layout). Previously the handler fell through to `npm install -g`, which hit `EACCES` because the runtime user is uid 1001 and couldn't write the global node_modules — and even when it could, the bot reads from `/app`, not the global root, so the new version was never picked up. The new branch detects the `/app` layout, `npm pack`s the latest tarball, overlays it onto `/app`, runs `npm install --omit=dev`, and exits so the orchestrator restarts the container on the new source. AgentSpace pods (which have `AGENTSPACE_POD_TOKEN` + `AGENTSPACE_API_URL`) still go through the control plane; nothing else changes for them or for npm-global installs.
- Dockerfile: `chown -R claudia:claudia /app` after the build steps so the runtime user can overlay new source during the in-place `/upgrade` above. Previously `/app` and its `node_modules` were root-owned because the COPY and `npm ci` ran as root before `USER 1001`.

## v2.2.7
- Docker image now ships `git`, `jq`, `python3`, `python3-pip`, and `build-essential` so spawned coding agents don't fall back to curling random binaries into userspace when a basic tool is missing.
- Symlink `/app/bin/cli.js` to `/usr/local/bin/open-claudia` so the CLI (used by agents for `send-file`, `task`, etc.) is on PATH from any cwd. Previously agents had to extract the packaged tgz to find it.
- Grant the `claudia` user passwordless sudo for `apt-get` / `apt` so the model can install additional packages at runtime via the normal path rather than ad-hoc binary downloads.

## v2.2.6
- Kazee inbound photos: V2 socket emits attachments as `msg.media` (single) or `msg.medias` (array), not `msg.attachments`. The adapter now reads all three so images sent from Kazee web/mobile clients reach the bot instead of being silently dropped as zero-attachment text.
- Kazee outbound files: rewrote `sendFile` to (a) upload via `POST /chat/media/:chatId` with field `media` (the actual route; the old code POSTed to `/upload` with field `file` and 404'd), then (b) post the message via the V2 socket event `message:send` with `mediaIds: [<uploadedId>]` instead of REST `sendMessage` with `media_url`. The REST path is currently broken upstream — `Chat/send_message` calls `Media.create` without the required `chat`/`bucketName`/`fileName`/`minioPath` fields and 500s with `Failed to create media record`. Going via the socket handler reuses the already-saved Media doc and avoids the duplicate create.
- New `_socketEmit(event, payload, timeoutMs)` adapter helper that promisifies socket.io acks with a timeout, used by the new outbound path.

## v2.2.5
- Fix `/codex` resume crash: `buildCodexArgs` no longer appends `--add-dir <transcripts-dir>`, which the Codex CLI does not accept and which caused every `codex exec resume` invocation to exit 2 with `error: unexpected argument '--add-dir' found`. Transcript pointer is still injected into the prompt via `promptWithTranscriptPointer`.
- Backend-aware empty-output failure message: when a non-Claude backend exits with no assistant output, the reply now labels it correctly (`Codex` / `Cursor`) and points at the right diagnostic commands (`/codex_auth_status`, `/codex_login`, `/codex_setup_token`, or `agent login`) instead of always saying "Claude exited" and recommending Claude-only commands.

## v2.2.4
- Bearer auth on `/api/*`: when `BOT_CONTROL_TOKEN` is set in the env, requests with `Authorization: Bearer <token>` are accepted in addition to the existing cookie session. Behaviour is unchanged when the env var is not set, so local installs are not affected.
- `/upgrade` detects an AgentSpace-managed pod (both `AGENTSPACE_POD_TOKEN` and `AGENTSPACE_API_URL` set) and delegates to `POST /pods/self/upgrade` on the control plane instead of running `npm install -g` locally. The control plane rolls the deployment with a fresh image pull. Local installs fall through to the existing npm upgrade path.
- Password change in the web UI fires a fire-and-forget `POST /pods/self/password-changed` callback so the control plane can flag the pod as having a user-set password and refuse to leak the stale initial value.

## v2.2.3
- Queue drain now batches: when multiple messages are received while a task is running, they're delivered as one combined follow-up turn instead of N isolated turns. The model sees them together (numbered, with HH:MM:SS queue timestamps), in context of what it just finished, so it can plan across them. A single queued message still delivers as before — no behavior change for the common case.

## v2.0.1
- Kazee owner detection: `envelope.channelId` is the chat-document id, but `KAZEE_OWNER_USER_ID` is the Kazee user id. `isChatOwner`/`isChatAuthorized` now also short-circuit when the inbound user id matches the configured transport owner, so the owner running `/auth` (or anything else) on a fresh Kazee install is recognized immediately instead of being queued as a non-owner request.
- `chatContext` now carries `userId` and `transport` in addition to `chatId`; new `currentUserId()` / `currentTransport()` exports.

## v2.0.0
- **Multi-channel**: the bot can now run on Telegram and Kazee Chat in the same process. `CHANNELS=telegram,kazee` selects which channels start up; each channel implements the new `ChannelAdapter` contract (send/edit/delete/upload/keyboard/typing/voice-fetch) and a `chatContext` AsyncLocalStorage routes replies, edits, and notifications back to the originating channel.
- Kazee adapter speaks V2 socket (inbound messages, edits, deletes, reactions, typing, presence) and uses REST for sending, editing, deleting, and uploads. Interactive button keyboards are sent as `type:"interactive"` with portable `interactive.buttons` so web and mobile clients render them natively.
- New env: `CHANNELS`, `KAZEE_URL`, `KAZEE_BOT_TOKEN`, `KAZEE_OWNER_USER_ID`, `KAZEE_BOT_USER_ID`, `KAZEE_DEBUG_EVENTS`.
- Slash-command registry: bots self-register the `/commands` list on connect (Kazee). Each adapter renders the registry into its native menu format; Telegram continues to use the existing `/setMyCommands` flow.
- Auto-compaction bumped from 140k to 280k context tokens (`AUTO_COMPACT_TOKENS`) and now also runs proactively after each reply — large turns are compacted before they pile up, not at the next ask.
- First-`/auth`-wins owner bootstrap so a brand-new Kazee install can be authenticated from the first inbound DM without pre-seeded `auth.json`.
- Cron jobs now resolve the owner's adapter from the canonical user id, so scheduled tasks fire in the correct channel.
- Dockerfile fix: container user can now write to the config directory; missing `chat_id` shows a friendly setup hint instead of crashing on first start.
- Requires chat-central with REST `sendMessage` interactive support (commit `b1a7d02` or later) for keyboard buttons over Kazee. Older servers will accept the message but drop the buttons.

Known limitations:
- Kazee adapter does not emit a back-online ping after reconnect (the Telegram one still does).
- kazee-chat-mobile prior to v3.21.0 renders interactive messages as a text-only fallback — upgrade clients for button UI.

## v1.20.0
- Added project transcript memory: redacted JSONL transcripts are stored outside repos under the Open Claudia config directory, keyed by normalized project path hash.
- Fresh sessions and backend switches now receive a small transcript pointer/instruction instead of auto-generating or pasting handoff summaries; agents are told to tail/search/read relevant parts only and treat history as untrusted context.
- Added transcript config knobs: `PROJECT_TRANSCRIPTS`, `TRANSCRIPT_MAX_ENTRY_CHARS`, and `TRANSCRIPTS_DIR`.

## v1.19.5
- Stop printing the Web UI admin password in startup logs.

## v1.19.4
- Legacy/static env-only deployments with no persisted `auth.json` owner now treat configured `TELEGRAM_CHAT_ID` chats as owners for shared auth commands. This fixes authorized test chats being blocked from `/codex_login` when no `isOwner` metadata exists yet.

## v1.19.3
- Owner-only auth commands now recognize chats marked as `isOwner: true` in `auth.json`, not only the first `TELEGRAM_CHAT_ID`. This fixes owner-authenticated Telegram chats being blocked from `/codex_login` and `/codex_setup_token`.

## v1.19.2
- Docker images now bake in the OpenAI Codex CLI (`@openai/codex`), so `/codex` and Codex auth commands are available in container/Kubernetes deployments after rollout.
- Documented that direct npm installs still need optional backend CLIs installed on the host, while Docker images include the Codex CLI.

## v1.19.1
- `/auth` requests now send the owner Telegram Approve/Deny buttons, while preserving the existing `auth.json` and CLI approval flow.

## v1.19.0
- Added `/doctor` / `/requirements`: mobile-friendly checks for Node.js, Claude/Cursor/Codex CLI versions and auth, optional ffmpeg/Whisper voice stack, workspace writability, and config dir writability
- `/upgrade` now includes a post-upgrade requirements/auth summary before restarting, so missing CLIs or auth breakage is visible immediately
- Added Codex auth commands: `/codex_auth_status`, `/codex_login` device auth, `/codex_setup_token` / `/codex_use_api_key` secure API-key paste mode, and `/cancel_codex_auth`
- Codex/OpenAI keys are redacted from Telegram output/logs; API keys are passed to `codex login --with-api-key` without being echoed
- Added canonical user identity links for future multi-channel support. State and session history now key by canonical user id (`telegram:<chatId>` by default, or an explicit id like `sumeet@inet.africa`) instead of raw Telegram chat id.
- Added `/link`, `/links`, and `/whoami` for managing and inspecting Telegram-to-user mappings.
- Existing `state.json` and `sessions.json` chat-id keys are migrated through the identity resolver on load.

## v1.18.0
- Auto-compacts high-context sessions before the next turn (`AUTO_COMPACT_TOKENS`, default 140k): summarizes the old session, seeds a fresh session, then continues the user's request there
- `/compact` now creates a fresh compacted continuation session instead of only adding another summary turn to the existing session
- `/continue` resumes the selected stored session ID with `--resume` instead of using cwd-most-recent `--continue`
- Stabilized the mobile system prompt: no timestamps, dynamic file lists, vault key names, or raw Telegram token curl examples
- Reply context is no longer redundantly injected when replying to the bot's own prior text from the active session

## v1.17.0
- **Multi-user / team mode**: a single bot can now serve multiple authorized users in parallel. Each user has their own conversation thread, project session, settings, model, backend, runningProcess, queue, and usage counters
- Per-user state lives in `userStates: Map<chatId, UserState>`; an `AsyncLocalStorage` chat context routes `send()`/`editMessage()`/typing indicators back to whoever triggered the work
- `state.json` and `sessions.json` are now keyed by chatId. Legacy single-user files are migrated under the owner's CHAT_ID on first load — direct-install users notice no change
- Owner-only commands: `/restart`, `/upgrade`, `/login`, `/setup_token`, `/use_oauth_token`, `/clear_oauth_token`, `/auth_code`, `/cancel_auth`, mode switch (they touch shared Claude OAuth or bot lifecycle)
- Shared across users: vault, soul, files, crons, Claude OAuth token, workspace files (no per-user worktrees yet — concurrent edits hit the same paths)
- `/stop` now stops only the calling user's running Claude, not everyone's
- Crons fire as the owner (the user who scheduled them); cron output reports to the owner's chat
- `isDuplicate()` is now keyed by `chat_id:message_id` so two users' messages can't collide on the dedup set

## v1.16.1
- `/model` is now a single unified picker — all Claude, Cursor, and Codex models in one keyboard, with section dividers per backend
- Tapping any model auto-switches the active backend if needed (e.g. picking gpt-5 from Claude flips you to Codex). No more `/codex` + `/model` two-step
- Only backends with a detected CLI binary appear in the picker (Cursor/Codex rows hidden if not installed)

## v1.16.0
- OpenAI Codex backend support: `/codex` to switch, `/backend` picker now has 3 options
- Auto-detects `codex` binary via `which codex` at startup; opt-in via `CODEX_PATH` in .env
- Codex sessions persist (`codexSessionId`) across restarts and resume via `codex exec resume <id>`
- Per-backend session IDs maintained independently — switch freely without losing context
- Stream parser handles Codex JSONL events: `thread.started`, `item.started`, `item.completed`, `turn.completed` (with usage)
- Plan mode on Codex maps to `--sandbox read-only`; outside plan mode uses `--dangerously-bypass-approvals-and-sandbox` (bot is the sandbox)
- `/model` picker shows Codex models (gpt-5, gpt-5-codex, o3, o4-mini) when on Codex
- Effort/budget gracefully reject on Codex (not exposed as CLI flags); Claude auth preflight scoped strictly to the Claude backend
- Codex auth errors in stderr produce an actionable "run `codex login`" recovery message

## v1.14.9
- Fix `/upgrade`: `latest` was const-scoped to the version-check try
  block, so the install step ran with `latest` undefined and threw
  `ReferenceError: latest is not defined`. Hoist the variable and fall
  back to npm's `latest` dist-tag if `npm view` fails. Regression
  introduced in 745c6cf — broke v1.14.7 and v1.14.8 upgrades.

## v1.14.8
- Fix: `/stop` and the 6h hard-timeout now walk the full descendant
  process tree, not just the Claude CLI's process group. Background
  bashes started with `run_in_background: true` (and other detached
  children) used to survive the kill and keep running indefinitely —
  e.g. a polling `until` loop with a broken grep would burn CPU and
  network for hours after the agent was already dead.
- Fix: on bot startup, sweep stale `claude login` / `claude setup-token`
  processes older than 30 minutes. These were left over from previous
  bot crashes (blocked on stdin forever) and could hold Keychain locks.

## v1.14.0
- Claude Code auth management from Telegram: `/auth_status`, `/login`, `/setup_token`, `/use_oauth_token`, and `/clear_oauth_token`
- Claude subprocesses now receive `CLAUDE_CODE_OAUTH_TOKEN` from config/env/vault when available, avoiding launchd/macOS Keychain auth failures
- Sensitive Claude tokens are redacted from Telegram output and logs
- Package lock metadata corrected to match `@inetafrica/open-claudia`

## v1.12.0
- Cursor Agent tool progress: Shell, Read, Edit, Write, Grep, Glob calls now show in real-time Telegram updates
- Plan output surfaced to Telegram: when Cursor creates a plan (`--mode plan`), the full plan markdown and task list are sent to the user
- Handles all Cursor tool_call event types with fallback for unknown tools

## v1.11.0
- Backend-aware plan mode: /plan passes `--mode plan` to Cursor Agent, `--permission-mode plan` to Claude
- New /ask command: read-only Q&A mode (Cursor Agent only, `--mode ask`)
- /effort and /budget now warn when on Cursor backend (unsupported flags)
- Worktree flag (`--worktree`) wired into Cursor Agent args
- buildCursorArgs now forwards plan/ask/worktree settings

## v1.10.0
- Cursor Agent backend: switch between Claude Code and Cursor Agent CLI
- New commands: /cursor, /claude, /backend with inline keyboard
- Separate session persistence per backend (Claude and Cursor sessions don't clash)
- Auto-discovers `agent` CLI in PATH if CURSOR_PATH not set
- /status shows active backend

## v1.9.2
- Fix: show what's new after upgrade
- Startup message shows version

## v1.9.1
- Fix: duplicate messages — progress message now edited instead of sending a second copy

## v1.9.0
- Force password change on first web UI login
- Password complexity requirements (12+ chars, uppercase, lowercase, number, symbol)

## v1.8.1
- Web UI accepts WEB_PASSWORD env var for managed deployments
- Config API whitelist (only safe keys editable)
- Stronger password entropy (32 chars)

## v1.7.4
- Run as non-root user in Docker (Claude Code security requirement)
- Numeric UID 1001 for K8s compatibility

## v1.6.0
- Agent mode: non-blocking side conversations while tasks run
- /mode command to switch between direct and agent modes

## v1.5.0
- Robust message delivery with retry on replyTo failure
- Adaptive rate limiting (2s -> 5s) to avoid Telegram 429 errors
- Global error handling with Telegram notification on crash
- editMessage handles rate limits gracefully

## v1.4.5
- Streaming progress with tool names and elapsed time
- Voice note transcription via whisper.cpp
- File and image handling
- Cron jobs for scheduled tasks
- Encrypted vault for credentials
