# Permissions system — production-readiness audit & redesign plan

Date: 2026-07-24 · Auditor: Open Claudia (session-reviewed, all fixes test-pinned)
Scope: the external-speaker guardrail stack — relationship classification, the
three policy layers, the enforcer (reply/action/triage vets), approvals +
owner cards, the tool deny-gate, mandates, and identity/people records.

---

## 1. Verdict

**Conditionally production-ready.** The core is sound: default-deny
classification, an independent model judge that fails closed, owner-only
approval, and a hardcoded security floor no note can override. This pass
closed the five gaps that were outright exploitable or operationally broken
(F1–F6 below). What remains open is not exploitable-by-message but is
*organizational* debt that will bite as more people use the bot: fragmented
identities, no group layer, and always-allow grants that are broader than the
owner probably intends. Those are designed (§6) but deliberately NOT built
until the plan is approved.

---

## 2. The system today (post-fix architecture)

Layers consulted for every guarded turn, narrowing in scope; union of grants,
any prohibition wins, default-deny:

1. **Security floor** — hardcoded in `GUARD_SYSTEM_PROMPT` (enforcer.js).
   Credentials/secrets/infra access can never be granted by any layer below.
2. **Global guardrails** — owner-authored, applies to every external.
3. **Channel rules** — per group chat / task thread.
4. **Person mandate** — per-person grants on their entity pack (≤4000 chars,
   deduped, slug-keyed; judge reads at most 6000 defensively).
5. **Turn grant** — one-off "handle this request" approval; reply vets only.

Enforcement points on a guarded (non-owner) turn:

- **Inbound triage** (router → `guardInboundTriage`): the request itself is
  judged BEFORE any agent work. Held → owner card ("should I handle it?"),
  external gets the holding line, no generation happens at all.
- **Provider file tools** (PreToolUse hook → tool-guard): Write/Edit/
  MultiEdit/NotebookEdit are denied outright for external speakers, with a
  redirect to write-tier tools (which escalate to the owner).
- **Bash deny-gate** (unchanged): prod surfaces require the covering
  open-claudia tool; write/destructive tool runs escalate with exact-command
  approval, one-shot consume, byte-bound.
- **Outbound reply vet** (`guardOutboundReply`): the authoritative gate on
  anything the external actually receives. Triage is advisory; this is final.

Identity: `speakerFor` fails closed to "external"; owner status comes only
from `people.isOwner` via canonical identity (a note can never confer it).
Unknown authed speakers get an auto-contact (fresh record, hashed entity slug,
never name-linked). Meet-greet lane covers UNVERIFIED first contacts.

Fail direction map (deliberate asymmetries):

| Gate | On internal error | Why |
|---|---|---|
| speakerFor | closed (external) | never widen on failure |
| Enforcer verdict | closed (escalate), never cached, 1 bounded retry on transient | transport blips ≠ verdicts |
| Triage (guarded speaker) | closed (held) | reply vet still downstream |
| Triage (unresolvable speaker) | open (proceed) | owner UX; output vet authoritative |
| File-mutation gate, no channel env | dormant (allow) | owner-local turns untouched (R9/R14) |
| File-mutation gate, env + resolve error | closed (deny) | env present ⇒ real channel turn |
| Bash deny-gate | open (allow) | availability guard, not identity guard; logged |
| Turn grant validation | closed (no grant) | any mismatch/expiry/read error |

---

## 3. Findings fixed in this pass

| ID | Sev | Finding | Fix (all uncommitted, test-pinned) |
|---|---|---|---|
| F1 | High (ops) | Guard model cold-starts overran the old timeout, failing closed on legitimate replies → spurious owner cards | 60s timeout (`ENFORCER_TIMEOUT_MS`) + exactly one retry on transient transport errors; non-transient never retried; persistent failure still closes |
| F2 | **Critical** | Only Bash was hooked: an external speaker could have the agent write/edit files directly via provider Write/Edit tools — bypassing the whole approval path | Second PreToolUse matcher `^(Write\|Edit\|MultiEdit\|NotebookEdit)$` → same deny-gate CLI → speaker-gated denial; deliberately outside relaxed tooling-mode (it's a security control, not a workflow preference); audited as `external-file-write` |
| F3 | Med (UX/authz clarity) | Approval cards led with the bot's output, not the person — the owner was approving text, unclear they were authorizing a PERSON's request | Cards reworded: "**X** asked me something — should I act on it?" / "wants me to handle a request — should I?" / "wants me to run a **write** action — should I?"; ask shown first |
| F4 | High | Work happened BEFORE any approval: the agent fully processed an external request, only the reply was vetted (cost, prompt-injection surface, side-effect risk) | Up-front triage gate in the router: inbound judged first; held turns generate nothing; approval re-runs enter via the scheduler so they can't re-triage into a loop |
| F5 | High (friction→risk) | Without a grant concept, every triage-approved request structurally produced a SECOND approval card for its reply — double-approval fatigue trains blanket-allow habits | Turn grant: the approval id rides the scheduler job → bot-process AsyncLocalStorage (agent subprocess can never forge it) → honored in reply vets ONLY, validated fail-closed (approved + triage-kind + adapter/channel match + <24h), quotes one request, never actions |
| F6 | High (latent bug) | `appendMandate` upserted by NAME: an always-allow for auto-contact `sharon-a1b2c3` could land on a different `sharon` entity — cross-person grant leakage; mandates also grew unbounded with duplicate lines | Slug-keyed upsert; whitespace-normalized dedup; 4000-char write cap (refused, never truncated — it's security text); 6000-char defensive read clip in the judge; same hygiene in the /guardrails person-rule path |

## 4. Known gaps — documented, not yet closed

| ID | Sev | Gap | Position |
|---|---|---|---|
| G1 | Med | Codex `apply_patch` file writes: hook matcher for Codex remains Bash-only; whether Codex file-mutation surfaces are covered is UNVERIFIED | Left unchanged rather than guessing hook names; verify against a live Codex fleet bot before relying on F2 parity there |
| G2 | Med | Media turns (voice/photo/document) from externals bypass inbound triage (text-only hook) | Output vet still covers everything they receive; close alongside G1 in a follow-up |
| G3 | ~~Low~~ CLOSED (P3) | ~~The `apr:` button handler (actions.js) has no test harness~~ Closed with P3: test-approval-cards.js pins the handler end-to-end (real approvals + entities + mandates, stubbed transport) — reply Always-allow delivery + scoped grant + ack ids, double-tap idempotence, deny hold, owner-only gating, triage approve wake job + deny decline, make-a-rule scope buttons, expired-card ack | Was: pre-existing posture; grant validation, card content, and record shapes were already pinned in test-enforcer.js |
| G4 | High (org) | Identity fragmentation: auto-contacts key on (adapter, userId) — the same human on Telegram + Kazee + Spaces mints separate records with separate mandates. An owner grant lands on one-third of a person (mechanism verified in code; live counts to be confirmed on the deployed store during the repair pass) | Redesign P1 |
| G5 | Med | No group layer: nothing between "this channel" and "this person" — can't say "the dev team may deploy to staging" | Redesign P2 |
| G6 | High (authz creep) | "Always allow" writes a verbatim-ask mandate line, forever, with no scope or expiry — over time a person accretes broad standing permissions the owner never reviewed as a set | Cap+dedup (F6) bound the bleeding; real fix is scoped grants, Redesign P3 |

## 5. Test coverage after this pass

All green locally (`npm test`; one known loaded-runner flake class in
test-utility-provider-policy re-verified 3× clean):

- **test-enforcer.js** (+137 lines): triage kind end-to-end (owner no-op,
  allow, block→card with handle/decline buttons + pinned record, fail-closed),
  turn grant (reply-only, quotes the ask, cache-key isolation, and rejection
  of pending / wrong-channel / wrong-kind ids), retry semantics, fail-closed
  + never-cache, audit hygiene (no payloads/mandates in logs).
- **test-tool-guard.js** (+60): file gate dormant without channel env, denies
  unknown-external Write/Edit/NotebookEdit, reads+benign Bash stay free,
  owner turns untouched, audited rule, settings carry both matchers on one
  hook command.
- **test-mandate-hygiene.js** (new): slug-keying (two Sharons never share),
  dedup, cap-refusal leaves bytes identical, name-fallback for legacy records.
- **test-provider-gateway.js**: router consults triage with the routed text +
  provider/project/dir pinning on every text turn.
- Plus the previously-passing enforcer/scheduler/hook/approval suites.

---

## 6. Redesign plan (approval requested — nothing below is built)

### P1 — One person, many surfaces (identity unification)

Anchor: **people.id stays the canonical key; Kazee userId becomes the
preferred anchor handle** (it is unique across all Kazee tasks, so every task
thread resolves to the same person — no per-task duplicates).

- **Agent-proposed, owner-confirmed merges.** When a new identity appears (or
  at repair time), the agent gathers signals — same display name, same email
  in profile, cross-references in conversation ("this is Sharon from
  accounts") — and posts a card: *"Is `telegram: Sharon N. (…111)` the same
  person as `kazee: sharon@… `? [Same person] [Different] [Not sure]"*.
  NEVER auto-merge, and never merge on display name alone (names are
  spoofable — that invariant already exists in ensureContact and stays).
- **Merge = handles fold into one people record**; entities merge Mandate
  (union, deduped, capped — over-cap requires owner trim), Notes/Log append.
  The losing record is tombstoned (`mergedInto: <id>`) so old approval
  records/audit lines still resolve. Merges are reversible from the tombstone.
- **Repair pass** for existing stores: one-time sweep proposes merge cards for
  probable duplicates (e.g. the fragmented Sharon records on the deployed
  bot); owner confirms each. Nothing merges silently.
- Guard changes: `speakerFor` unchanged (already person-keyed); the win is
  that one mandate follows the human across surfaces.

### P2 — Groups as a policy layer

Precedence becomes: **floor > global > channel > group > person > turn-grant**
(same rule: grants union, any prohibition wins, floor absolute).

- A group is an owner-curated named set of person ids (`groups.json`), e.g.
  `dev-team`, `family`. People can be in several.
- `/guardrails` gains group scope; the "Make a rule…" card offers *this
  person / this channel / this group / everyone*.
- The enforcer prompt gains one GROUP RULES section (union of the speaker's
  groups, clipped like mandates). Cache key includes it.
- No nesting, no inheritance between groups in v1 — flat sets only.

### P3 — Capability-scoped approvals (kills the blanket)

Replaces free-text "Always-allowed by owner: <verbatim ask>" with a
structured grant the owner explicitly shapes at approval time:

```
- grant: {capability} · scope: {this-exact-thing | this-kind-of-thing}
  · for: {person | group} · expires: {7d | 30d | never} · granted: 2026-07-24
```

- The card's "Make a rule…" flow asks two more taps: *what* ("just this" vs
  "requests like this — <agent-drafted capability line the owner can edit>")
  and *how long* (7d / 30d / no expiry). Default is the NARROW option.
- Stored in the same Mandate section (human-readable lines, same cap), so the
  judge consumes them without a format break; expired lines are skipped at
  read time and swept by the nightly dream pass with a chat notice.
- `/guardrails` lists a person's active grants with per-line revoke — the
  owner finally sees the accumulated permission set in one place.
- Rationale: an approval should carry WHO + WHAT KIND + HOW LONG, not become
  an eternal verbatim precedent (this is the "blanket approve" concern raised
  on the always-allow flow, solved at the model layer the judge already reads).

### P4 — Rollout order

1. **P3** first (smallest blast radius, immediate authz-creep fix; pure
   card-flow + mandate-format change, judge prompt already compatible).
2. **P1** second (unification + repair pass; needs the merge card UI and the
   tombstone plumbing; highest data-migration care — every step
   owner-confirmed, reversible, backed up like dream merges).
3. **P2** last (new layer; cleanest once identities are unified so group
   membership means one person, not one fragment).
4. Close G1/G2 (Codex file-tool verification, media-turn triage) alongside.
   (G3, the `apr:` handler harness, shipped with P3 — see §4/§9.)

Risks: P1 merge cards interrupt the owner (batched, max a few per day);
P3 makes approvals two taps longer (only on "Make a rule", never on
approve-once); all three change nothing for the owner's own turns.

---

## 7. Pass 2 — 360° dangling-pathway sweep (2026-07-24, post-go)

After the go, one full pass enumerating every EXISTING entry point that
reaches the same concerns as the new gates (the /intros lesson: a new system
must reconcile every old pathway, not just the one it was built on). Findings,
all fixed and test-pinned this pass:

- **D1 · CRITICAL — model output could write guard policy.** The per-turn
  pack reviewer and the nightly dream both passed `relationship` + `mandate`
  from model JSON into `upsertEntity`. Both run over mixed-provenance content
  (including external chatter), so a prompt-injected turn could have granted
  itself a mandate. Writers now strip both fields (owner actions remain the
  only writers), prompts/schemas updated, attack fixtures added
  (test-provider-pack-review, test-provider-dream). Introspection's
  entity_edits applier verified notes-only.
- **D2 — media turns bypassed inbound triage (closes G2).** Voice, audio,
  photo, and document handlers ran the agent without `guardInboundTriage`;
  only text was triaged. All five inbound shapes now enter through one
  `admitGuardedTurn` helper; media asks carry saved file paths so an
  owner-approved re-run can still find them, and the resolved ask is
  back-filled into the turn scope (`setCurrentInboundText`) so owner cards
  and F5 turn-grant validation see the real request, not an empty string.
- **D3 — dream auto-merges could delete guard policy.** An entity merge
  could remove a mandate-carrying entity or orphan a people-record
  `entitySlug` pointer (silent speaker-resolution break). Merges touching
  either are now skipped (`guardPinned`) — owner-confirmed merges only
  (arrives with P1).
- **D4 — Bash was an open side door around F2/F6.** On a guarded turn one
  heredoc could edit config, packs, crons, tool sources, or the installed
  package. New speaker-gated control-plane gate in the PreToolUse hook:
  any command referencing the config dir, package root, or `.open-claudia`
  is denied for externals (reads too — read-then-quote leaks policy), ahead
  of the tooling-mode switch, audited as `external-control-plane`.
  Authorised channels stay exempt (bin/tool.js gates externals itself).
- **D5 — Codex file edits were unhooked (closes G1, live-probed).** Probed
  real codex 0.144.0 with a log-everything hook: shell arrives as tool_name
  `Bash`, file edits as `apply_patch` — and normal runs bypass the sandbox,
  so the `^Bash$` matcher left externals able to write ANY file under Codex.
  Matcher widened to `^(Bash|apply_patch)$`, `apply_patch` added to
  FILE_MUTATION_TOOLS, fixture + pins updated.
- **D6 — two mandate writers, drift risk.** The always-allow grant writer
  (actions.appendMandate) and the rule-card writer (router) each implemented
  dedup/cap/slug-keying. Unified on one `appendMandateLine` so hygiene can't
  drift.
- **D7 — keyring CLI was speaker-blind.** `keyring exec` runs an arbitrary
  command with the FULL keyring in env, `get` prints raw values — and the
  control-plane gate exempts `keyring exec` as an authorised channel, so on
  an external turn it was a creds-plus-config-write side door around D4. The
  keyring CLI now refuses guarded external speakers outright (every verb,
  audited as `external-keyring`), dormant owner-local, denied on resolution
  failure.

Verified-OK pathways (no change needed): scheduler re-runs stamp
`OC_CHANNEL_USER_ID` (file gate stays live on approved external re-runs);
Spaces turns flow through the common router (triage covered); sub-agents
spawn read-only and inherit channel env; codex transport confinement
(hooks on, plugins off, untrusted trust) is mandatory-or-refuse. unified_exec is
no longer forced off (2026-09-05): Codex 0.153.x locks it on and its shell calls
still reach the PreToolUse hook as "Bash".

G3 closed with P3 (test-approval-cards.js pins the `apr:` handler). `tool run`
remains the one authorised channel reachable on external turns — by design:
bin/tool.js vets write/destructive tiers per speaker (mandate-allow or
owner escalation; vet logic pinned in test-enforcer.js /
test-provider-enforcer.js), and read-only verbs answer without side effects.

---

## 8. Ship state

Go received 2026-07-24. All §3/§5 fixes plus the §7 sweep are committed and
pushed (v3.1.25, 3025f53); fleet roll/publish is confirmed separately
(restarting the bot kills the live session). §6 starts next: P3 → P1 → P2.

---

## 9. P3 — capability-scoped approvals (2026-07-24)

The owner's driving concern, verbatim: "the approval all is a blanket for all
users or how can we get it to be for the user plus maybe what we want them to
be able to do as a guardrail vs a clanket approve everything going forward for
all or event that person." A mandate line is now a scoped, expiring, revocable
GRANT — not a permanent blanket.

New module `core/mandates.js`; the three policy layers now share one
lifecycle (id → list → revoke):

- **Ids.** Every mandate bullet has a stable 8-char id =
  sha1(slug + normalized text) — slug-scoped, so the same capability granted
  to two people revokes independently. Ids ignore case, whitespace, and the
  expiry marker.
- **Expiry.** A trailing ` [until YYYY-MM-DD]` marker (UTC, inclusive).
  Tap-grants ("Always allow" on an approval card) default to a bounded TTL —
  `MANDATE_GRANT_TTL_DAYS`, 30 by default, 0 = permanent. Owner-authored rule
  cards stay permanent unless the text ends "for Nd" / "until YYYY-MM-DD"
  (`extractExpiry`). Filtering is deterministic and read-side:
  `relationship.speakerFor` (the single choke point every consumer of
  `speaker.mandate` goes through — enforcer, external-mode prompt, tool vets)
  drops expired bullets via `activeMandateText`, so the judge never sees a
  lapsed capability and never does date math. The enforcer verdict cache keys
  on the filtered text, so expiry busts it naturally.
- **Refresh, not stacking.** `appendMandateLine` dedups on the
  marker-stripped text: re-granting the same capability replaces the line
  with the new expiry (`refreshed: true`). Successful writes also prune
  expired bullets and collapse legacy duplicates; refused/duplicate writes
  mutate nothing (cap-refusal pin unchanged). Non-bullet hand-written text
  survives rewrites.
- **Revoke.** `/guardrails` now lists per-person grants (`👤 name (slug)`,
  ids + until-dates) alongside global + channel rules; `remove|rm|revoke <id>`
  resolves across all three layers (global → channel → person). Revoking the
  last line empties the section via direct `writeEntity` (upsertEntity treats
  "" as no-op). Approval-card acks quote the id + expiry and the revoke
  command.

Pathways reconciled (the §7 discipline applied to the new system):
`bin/entity.js persona --mandate` is a deliberate owner-only WHOLESALE editor
(same class as hand-editing the entity file) — the parser tolerates whatever
lands there and the read filter still governs what the judge sees. Dream and
pack-review remain stripped of mandate writes (D1). All other writers funnel
through `appendMandateLine` (D6).

Pinned in test-mandate-hygiene.js (expiry marker, refresh, prune-on-write,
id scoping, extractExpiry, list/revoke, last-line clear) and
test-external-rule-scopes.js (speakerFor drops expired lines while the file
still holds them). G3 closes here too: test-approval-cards.js exercises the
`apr:` button handler end-to-end with the real approvals + entities +
mandates modules and stubbed transport — Always-allow delivers AND writes the
scoped grant (TTL marker, ack quotes id + revoke command), double-tap can't
flip a decision or duplicate a grant, deny holds with no delivery and no
mandate write, non-owner taps bounce, triage approve wakes the agent in the
origin channel with the approval id as a one-off turn grant (request framed
strictly as data), triage deny sends the polite decline without waking the
agent, "Make a rule" delivers and offers person/channel/global scope buttons,
and stale cards get a clean ack.
