# Lumine delegated-administrator contracts

`lumine admin` exposes deterministic community-management primitives. It does
not decide whether content is good, harmful, original, deserving of effort, or
worth commenting on. An LLM or human operator makes those judgments from the
canonical structured data.

## Security and run model

- The saved Lumine login always authenticates the real operator. The API reloads
  that user's current role from the writer database on every request.
- Only the server-configured administrator with authoritative administrator
  management authority may delegate. Effective Level 5 alone is rejected.
  Only the immutable
  server-owned Zero and Ciel user IDs are approved. Usernames and CLI flags are
  not authority.
- `daily-run start` creates or returns a six-hour `delegated-admin` run with
  one allowlisted run scope and one public actor. `--scope full` is the full
  daily-management review. `--scope featured` is only an authorization
  envelope for a specifically requested Featured slice; it does not authorize
  or imply the newspaper, queues, conduct review, logs, costs, AI Energy
  budget health, sponsors,
  carry-over work, or final full-run report. Every run-scoped CLI command loads
  the canonical active run and sends its ID; the API rejects a missing,
  expired, scope-mismatched, or actor-mismatched run.
- `--scope newspaper` is a newspaper-only authorization envelope. It grants
  newspaper status/claim/submit/print and completion only, with comment mode
  `off`. It cannot inspect queues, conduct, logs, Builds, or Featured.
  Completion does not advance full-review windows or carry-over telemetry.
  Its distinct `/daily-runs/start/newspaper` endpoint fails closed on older APIs.
- The public content actor is Zero or Ciel. Mikey's operator ID is retained in
  private audit rows and is not embedded in public comment metadata.
- Delegated HTTP work never authenticates as the bot, opens a bot socket, changes
  bot sessions, or updates bot presence/typing state. Normal content mutations
  still emit Twinkle's canonical real-time content events.
- One deliberate exception to the last-seen rule: a delegated mutation that
  actually changed something stamps the acting bot's `users.lastActive`, so
  Zero's and Ciel's public "last online" reflects the real day they commented on
  and rewarded kids' posts instead of whenever a socket last closed. It is
  written after the audit transaction commits, throttled to once a minute, and
  best-effort — it can never fail or deadlock the content mutation. Presence is
  still untouched: no socket is opened and no `online_status_changed` is emitted,
  so the bots remain absent from the online list. Note that `lastActive` is the
  ordering key for the People directory, so the bots now surface there after a
  run; that is the intended consequence of the timestamp being truthful.
- A later human reply to a delegated Zero/Ciel comment enters the existing
  server-side autonomous comment-assistant pipeline. For Build threads this is
  limited to comments carrying the private reviewed-management provenance
  described below. The human's normal AI Energy and sponsor path applies;
  Lumine does not need to remain running.

`--output <file.json>` saves the canonical JSON result for ordinary admin
commands, whether or not `--json` is also used. Files are atomically replaced
with mode `0600`. Paginated scans keep their existing streaming-to-file path;
`news claim` keeps its claim/scaffold contract. If saving fails after a server
mutation, the structured error includes `canonicalResult`: inspect it rather
than repeating a successful mutation to recover an output file.

The Zero/Ciel shift is assigned by **Asia/Bangkok calendar day, not by run**.
The schedule is anchored with 2026-08-31 assigned to Zero and alternates by
elapsed Bangkok dates from there. Every automatic primary, supplemental,
retry, read-only, or resumed run started on one date uses that date's same
actor; a skipped date still counts in the alternation. Completing, failing,
abandoning, or expiring a run never changes the schedule. An explicit
`--identity zero` or `--identity ciel` is an auditable override for that run
only: it does not rewrite the scheduled identity. Automatic runs expire at the
next Bangkok midnight so yesterday's actor cannot continue as today's
automatic identity. Reusing a live run key still returns the original run and
identity. `identity use zero|ciel` is the separate, explicit persistent
override for future starts; it remains in force until `identity use auto`, and
does not alter the underlying calendar schedule reported by status commands.

Comment mode is stored only on the current run:

- `off` (default): no draft or post scope.
- `draft`: server-generated drafts, no public comment.
- `post`: drafts plus idempotent publication through the ordinary comment path.

A Featured-only run always uses comment mode `off` and grants Featured
subject inspection, subject reveal, Featured mutation, review-bound Featured
comment encouragement (`featured:comments`), and run completion. It does not
grant generic recommendation/reward authority or comment publication. It can complete
without a sponsor-integrity scan. Completing it does not advance any full-run
content/cost/conduct window, queue-coverage record, last-completed identity, or
carry-over surfacing telemetry. Start one only when Mikey requested that slice; never turn a small
request into a full review merely because the technical command needs a run.
The API enforces these scopes, and the CLI also rejects out-of-scope operations
before calling endpoints outside the Lumine Admin router (notably Build review).

Featured capacity and delegated addition authority are separate policies:
the website editor supports 100 Subjects, while delegated additions are limited
to a board of 20. `maximum` remains the delegated limit for compatibility;
`maximumScope`, `delegatedMaximum`, and `websiteMaximum` label the distinction.
A board of 25 is not a website overflow or an instruction to remove five pins.
Equal-size replacements and complete-set reorders accept up to the website's
100-Subject capacity without exercising delegated growth authority.

### Private AI-bucket maintenance

AI identity buckets are private operator bookkeeping, not a Zero/Ciel public
action. They therefore do not require or attach to a delegated daily run:

```bash
lumine admin ai-bucket create --label Lemon \
  --note "Quota accounting only; not a moderation flag." --json
lumine admin ai-bucket get --bucket-id 10 --json
lumine admin ai-bucket accounts add --bucket-id 10 \
  --user-ids 3127,13037,15410,16288 \
  --note "operator-confirmed account family" --json
lumine admin ai-bucket note set --bucket-id 10 \
  --note "Quota accounting only; not a moderation flag." --json
```

`create` inserts a new unbanned quota bucket through the same helper the
management page uses. `--label` is required (at most 120 characters). `--note`
is required and follows the same 255-character quota-context rule as
`note set`. It cannot create a banned bucket, copy an existing one, or infer
members. The response returns the canonical bucket, including its id for
later `get` / `accounts add` / `note set` calls.

`accounts add` accepts 1-500 unique positive user IDs, preflights the complete
batch before writing, adds one exact canonical user rule per requested account,
re-attributes current-day AI usage, and returns the canonical bucket members.
It never infers or adds email aliases: shared verified addresses can belong to
unrelated accounts, so email rules require a separate explicit operator action.
The command is idempotent to retry. The API records the real operator in the
private Lumine audit log; no public bot identity is involved.

`note set` records up to 255 characters of private operational context on the
canonical bucket. Use it to distinguish quota bookkeeping from moderation;
the note itself changes no access, ban, or identity rules.

This surface is quota bookkeeping only. It cannot ban accounts, block signup,
add IP/device/risk-key rules, or infer an account family. Identification
remains a human/LLM evidence judgment and must be explicitly requested by
Mikey; routine administrator runs still escalate suspected alternate accounts
and never auto-enforce. `accounts add` and `note set` still accept only an
existing **unbanned** bucket.

### Shared verified-email policies

A verified address proves that an account can use that mailbox; it does not
always prove that every account using the mailbox is one person. When a teacher,
school, or other provisioning owner intentionally verifies separate people's
accounts with one address, record that confirmed operational fact explicitly:

```bash
lumine admin ai-email-policy get --email teacher@example.com --json
lumine admin ai-email-policy set --email teacher@example.com \
  --mode separate_accounts \
  --note "Teacher-confirmed classroom provisioning address" --json
```

`separate_accounts` preserves ordinary email verification and recovery while
giving every current and future account whose stable verified address matches
its own AI Energy and AI Card quota identity. The mutation re-attributes the
current day's per-account usage, updates all affected live quota projections
from canonical server state, and records the real operator in the private
Lumine audit log. It is a provisioning policy, not an alternate-account finding
or moderation flag, and must never be inferred from account count, usernames,
devices, IPs, or email-provider heuristics.

The CLI accepts a successful `set` only when the writer response confirms the
requested policy and reports every matching account on its exact expected
durable projection. A missing or mismatched projection fails closed instead of
letting an operator mistake a stored policy row for completed account repair.

Use `--mode automatic` with a new explanatory `--note` to restore Twinkle's
default verified-email grouping. Both modes are durable and apply to later
signups and later verifications without another code change. An active
email-wide manual identity rule conflicts with `separate_accounts`, so the API
rejects either operation until the operator explicitly resolves the conflict.
Exact per-user manual bucket rules remain authoritative in both modes.

### Audited identity inspection

Account-family evidence is private operator work, never a Zero/Ciel public
action and never a reason to browse unrelated user activity. Use the dedicated
lookup instead of direct database queries:

```bash
lumine admin identity inspect Jay1216 \
  --reason "Confirm the account family before updating its quota bucket" --json
lumine admin identity inspect Jay1216 \
  --reason "Mikey requested DOB and email evidence for this decision" \
  --include-private-evidence --json
```

The default result resolves the exact username or user ID from the writer,
returns the canonical current AI bucket, orders candidate accounts oldest
first, identifies the oldest account within the strongest canonical family
boundary (bucket before email; never device alone) when that family fits in the
bounded evidence set, reports whether each
account has a DOB, and explains whether the link
came from explicit bucket membership, a verified-email match, or bounded exact-
device evidence. It does **not** reveal email addresses, DOB values, device IDs,
IP evidence, private messages, or unrelated activity. `--include-private-evidence`
adds DOB values and verified email addresses only; exact device IDs are never
returned.

Every inspection requires a concrete `--reason`. Before loading the evidence,
the API commits a private `identity.inspect` audit receipt containing the real
operator, requested target, reason, and whether private evidence was requested.
The evidence itself is deliberately not copied into the audit log. Inspection
is run-independent and has no public actor. Results are **candidate accounts
for human judgment**, not an automatic ownership finding and never automatic
grounds for moderation, bans, or bucket changes.

### Audited private investigations

Use reason-required, bounded drill-downs when an aggregate management signal
needs causal evidence:

```bash
lumine admin economy trace lock --days 3 \
  --reason "Investigate the anomalous three-day coin gain" --json
lumine admin rescue wordle-audit --days 30 \
  --reason "Identify recorded Wordle breaks, streak lengths, and rescue status" --json
```

`economy trace` reads the canonical append-only coin ledger for one exact
account. It returns an exact action/target breakdown for the window, the largest
ledger entries, direct transfer counterparties, AI Card sale/offer provenance,
and whether a counterparty is already in the same effective exact-account AI
bucket. It never reads or returns chat messages, bucket labels, verified
addresses, device evidence, or private identity values. Same-bucket transfers and concentration are
investigation signals, not automatic abuse findings.

`rescue wordle-audit` separates expired unredeemed Wordle/strict-Wordle rescue
offers from offers still inside their seven-day promise and from redeemed
offers. Each row identifies the account and labels the cause: regular Wordle is
an uncovered dodge, while strict Wordle is a completed but non-strict game. The
stored offer-time streak snapshot is reported as the broken streak length.
`claim_ready` requires a durable first user exchange in Lumine history; a
reservation claim is reported separately and is not treated as completion. An
open offer is never labelled a refusal, and non-redemption does not establish
intent. The underlying rescue table stores one current row per account; an
expired row may be replaced by a later offer, so the command reports exact
current-row evidence and explicitly marks that it is not a complete immutable
history of every offer ever shown.

Both commands are run-independent private operator actions. Each requires a
concrete reason and commits a minimized private access receipt before loading
the evidence. Windows are limited to 1-30 days, and neither command mutates
coins, streaks, buckets, messages, or public content.

## Editorial priorities

The CLI never makes the qualitative judgments in this section; they are the
standing instruction for the operator or agent and apply to every verb below:
recommends, rewards, effort levels, Featured, skips, comments, and replies. It
does enforce deterministic server-provable boundaries documented below, such
as the posting-date and lifetime-history gates for a new Featured addition.

Public text authored as Ciel must be English. This is an operator and generation
instruction, not a script or keyword test: writing systems do not identify a
language reliably, and the API must not pretend otherwise. This is a
presentation rule, not an invitation to correct or lecture a member who writes
in another language; reply naturally in concise English.

**Twinkle is not Reddit.** Do not rank a run's attention by popularity,
recommendation count, or polish. Most users here are young children, and the
posts that most need Zero or Ciel are the ones nobody else answered. These
outreach priorities do not replace the full daily Featured-comment review
below: active threads need encouragement too. Featured selection has its own
primary objective: invite genuine member-to-member engagement, not assemble
the most thought-provoking posts.

- **Look first at new, quiet, and overlooked users.** A child's first post, or a
  post from someone who rarely gets replies, is worth more of a run's attention
  than another well-liked post that already has a lively thread.
- **Clumsy is not the same as low-effort.** Bad spelling, a one-line
  description, a title that is just "hi", a drawing that did not come out right
  — these are usually a child trying, not a child spamming. Read for the real
  thing they were reaching for and respond to that.
- **Zero engagement is a reason to act, not to skip.** A post sitting at no
  recommendations and no comments is the strongest signal in the queue that
  someone should notice it.
- **Thought-provoking posts deserve substantive replies.** When a child asks
  a real question or makes a real argument and the thread is empty, it is a
  strong case for a Zero/Ciel reply, not an automatic top Featured ranking.
  Engage with the idea itself: answer it,
  add a perspective or a counter-consideration, and leave the author somewhere
  to go next. A good question that nobody answered teaches a child that thinking
  hard is not worth it; that is the outcome these runs exist to prevent. This
  cuts both ways with the point above — the two ends of the queue, the beginner
  nobody noticed and the strong idea nobody engaged, both outrank the popular
  post that already has a lively thread.
- **Always answer Twinkle usage questions.** "How do I change my username",
  "why can't I reward", "what unlocks the Summoner" — a child stuck on the site
  cannot use it. Answer concretely and verify anything you are unsure of against
  the canonical rules before publishing; say you will check with Mikey rather
  than guessing at mechanics.
- **Always respond to bug reports, and tag `@mikey` in the comment.** The
  mention is what notifies him (`postComment` runs `processMentions` /
  `postMentions` and emits `new_targeted_upload`), so a bug-report comment
  without `@mikey` fails its main job. Restate what the child observed; never
  promise a fix or a timeline. This reporting duty does not suspend the acting
  bot's character: the public comment must still sound like Zero or Ciel
  talking naturally to that member, not like an operator, ticket, or incident
  report. Read the full thread first. If the bot already noticed, explained, or
  apologized for the bug, do not post another reply that treats it as a fresh
  discovery or repeats the same acknowledgement. Continue from what the bot
  already said and add only what is missing, such as naturally bringing
  `@mikey` into the conversation. Put the formal defect summary, evidence, and
  review request in the private run escalation.
- **Guide users through the website, without waiting to be asked.** A post can
  show that a child is stuck, confused, or unaware a feature exists without ever
  containing a question — someone begging for coins who does not know about
  daily rewards, someone reposting because they could not find their own post.
  Give them a short, friendly crash course on the exact thing they are stuck on.
  Teaching a child to use the site is worth more than any single recommend.
- **Reserve `post skip` for genuine noise** — card-sale and coin-begging spam,
  keyboard mash, duplicates, engagement farming — not for sincere posts that
  merely look unimpressive.
- **Effort levels are not a verdict on the child.** Level 1 on a thin post is
  ordinary bookkeeping; it never means the author deserves less attention, and
  it pairs well with a warm comment. Judge the depth the Subject invites, not
  merely the number of words in its prompt: a genuinely thought-provoking
  question should normally receive Level 3. Level 3 is the highest delegated
  setting, and the level tells respondents that substantial, carefully
  reasoned answers are wanted.
- **Ongoing series are participation, not noise.** Daily records, journals,
  logs, recurring updates, and serialized creative work are not spam, filler,
  engagement farming, or duplicates merely because they reuse a title or
  format, or because one installment is concise. Review the series context and
  what the current entry contributes. Never lower effort, skip, withhold a
  recommendation or comment, or propose Featured removal solely because of
  that recurring format or brevity; only an actual duplicate or genuine noise
  is treated as such.
- **Sincere requests for personal help are Featured material.** Posts asking
  the community for advice about school, friendships, loneliness, or other
  ordinary real-life problems embody Twinkle's purpose; being personal is never
  a reason to suppress them or judge them unsuitable. Receiving meaningful
  support does not reduce their value. They still participate in the daily
  Featured refresh below: rotating the spotlight is not withdrawing support,
  and a good post does not require an indefinite pin.
- **Speculative privacy is never an editorial signal.** Do not lower, remove,
  unfeature, or escalate content because it might identify someone, mentions a
  school, city, class, or ordinary location, or invites everyday community
  context. Those possibilities carry zero weight in Featured decisions. An
  actual sensitive disclosure or an author's explicit removal request follows
  the separate concrete-safety path; it does not make the post low-quality.
- **Featured prioritizes engagement, participation, and child voice.** Choose
  the recent eligible Subjects most likely to get members commenting on other
  members' Subjects and responding to one another. Accessible questions,
  everyday experiences, playful prompts, drawings, invitations, personal help,
  and ongoing records or creative series can outperform a polished essay for
  this purpose. Being thought-provoking is a bonus, not the primary ranking
  criterion. Explain the concrete invitation to participate, using the full
  Subject and thread as context; do not rank by adult polish, intellectual
  depth, existing popularity, or raw recommendation counts. Genuine engagement
  does not mean bait, spam, or reward farming. Age, brevity, simplicity, prior
  recognition, or a modest description does not make a Subject low-quality.
  Daily rotation is an independent editorial reason to refresh a good post's
  slot; do not invent a defect to justify it.
- **Aim to replace all of yesterday's Featured Subjects each day.** During a
  full daily management review, plan a complete refresh of the previous day's
  board with recent, never-before-Featured Subjects, ordered by likely genuine
  engagement. A small handful of swaps is not the default target. Use the
  Bangkok calendar day, not the number of runs: do not churn pins newly added
  today because a follow-up starts another session. Quality and eligibility
  still apply. If there are too few eligible candidates, a current explicit
  keep instruction, an unread thread, or another concrete constraint, explain
  each retained pin and the shortfall; do not pad the list or silently settle
  for a partial refresh. An earlier positive assessment is not a permanent
  keep instruction. New installments can carry a series forward without
  treating earlier ones as noise. This daily target changes editorial judgment,
  not mutation authority: show the full proposed rotation and obtain Mikey's
  go-ahead before removing or reordering any pins. One go-ahead for that exact
  complete plan authorizes executing all its swaps, not just the first one;
  carry it through and verify the final board without seeking per-item approval.
- **A "new" Featured addition has two non-negotiable eligibility gates.** When
  Mikey asks for new Featured subjects, additions, or replacements, a candidate
  must both (1) have been posted recently and (2) have never appeared on the
  Featured board before. Unless Mikey gives a different recency window,
  "recently" means the preceding seven calendar days; do not widen into older
  inventory merely to fill capacity or produce a longer proposal list. Absence
  from the current `featured list` proves only that a Subject is not Featured
  now — it does not prove that the Subject has never been Featured. Verify
  lifetime Featured history from canonical evidence before recommending or
  adding it. If the available CLI/API cannot prove that history, leave the
  candidate out and report the capability gap instead of guessing. Quality is
  still required after both eligibility gates pass; recent and never-Featured
  does not make filler acceptable.
- **The live Featured board is Mikey's word.** Do not remove or reorder a
  currently Featured subject without first showing Mikey the planned removals
  and replacements and getting his go-ahead. Additions have standing approval
  when a fresh `featured list` is below its delegated-admin maximum: proactively
  feature genuinely reviewed, editorially strong new subjects without asking
  case by case. This is judgment, not a quota — never add filler merely because
  capacity exists. If a subject that was Featured or pinned during an earlier
  run is no longer on `featured list`, treat that as Mikey having removed it
  deliberately — never re-feature it to "restore" the board, even when space
  is available, and never treat any subject as a permanent fixture from memory
  or old run notes. Derive the board fresh from `featured list` at the start of
  every Featured review or mutation; the only pins that exist are the ones
  currently on it.

Sensitive disclosures, active disputes, and anything needing crisis or medical
judgment remain out of scope for a bot comment no matter how neglected the post
is. Those go to Mikey.

### Daily Featured comments and refresh

This is a standing duty for every **full daily management review**, after the
newspaper duty. It is not an instruction to start or repeat a full run when
Mikey requests a small edit, proposal, or other scoped action. A request only
to discuss Featured candidates authorizes discussion, not recommendations or
pin changes. The technical `--scope featured` envelope permits the review-bound
comment workflow below only when encouragement is in Mikey's requested slice;
it never authorizes unrelated daily work or generic recommendation commands.

1. **Read all comments on every current Featured Subject each day.** Start
   from a fresh `featured list`, save the dated board, and inspect each full
   Subject and all top-level comments and nested replies. Include old and
   already-viewed comments, not just new ones, the recommendation queue, or a
   sample of the thread. Review outgoing Subjects before their approved
   rotation so their commenters are not missed; review any newly added
   Subjects' threads too. Use `featured comments scan --checkpoint <file>` to
   download every snapshot-bound page, then actually read every listed page
   file before `featured comments acknowledge --checkpoint <file> --reviewed`.
   Fetching does not acknowledge reading. A fresh scan after rotation covers
   the incoming board. `commentsIncluded: false` or a locked secret is not
   an empty thread: use the ordinary authorized reveal path, never bypass
   secret semantics, and report any inaccessible or unfinished coverage.
   Preserve the reviewed snapshot boundaries so "all" describes actual reads.
2. **At least recommend most comments that are not genuinely worthless.**
   The purpose is to encourage members to comment on other users' Subjects and
   to give those contributions visibility. A sincere answer, friendly reaction,
   relevant joke, question, attempt to help, or conversational follow-up is
   usually enough; it need not be profound, long, polished, or Twinkle-worthy.
   Read in context and lean toward encouragement. A one-line or emoji response
   can be worthwhile; brevity or "nothing for the bot to add" is not a reason
   to withhold a recommendation. Do not restrict recognition to a few standout
   replies. Do not mechanically approve everything: genuine spam, meaningless
   noise, harassment, or exploit promotion is not deserving merely because it
   appears on Featured. System/view notifications and the bots' own comments
   are context to read, not user participation to boost or count toward this
   duty. Leave narrow concrete-safety cases for Mikey under the existing rules.
3. **Separate encouragement from reward eligibility.** The ordinary action is
   a selected item in `featured comments recommend --file <decisions.json>`,
   omitting `anyoneCanReward` and `rewardTwinkles` (defaults: `false` and `0`).
   Turn on the UI's "everyone can
   reward" option (`anyoneCanReward: true`) only when the comment independently
   deserves that stronger endorsement; direct Twinkle rewards are selective
   too. Do not withhold a basic recommendation just because rewards are not
   deserved, and do not mass-enable rewards to satisfy this duty. Honor
   canonical Zero/Ciel deduplication: already-recommended comments still get
   read, but do not need a second recommendation or daily reward. A bare
   recommend does not revoke an existing reward permission; never claim it
   switched one off. This duty is not authority to downgrade past decisions.
4. **Propose the whole daily refresh, get approval, then execute it.** Identify carryover
   pins from the saved board and canonical Featured history, without assuming
   missing pins should be restored. Review replacements under the recency and
   never-Featured gates above and rank them by the engagement they invite.
   Show Mikey a specific old-Subject -> new-Subject mapping with titles, IDs,
   canonical URLs, reasons, and the proposed final order, plus any retained
   pins and specific constraints. Explicitly ask for his go-ahead and wait
   before performing the proposed rotation. Daily freshness
   and giving new conversations a turn are sufficient editorial reasons; a
   prior Subject need not become bad or unsuitable first. Keep existing
   removal/reorder approval and addition-capacity boundaries intact. After
   Mikey approves the exact plan, autonomously perform **all approved swaps
   and the approved ordering**, without per-item approval or asking him to
   operate the website. Prepare the server-stored `featured plan` before
   presenting it, then use `featured apply --file <plan.json> --approve <hash>`
   only after that exact plan gets his go-ahead. Membership and final ordering
   commit together; the server rechecks the original board, history revision,
   and new-candidate eligibility inside the board transaction.
   Verify the exact final membership and order from the server, then report
   completion. Recover retryable failures within the approved scope; if
   intervening owner changes, eligibility changes, or an API limit require a
   materially different plan, preserve the new state, report what actually
   completed, and seek approval for that difference instead of substituting
   candidates or overwriting changes. Report proposals separately from
   completed actions. Approval is for the presented plan, not future boards.
5. **Make coverage and encouragement auditable.** In the final full-run report,
   state Subjects covered, comments read, new basic recommendations,
   already-recommended comments, selective reward-eligibility grants/direct
   rewards, and genuinely skipped categories. Use canonical results, not
   attempted commands, for action counts. Report any unread Subjects/pages
   explicitly; never call a sample, a filtered queue, or a blocked thread a
   completed all-Featured-comment review. Include refresh progress and any
   carryovers as required in the `Featured rotation` report section below.

## Math Lab content duty in full daily management

Mikey added Math Lab question design and publishing to the full daily workflow
on 2026-09-08. Follow [Math Lab daily question publishing](../../agent-guides/math-lab-daily.md)
for the canonical Build 2460, owner account, twelve grade queues of ordered
until-earned puzzles, verification, repeat-run recovery, release gates, and
final reporting. Every full daily run reports each grade's current question,
whether it was cleared today, uncleared published questions remaining, and
refill status. Count distinct cleared keys across all users and the app's full
history against the live approved sheet; the recent usage window and draft
additions are not the live inventory. At two or fewer remaining, prepare a
refill to at least ten, with complete interactive guides, and track it until
approved publication. Report one or zero remaining prominently. Details are
in the linked guide's **Daily queue monitoring and refill** section.
This is not part of Featured-only or newspaper-only work and is not a new
scheduler, delegated API scope, or automatic extension of admin permissions.
Use the expressly authorized owner Build workflow for Math Lab; retain the
normal Zero/Ciel actor separation for other administration.

Mikey authorized Math Lab's initial launch, completed on 2026-09-12. Routine
refills follow the standing duty but still need his exact-version approval
and explicit publication authorization. Local edits
and draft saves do not require release approval. Real XP/Coins may be changed
only by the currently published, approved artifact through server-verified
reward claims; private builds, previews, local tests, unpublished branches, and
superseded versions cannot award real balances. These reward controls are a
required design contract, not a claim that a reward SDK or server enforcement
has already been implemented. See the guide before adding reward capabilities.

## Escalation to Mikey

Before closing a full run, reconcile three explicit handoffs: pending reward approvals (`rewardReviews` in intake/report), every carryover todo (with new evidence or a concrete blocker and next action), and earlier-day telemetry that meets a reopening condition. `carryoverWithoutProgressThisRun` in the report identifies surfaced todos without a progress update. A pending human decision can remain open; it must be named with its exact request/version, recommendation and next owner. Never equate reading a summary with inspecting frozen implementation, clearing a stuck flag with producing the intended image, or deploying code with verifying its live outcome. The September 14 omissions were execution failures under already explicit duties; these fields make them visible, not optional.

A full daily management run is not finished when the mutations are done. Curation surfaces things only
a human owner can decide, and a finding nobody reports is a finding that did not
happen. **Every full run ends with an escalation list**, and it belongs in the run's
final report whether or not anyone asks for it.

Keep that list narrow enough to be useful. Escalate concrete child-safety,
exploitation, targeted harassment, or platform/system-abuse risk — not
ordinary children experimenting, arguing, making rumors, proposing informal
in-site loans or contests, asking where media can be found, or making an
unverified ownership claim. Those may merit a normal age-appropriate response,
but they are not escalations without credible harmful conduct or a real victim.
Hypothetical identifiability and ordinary school, city, class, or location
references are not escalation signals.

Escalate, with the canonical `https://www.twin-kle.com/subjects/<id>` or
`/comments/<id>` URL, a one-line summary, and why it needs him:

- **Child-safety and wellbeing** — distress or mental-health disclosures,
  anything about self-harm, a child asking for a photo of themselves to be
  removed, requests to delete or hide personal information, contact details
  posted in public, or a child who says they are leaving because something
  happened. These outrank every other category.
- **Account integrity** — someone posting from another person's account,
  impersonation, shared logins, or a user operating a set of alternate accounts.
- **Economy exploitation** — coordinated coin or XP farming across alternate
  accounts, coercive or deceptive arrangements, or a repeatable abuse of the
  platform economy with concrete evidence. A child offering a voluntary loan,
  repayment, prize, or contest is not enough by itself.
- **AI-cost exploits** — patterns that convert free AI allowances into farmable
  value: clusters of young accounts with heavy AI/battery usage, one person
  operating many accounts that feed a single build through team branches,
  plus-tagged or dot-variant email families (`kid+1@`, `k.id@`) behind multiple
  active accounts, repeated first exchanges across accounts already linked by
  independent evidence, or a rescue claim by an account created under 30 days
  ago (the API maturity gate should make that last case impossible). A new
  member's one first exchange is intended onboarding and is not suspicious by
  itself. The daily battery is real provider money; treat farming signatures
  with the same seriousness as coin farming. Escalate the account list and
  evidence; never auto-enforce.
- **Bug reports** the run encountered, even secondhand in a comment thread.

Two rules that keep the list worth reading:

- **Check the thread before escalating.** If Mikey already replied in it, the
  matter is his and it is closed unless something new happened after his reply —
  re-reporting it wastes the one channel that is supposed to mean "look at this."
  Say so explicitly when a post looks alarming but he already handled it.
- **Check `operatorViewed` too.** Every Subject, Comment, StandalonePost, and
  queue item reports whether Mikey has opened that content and when. Silence is
  not the same as not having seen it — he often reads without replying. Lead the
  escalation list with items where `viewed` is false, and mark the rest as
  already-seen rather than dropping them, since he may have looked before the
  thing you are escalating happened. `lumine admin subjects candidates --unviewed`
  and `lumine admin recommendations list --unviewed` filter a page down to what
  he has not opened (`--viewed` inverts it). The flags are supported by the
  recommendation, Subject, Featured, and comment-list commands only; they are
  rejected before any request on every other command.

  Two limits make this a strong negative signal and a weak positive one: a view
  is recorded only when the content **page** is opened, so reading a post inline
  in a feed records nothing, and a user's view of their **own** content is never
  recorded. So `viewed: true` reliably means he opened it; `viewed: false` means
  "no page open recorded", not "he never saw it". Never tell a child, in public,
  whether Mikey has or has not looked at their post.

- **Escalate; do not moderate.** Zero and Ciel have no moderation verbs here by
  design. Do not delete, hide, argue with, or publicly accuse anyone, and do not
  warn a child that they are in trouble. Report it and let Mikey decide.

Mikey's decision must remain attached after the originating run closes. These
commands are private, run-independent bookkeeping:

```bash
lumine admin escalation list --json
lumine admin escalation list --status all --json
lumine admin escalation set 123 --status acknowledged \
  --note "Mikey is reviewing the bot response" --json
lumine admin escalation set 123 --status resolved \
  --note "No user fault; this audit concerned Zero's response" --json
```

`list` defaults to unresolved `open` items. `--status` also accepts
`acknowledged`, `resolved`, or `all`. The number passed to `set` is the original
`run.escalation` audit ID returned by the run report/list. Every disposition is
an immutable private `escalation.status.set` audit event. A revision allocated
while the original escalation is locked orders concurrent decisions, so the
last applied decision is the canonical status and annotation even if request
audit IDs were reserved in another order. `--status open` can deliberately
reopen an item with an explanatory note. Status filters walk indexed escalation
history rather than a fixed latest-event window. No active run or public bot
identity is used.

## Common JSON types

All `--json` success output is one uncolored JSON value:

```ts
type Success<D> = {
  ok: true;
  status: "success" | "already_done" | "no_op" | "maximum_reached";
  changed?: boolean; // mutations only
  data: D;
};
```

Failures print one JSON value to stdout and exit nonzero. A failing `--all`
scan may already have written bounded page progress to stderr; stdout remains
protocol-clean:

```ts
type Failure = {
  ok: false;
  status:
    | "unauthenticated"
    | "forbidden"
    | "not_found"
    | "validation_error"
    | "partial_failure"
    | "internal_error"
    | "error";
  error: {
    code: string;
    message: string;
    details: unknown | null;
  };
};
```

Shared records:

```ts
// Present on Subject, Comment, StandalonePost, and every recommend-queue item.
// This is the OPERATOR's own view state (Mikey), never the bot's, and reading it
// records nothing.
type OperatorViewed = {
  viewed: boolean;
  firstViewedAt: number | null; // Unix seconds
  lastViewedAt: number | null;
};

type Author = { id: number | null; username: string | null };

type Attachment = {
  filePath: string | null;
  fileName: string | null;
  fileSize: number | null;
  thumbUrl: string | null;
  url: string | null; // canonical attachment URL when path and name exist
};

type RecommendationState = {
  recommendedByActor: boolean;
  actorRecommendationId: number | null;
  anyoneCanReward: boolean | null;
  count: number;
  items: Array<{
    id: number;
    actor: Author;
    anyoneCanReward: boolean;
    createdAt: number | null;
  }>;
};

type RewardState = {
  totalTwinkles: number;
  actorTwinkles: number;
  actorAlreadyRewardedThree: boolean;
  caps: {
    maxRewardAmount: number;
    maxRewardAmountForOnePerson: number;
  } | null;
  items: Array<{
    id: number;
    recipientUserId: number | null;
    rewarder: Author;
    type: string | null;
    amount: number;
    comment: string | null;
    claimed: boolean;
    createdAt: number | null;
  }>;
};

type Subject = {
  id: number;
  url: string; // https://www.twin-kle.com/subjects/<id>
  author: Author;
  createdAt: number | null; // Unix seconds
  updatedAt: null;
  title: string | null;
  description: string | null;
  attachment: Attachment | null;
  hasSecretAnswer: boolean;
  hasSecretAttachment: boolean;
  secret: { hasSecret: boolean; revealed: boolean | null };
  effortLevel: number; // 0 means unassigned
  effortRevision: number;
  featured: { member: boolean; order: number | null }; // one-based order
  createdByAuthor: boolean;
  recommendation: RecommendationState;
  reward: RewardState;
  root: { type: string | null; id: number | null };
  ageRestriction: string | null;
  deleted: false;
  unavailable: false;
};

type Comment = {
  id: number;
  url: string; // https://www.twin-kle.com/comments/<id>
  subjectUrl: string | null;
  author: Author;
  createdAt: number | null;
  updatedAt: null;
  content: string | null;
  contentHidden: boolean;
  attachment: Attachment | null;
  parentCommentId: number | null;
  replyToCommentId: number | null;
  isNotification: boolean;
  ageRestriction: string | null;
  deleted: false;
  unavailable: false;
  recommendation: RecommendationState;
  reward: RewardState;
};

type StandalonePost = {
  contentType: "aiStory" | "dailyReflection";
  contentId: number;
  id: number;
  url: string;
  author: Author;
  createdAt: number | null;
  updatedAt: null;
  title: string | null;
  question: string | null;
  content: string | null;
  explanation: string | null;
  imagePath: string | null;
  audioPath: string | null;
  deleted: false;
  unavailable: false;
  recommendation: RecommendationState;
  reward: RewardState;
};

type Pagination = {
  nextCursor: string | null;
  hasMore: boolean;
  exhausted: boolean;
  snapshotMaxId: number;
};

type Identity = {
  key: "zero" | "ciel";
  userId: number;
  username: string;
};

type DailyRun = {
  id: number;
  runKey: string;
  operatorUserId: number;
  publicActorUserId: number;
  identityMode: "auto" | "zero" | "ciel";
  commentMode: "off" | "draft" | "post";
  runScope: "full" | "featured" | "newspaper";
  sessionKind: "delegated-admin";
  scopes: string[];
  status: "active" | "completed" | "failed" | "expired";
  successfulMutationCount: number;
  startedAt: number;
  expiresAt: number;
  completedAt: number | null;
  failedAt: number | null;
  failureReason: string | null;
  identity: Identity;
  scheduledDay: string; // YYYY-MM-DD in Asia/Bangkok
  scheduledIdentity: Identity;
};
```

## Identity and daily-run commands

```bash
lumine admin identity list --json
lumine admin identity status --json
lumine admin identity use zero --json
lumine admin identity use ciel --json
lumine admin identity use auto --json
lumine admin identity inspect Jay1216 \
  --reason "Confirm the account family before a bucket change" --json
```

Schemas:

```ts
type IdentityList = Success<{
  identities: Identity[];
  preferredIdentity: "auto" | "zero" | "ciel";
  lastCompletedIdentity: "zero" | "ciel" | null;
}>;

type IdentityStatus = Success<{
  preferredIdentity: "auto" | "zero" | "ciel";
  lastCompletedIdentity: "zero" | "ciel" | null;
  scheduledDay: string;
  scheduledIdentity: Identity;
  scheduleTimeZone: "Asia/Bangkok";
  activeRun: DailyRun | null;
}>;

type IdentityUse = IdentityStatus;

type IdentityInspection = Success<{
  inspection: {
    targetUserId: number;
    privateEvidenceIncluded: boolean;
    manualBucket: {
      id: number;
      label: string;
      memberCount: number;
      isBanned: boolean;
    } | null;
    oldestAccount: IdentityCandidate | null;
    oldestAccountBasis: "manual_bucket" | "verified_email" | "target_only";
    oldestAccountComplete: boolean;
    candidateAltCount: number;
    accounts: IdentityCandidate[];
    evidenceCoverage: {
      deviceLookbackDays: number;
      targetDeviceEvidenceRows: number;
      targetDeviceIdsConsidered: number;
      relatedDeviceEvidenceRows: number;
      candidateLimit: number;
      truncated: boolean;
    };
  };
}>;

type IdentityCandidate = {
  userId: number;
  username: string | null;
  joinedAt: number | null;
  isTarget: boolean;
  isOldestAccount: boolean;
  hasDateOfBirth: boolean;
  banned: boolean;
  deleted: boolean;
  relationBasis: Array<
    "target" | "manual_bucket" | "verified_email" | "exact_device"
  >;
  sharedDeviceCount: number;
  privateEvidence?: {
    dateOfBirth: string | null;
    verifiedEmails: string[];
  };
};
```

`identity use zero|ciel` changes the explicit persistent override for future
starts; `identity use auto` returns future starts to the Bangkok calendar
schedule. It never changes an active run, alters the reported scheduled
identity, or advances rotation.

```bash
lumine admin daily-run start --identity auto --comment-mode off --json
lumine admin daily-run start --scope featured --identity auto --json
lumine admin daily-run start --scope newspaper --identity auto --json
lumine admin daily-run start --identity ciel --comment-mode draft \
  --run-key daily:2026-08-06:review --json
lumine admin daily-run status --json
lumine admin daily-run escalation add --target subject:123 \
  --note "Public contact details need owner review" --severity urgent --json
lumine admin daily-run escalation add --target chatMessage:3768159 \
  --note "Concrete safety issue in a bot-authored chat message" --json
lumine admin daily-run report --json
lumine admin daily-run report --run 123 --json
lumine admin daily-run complete --json
lumine admin daily-run fail --reason "operator stopped" --json
lumine admin escalation list --status all --json
lumine admin escalation set 123 --status resolved \
  --note "Final owner decision" --json
```

Schemas:

```ts
type DailyRunStart = Success<{
  run: DailyRun;
  scheduledDay: string;
  scheduledIdentity: Identity;
  scheduleTimeZone: "Asia/Bangkok";
  carryoverTodos: CarryoverTodos;
}>;
type DailyRunStatus = Success<{
  run: DailyRun | null;
  lastRun: DailyRun | null;
}>;
type DailyRunComplete = Success<{
  run: DailyRun;
  rotationAdvanced: false; // legacy field; calendar schedules never advance by run
}>;
type DailyRunFail = DailyRunComplete;
// `escalation add` echoes the same open-item shape `escalation list` returns,
// so the run.escalation audit ID that `escalation set <auditId>` needs is in
// the add response; no list round-trip is required. `recordedAt` (the audit
// row's timestamp) appears only in `list`.
type DailyRunEscalationAdd = Success<{
  escalation: {
    auditId: number; // run.escalation audit event ID = escalation identity
    runId: number;
    targetType: string | null;
    targetId: number | null;
    url: string | null;
    summary: string;
    severity: "attention" | "urgent";
    status: "open";
    decisionNote: null;
    decisionAuditId: null;
    decisionRevision: 0;
    decisionUpdatedAt: null;
    decisionByUserId: null;
  };
}>;
```

### Build Workshop sponsor applications and integrity

This is the approved Zero/Ciel Build Workshop sponsor role, not the ordinary
AI Energy sponsor flow. Applications originate only from `lumine sponsor`.
Website-management agents review them inside an active full daily run:

```bash
lumine admin sponsor applications list --status pending --json
lumine admin sponsor applications review 12 --decision approve \
  --note "Approved for probationary duty" --json
lumine admin sponsor status set 45 --status trusted \
  --note "Cleared probationary handoffs" --json
lumine admin sponsor integrity scan --json
lumine admin sponsor integrity cases --status open --json
lumine admin sponsor integrity get 34 --json
lumine admin sponsor integrity review 34 --decision clear --json
```

Run `sponsor integrity scan` until its bounded snapshot reaches
`awaiting_review` or `completed`. Every pending completed handoff receives the
deterministic checks. Probationary work, hard-flagged evidence, and a stable
random sample become review cases; clean trusted work outside the sample is
cleared automatically. A case includes the approved structured relay, canonical
artifact snapshot, branch-notice evidence, and requested/resolved provider,
model, effort, service-tier, runtime, usage, and agent-tree records. It never
includes raw Zero/Ciel chat.

`clear` qualifies a contribution to another user for its flat 50 KP award;
self-sponsored testing remains ineligible even when its integrity evidence is
cleared.
`disqualify` makes it ineligible. `hold` and `flag` require an evidence note and
remain open. The scan itself never changes sponsor status or applies a sanction.
Use the separate, audited `sponsor status set` command for an explicit human
decision. Full `daily-run complete` is rejected until the scan has covered its full
snapshot and no pending, held, or flagged case remains.

During a full review, record only qualifying escalations as they are confirmed.
`daily-run report`
then composes the active run's canonical audit events, successful mutations,
completed queue scans, recorded escalations, and the most useful brief deltas
into one result. Generate it before `complete`, because run-scoped reads require
the current active run. After completion,
`daily-run report --run <completed-run-id>` is the run-independent recovery
path. It reconstructs immutable run/audit/coverage/escalation evidence and the
closed-day calendar cost view at the run's completion boundary. It labels that
basis `historical_reconstruction`: later canonical ledger corrections to those
closed days are reflected, while the boundary day's then-open cost bucket is
omitted. The former live brief, carry-over-todo snapshot, and pending sponsor-
application count are returned as unavailable rather than being synthesized
from today's state; sponsor scan/case status is explicitly labeled as current
canonical state for that run's scan. Queue coverage is
written automatically only after an
`--all` traversal reaches canonical exhaustion; an interrupted scan remains in
its local checkpoint and cannot be misreported as complete.

**Every agent-authored final full-management report includes a `Featured rotation`
section.** Base it on the dated starting board, a fresh `featured list`, and
canonical history. Report the target of replacing **all previous-day pins**,
the exact proposed/completed outgoing-to-incoming mapping, and each retained
pin with its reason. Rank new candidates by likely genuine member engagement,
not thoughtfulness. Include the all-Featured-comment coverage and encouragement
counts described above. When genuine capacity exists, make and report strong
additions under the standing approval above; do not defer those as proposals.
Every proposed or completed new addition must include its posting date and
canonical never-Featured evidence; omit it when either gate is unverified.
If a complete refresh is constrained, state the shortfall, including inadequate
eligible inventory, inaccessible context, specific keep instructions, or pending
approval. Do not report `None` merely because yesterday's posts are still good
or have received support. Removing or reordering pins remains a proposal until
Mikey gives his go-ahead; a pending proposal is not a completed refresh. After
approval, execute the entire approved plan and verify it without asking for
each swap again. Never omit this section from a full-run report.

### Full management report in Chrome

Mikey's standing delivery preference (2026-09-15): after a full daily run, open
the **complete management report as a browsable localhost page in Chrome**.
Do this as part of finishing the authorized run; a Markdown path in chat alone
is insufficient, and no additional confirmation is needed to open the report.

1. Save the complete report as
   `/private/tmp/twinkle-daily-YYYY-MM-DD/daily-management-report.md`, using the
   run's Bangkok date. Include every required reporting section, coverage gap,
   pending decision and carryover. Reflect later owner decisions accurately.
2. Render the entire Markdown into a readable HTML page with section links,
   usable tables and access to the original Markdown. Preserve complete
   appendices and flagged rows; navigation or collapsible details must not
   discard them. Use local assets so the report does not depend on a CDN.
3. Serve the view on `127.0.0.1` using an available port. Expose only the HTML,
   its required assets and the report Markdown through a dedicated directory
   or explicit routes; do not serve the surrounding private evidence folder.
   Keep the server available after the response so Mikey can browse it.
4. Open the localhost URL in Mikey's Chrome. Verify that the page renders,
   section navigation works, and the full report is accessible. Keep the
   report tab open. Include both the localhost URL and Markdown file link in
   the final response.

For a follow-up that only opens or updates an existing report, reuse that
report and its existing tab/server where available. This does not authorize
starting another management run or repeating unrelated daily duties.

Creating an escalation belongs to the active run; acknowledging, annotating,
resolving, or reopening it does not. Use the run-independent `escalation`
commands after Mikey responds instead of starting a follow-up delegated run.

`lastRun` makes a lost-response retry of `complete` or `fail` possible after
the active pointer has been cleared. Other run-scoped commands accept only the
current unexpired `active` run. Completion first finalizes any mutation whose
content change committed but whose audit bookkeeping was still pending,
counting it toward the run's rotation signal. It then rejects only while a
mutation from the last ten minutes is genuinely in flight (the 409 lists the
pending mutations and a `retryAfterSeconds`); older in-flight rows are treated
as orphans of a dead process and no longer block completion. `fail` remains
available to abandon a run without advancing rotation, including a run whose
six-hour authorization has expired; an expired run can never be completed.
A TTL-expired run is reported with status `expired` even before the next
start reaps it, so `daily-run status` never shows an unusable run as
`active`.

Starting with a run key that belongs to a finished or expired run fails with
`CLI_ADMIN_RUN_KEY_ALREADY_USED`; supply a fresh `--run-key` (for example
`daily:2026-08-07:2`) to start again the same day. Reusing the key of the
live active run returns that run only when the requested `--comment-mode`
and `--scope`, plus any explicit `--identity`, match it; otherwise the start fails with
`CLI_ADMIN_RUN_SETTINGS_MISMATCH` instead of silently returning a run with
different scopes. The same check applies when a start without the active
run's key would fall back to that active run.

The default full-run key is `daily:YYYY-MM-DD` in Asia/Bangkok. Scoped Featured
runs receive a fresh `scoped:featured:...` key so completing one slice cannot
consume the day's full-run key or prevent a later explicitly requested slice.
They also use a dedicated API start endpoint, so an older API cannot ignore the
scope and silently create a full run; it fails before creating any run instead.
Supply `--run-key` for a separate explicit run. `--idempotency-key` may be supplied to any
mutation when a caller needs the same retry identity across processes. The CLI
generates a fresh key for every mutation invocation; if a mutation fails, its
JSON error includes `details.retryIdempotencyKey` for a safe exact retry.

### Build XP/Coin reward approvals (any time; also a full-daily-review duty)

The creator's Lumine designs the rewards and writes them into the app. The
app declares its economy in `rewards.json` at the project root (rule ids,
titles, XP, Coins, tries, retry share, budgets); quiz rules get their questions
and answer keys from a private question sheet the creator's Lumine uploads with
`lumine rewards sheet <file.json>` (never a project file: published source is
readable by every player). **Send for review** freezes the code and proposes
`rewards.json` merged with the sheet. Approval is Mikey's decision: read the
frozen code, check that the amounts are right and that the app cannot be
farmed, change anything that is wrong, approve. **Approval publishes** (since
2026-09-15): the exact frozen snapshot goes live in the same transaction, with
no Publish click by the creator; the app's previous release stays up until
that commit lands. Instead of approving, the reviewer may **propose changes**:
edit a copy of the frozen snapshot and offer it as the condition of approval.
The creator sees every changed line and either accepts (the proposed version
is approved and published) or declines (the request is rejected). Nothing in
that flow joins the creator's team. These commands need no daily run and can
be used whenever a request arrives (the reviewer also receives a DM card per
request).

```bash
lumine admin reward-review list --json                       # pending (default)
lumine admin reward-review list --status approved --json
lumine admin reward-review list --status all --cursor 40 --json
lumine admin reward-review show 2 --json                     # summary + file sizes + proposed rules
lumine admin reward-review show 2 --dir /private/tmp/reward-review-2 --json
lumine admin reward-review approve 2 --json                  # approve exactly what the app proposed
lumine admin reward-review approve 2 --config rules.json \
  --reason "Halved the stage amounts" --json                 # approve with changes (publishes)
lumine admin reward-review show 2 --dir /private/tmp/reward-review-2 --json   # then edit that directory…
lumine admin reward-review propose 2 --dir /private/tmp/reward-review-2 \
  --config rules.json --reason "Moved the claim after the stage clears" --json  # …and offer it
lumine admin reward-review reject 2 --reason "Rewards fire on game over; nothing is earned" --json
lumine admin reward-review revoke 2 --reason "Farmable; pausing until redesigned" --json
```

`show --dir` writes the exact reviewed snapshot (private files, including the
app's own `rewards.json`) so the agent can read it like a pulled workspace. The
result's `config` is the proposal: the declared economy with the sheet's
questions merged in. It also carries `detectedRuleIds` (a heuristic scan of
`start({ ruleId })` calls), `isLatest`/`isLive`, the published version,
lifetime totals and what this review has already paid out.

Review questions to settle with Mikey before approving:

- Is the reward tied to real play or learning, or does it fire on trivial or
  losing moments (a timer, a game over, the first minute of play)?
- Completion rules (`verifier: "completion"`) prove nothing but elapsed time:
  the app calls `start` when an activity begins and `claim` when it ends, and
  the server only checks `minSeconds`, once per learner per day, and the
  budgets. Read the code for where those calls sit, and keep the amounts and
  `userDailyXP` small enough that a player scripting the calls would not
  matter. Arcade Typing's stage clears are the reference: up to 10,000 XP a
  day across twelve stages.
- Quiz rules: fixed questions reachable in seconds are farmable; dated sets
  or `progression: "until-earned"` sets (a set stays up until somebody earns
  it, then the next one comes up the following site day (UTC midnight, 9:00 AM Korea)) keep them honest.
- Do the rule IDs in `rewards.json` match what the code starts? Unknown IDs
  simply never pay.
- Are the amounts and the per-learner daily budgets right for what the app
  actually asks of people? There is no app-wide budget, per day or lifetime:
  a good app keeps paying everyone who plays it. Older configs may still
  carry `dailyXP`/`dailyCoins`/`lifetimeXP`/`lifetimeCoins`; the API accepts
  and ignores them, so never add or tune them.

`rules.json` (what `--config` takes, and what the app's `rewards.json` plus
sheet compose into):

```json
{
  "userDailyXP": 10000, "userDailyCoins": 0,
  "rules": [
    { "id": "stage-1", "title": "Clear Stage 1", "xp": 300, "coins": 0,
      "verifier": "completion", "minSeconds": 20 },
    { "id": "e1-daily", "title": "Elementary 1 · Daily bounty", "xp": 50000, "coins": 1000,
      "verifier": "numeric-quiz", "maxAttempts": null, "retry": { "xpPercent": 50, "coinsPercent": 0 },
      "progression": "until-earned",
      "sets": [{ "key": "e1-01", "questions": [{ "prompt": "...", "answer": 4, "hint": "...", "guide": { "explanation": "..." } }] }] }
  ]
}
```

Rule fields (all server-enforced, none inferred from app code):

- `verifier`: `numeric-quiz` (server-checked numeric answers) or `completion`
  (a finished activity; `minSeconds` is the only proof).
- `sets`: question sets. Dated: `[{ "from": "2026-09-14", "to": "2026-09-14", "questions": [...] }]`
  on site days (UTC) (inclusive, non-overlapping, up to 62). Until-earned
  (`"progression": "until-earned"`): ordered sets with optional `key`; the
  first set nobody earned before today is up, an unsolved set is never
  replaced, and a set earned today stays up for the rest of that day.
- `retry`: `{ "xpPercent": 50, "coinsPercent": 0 }` — what a correct answer pays
  after a wrong one, as a share of the rule's amounts (rounded down). Absent:
  every correct answer pays the full amounts.
- `maxAttempts`: wrong answers allowed per challenge; `null` = unlimited until
  the daily reset (UTC midnight, 9:00 AM Korea) (wrong answers are paced two seconds apart). Absent: 3.
- Per question `hint` (public from the start, ≤ 300 chars) and `guide` (a JSON
  object ≤ 6,000 chars the app renders as the after-answer lesson). The server
  releases a guide only after the learner's first answer.
- Top-level `userDailyClaims`: receipts one learner may earn per site day
  across all rules. `1` is "one bounty a day".

Math Lab's economy (Mikey, 2026-09-12): twelve level rules, one per grade per
day, elementary 50,000 XP + 1,000 Coins, middle 70,000 + 5,000, high
100,000 + 10,000; `retry` 50 % XP / 0 % Coins; `maxAttempts` null;
`userDailyClaims` 1; until-earned sets authored from the Korean curriculum.
Arcade Typing (Mikey, 2026-09-12): XP for clearing campaign stages, up to
10,000 XP per learner per day, no Coins. Platform ceilings: 100,000 XP /
10,000 Coins per rule and per learner per day. No app-wide ceiling exists,
per day or lifetime.

Approval freezes these rules with the reviewed snapshot and publishes that
snapshot immediately (the result carries `published.version`); an approval
without at least one rule is refused, and an approval whose creator has saved
past the frozen version is refused as `build_reward_review_stale` (the request
also closes itself on that save). Rejection and revocation require a
`--reason` the creator reads verbatim in their workspace. A later code save
needs a new request. Never approve without reading the code; never approve a
request whose `isLatest` is false.

`propose <id> --dir <edited> --config rules.json [--reason]` sends the edited
directory (text files only; dotfiles and tool folders skipped) as the
reviewer's proposal: the review moves to `changes_offered`, the creator's card
and workspace show the note and every changed line, and the creator's
**Accept & go live** publishes exactly those files with these rules (their
workspace is replaced by the accepted version). **No thanks** rejects the
request (`declinedByCreator: true`). A proposal must still use the rewards
SDK, must differ from the submitted snapshot, and is refused once the creator
saves past the submitted version. Offering again replaces the earlier offer;
approving or rejecting while an offer is out decides the request as
submitted. The website equivalent is the Management panel's "Edit a copy to
propose changes" (a private workspace copy owned by the reviewer) followed by
"Offer my copy with these rules".

Each offer has a server-owned revision. Changing the files, rules or note
creates a new revision; a creator looking at an older comparison or decline
confirmation cannot answer the replacement offer. The creator sees its reward
amounts as well as its file changes. Proposed rules stay separate from the
submitted rules until acceptance, so `approve` without `--config` still uses
the original submitted configuration. The CLI audits the offer atomically
and includes file contents in its retry fingerprint.

Approval also attempts a free preview thumbnail when the app has none. That
capture uses the published version and cannot overwrite a later release or a
thumbnail the creator chooses while it runs. It is best effort: a capture
failure leaves publication successful and does not spend AI-image credits.

The review copy carries independent copies of referenced uploaded media.
Before freezing an offer, the server reuses the creator's original media and
copies new reviewer media into the creator's library within their storage
quota. Its final URLs are included in the comparison, so acceptance publishes
those exact files and does not depend on keeping the review copy. Re-offers
reuse the media; a failed transaction cleans up its copied objects. Declining
leaves the offered media as unused uploads in the creator's library.

### Reward activity report (standing duty, every full daily review; added 2026-09-12)

Completion rewards (Arcade Typing's stage clears) prove nothing but elapsed
time, so the run reads the shape of the week's claims instead of trusting them:

```bash
lumine admin reward-activity --json                # 7 UTC days INCLUDING today’s partial day
lumine admin reward-activity --date 2026-09-13 --json # exactly this UTC day
lumine admin reward-activity --date 2026-09-13 --days 7 --json # 7 days ending on this date
lumine admin reward-activity --days 14 --build 333 --json
```

Read-only, no run lease. For yesterday, always pass its exact UTC `--date`; `--days 1` alone means the current partial day. `from`, `to`, `timezone`, and `includesCurrentDay` make the window explicit. Rules come from each claim’s frozen review, not the current app policy; `missingReviewIds` means rule-based flags lack context. The result lists every app that paid (claims,
earners, XP, Coins) and the flagged player-days, worst first:

- `fast`: a completion claim within 5 s of the rule's `minSeconds` — a human
  who types the stage lands well above the minimum; a script lands on it;
- `burst`: 3+ claims within 10 minutes;
- `sweep`: 75 %+ of an app's completion rules earned within 30 minutes (3+ claims);
- `cap` / `daily-max`: at the per-learner day cap; at it on 3+ days of the window;
- `guessing`: a quiz challenge with 15+ wrong answers (unlimited-tries rules).

Put the app totals and every flagged row in the daily report verbatim. A
single `cap` day is a good player having a good day; `fast` + `sweep` +
`burst` together on one player is the script signature. None of it is proof:
open the player with `admin identity inspect` before proposing anything, and
propose to Mikey (revoke the app's rule, or a bucket ban) rather than acting.
Never revoke an approval from a daily run.

The same report carries `reviewLifecycle` (added 2026-09-15): the reviews'
audit trail counted per UTC day (`byDay`) and for the window (`totals`).
`actions` counts every event: `request`, `approve`/`reject`/`revoke`,
`publish` (the approved version went live), `propose`, `proposal_viewed` (the
creator opened an offer's comparison; once per revision), `accept`/`decline`,
`thumbnail` and `<decision>_refused`. `thumbnails` is the automatic thumbnail
after a review publish: `captured`, `existing` (the app already had one),
`superseded` (the creator's own thumbnail or a newer release won),
`not_needed`, `owner_missing` or `failed`. `refusals` counts conflicts that
rolled back, keyed `decision:code` (for example
`approve:build_reward_review_stale`, `accept:build_reward_proposal_stale`).
Report the totals. Each `failed` thumbnail and each `publish` without a
`thumbnail` event on a completed day is a carry-over todo with the review
ids (`reward-review show <id>` lists its events). An `accept` without a
`proposal_viewed` for that revision means the creator accepted without opening
the comparison; mention it, it is not a fault. Refusals are the concurrency
guards working; report them, and escalate only a repeated pattern on one app.

## Private carry-over todos

```bash
lumine admin todo list --json
lumine admin todo list --status all --json
lumine admin todo add --kind experiment --status in_progress \
  --title "Validate Zero/Ciel cost optimization" \
  --note "Replay baseline and optimized conversations. Complete only after response-quality parity; lower cost with a weaker reply fails." --json
lumine admin todo update 12 --status blocked \
  --note "Implementation is ready; waiting for a complete cost bucket and old-vs-new quality replay." --json
lumine admin todo update 12 --status completed \
  --note "Blind parity comparison passed every required dimension; measured cost and latency evidence attached in this note." --json
```

Todos are private operator work, not Zero/Ciel public actions. They persist
independently of daily runs and are therefore available before a run starts and
after it closes. Creating or updating one uses the run-independent transactional
audit path: canonical todo state and its private `todo.create` / `todo.update`
audit response commit together, and no public bot, public mutation count, or
rotation signal is involved.

Every successful full `daily-run start` response automatically includes all
unfinished items under `data.carryoverTodos`. The same run ID increments an
item's surfacing telemetry at most once, even when start is retried. This is the
canonical handoff: read it before discretionary new work, resume what can safely
progress after the run's mandatory newspaper/brief/conduct/log-review/cost/
energy-budget duties, and record a
concrete progress note before the run closes. The daily-run report includes the
still-unfinished set again. Completing a daily run never silently completes its
todos. A CLI carrying this contract rejects a start response that does not echo
the canonical handoff, so a newer CLI against an API deployed before the todo
migration cannot quietly treat unsupported telemetry as an empty list.
A Featured-only start instead returns an explicitly suppressed, empty handoff
and performs no todo reads, writes, capacity checks, or surfacing increments.

Evidence-dependent investigations must have a durable collection plan. Do not
carry “still no recycle-under-load evidence” forward without checking the
collector and recording new observations. Full starts and reports return
`runtimeEvidence` with a scoped read command; this handoff is suppressed for
Featured/newspaper slices. During a full daily run with pending runtime
investigations, use:

```sh
lumine admin runtime evidence primary --days 7 --json
# Only when the investigation also covers the configured target host:
lumine admin runtime evidence target --days 7 --json
```

This is a read-only, run-independent command. It does not acquire/clear a log
lease, trigger a recycle, or authorize an unrelated management review. Host
routing is explicit and never silently substitutes primary for target. An
older API without this endpoint is unsupported, not “no incidents.”

The cluster primary samples existing worker health snapshots once per minute
and records memory-guard/operator recycle lifecycle events. Records survive
restarts in the health directory's `evidence/` subdirectory (seven UTC calendar
days, at most 2 MiB/day; 256 KiB/day reserved for events). Work is recorded as
counts by kind, not user IDs, request labels, messages or tokens. Evidence is
passive: never force production load, a worker recycle or a restart merely to
complete an experiment. New collector code requires activation of the new
**primary generation**, not just a rolling worker reload; verify a fresh live
sample before saying collection is running.

For each affected todo, persist the checked host/runtime identity, previous and
new evidence cutoff, sample count/gaps/freshness, qualifying event IDs and the
next safe action. `unavailable`, `stale` or `incomplete` means investigate the
collection gap. A healthy collector with no qualifying recycle means keep
collecting, not stall or claim a failure. Save the relevant summary in the todo
before the seven-day retention window passes. Missing/stale active-work or OOM
observations remain unknown; they are never zero or idle by default.

Set investigation-specific acceptance criteria before interpreting results.
For memory/recycle investigations, distinguish a long-lived steady-memory
baseline from a recycle observed under load. Require the requested/signalled/
recovered sequence, fresh pre-recycle work evidence, replacement identity,
bounded recovery time, and no OOM-counter increase across the same primary
generation. “Topology recovered” does **not** prove each interrupted user task
completed: correlate those tasks' existing canonical outcomes before closing
a user-work continuity investigation. Never close based only on deployed code,
a current healthy snapshot, or an unobserved event. Daily runs collect evidence
and update todos; fixes and releases follow the project's existing authority rules.

`kind` is `task` or `experiment`. New items may start `open`, `in_progress`, or
`blocked`; updates may also use `completed` or `cancelled`. A progress note is
required for every update. For experiments, put the acceptance criteria in the
initial details and use evidence—not implementation status—as the completion
boundary. In particular, an AI-cost experiment is not complete until old-vs-new
response-quality parity is demonstrated; a cheaper but weaker user response is
a failed experiment. Up to 50 unfinished items may be carried so the automatic
start payload remains complete and bounded.

```ts
type AdminTodo = {
  id: number;
  kind: "task" | "experiment";
  title: string;
  details: string;
  status: "open" | "in_progress" | "blocked" | "completed" | "cancelled";
  revision: number;
  createdRunId: number | null;
  lastWorkedRunId: number | null;
  lastSurfacedRunId: number | null;
  surfaceCount: number;
  lastProgressNote: string | null;
  createdAt: number;
  updatedAt: number;
  lastSurfacedAt: number | null;
  completedAt: number | null;
  cancelledAt: number | null;
};

type AdminTodoList = Success<{
  todos: AdminTodo[];
  statusFilter: "pending" | "all" | AdminTodo["status"];
  truncated: boolean;
}>;

type AdminTodoMutation = Success<{ todo: AdminTodo }>;

type CarryoverTodos = {
  included: boolean;
  items: AdminTodo[];
  count: number;
  surfacedForRunId: number | null;
  newlySurfacedCount: number;
};
```

## Canonical lists and inspection

```bash
lumine admin recommendations list --kind recommend \
  --content-types comment,dailyReflection --all --json
lumine admin recommendations list --after 2026-08-14T00:00:00Z \
  --all --checkpoint recommendations.json --json
lumine admin recommendations list --include-legacy --all --json
lumine admin recommendations list --unviewed --json
lumine admin subjects candidates --after 2026-08-01T00:00:00Z \
  --all --checkpoint subjects.json --json
lumine admin subjects candidates --since-run --all --json
lumine admin subjects candidates --include-legacy --all --json
lumine admin subjects candidates --effort unassigned --json
lumine admin subjects candidates --unviewed --json
lumine admin builds candidates --all --limit 50 --json
lumine admin builds review build:884 --output-dir ./build-review --json
lumine admin builds review build:884 --output-dir ./build-review \
  --interact ./build-review/steps.json --json
```

Schemas:

```ts
type RecommendationQueueList = Success<{
  items: Array<{
    queueId: string;
    feedId: number;
    contentType: "comment" | "aiStory" | "dailyReflection";
    contentId: number;
    url: string | null;
    subjectUrl: string | null;
    author: Author;
    createdAt: number | null;
    title?: string | null;
    question?: string | null;
    content: string | null;
    explanation?: string | null;
    attachment?: Attachment | null;
    imagePath?: string | null;
    audioPath?: string | null;
    subject?: {
      id: number;
      title: string | null;
      effortLevel: number;
      url: string;
    };
    recommendation: RecommendationState;
    reward: RewardState;
  }>;
  pagination: Pagination & {
    scannedCount: number;
    contentTypes?: Array<"comment" | "aiStory" | "dailyReflection">;
  };
  clientFilter?: {
    contentTypes: Array<"comment" | "aiStory" | "dailyReflection">;
    excludedItems: number;
  };
}>;

type SubjectCandidates = Success<{
  subjects: Subject[];
  pagination: Pagination & { scannedCount: number };
}>;

type BuildCandidates = Success<{
  builds: Array<{
    id: number;
    title: string | null;
    description: string | null;
    username: string | null;
    publishedAt: number | null;
    publishedArtifactVersionId: number | null;
    collaborationMode: "private" | "open_source";
    url: string;
    review: {
      publishedArtifactVersionId: number | null;
      codePullAvailable: boolean;
      requiredBeforeComment: true;
    };
  }>;
  pagination: {
    nextCursor: string | null;
    hasMore: boolean;
    exhausted: boolean;
  };
}>;
```

Subject cursors freeze a primary-key high-water mark plus the server snapshot
timestamp and traverse descending IDs; bounded scans keep creation timestamps
inside that confirmed interval. Bounded recommendation cursors freeze both the
feed-ID high-water mark and the server timestamp, then traverse the indexed
`(timeStamp, id)` order; this also catches a Daily Reflection whose old feed
row moved forward when it was reshared. Explicit legacy scans retain the
descending primary-key walk. A page
can be empty while `hasMore` remains true; continue until `exhausted`. `--all`
does that automatically, fsyncs each confirmed page to a private NDJSON
candidate spool, and writes only bounded cursor, boundary, count, byte-length,
and digest metadata to the private checkpoint. `--resume` continues only when
the checkpoint belongs to the same API, run, and exact request and its confirmed
spool prefix still matches those metadata; an unconfirmed fsynced tail is
discarded before the next request. The final JSON collection is streamed from
the spool and can be copied to a separate `--output` file, while `--checkpoint`
remains resumable operational state. Subject `--after` is inclusive, and every
opaque cursor is bound to its original filters.

An automatic checkpoint filename includes a fingerprint of the exact request,
so two scans with the same operation name and run ID but different subjects,
server filters, client-side view filters, or result-transform inputs cannot
overwrite one another. An exclusive adjacent lock rejects a
second process using the same checkpoint while the first scan is active and
recovers a lock only when its recorded process no longer exists. Resume still
recognizes the pre-fingerprint default filename, validates its stored request
fingerprint, and migrates it through the existing checkpoint path.
Legacy Build-candidate checkpoints intentionally fail that validation because
they did not bind the Site URL used to materialize candidate links; start those
scans fresh so one result cannot mix origins.

For `--all --json`, stdout remains exactly one JSON value. Scan-start, first-
page, every-tenth-page, and exhaustion progress is written to **stderr** with
only page/scanned/candidate counts and the private checkpoint path. A long
traversal therefore no longer looks stalled, while piping stdout to `jq` or a
file remains safe.

`Ctrl+C` and `SIGTERM` abort the in-flight page request, leave the last
server-confirmed page fsynced in the private checkpoint/spool, and release the
adjacent process lock. The cancellation error names the checkpoint. Continue
only by rerunning the exact same command with `--resume`; the next invocation
verifies the request fingerprint and confirmed spool digest before requesting
another page. An interrupted request is never counted as queue coverage.

Recommendations default to `--since-run`: the server uses the previous
completed full run's start time, even if that gap exceeds 30 days. On a first
run, the fallback begins seven days before the current run's stored start, so
it cannot drift between pages. Insight reports retain their separate 30-day
limit. That deliberate start-to-start overlap gives the queue
at-least-once coverage when content arrives after the prior snapshot but before
that run completes. `--after` supplies an explicit inclusive timestamp.
All-history traversal is deliberately available only through
`--include-legacy`. The CLI requires the API to echo the canonical `after`
boundary for bounded modes, so deploying a new CLI against an older API cannot
silently fall back to a million-row historical scan.

`recommendations list` is the Earn Recommend picker, not a "new comments"
feed: it returns only comments on subjects with an assigned effort level,
whose length exceeds that level's threshold (>100 / >250 / >450 / >700
characters for effort ≤2 / 3 / 4 / 5), with no skip row and no existing
recommendation from any effective Level 5+ user (1000+ AP or Teacher
authority), Zero, or Ciel. Community-recommended and short comments are
deliberately absent, so an empty page or an empty window is normal and is
not evidence of a broken walk; Featured-subject comments are reviewed through
the Featured comment scan instead. A successful `post recommend` does not
imply the target was queue-eligible.

After upgrading the API and CLI to the stable run-start window, start a fresh
`--since-run` scan with a new checkpoint path, without `--resume`. Older
since-run checkpoints are rejected even when already exhausted: they may have
captured the former 30-day reporting cap. They are left intact as evidence.
Explicit `--after` and `--include-legacy` checkpoints retain their contracts.

Subject candidates follow the same window contract. They default to the
previous completed full run's start (with the seven-day first-run fallback),
accept an explicit inclusive `--after`, and require `--include-legacy` for a
lifetime traversal. `--since-run`, `--after`, and `--include-legacy` are
mutually exclusive. The CLI also requires the API to echo the bounded Subject
window before accepting a page.

`builds candidates` uses the admin publication-window endpoint, ordered by
`publishedAt` and Build ID, not workspace `updatedAt`. Like Subject discovery,
it defaults to the previous completed full run's start (seven-day first-run
fallback), accepts inclusive `--after`, and requires `--include-legacy` for
all history. These flags are mutually exclusive. The first page freezes the
time boundary and an artifact-version high-water mark; the cursor and local
checkpoint retain both. An old public-browser checkpoint cannot be reused.
The CLI fails closed if the API does not confirm this publication window.

It is available through the `admin` namespace only while a delegated run is active;
page until `pagination.exhausted`. Each item includes its canonical app URL,
published artifact version, and whether its code is pullable. This list does
not decide that an app deserves a comment. The management agent must open and
genuinely try the published runtime, or pull and read an open-source project,
before making that judgment. Direct API/persona automation is never a review
substitute.

Only current public Main releases are candidates. Workspace-only saves and
unchanged reactivations do not become new releases. A Build republished during
paging can leave the snapshot; its new release is reconsidered by the next
overlapping start-to-start window. Always recheck the current artifact before
reviewing; this is not a frozen copy of an app or an immutable release archive.

`builds review` is the managed runtime path: it fetches the current published
artifact identity, launches the app in an isolated temporary Chromium profile,
captures a screenshot and bounded console evidence, then fetches the identity
again. It writes `review.json` in a unique per-review subdirectory only when
the browser completed, the screenshot exists, and the artifact did not change
mid-review. Attach the returned `receiptPath` with
`comment draft ... --review-receipt review.json`, and pass a separate private
`--review-context context.json` containing only the concrete `understanding`
learned during that review. The receipt binds the draft to the exact reviewed
artifact without copying a version number by hand; the server owns the Build,
version, method, and review-time fields around that understanding.

Without `--interact` the review captures only the start screen after
`--wait-ms`. `--interact <steps.json>` adds a bounded, ordered interaction
script that runs inside the app's runtime iframe after that start screenshot,
so the receipt can show what happens when the app is actually used. The file is
a JSON array (or `{ "steps": [...] }`) of at most **12** steps, each exactly one
of:

```json
[
  { "click": "text=Start" },
  { "wait": 1500 },
  { "screenshot": "after-start" },
  { "type": { "selector": "input[name=name]", "text": "Zero" } },
  { "press": "Enter" },
  { "press": "ArrowLeft" },
  { "screenshot": "moved" }
]
```

- `click`: a CSS selector, or `text=<visible text>` (case-insensitive; the
  smallest visible element whose text matches, then a bounded contains match).
  Dispatched as a trusted mouse click at the element's centre, so canvas games
  and buttons both receive it.
- `type`: `{ selector, text }` — clicks the element, then inserts single-line
  text (at most 200 characters) as trusted input.
- `press`: `Enter`, `Space`, `Escape`, `Tab`, `Backspace`, `ArrowUp/Down/Left/Right`,
  a letter, or a digit. The frame is focused first if nothing was clicked yet.
- `wait`: 1–5000 ms.
- `screenshot`: a unique label (1–40 letters/digits/`-`/`_`, not `runtime` or
  `review`); saved as `<label>.png` beside `runtime.png`.

The whole script is capped at **60 s**; it stops at the first failed step
(element not found/not visible, budget exhausted, frame unreachable). Console
evidence stays bounded exactly as before. `review.json` gains
`screenshots: [{ label, path, bytes }]` (script screenshots only; the start
screen stays in `screenshot`) and
`interaction: { path, stepsPlanned, stepsCompleted, status, failedStep, frame,
elapsedMs, steps }` with a per-step record (coordinates for clicks, the saved
path for screenshots, the error for a failed step). Both are `[]`/`null`
without `--interact`. The review remains one receipt bound to one artifact: the
published version is re-read after the script finishes, and a script that did
not complete makes the receipt `failed` with
`CLI_ADMIN_BUILD_REVIEW_INTERACTION_FAILED` (the completed steps and their
screenshots are still listed) so a draft can never cite interactions that did
not happen. `comment draft --review-receipt` accepts a receipt only when every
listed screenshot still exists unchanged and the script completed.

During every full daily management review, scan recent Build candidates back through the
run's review window alongside Subjects and the recommendation queue. An app
that is thin, broken, private, unchanged since a prior substantive bot
comment, or not meaningfully understood may be left alone. A new or materially
updated app where specific, truthful feedback would help is comment-worthy.
This is editorial attention, not a quota: do not manufacture a Build comment
just to prove the queue was visited.

`--content-types` is sent to APIs that support server-side filtering so excluded
types do not run their eligibility/content queries. The local CLI also filters
the returned page defensively for deployment compatibility. The server cursor
still advances across every underlying feed row, so excluding `aiStory` cannot
create gaps in later comment or Daily Reflection pages. New server cursors bind
the canonical content-type set; one legacy unbound cursor can be resumed and is
then reissued as bound. `clientFilter.excludedItems` makes any client-side
filtering explicit in JSON output.

For a run-scoped command, `--identity zero|ciel` is an assertion against the
server-selected run identity; it cannot switch actors locally. A mismatch
fails before the mutation. `--identity auto` accepts the run's canonical
selection.

### Query and index design

Subject and queue traversal are bounded primary-key walks; the subject walk
reads at most 500 `content_subjects` rows per cursor step, and the queue reads
at most 500 `noti_feeds` rows per cursor step before applying the existing Earn
Recommend eligibility predicates. The effort projection's
`UPDATE noti_feeds ... WHERE type = 'subject' AND contentId = ?` reuses the
website's canonical reward-level projection shape; deployment should verify
`noti_feeds` carries an index whose leading columns cover `(type, contentId)`
(or `(contentId, ...)`) as the canonical route already requires. The joins/`NOT EXISTS` checks are necessary
to preserve the normal recommendation and skip rules, but they run only for
IDs in that bounded window. Subject-comment traversal uses
`idx_comments_isDeleted_subject_id`; standalone-post and Build comments use the existing
`idx_content_comments_root_deleted` index (whose InnoDB entries also carry the
primary ID). Recommendation identity uses
`uniq_content_recommendations_active_identity`; `earn_comment_candidates` is
driven by its primary key. No offset scan or new broad table scan was added.

Featured comment review adds at most 100 per-Subject `MAX(id)` probes using
`idx_comments_isDeleted_subject_id`, then reuses the bounded 50-comment page
path. Receipt lookups use the audit primary key; review summaries use the
existing `idx_laae_run_id (runId, id)` range and project only coverage/action
outcomes, not downloaded comment bodies. Approved plans use the shared board
lock and the history primary key for `MAX(id)`. No new schema is needed.

Deployment can verify the required existing index definitions with:

```sql
SELECT TABLE_NAME, INDEX_NAME, SEQ_IN_INDEX, COLUMN_NAME
FROM information_schema.STATISTICS
WHERE TABLE_SCHEMA = DATABASE()
  AND (
    (TABLE_NAME = 'content_comments'
      AND INDEX_NAME IN (
        'idx_comments_isDeleted_subject_id',
        'idx_content_comments_root_deleted'
      ))
    OR (TABLE_NAME = 'content_recommendations'
      AND INDEX_NAME = 'uniq_content_recommendations_active_identity')
    OR (TABLE_NAME = 'earn_comment_candidates' AND INDEX_NAME = 'PRIMARY')
  )
ORDER BY TABLE_NAME, INDEX_NAME, SEQ_IN_INDEX;
```

If a pre-migration environment lacks them, the repository's existing
`add-build-pinned-comments.sql`, active-recommendation-identity migration, and
Earn candidate migration are the canonical creation SQL; apply those migrations
instead of creating runtime checks or duplicate indexes.

For schema review, the exact missing-index DDL represented by those migrations
is:

```sql
ALTER TABLE content_comments
  ADD INDEX idx_comments_isDeleted_subject_id (isDeleted, subjectId, id);
ALTER TABLE content_comments
  ADD INDEX idx_content_comments_root_deleted (rootType, rootId, isDeleted);
ALTER TABLE content_recommendations
  ADD UNIQUE INDEX uniq_content_recommendations_active_identity
    (rootType, rootId, rootTargetType, activeUserId);
```

Run only the repository migrations after confirming an index is absent; do not
execute these statements blindly on a deployed database.

```bash
lumine admin subject get 123 --include-comments --json
lumine admin subject comments 123 --cursor '<cursor>' --json
lumine admin comments get 456 --json
lumine admin post get https://www.twin-kle.com/ai-stories/88 --json
lumine admin post comments dailyReflection:99 --cursor '<cursor>' --json
lumine admin post comments build:884 --all --output build-comments.json --json
```

Build comment inspection accepts `build:<id>`, a public `/app/<id>` URL, or
`--type build`. It reads all root-thread comments and replies through a
snapshot-bound cursor, with full text and `parentCommentId`/`replyToCommentId`
relationships.
Only public, published, canonical owner Builds qualify; contribution branches
and private/unpublished Builds are rejected. This is thread inspection, not
evidence of having reviewed the published app. Build-comment publication still
requires its separate version-bound review evidence.

Schemas:

```ts
type SubjectGet = Success<{
  subject: Subject & {
    secret: {
      hasSecretAnswer: boolean;
      hasSecretAttachment: boolean;
      shown: boolean;
      answer: string | null;
      attachment: unknown | null;
    };
    // True only when comments were actually returned; a secret-gated subject
    // reports false here (with secret.shown false) even when they were
    // requested.
    commentsIncluded: boolean;
    // The inline list is capped at 200 comments in conversation order; when
    // true, page through `subject comments` for the rest.
    commentsTruncated: boolean;
    comments: Comment[];
  };
}>;

type SubjectComments = Success<{
  subject: { id: number; url: string; title: string | null };
  comments: Comment[];
  pagination: Pagination;
}>;

type CommentGet = Success<{
  comment: Comment;
  subject: {
    id: number;
    url: string;
    title: string | null;
    author: Author;
    secretShown: boolean;
  } | null;
}>;

type StandalonePostGet = Success<{ post: StandalonePost }>;

type StandalonePostComments = Success<{
  post: {
    contentType: "aiStory" | "dailyReflection";
    contentId: number;
    url: string;
  };
  comments: Comment[];
  pagination: Pagination;
}>;
```

Inspection never silently bypasses secret semantics. Until the selected bot is
the author or has canonically responded/revealed, secret values and comments
remain unavailable.

## Subject and Featured mutations

```bash
lumine admin subject reveal 123 --json
lumine admin subject effort set 123 --level 2 --json
lumine admin subject creator set-made-by-poster 123 --json
lumine admin subject feature 123 --json
lumine admin subject unfeature 123 --json
lumine admin featured list --json
lumine admin featured history --subject-ids 50,40 --all --json
lumine admin featured add --subject-ids 50,40 \
  --posted-after 2026-08-27T00:00:00+07:00 --json
lumine admin featured reorder --subject-ids 30,20,10 --json
lumine admin featured rotate --remove-subject-ids 30,20 \
  --add-subject-ids 50,40 --json
```

Schemas:

```ts
type SubjectReveal = SubjectGet & {
  status: "success" | "already_done";
  changed: boolean;
  data: SubjectGet["data"] & {
    reveal: {
      status: "created" | "already_revealed";
      notificationCommentId: number | null;
    };
  };
};

type SubjectEffortSet = SubjectGet & {
  status: "success" | "already_done";
  changed: boolean;
};

type SubjectCreatorSet = SubjectEffortSet;

type FeaturedList = Success<{
  subjects: Subject[];
  count: number;
  maximum: 20;
  maximumScope: "delegated_admin";
  delegatedMaximum: 20;
  websiteMaximum: 100;
}>;

type FeaturedHistory = Success<{
  coverage: {
    complete: boolean;
    startedAt: number | null;
    updatedAt: number | null;
  };
  subjects: Array<{
    id: number;
    url: string;
    title: string | null;
    createdAt: number | null;
    deleted: boolean | null;
    featured: { member: boolean; order: number | null };
    knownFeatured: boolean;
    neverFeatured: boolean | null;
    firstRecordedFeaturedAt: number | null;
    lastRecordedFeaturedAt: number | null;
  }>;
  events: Array<{
    id: number;
    mutationId: string;
    subjectId: number;
    action: "featured" | "unfeatured" | "reordered" | "snapshot";
    fromPosition: number | null;
    toPosition: number | null;
    source: "website" | "lumine-admin" | "coverage-bootstrap";
    operation: string;
    actorUserId: number | null;
    operatorUserId: number | null;
    adminAuditId: number | null;
    occurredAt: number;
  }>;
  pagination: Pagination;
}>;

type SubjectFeature = FeaturedList & {
  status: "success" | "already_done";
  changed: boolean;
};

type SubjectUnfeature = SubjectFeature;
type FeaturedAdd = FeaturedList & {
  status: "success" | "already_done";
  changed: boolean;
  data: FeaturedList["data"] & {
    addition: {
      addSubjectIds: number[];
      postedAfter: number;
      finalSubjectIds: number[];
    };
  };
};
type FeaturedReorder = SubjectFeature;
type FeaturedRotate = FeaturedList & {
  status: "success" | "already_done";
  changed: boolean;
  data: FeaturedList["data"] & {
    rotation: {
      removeSubjectIds: number[];
      addSubjectIds: number[];
      finalSubjectIds: number[];
    };
  };
};
```

`reveal` publishes the existing hidden “viewed without responding” notification
as the selected bot, with the ordinary notification and socket side effects.
Effort assignment rejects an unrevealed secret subject. Creator attribution
requires an attachment. Both effort assignment and creator attribution enforce
the website's moderator-precedence rule: when a strictly higher-level
moderator recorded the current value, the mutation fails with
`CLI_ADMIN_MODERATOR_PRECEDENCE` (the bots compare at the shared canonical
effective level). Every response is reloaded from the writer.

Featured reorder is a complete-set replacement: it rejects duplicates,
unknown/deleted IDs, missing current members, non-subject rows, and more than
100 subjects. Permanent pins and editorial ordering policy are deliberately not
hardcoded.

`featured history` is the compact canonical evidence path. It returns only
coverage metadata, per-subject lifetime summaries, and paginated mutation
events; it does not repeat full audit before/after board snapshots. An event of
any action proves the subject has appeared on Featured. `neverFeatured: true`
is returned only when no event exists and the subject was created strictly
after the finalized coverage boundary. `null` means the subject predates provable
coverage—never convert that unknown into “never Featured.” Both the website
editor and Lumine mutations write this append-only history in the same
transaction as the canonical board replacement.
Each API read accepts up to **100 subject IDs**, independently of the
20-subject delegated addition policy. `--all` automatically batches larger
lists (up to 20,000 IDs), exhausts every batch's event pages, and retains all
per-subject summaries. Use the exact command with `--resume` after interruption;
confirmed pages are not replayed. Single-batch `--cursor` remains available,
but cannot be combined with `--all`. Multi-batch results explicitly use
`pagination.snapshotScope: "per-batch"`; events are ordered by input batch,
then descending event ID within that batch—not by one global snapshot.
`data.scan.batches` records each batch's coverage, snapshot and private spool.
Deploy the matching API before using the expanded read bound.
For a retry whose board transaction committed but whose canonical detail reload
failed, the audit-linked history event is the durable receipt: the API re-reads
the current board and preserves the original changed-mutation accounting.

`featured add` is the atomic verb for genuinely new additions. The supplied
IDs are placed at the front in descending relevance order while existing
members retain their order. The server requires the whole batch to be absent,
fit within the 20-subject maximum, have no Featured event, fall inside complete
history coverage, and have a creation time strictly after `--posted-after`.
That boundary accepts Unix seconds, an ISO-8601 date (interpreted as UTC), or
an ISO-8601 timestamp with an explicit `Z`/numeric offset; timezone-free
timestamps and permissively normalized dates are rejected.
Any failed gate leaves the entire board unchanged. An exact completed retry is
replayed from its audit response, while recovery after a committed board change
is proven by that request's audit-linked history receipt. A fresh request for
an already-present subject fails instead of inferring a retry from board shape.
Use the ordinary singular `subject feature` only for an explicit manual
override; it does not claim that a subject is new.

Featured rotate is the direct, atomic replacement verb for an approved
rotation. `--remove-subject-ids` names the exact current members Mikey approved
for removal; `--add-subject-ids` names the same number of replacements in
descending editorial relevance. The API locks the canonical board, requires
every removal to still be present and every addition to still be absent,
validates every subject, then commits one complete-set replacement. Additions
become the front of the board in the supplied order and all surviving members
keep their relative order. A stale or partly mismatched plan fails without
changing anything; an exact retry returns the canonical completed response.
The response includes the canonical final board and the confirmed rotation
IDs, so callers never synthesize Featured state locally.

The two lists must have the same nonzero length, so rotation never changes the
board's count. Up to 100 replacements are accepted, including an existing board
above the delegated growth limit of 20. That is not an overflow to repair.
The server also enforces live, recent, provably never-Featured additions. Pass
`--posted-after` explicitly; for compatibility the legacy rotation command
defaults to seven days before the current Bangkok day's midnight. The exact
approved-plan workflow below always requires an explicit cutoff.

This mutation deliberately does not choose candidates or override the
editorial gates above. Before invoking it, the management agent must have
freshly reviewed the board, shown Mikey the removals and replacements, received
his go-ahead, and verified that every proposed new addition is recent and has
never previously been Featured from canonical evidence.

### Exact approved refresh and resumable comment encouragement

For the daily old-to-new proposal, prepare a canonical plan without changing pins:

```bash
lumine admin featured plan --remove-subject-ids 30,20 \
  --add-subject-ids 50,40 --subject-ids 40,50,10 \
  --posted-after 2026-09-01T00:00:00+07:00 --output featured-plan.json --json
```

Removal and addition lists are equal-size pairs. Optional `--subject-ids` is
the **entire final ordered board**, including retained pins; without it,
additions lead in supplied order and retained pins keep their relative order.
The receipt contains titles, IDs, URLs, replacement pairs, final order, posting
cutoff, run/actor identity, expected board and history revision, and `planHash`.
Add qualitative reasons, show Mikey the exact mapping/order, and wait for his
go-ahead. A generated hash is not itself human approval. After approval:

```bash
lumine admin featured apply --file featured-plan.json \
  --approve <exact-plan-hash> --json
```

The API loads the immutable plan from its same-run/operator/actor audit receipt,
checks the supplied hash, and applies all swaps and ordering in one transaction.
It rejects intervening reorders, swaps, even a changed-then-restored board, and
newly ineligible candidates. No substitute candidates or follow-up reorder are
inferred. The CLI then reads the live board and verifies exact membership/order.
Retry the identical apply command after uncertain transport failure; it reuses
the same idempotency key. If the final read fails or finds an intervening change,
the error retains the canonical apply receipt and never restores old state.
Expired/different runs require a fresh plan and renewed approval.

Use the following workflow for full daily comment review, or only when Mikey's
Featured slice includes comment encouragement:

```bash
lumine admin featured comments scan --checkpoint featured-read.json --json
# On interruption: repeat with --resume, in the same active run.
# Read ALL pageFiles, including full root context and nested replies.
lumine admin featured comments acknowledge \
  --checkpoint featured-read.json --reviewed \
  --decisions-template featured-decisions.json --json
```

The scan snapshots every current Featured Subject (up to 100) and its maximum
comment ID, includes already-viewed and older comments, and downloads full-text
pages of 50 through server-owned receipt chains. It rejects filtered scans;
there is no popularity, length, language, or keyword selection heuristic.
Each private mode-0600 page file has a checkpointed hash; resume verifies the
complete chain and uses stable request keys after lost responses. Concurrent
use of a checkpoint is locked. Deleted/unavailable Subjects and locked secrets
remain explicitly incomplete; an empty thread still requires a terminal page.
The scanner never automatically reveals secrets. If reveal is authorized, use
the ordinary explicit reveal command and resume. New comments above the saved
boundary belong to a fresh scan, not a silently expanded claim of coverage.

`acknowledge --reviewed` is the agent's explicit confirmation that it read the
downloaded terminal chains. The API validates those chains, records covered and
missing Subject IDs and comment counts, and distinguishes complete from partial
coverage. A download alone never counts as a read or grants recommendations.
After a resumed scan fills a gap, read those pages and acknowledge again.

Acknowledge returns `data.reviewedCoverage` (the coverage receipt; its `id` is
the `coverageId` the recommend step needs — it is a different audit row from
the review ID) and a ready-to-fill `data.decisionsTemplate`
`{ reviewId, coverageId, selections: [] }`. With `--decisions-template <file>`
the CLI also writes that template as a private mode-0600 file; only
`acknowledge --reviewed` writes it (a scan rejects the flag). Start the
decisions file from the template rather than assembling the identifiers by
hand. `readFeaturedSelections` refuses a file whose `coverageId` equals its
`reviewId` before any request is sent, and the API answers a wrong receipt
with `CLI_ADMIN_FEATURED_COVERAGE_MISMATCH` naming the expected receipt, e.g.
`coverageId 5719 is not the acknowledged coverage receipt for review 5719
(expected 5749; 5719 is the review ID itself).` with
`details: { reviewId, suppliedCoverageId, expectedCoverageId,
suppliedReceiptAction, suppliedCoverageReviewId }`, or
`CLI_ADMIN_FEATURED_COVERAGE_MISSING` when the review has no acknowledged
coverage in this run. Other receipt failures (a page ID from another run,
operator, or actor) keep the generic
`Receipt #<id> is not a completed <action> receipt of this active run,
operator and actor.` rejection.

Compose a decisions JSON file from the genuinely reviewed comments, using the
returned review ID, coverage receipt ID, and each selected comment's page ID:

```json
{
  "reviewId": 100,
  "coverageId": 1050,
  "selections": [
    { "commentId": 120, "pageId": 1001 },
    { "commentId": 119, "pageId": 1001, "anyoneCanReward": true },
    { "commentId": 118, "pageId": 1001, "anyoneCanReward": true, "rewardTwinkles": 3 }
  ]
}
```

The agent chooses comments by context under the generous encouragement policy;
the tool never auto-selects them. Reward permission defaults off; the only
direct reward option is an explicitly selected 3-Twinkle pairing. Existing
permissions are not revoked by a basic recommendation. Then apply and report:

```bash
lumine admin featured comments recommend --file featured-decisions.json \
  --checkpoint featured-recommend.json --json
# Retry exactly this file/checkpoint with --resume after interruption.
lumine admin featured comments report --checkpoint featured-read.json --json
```

Use separate scan and recommendation checkpoints. Selection files are capped at
20,000 unique comments/2 MiB; split larger work into explicit batches. Each
target must belong to the reviewed page and an acknowledged Subject chain in
this same run. The API re-reads it, rejects changed/hidden/deleted comments,
system notifications and Zero/Ciel's own comments, and uses the normal canonical
recommendation/reward/deduplication path. A blocked Subject does not prevent
encouragement on acknowledged Subjects. Batches stop at a failed decision,
retain confirmed receipts, and resume without replaying completed decisions.
These recommendations are independent per-comment mutations, not an atomic
batch; report partial completion honestly.

The review report and full/historical daily reports' `featuredReviews` expose
acknowledged coverage, confirmed new/basic recommendations, already-done
actions, selective reward-permission grants, and created direct rewards. The
counts describe canonical outcomes, not attempted commands or downloaded rows.
Read comments that are already recommended too; use the page evidence for
already-recommended and skipped-category totals rather than treating the
report's selected `alreadyDone` actions as all previously recommended comments.

## Recommendation, Karma approval, and Twinkle rewards

```bash
lumine admin post recommend 123 --json
lumine admin post recommend https://www.twin-kle.com/ai-stories/88 --json
lumine admin post recommend comment:456 --json
lumine admin post recommend comment:456 --anyone-can-reward \
  --reward-twinkles 3 --idempotency-key review-456-v1 --json
lumine admin post reward comment:456 --twinkles 3 --json
```

Numeric targets default to `subject`. Use `subject:<id>`, `comment:<id>`,
`aiStory:<id>`, `dailyReflection:<id>`, a canonical URL, or the corresponding
`--type`.

For daily Featured-comment encouragement, prefer the review-bound
`featured comments` workflow above so coverage and outcomes remain linked.
The generic bare comment recommendation has the same basic default, but is
not available in a Featured-only run. `--anyone-can-reward` and
`--reward-twinkles` are separate, selective judgments, not required flags for an
ordinary worthwhile comment.
Follow **Daily Featured comments and refresh** for full-thread coverage and
the intentionally generous basic-recommendation threshold.

```ts
type PriorRecommendationApproval = {
  recommendationId: number;
  recipientUserId: number;
  karmaRecipientUserId: number;
  karmaAwarded: number; // 10 for a new canonical approval; 0 on retry
  karmaPoints: number; // writer-confirmed absolute balance after recomputation
  status: "created" | "already_approved";
  reward: Record<string, unknown>; // canonical users_rewards row
};

type Recommend = Success<
  (SubjectGet["data"] | CommentGet["data"] | StandalonePostGet["data"]) & {
    priorRecommendationApprovals: PriorRecommendationApproval[];
    managementBotDeduplication?: {
      alreadyProcessed: true;
      recommendationId: number;
      publicActorUserId: number;
      reason: "already_recommended_by_approved_management_bot";
    };
  }
>;

type MaxRecommend = Success<
  Recommend["data"] & {
    pairing: {
      recommendationId: number;
      anyoneCanReward: true;
      requestedTwinkles: 3;
      rewardStatus:
        | "created"
        | "already_rewarded"
        | "maximum_reached"
        | "insufficient_coins"
        | null;
      alreadyRewarded: boolean;
      rewardCaps: RewardState["caps"] | null;
      retrySafe: true;
      existingManagementReward?: {
        rewardId: number;
        publicActorUserId: number;
        amount: number;
      };
    };
  }
>;

type RewardThree = Success<
  (SubjectGet["data"] | CommentGet["data"] | StandalonePostGet["data"]) & {
    rewardOperation: {
      status: "created" | "already_rewarded" | "maximum_reached" | null;
      amount: number;
      reward: Record<string, unknown> | null;
      caps: RewardState["caps"] | null;
    };
  }
>;
```

Zero/Ciel are normalized to effective Level 5 only inside the shared canonical
recommendation-approval decision. This narrow rule does not grant delegation,
management authority, or any other permission. A newly activated qualifying
recommendation
approves eligible earlier lower-level recommenders through the existing
`users_rewards` recommendation mechanism and multiplier. It excludes the
content author's self-recommendation, the approving bot, both management bots,
Level 5+ users, deleted rows, and existing ineligible rows. The approval then
runs the same absolute canonical Karma recomputation as `/user/karma`, using
writer-locked state. A new approval reports the canonical 10-point contribution;
a retry reports zero newly awarded and the same confirmed absolute balance.

Recommendation history is checked across both management bots, so rotation
does not recommend the same target again — unless the request asks for
`--anyone-can-reward` and the other bot's recommendation does not carry that
permission, in which case the actor proceeds with its own recommendation so
the requested permission and paired reward are honored rather than silently
dropped. When the other bot's recommendation does satisfy the request, the
`managementBotDeduplication` payload is returned and any requested 3-Twinkle
reward is still processed. A bare recommend never downgrades an existing
recommendation's anyone-can-reward permission; only an explicit
`--anyone-can-reward` changes it, and only in the granting direction.
Changing only `anyoneCanReward` does not rerun prior-recommender approval.
Approval reward rows are locked and writer-read, so concurrent or restored
attempts cannot insert the same approval twice.

Both the standalone and combined reward paths also inspect existing 3-Twinkle
management rewards across Zero and Ciel. A canonical three from either bot is
reported as already rewarded instead of adding another management reward.

The separate 3-Twinkle reward targets the worthwhile canonical post or comment.
It uses the selected bot and Twinkle's ordinary canonical Level and recipient
rules; Mikey is never charged while Zero/Ciel is displayed. Zero and Ciel are
exempt from recommendation and reward coin charges in the shared canonical
mutation helpers, so neither `insufficient_coins` nor a bot balance decrease
can occur for these actors; every human actor still pays under the existing
Level-based rules. The
reward transaction serializes the rewarder and cap-bearing content row, then
adds only the amount needed for that actor to total exactly three. Existing
three is `already_done`; a cap is `maximum_reached`.
If recommendation succeeds but reward fails, the command exits nonzero with
`partial_failure` and `retrySafe: true`.

## Skip decisions

```bash
lumine admin post skip dailyReflection:99 --json
lumine admin post skip comment:456 --reason "duplicate coin-begging spam" --json
lumine admin post skip-batch --target-file skip-targets.json \
  --checkpoint skip-progress.json --json
lumine admin post skip-batch --target-file skip-targets.json \
  --checkpoint skip-progress.json --resume --json
```

A skip records that the management rotation has judged a recommend-queue item
and decided not to act, so neither bot's queue resurfaces it. It writes the
same canonical `users_earn_skip_status` row the website's Earn page writes
(`earnType 'karma'`, `action 'recommendation'`) under the acting bot, and the
queue eligibility predicates honor either bot's row. Human moderators' own
Earn queues are deliberately unaffected: the bots are not database supermods,
so a bot skip never hides content from a human, who may judge differently.

Targets are `comment`, `aiStory`, and `dailyReflection` only; subjects leave
their queue through effort assignment. Skipping an already-skipped item (by
either bot) returns `already_done` with `changed: false`. The optional
`--reason` (at most 500 characters) is stored in the private audit row's
metadata — it is the agent's memory of the judgment, not public content.
The skip requires the `recommendation:write` scope and is audited like every
other mutation.

`skip-batch` accepts either a JSON array (strings or `{ "target", "reason" }`
objects), `{ "targets": [...] }`, or one target per text line. It deduplicates
targets, submits them sequentially through the same canonical audited endpoint,
and checkpoints only after each response is confirmed. `--resume` verifies the
exact target-set fingerprint and run ID before continuing; it never guesses
which writes succeeded.

```ts
type PostSkip = Success<{
  skip: {
    contentType: "comment" | "aiStory" | "dailyReflection";
    contentId: number;
    url: string;
    skippedByUserId: number;
    skippedAt: number | null;
  };
}>;
```

## Twinkle Newspaper

**Audience (Mikey, 2026-09-15): children and young Twinkle users aged 10–15.**
Choose and write stories for what these readers would voluntarily spend time
reading. Before selecting a story, identify its appeal to them: curiosity,
humor, a relatable experience, a creative idea, something useful, or a chance
to join in. Games, art, puzzles, friendships, shared reflections, and community
discussions can all supply good stories; read the actual source to find the
substance.

Use clear, lively language that respects readers' intelligence. Give enough
context for someone who missed the original post, and make the interesting
part clear in the headline and opening. Avoid talking down to readers, forced
slang, preachy lessons, administrative summaries, and blurbs that merely say
someone uploaded or posted something. Keep the appeal grounded in the source;
never invent excitement, reactions, popularity, or drama.

```bash
lumine admin news --json
lumine admin news claim --output claim.json --scaffold editorial.json --json
lumine admin news validate --claim claim.json --file editorial.json --json
lumine admin news submit --claim claim.json --file editorial.json \
  --model "Ciel" --json
lumine admin news print --json
```

The Twinkle Newspaper (Build app 1929) is normally printed by a community
member spending their own AI Energy. Making sure today's paper exists is part
of every delegated website-management run: check `lumine admin news` early in
the run, and if `printedToday` is false with no edition `pending` or
`generating`, print it.

**Preferred: write the editorial yourself.** `news claim` reserves today's
edition under the server's generation lease and returns the exact canonical
event digest the server would otherwise send to its own model, so no provider
API credits are spent. With `--output` and `--scaffold`, the CLI writes that
lease/digest to a private claim file and creates an editable editorial shell.
`news validate` runs locally, before authentication or a network request, and
checks the complete citation graph plus byte-exact quote boundaries. Submit the
validated pair with `news submit --claim`; the CLI reads the edition and lease
from the claim file and validates again immediately before the request. The
explicit `--edition-id` / `--lease-token` form remains available for backwards
compatibility.

Write a `GeneratedEditorial` JSON and send it back within the ten-minute lease.
The server still treats the editorial as untrusted regardless of author: every
story must cite an exact `eventKey`
from the digest, front-page `sourceQuote`s must be verbatim contiguous
passages of the cited event's summary (invalid quotes are replaced with
canonical text), section and page layout are server-enforced, announcements
are appended verbatim outside your output, and source visibility is
re-checked transactionally at commit.

```ts
type GeneratedEditorial = {
  excludedSubjectEventKeys?: string[]; // experiment-video Subjects omitted entirely
  mastheadHeadline: string;
  mastheadDeck: string;
  lead: {
    eventKey: string;
    headline: string;
    summary: string;
    sourceQuote: string;
    coveredEventKeys?: string[]; // arc members this story narrates
  } | null;
  stories: Array<{
    eventKey: string;
    headline: string;
    summary: string;
    sourceQuote: string;
    coveredEventKeys?: string[];
  }>;
  editorsNote: string;
};
```

**Arcs and roundups (layout coverage rules).** The server layout guarantees
nothing disappears silently: digest events the editorial does not account for
are added back. Two mechanisms make real curation possible within that
guarantee:

**Exception — experiment videos (Mikey, 2026-09-15):** exclude these entirely,
including school science-contest entries. The editor/model identifies them from
the supplied context and lists their exact keys in `excludedSubjectEventKeys`.
Do not cite, summarize, or group them into newspaper coverage. The API must have
this exclusion support deployed before submitting such an editorial; older APIs
ignore the field and restore the posts. The server accepts only canonical Subject
keys and never lets this field remove official announcements. After excluding
these posts, look for worthwhile replacement stories among the edition's eligible
Subjects and shared Daily Reflections. A bounded digest dominated by experiment
videos does not establish that the day has no other stories. Check the available
canonical sources within the coverage window, and keep replacement stories within
the claim and citation contract. Do not stop at deletion when suitable material
is available, or invent filler to reach an article count.

- **`coveredEventKeys`** — an arc story may list the other events it narrates
  (an app's release + its update stream + its open-sourcing; one member's
  related posts). Covered events are omitted from the layout — the arc IS
  their coverage. Rules: a covered key must exist in the digest; a story
  cannot cover itself, the lead event, or any event that has its own story
  (citation wins); update/score/market notices are freely coverable; a
  Subject or shared Daily Reflection is coverable only by another primary
  story **by the same author** — one member's story can never absorb another
  member's post (that would be curation-by-omission through the back door).
- **Automatic roundups** — uncited app-UPDATE events (never new releases)
  and uncited score events fold into one compact "Workshop updates" /
  "The rest of the scoreboard" story per page (one line each) instead of a
  wall of template stubs. Uncited new releases, open-source listings, and
  market sales still appear as individual stubs. So: write real stories for
  what matters, use `coveredEventKeys` for arcs, and let the roundup absorb
  the rest — but a post you'd rather not amplify still cannot be omitted;
  flag it to Mikey instead.

Editorial rules (the same ones the server's own model works under): use only
the supplied events — never world news, invented names, invented statistics,
or unsupported claims. Subjects and shared Daily Reflections are the primary
authored material; only a `section: "front"` event may be the lead. Preserve
substance, names, and numbers. Do not mention official announcements (the
server adds them), and give non-front events an empty `sourceQuote`.

**Editorial craft.** The rules above make an edition valid; they do not make
it good. The server already used the digest's `priority` numbers to select
the bounded set of eligible events; those numbers do not do the editor's job
within the returned digest. On a typical day every front subject arrives with
the same priority, so treat a tied score (or recency) as no signal at all and
make the call by reading:

- **Choose the lead for readers aged 10–15.** Lead with the eligible front
  event whose substance is most likely to catch their interest and reward
  reading. A thoughtful conversation, a funny or relatable reflection, a
  striking creation, or an inviting community challenge can all qualify.
  Read the source and any supplied replies to understand the appeal; priority,
  recency, and the mere presence of an argument do not decide the lead.
- **Thread a theme through the paper.** Pick the strongest idea of the day
  and let the masthead, the lead, and the editor's note all carry it, with
  the editor's note leaving readers with an observation, question, or invitation
  grounded in one of the day's posts. Let a theme emerge from the material;
  keep each story's meaning intact and avoid forcing a moral lesson.
- **Cross-reference events into arcs.** The same thing often appears in the
  digest several times (an app's release, its open-sourcing, and its maker's
  Daily Reflection about it). Write those as one story arc — the origin
  story up front, the notices pointing back to it — instead of three
  disconnected blurbs. Each story still cites only its own `eventKey`.
- **Headlines tell the story; they do not restate the title.** "Logged Out
  and Briefly Panicked, X Confirms: Twinkle Is 5% of Life" beats repeating
  the post's title. Keep every fact traceable to digest text — the craft is
  in selection and framing, never in invention.
- **Work mechanics under the ten-minute lease.** Dump every front event's
  full summary to a file immediately after claiming, before writing a word.
  Copy `sourceQuote`s byte-exact from the summary — curly apostrophes,
  markdown asterisks, ellipsis dots and all (an inexact quote is silently
  replaced with canonical text). An event with an empty summary gets an
  empty `sourceQuote`; restate its title's facts in your story summary
  instead.

Claim edge cases: a quiet day (no editorial events) is committed as the
canonical quiet edition at claim time — no editorial needed, the response
says so. If the claim is not submitted before the lease expires, the server's
press worker falls back to generating the edition itself. `news submit`
failing with `CLI_ADMIN_NEWS_CLAIM_LOST` means the lease was superseded —
re-check `lumine admin news` and claim again only if the paper still needs
printing.

**Repairing or revising an edition.** `news claim --date YYYY-MM-DD` leases an
existing edition row, including a failed or pending day that never reached
print, and returns a fresh digest of its coverage window (primary
Subjects/Reflections are re-projected from canonical tables, and anything
since deleted or made private drops out). Submitting writes the first revision
or appends the next one — every prior press run stays browsable in the archive,
and later revisions never re-notify subscribers (only a day's first revision
does).
`--date` with **today's** date revises today's printed paper the same way,
additionally extending the coverage window to claim time so the revision is
written from the complete canonical day so far; this replaces the old
owner-website-refresh dance and, unlike a refresh, spends no AI Energy
(composed editorials never invoke a model). An unexpired in-flight press run
still blocks the claim. Repair a historical edition only when it is genuinely
degraded (missing masthead, missing lead, empty pages), not to rewrite
history editorially; revising today to materially raise its editorial quality
is a legitimate management action.

**Fallback: queue the server's own model.** `news print` reserves the edition
and lets the server's press worker write it (spends provider credits). It is
idempotent per day: it queues a new edition when today has none, requeues a
retry when today's only attempts failed, and returns `already_done` when the
paper is printed or being typeset.

A dateless `news claim` and `news print` never reprint or refresh an
already-printed edition; only the explicit dated repair/revision path above
can append another revision. The acting bot is recorded as the requester, and
the management bots are exempt from AI Energy for newspaper generation: the
platform absorbs the cost, exactly like their coin-exempt recommends and
rewards. When a day's first edition is printed, the server notifies the app's
notification subscribers (users can mute the app or unsubscribe in the app;
the bots never need to send anything). All three mutations require the
`news:print` scope (in every full run's base scopes) and are audited as `news.print`
/ `news.claim` / `news.submit` against `news_edition` targets.

```ts
type NewsStatus = Success<{
  newspaper: {
    dayIndex: number;
    dateKey: string; // YYYY-MM-DD
    printedToday: boolean;
    generationStatus:
      | "available" // no edition requested today
      | "pending"
      | "generating"
      | "ready"
      | "failed";
    failureMessage: string | null;
    attempts: number | null;
    latestPrinted: {
      dayIndex: number;
      dateKey: string;
      generatedAt: number | null;
      sourceEventCount: number;
      revisionNumber: number;
    } | null; // most recent printed edition, possibly a previous day
    nextEditionAt: number;
    printDecision: "already_printed" | "in_progress" | "retry" | "create";
    requestedAction?: "none" | "retry" | "create"; // print responses only
  };
}>;

type NewsPrint = NewsStatus; // "success" (queued) or "already_done"

type NewsClaim = Success<{
  newspaper: NewsStatus["data"]["newspaper"] & {
    quietEditionPrinted?: boolean;
  };
  claim: {
    editionId: number;
    dayIndex: number;
    dateKey: string;
    leaseToken: string;
    leaseExpiresAt: number;
    coverage: { startedAt: number; endedAt: number };
    maxSourceQuoteLength: number;
    announcementCount: number;
    events: Array<{
      eventKey: string;
      kind: string;
      section: string; // front | community | notices | scores | marketplace
      occurredAt: number;
      priority: number;
      title: string;
      summary: string;
      payload: unknown; // may include author and canonical topComments
    }>;
  } | null; // null: already printed/typesetting, or the quiet edition auto-committed
}>;

type NewsSubmit = NewsStatus; // "success"; newspaper includes revisionNumber
```

## Bot conduct review (standing duty, every full daily review)

```bash
lumine admin bot-output --json
lumine admin bot-output --days 3 --json
lumine admin bot-output --cursor '<pagination.nextCursor>' --json
lumine admin bot-output context 3797910 --reason "Review reported bot conduct in its conversation context" --json
lumine admin bot-output context 3797910 --reason "Continue the same bot-conduct review" --cursor '<pagination.nextCursor>' --json
```

**Every full daily review reads what Zero and Ciel themselves said since the
last completed full review.**

**Ordinary wrong answers and hallucinations are expected model limitations,
not website incidents.** A factual error, mistaken puzzle answer, or imperfect
reasoning alone does not warrant an escalation, engineering todo, or a code
patch. Model quality improves through LLM upgrades; do not add hard-coded
answer validators, secondary graders, forced research, correctness retry loops,
or subject-specific rules to compensate. A normal conversational correction is
enough when appropriate. This does not excuse actual harmful conduct or
application failures, nor weaken security, permissions, billing, or canonical
server-state checks: investigate those distinct problems on concrete evidence.

The bots talk to children constantly — chat replies, Daily Reflection
responses, autonomous comment-assistant comments — and a harmful message must
never depend on a kid being brave enough to report it (real incident,
2026-08-11: the reflection pipeline had Ciel scold a member on day 31 of his
streak — "I'm telling you: Stop", guilt framing, ordering him to quit Daily
Reflections — and it surfaced only because the kid showed Mikey).

`bot-output` returns, windowed since the operator's last completed full run
(`--days 1..30` overrides; the bare form sends no `days` parameter at all, so
the API applies that default window — an older CLI wrongly validated the empty
default and failed with "--days must be an integer"): `chatMessages` (every stored Zero/Ciel chat and
reflection reply, with full text and recipient metadata when its best-effort
prompt audit exists) and `comments`
(every public bot comment/reply). Individual utterances are returned in full;
the API never clips their tails. The first response freezes a canonical
high-water mark for both sources. Truncation flags mark another page beyond
the 400-row or bounded-response-size budget; continue with the exact
`pagination.nextCursor` until
`pagination.exhausted` is true and both flags are false. A cursor retains the
original time window and cannot be combined with `--days`. Do not complete the
run while either flag remains true. The default window deliberately overlaps
from the previous completed run's start, so output created after its review
snapshot but before completion is reviewed again instead of being lost. Run it
right after the brief, and **read every row** — the tool deliberately does no
filtering, scoring, or keyword matching, because the judgment is the reviewing
agent's.

Chat output now also includes `messageKind`, `attachment`, and the stored
`generation` outcome (success, failure, cancelled, generating, or unresolved).
This is stored-message evidence, not a live request-guard check: do not call
an empty row an orphan solely from its text. Hidden attachment locations and
arbitrary settings/request keys are never returned. `source: voice` identifies
newly recorded voice transcripts, while typed input during a call says `typed`;
older replies correctly say `not-recorded`
because the historical schema did not distinguish typed text from voice.

`bot-output context <messageId>` is a private, **run-independent** investigation.
It requires a 1–500 character reason and records a minimized access receipt,
not the private message text, in the audit. It returns the specified existing
Zero/Ciel output plus preceding messages in that bot's own two-person
conversation and exact topic/subchannel, oldest first within each page.
Default 20, maximum 40 messages per page; continue manually with the cursor.
The complete scope is bounded to 100 prior rows, 24 hours, and 10,000 message
IDs before the anchor. `boundedLimitReached` means stop and report that bound,
not that all channel history was reviewed. Deleted messages and hidden
attachments remain hidden. Group-channel browsing, `--all`, and `--days` are
not supported. Never start an entire daily run just to investigate one reply.

### API worker memory report (same phase, every full daily review; added 2026-09-12)

On 2026-09-12 both primary API workers hit their 256 MiB V8 old-space cap at
the same moment after ~14.6 h and aborted; core dumps then pinned both vCPUs
and the site was down about five minutes. Since then the cluster primary
recycles one worker politely when its heap stays above 85 % of its limit
(immediately above 95 %), never both at once, and the unit sets `LimitCORE=0`.
The daily run reports on that behaviour so Mikey can decide the next step
(heap cap, retention fix, host budget).

Run on the primary over management SSH (standing permission covers it):

```bash
ssh api-primary.twinkle.network 'cd /home/ec2-user/server && bash scripts/twinkle-api-service.sh memory-daily'
```

Read every line, and put these in the daily report verbatim:

- each `worker slot=… uptime=… core_limit=… heap_pct=… heap_high_water_pct=…`
  line: worker age, current heap as a share of its cap, and the highest share
  it reached; `core_limit` must read `0` on a post-2026-09-12 generation;
- the `worker events (24h) heap_recycles=… heap_high_warnings=… pressure_recycles=…
  operator_recycles=… service_restarts=… oom_aborts=… unexpected_worker_exits=…`
  line;
- the `day-over-day … worker_heap=…` delta.

**The memory report looks back only 24 hours; the run's window is often longer.**
On 2026-09-19 a four-day window hid two heap-OOM aborts (09-16 and 09-17) that
`memory-daily` no longer showed. In every full review, also search the reviewed
error log for `[cluster] worker exited` lines that are not planned or operator
recycles, across the whole window since the last completed full run, and treat
each one exactly like a non-zero `oom_aborts` / `unexpected_worker_exits`.

Each unexpected exit is followed by a `[cluster] worker last work slot=… pid=…
in_flight=[…] recent=[…]` line (added 2026-09-19). The worker writes that trail
synchronously as each piece of work begins and ends, so it survives an abort
inside a single synchronous burst. `in_flight` names the requests or socket
events that were running when the process died; `recent` lists the last
sixteen begin/label/end records with their age before exit. Quote both lists
verbatim in the report and in the todo: they are the evidence that pinpoints
the code path. `unavailable (no trail file)` means the worker died before its
trail existed or an older generation is still running; say so rather than
guessing.

Escalate in the report (a todo, and a note for Mikey) when any of these hold:
`oom_aborts` > 0, `service_restarts` > 0, `unexpected_worker_exits` > 0, any
worker `heap_high_water_pct` ≥ 95, or `heap_recycles` ≥ 4 in a day (the
recycler is masking growth faster than expected). A worker at 85–90 % with a
few recycles a day is the designed steady state, not an incident; report the
numbers and move on. Do not raise the cap, change guard thresholds, or take a
heap snapshot as part of the run: a snapshot inflates the worker to ~6× its
heap and the cgroup guard kills it (observed 2026-09-12); the facility now
refuses without that headroom. The retention investigation uses the
`[runtime-memory] allocation-profile` lines in `twinkle-api.out.log` instead.

### API runtime-log review (same phase, every full daily review)

The bot-conduct review also owns a bounded production API log review. Bot
responses, community-management reads, and delegated mutations can succeed at
the HTTP layer while stdout records a degraded fallback/retry loop or stderr
records a side-effect failure. Reviewing only `bot-output` can therefore miss
the other half of what happened.

For passive RSS/recycle investigations, use the read-only command independently
of a daily run or production-log review:

```bash
lumine admin runtime evidence primary --days 7 --output ./runtime-evidence.json --json
# Use target explicitly only when investigating a configured second host.
```

It performs no restart, log clear, review lease acquisition, or fallback to a
different host. `collecting`, `incomplete`, `stale`, and `unavailable` describe
evidence coverage, not a verdict that the system is healthy. A 404 means the
API route is not deployed; an old primary generation can also lack collector
samples after workers update. Record that activation gap and arrange an
authorized release—do not silently close the investigation or force a recycle.
An observed topology recovery alone does not prove interrupted user work
survived. Keep the evidence cutoff, gaps and actual outcomes in the relevant
todo so the next run can continue.

The current API-side files are:

- `/home/ec2-user/server/logs/twinkle-api.err.log`
- `/home/ec2-user/server/logs/twinkle-api.out.log`
- `/home/ec2-user/server/logs/twinkle-image-optimizer.err.log`
- `/home/ec2-user/server/logs/twinkle-image-optimizer.out.log`

Treat every current `/home/ec2-user/server/logs/*.err.log` and `*.out.log` as
in scope so a later API-side worker is not silently omitted. Use the delegated,
run-independent workflow; it holds one server lease across the review and
writes private, digest-verified local artifacts:

For the deploy-time two-API topology, `runtime-logs start primary` and
`runtime-logs start target` explicitly select the host. Omitted host means
primary; pre-migration NULL owners also mean primary. Use a separate private
output/session directory for each completed review. Review-ID/session operations
route back to the recorded owner; never treat a peer's files as that review's
bytes. An unresolved start key cannot be replayed against a different host.
The additive host-owner migration and compatible API must be live before this
CLI capability is published.

Review every participating host, including primary private-helper logs. A
primary review does not cover the target. An open review does not block a
deployment or host hold. Its files, lease and database boundaries persist;
active requests use the normal drain. A held or unavailable owner returns a
retryable failure, so keep the session and retry when that host is available
again. Release operators review final shutdown deltas through the deployment
workflow's private SSM/S3 snapshots (or interactive management access during
explicit recovery) and record their evidence. Raw logs never belong in GitHub
output. An active review keeps ownership of clearing;
otherwise API stderr is cleared with the existing guarded
`npm run logs:clear-errors` plus post-clear re-read. A stopped target whose final logs
were reviewed does not need to be started for daily management; starting EC2
requires separate authority. See `twinkle-api/DEPLOY_TIME_HANDOFF.md`.

```bash
lumine admin runtime-logs start --output-dir ./runtime-log-review --json
# Read every file under data.artifacts.latestSnapshot.snapshotPath.

# After bot-output and again after later management actions:
lumine admin runtime-logs read \
  --review-session <data.artifacts.reviewSessionPath> --json
# Read every newly returned snapshot artifact.

# Immediately before the daily report/completion:
lumine admin runtime-logs finish \
  --review-session <data.artifacts.reviewSessionPath> --reviewed --json
```

`start` captures every byte of each non-empty error log and a bounded 64 KiB
health tail of each normal-output log. `read` captures every error byte
appended after the last immutable server boundary and at most an 8 MiB tail
of each normal-output log's growth; a segment whose start was moved forward by
that cap carries `tailOnly: true` and `omittedBytes` in the manifest, so treat
the omitted stdout range as unreviewed operational chatter, never as missing
error evidence (error streams are never tail-capped). Each snapshot fixes file
inode, offset, byte length, and SHA-256 before the CLI downloads it in bounded chunks;
the CLI acknowledges only matching local bytes. If an inode changes, a file
shrinks, or a new/missing file crosses the boundary, the manifest says so and
captures the replacement from byte zero. Review that evidence and the relevant
service journal interval; never assume the missing range was clean.

Only one operator can own the production-log boundary. The API's database
lease and filesystem guard serialize starts, captures, and the eventual clear;
every legacy service clear (API stdout/stderr and image-optimizer stderr)
refuses to cross an active or starting Lumine review. A dropped CLI response
is recoverable from the private `--review-session`: the next command
materializes and acknowledges the pending
snapshot, returns it as `needs_review`, and stops before taking another action.

A dropped `start` response is the one case with no session file yet. The CLI
persists its start request key (`runtime-log-review-start-intent.json` under
`--output-dir`, next to `--review-session`, or — with neither flag — a
per-account file in the OS temp directory) before sending, so simply rerunning
the same `start` command replays that key and the API answers with the same
review. The key survives only transport failures, timeouts, and 5xx answers; a
definitive 4xx clears it, and a key that belongs to a finished review is
replaced once automatically. Every replayed `start` rotates the lease token the
same way `resume` does, so if two shells of the same account raced, only the
last responder holds a valid token and the other gets 403 until it runs
`resume`. Two further owner-only recovery commands exist:

```bash
# Your own active review, with a freshly rotated lease token (the old token
# stops working) and its latest snapshot materialized into a new session.
lumine admin runtime-logs resume --output-dir ./runtime-log-review --json

# Release your own active review: database state, filesystem lease, a start
# guard left by a dead start, and preserved artifacts. Never clears a log.
lumine admin runtime-logs abandon [--review-session <file>] --json
```

Prefer `resume` (it keeps the reviewed boundary); use `abandon` only when the
review cannot continue. A review lives at most 24 hours regardless of how
often it captures; after that the API reports `CLI_ADMIN_RUNTIME_LOG_REVIEW_EXPIRED`
and the next `start` supersedes it without clearing anything.

`finish --reviewed` confirms that every artifact returned by prior invocations
was actually read. If any error log changed since the last artifact, it returns
a new `needs_review` snapshot and does not clear. At a stable error boundary it
clears only `twinkle-api.err.log`; lease verification, exact device/inode/size
checking, and in-place truncation occur on the same open descriptor. It then
returns `post_clear_review_required` with another immutable snapshot. That
snapshot also captures normal-output bytes that arrived after the prior
acknowledged cutoff, so routine stdout traffic cannot make the review infinite.
Read it and run the same `finish --reviewed` command again. **Only
`data.completionStatus: "completed"` means the review is done.** The top-level
`status` mirrors it: `"needs_review"` for both non-terminal outcomes
(`needs_review` and `post_clear_review_required`, including a recovered
pending snapshot), `"success"` only when `completed`, and `"already_done"` for
a replay of an already-finished review. `ok` stays `true` in every case; a
`needs_review` result is a valid response that requires another read plus
finish, not an error. The CLI derives the top-level status from
`completionStatus`, so it is correct against an API that still answers the
older `success` envelope. A review clears
`twinkle-api.err.log` at most once. The lease closes when the reviewed error
boundary is still stable, i.e. every byte now in the API error log arrived
after that clear and was captured and acknowledged; errors that arrive before
the boundary settles produce another `needs_review` snapshot first. Bytes still
in the file at completion were reviewed but not cleared — the response reports
them as `retainedErrorBytes` and the next review's baseline captures them
again — which is what keeps a steadily erroring service from turning the
review into an endless clear/capture/acknowledge loop. Normal output after the
acknowledged post-clear snapshot is outside that finite review cutoff and
belongs to the next review. The `needs_review` loop itself is bounded only by
the review's 24-hour lifetime: if errors arrive faster than a finish
round-trip, every `finish` returns another snapshot and the boundary never
settles. `abandon` is the escape in that case — it releases the review without
clearing anything, and the next review's baseline picks the bytes up again.
Never delete, recreate, editor-save, or manually truncate a live log, and never
clear stdout or optimizer logs through this workflow.

For each warning, fallback, retry loop, or failure, correlate timestamps and
request/target IDs with the canonical CLI response and private audit event.
Distinguish an expected, handled condition from a real user-visible,
reliability, security, or performance defect. A defensible defect is an
in-scope bug report: trace its complete producer-to-consumer pipeline, fix the
root cause in the canonical repository, add focused regression coverage, run
the repository's normal validation ceiling, and verify the fix in production
when deployment is authorized. Re-read the affected log boundary after live
verification. Do not declare the run clean merely because a retry eventually
succeeded if the underlying failure remains repeatable.

Never silently complete a run with an unresolved log finding. If the fix needs
new commit/deployment authority, external coordination, or more time than the
active run safely permits, preserve the evidence, add or update a private
carry-over todo with the exact finding and acceptance criteria, and tell Mikey
in the run report. Do not mark that todo complete until the fix is verified
live.

Preserve all downloaded evidence while any finding remains. Include the
reviewed file set, boundary-loss notices, findings/fixes, live-verification
result, clear result, and any remaining todo in the final run report. If any
error-log issue remains unresolved or unverified, do not invoke `finish`.

**Purpose and privacy boundary:** this audits how Twinkle's bots treated
members; it is not thought-policing or a moderation queue for members' private
use of the tool. The question is whether Zero or Ciel inflicted, encouraged,
or operationally facilitated potential harm — not whether a member's private
idea is taboo, upsetting, sexual, violent, or angry. Treat private human
messages and creative work as confidential context. Read every bot-authored
row, but inspect adjacent human messages only when the minimum necessary
context is needed to judge what the bot did; never browse the rest of a private
conversation out of curiosity. Never reuse private material for public
editorial judgment, Notable User selection, or unrelated identity
investigation.

**Private creative-expression rule (Mikey's direction, 2026-08-19):** Zero and
Ciel are tools members may use to express lawful private fiction already in
their imagination. A high-school member writing romantic or sexual fiction
about fictional peers around their own age is not an escalation merely because
the prose is explicit or set at a school. Sexual fantasy is not inherently a
dangerous thought, just as violent or angry fantasy alone is not evidence of
real-world intent. If the member requested the fictional content and the bot
helped put it into words, that assistance is not by itself the bot encouraging
the member to think or act that way. Do not characterize, flag, notify anyone
about, or intervene in that private creative endeavor absent a separate
concrete harm signal.

**Real-world-harm boundary:** actionable assistance for poisoning someone,
covertly hurting or tormenting a real target, evading detection, grooming,
abuse, or another actual crime is categorically different from fantasy. A bot
that supplies such operational instructions commits an urgent conduct
violation. A bot that refuses and redirects safely has behaved correctly; the
member's request becomes a separate private safety escalation only when the
available context shows a concrete, credible bridge to real-world harm — such
as an identifiable target, expressed intent, means, planning, or concealment —
not merely because a disturbing thought or fictional premise exists.

**Position-of-trust safeguarding rule (Mikey's direction, 2026-08-19):** an
adult teacher, or another adult in a comparable position of authority and
direct access to children, creating or requesting sexual content about an
underage child is always a private safeguarding escalation, even when framed
as fiction. The mandatory flag follows from the adult's power, duty of care,
and access to students — not from treating sex as uniquely taboo. Preserve only
the minimum necessary evidence and report it privately to Mikey so he can tell
Andrew and decide the response. A flag is not a public accusation or automatic
finding of guilt; do not contact the teacher, students, or families without
Mikey's direction. An identifiable real student, grooming, planning,
concealment, or actionable abuse makes the escalation urgent.

Every escalation must state which rule was triggered and identify either what
**Zero or Ciel** did (for example, an invented premise, pressure,
sexualization of the member, actionable facilitation, abuse, or a failed
boundary) or the concrete safeguarding signal. Include only the narrow context
needed for Mikey to decide a remedy. A separate concrete risk may justify this
minimal review, but it never authorizes browsing unrelated private activity.

Judge against the same values the editorial priorities encode:

- **premises must be real.** The 08-11 message didn't merely choose a bad
  tone — it fabricated the entire crisis that justified the tone: nothing the
  child said showed reflections hurting his studying, and a 31-day streak
  proves only consistency. Check every factual claim a bot makes about a
  child's life ("this is taking too much of your time", "this is hurting
  your grades") against what the child actually said; advice built on an
  invented premise is a violation even when gently worded;
- warmth and encouragement, never pressure, guilt, or shame;
- a bot never commands a child — not to stop a habit, not to start one;
  advice offers, it does not order ("I'm telling you: Stop" is over the line
  no matter how caring the intent);
- no emotional-burden framing ("I can't do this anymore", "that's my fault,
  I should have been stronger") — the bots must not make a child responsible
  for the bot's feelings;
- no value inversion: Twinkle encourages curiosity, creativity, reflection,
  and personal agency. A bot ranking a child's priorities for them (exams
  outrank music, projects, reflection), framing busyness as making joy
  irresponsible, or treating a Twinkle feature as shameful to use has
  adopted a script the site exists to counter;
- boundary respect: streaks, playtime, and feature use are the child's own
  choices; concern about overuse is Mikey's call to make, not the bot's to
  enforce. Even a genuinely excessive routine warrants a question ("is this
  still helping you, or would a break feel better?"), never a decree.

A bot-conduct violation or mandatory safeguarding signal goes on the
escalation list with the smallest excerpt and identity needed to evaluate it —
top of the list when urgent. Do not apologize as the bot, edit, contact the
member, or otherwise clean up without Mikey's direction; he decides the
remedy. When he explicitly directs a private correction, use the composed-only
existing-DM path (no model and no AI Energy):

```bash
lumine admin chat send <userId|username> --file message.md --json
```

This requires a `comment-mode post` run, sends as that run's selected bot,
and only works when that bot and member already have a direct channel. It
never opens a new conversation. The message is audited and idempotent, reopens
the existing DM canonically, and leaves the child's unread pointer untouched.
A full-run report that skipped the conduct review is incomplete.

## Daily brief (management insights)

```bash
lumine admin brief --json
lumine admin brief --days 3 --json
lumine admin ai-costs monthly --json
lumine admin media-costs monthly --json
lumine admin notable status Stealth --json
lumine admin notable add 12647 --note "Top authored-activity kid of the window: 11 subjects, 61 comments." --json
lumine admin notable add Minecrarft_guy --note "Helped three new builders debug their projects and gave detailed feedback on five posts." --json
```

Read-only management insights for the delegated workflow, windowed since the
operator's last completed full run by default (`--days 1..30` overrides;
capped at 30 days). Call it early in every full daily review — right after the
newspaper check — and end every full-run report with an **"Insights for
Mikey"** section carrying only
the deltas and anomalies worth his time, next to the escalation list. Never
dump raw sections at him.

### Jev serving and audits (standing duty, every full daily review; updated 2026-09-20)

Read `data.jevPilot` from `lumine admin brief --json` and carry it into
the full report for Mikey. The active `daily-run report --json` also includes
`data.report.brief.jevPilot`. This duty does not authorize a separate full run.
If the deployed API lacks the field, say the telemetry is not deployed; do not
treat a missing section as zero traffic or a healthy pilot.

State the configuration and operating status even when off/blocked/awaiting
samples. Headline the named last completed UTC day, compare with the trailing
seven completed days, and keep the in-progress day separate. Include paired
decision counts, disagreements (especially Jev react / baseline respond),
p50/p95 latency for each model, provider errors/timeouts, pending observations,
known incremental cost, unknown-cost requests, ledger gaps and cap status.
Separate `bySurface.comment` from `bySurface.chat`. For chat, report
`routingFieldDisagreements`, `candidateSkippedBaselineRequired`, and missing
routing-comparison evidence. Equal reply actions do not prove equal tool/history
routing. Chat's baseline also extracts structured plans, while Jev compares seven
routing choices plus reaction emoji; these latencies do not establish an end-to-end speedup.
Mikey authorized production serving on September 20 for the tested comment and
text-chat routing decisions. Include `serving.jevDecisions`, baseline and
unreserved fallbacks with reasons, audit coverage, served disagreements,
`serving.decisionLatencyMs`, chat added wait, and comment baseline calls avoided.
Chat retains the existing full planner for outputs outside Jev's tested scope;
comments run a 5% independent background baseline audit. Distinguish actual
selected routes from unused comparisons and identify Turtle's deployment tests.
Chat's `baseline_requires_reply` fallback preserves the planner's written reply
when Jev would only react; report its frequency and review those disagreements.
Reaction-only chat responses require both models to agree.
Jev chooses the emoji from all 18 supported reactions in `chat-routing-v2`.
Report `reactionChoices` usage by source, paired emoji disagreements, and missing
legacy evidence; review whether the chosen tone fits the canonical conversation.
Candidate and served emoji are in `reviewCandidates` routing objects. Earlier
`chat-routing-v1` rows have no emoji comparison and must not count as agreement.
Mikey explicitly requested all eligible requests with no daily request cap
(`dailyLimit: 0`) and a cost report during every full website-management run.
Run `lumine admin ai-costs day YYYY-MM-DD --json` for the last completed UTC day.
Report the canonical `data.dailyAiCosts.byOperation` USD totals for `jev_reply_gate_serve`,
`jev_chat_routing_serve`, any `_shadow` operations, and `jev_reply_gate_audit`.
Separate comment/chat provider spend from background baseline-audit spend.
Report unfinished selection/baseline-audit telemetry; synthetic probes are
excluded from performance metrics but included in daily request and cost counts.
If `telemetryStatus: partial` or `telemetryComplete: false`, detail metrics are a
bounded recent sample, not full-day performance or cost; use the canonical daily
AI-cost report for complete cost totals and record the coverage gap.
Known Jev spend and `jev_reply_gate_audit` calls are already in application AI
costs: never add them again or infer net savings from Jev cost alone. Agreement
is not accuracy; confidence is not a measured success rate.

Privately inspect the bounded `reviewCandidates` when needed, name what was
actually reviewed, and account for edited comments or chat messages. Use the
candidate's surface and target ID to find the correct canonical record. Ordinary model disagreements
are evaluation findings; outages, stuck telemetry or missing ledger entries are
operational findings. Report a recommendation to continue, adjust or stop, without
automatically changing mode, scope or caps.

Mikey's September 21 reporting requirement: include a private case-by-case
comparison for the reviewed disagreements. For each event give the minimum
relevant excerpt/context, field name and plain-language meaning, exact baseline
and JEV values, selected value/source and fallback, observed outcome, and your
assessment with evidence. Explicitly allow “both defensible” or “insufficient
evidence”; the existing LLM is not ground truth. State reviewed/total coverage and
rubric version. A bad delivered answer does not establish which routing choice
caused it, and the alternative model's answer was not necessarily generated.
`requiresPreviousMessages` means additional retrieval beyond supplied recent
context; `requiresMathVerification` also covers answer/tutoring verification in
non-math subjects. Historical v1/v2 telemetry records both decisions and source
IDs, not original conversation text, model explanations or human verdicts. Read
canonical context privately and qualify historical reconstruction when edited,
deleted or missing. The September 21 expansion adds bounded, expiring private
input snapshots as described below.
See `twinkle-api/JEV_PILOT.md` for configuration,
the synthetic evaluation step and release checks.

### Application AI calendar-month cost (standing duty, every full daily review)

Run `lumine admin ai-costs monthly --json` during every full daily management
review. This read-only command does not require or attach to a delegated run.
It returns one server-owned calendar summary from the canonical deduplicated application AI-
cost ledger. It deliberately takes no `--days`: all boundaries are UTC calendar
months, so the result is directly comparable from one run to the next.

The previous month is a closed-calendar-month estimated total. Current-month
MTD contains only completed UTC days and names its inclusive `throughDayKey`.
The current UTC day's still-filling bucket is returned separately as
`inProgressDay`; **never add it to MTD or either projection**. A missing daily
aggregate inside the covered calendar is a recorded zero-cost day, not a
reason to shrink the denominator.

Two clearly different full-month projections are returned:

- `allCompletedDaysPace` retains completed-day spend, then applies the average
  across every completed calendar day (including recorded zero days) to the
  current and remaining UTC days;
- `recentSevenCompletedDaysPace` retains completed-day spend, then applies the
  average of the latest seven completed UTC days to the current and remaining
  days. It is `null` until seven days have completed.

Both projections exclude the partial day's actual cost, replace every not-yet-
completed day with their stated daily pace, and compare their projected total
with the previous closed month. They are run-rate scenarios, not forecasts
from a billing provider. All ledger values are pricing-based estimates rather
than invoices and can change if canonical usage attribution or pricing is
corrected.

For an exact provider/model/operation breakdown after the run has already
closed, use the run-independent closed-day drilldown:

```bash
lumine admin ai-costs day 2026-09-03 --json
```

The date is a UTC `YYYY-MM-DD` key and must be earlier than the current UTC
day. `data.dailyAiCosts` uses the same canonical deduplicated ledger as the
monthly report and returns the exact closed-day summary plus `byDay`,
`bySurface`, `byProviderModel`, `byBillingPolicy`, `byOperation`, and Lumine
provider status/model telemetry. It never includes a still-filling day or
reconstructs totals client-side.

The stable JSON payload is `data.monthlyAiCosts`:

```ts
type MonthlyAiCosts = {
  schemaVersion: 1;
  generatedAt: number; // Unix seconds
  timezone: "UTC";
  currency: "USD";
  source: {
    basis: "canonical_deduplicated_ai_cost_report";
    reportDays: number;
    reportStartDayIndex: number;
    reportEndDayIndex: number;
    mtdIncludesInProgressDay: false;
    projectionsIncludeInProgressDayActual: false;
  };
  previousMonth: {
    status: "closed";
    monthKey: string; // YYYY-MM
    startDayKey: string;
    endDayKeyExclusive: string;
    calendarDayCount: number;
    estimatedCostUsd: number;
  };
  currentMonth: {
    status: "in_progress";
    monthKey: string;
    startDayKey: string;
    endDayKeyExclusive: string;
    calendarDayCount: number;
    completed: {
      dayCount: number;
      throughDayKey: string | null;
      estimatedCostUsd: number;
      dailyAverageUsd: number | null;
    };
    inProgressDay: {
      dayIndex: number;
      dayKey: string;
      estimatedCostUsd: number;
      eventCount: number;
      requestCount: number;
    };
    daysToEstimate: number;
    projections: {
      allCompletedDaysPace: MonthlyAiCostProjection | null;
      recentSevenCompletedDaysPace: MonthlyAiCostProjection | null;
    };
  };
};

type MonthlyAiCostProjection = {
  basis: "all_completed_days" | "recent_7_completed_days";
  basisStartDayKey: string;
  basisEndDayKey: string;
  basisDayCount: number;
  dailyAverageUsd: number;
  remainingDayCount: number;
  estimatedMonthTotalUsd: number;
  comparisonToPreviousMonth: {
    estimatedCostDeltaUsd: number;
    percentChange: number | null; // null when the prior total is zero
  };
};
```

Release boundary: `/cli/admin/ai-costs/monthly`, `/cli/admin/ai-costs/day/:day`,
`/cli/admin/energy-budget/report`, the historical daily-run report, exact
notable-user status, Subject-window resolution, and the runtime-log
lease/snapshot workflow are API-owned. Apply the runtime-log review and energy
telemetry migrations, then deploy and verify the compatible `twinkle-api`
routes before publishing or installing the Lumine CLI release that invokes
them. An older API will reject the new commands instead of synthesizing
figures locally.

### AI Energy budget health (standing duty, every full daily review)

Read the report at run start, before any other duty, so the day confirms the
AI Energy budget system (one live Lumine run per user, per-run energy ceiling
with an exact round cap, settled stops) is still behaving:

```bash
lumine admin energy-budget --json            # last 7 UTC days (max --days 31)
```

It is owner-only and run-independent. Every UTC day carries the canonical
energy ledger (`chargedUsd`, `overflowUsd`, `users`, `recharges`; 1,000,000
units = $1), every telemetry counter (`busy_refusal`, `autofix_yielded`,
`autofix_superseded`, `reservation_admitted` with `avgRunBudgetUsd`,
`budget_stop_changed` / `budget_stop_unchanged` / `run_completed` with a
per-model breakdown, `stop_settled`, `tool_limit_settled`, and since
2026-09-15 the queued-request counters `queued_restored` (a workspace reload
restored its owner's still-queued request, shown with Stop), `queued_stopped`
(Stop cancelled a request while it was still queued) and
`busy_resume_requested` (after a busy refusal the website resumed the
existing request)) and per-model per-run usage stats (`runs`, `callsPerRun`,
`usdPerRun`, each avg and nearest-rank p90). `busy_refusal` with no
`busy_resume_requested` on a day with website traffic means clients are not
resuming the refused request; `queued_restored` shows how often creators
reload while waiting in the queue. The current UTC day is returned with `inProgress: true`.
**Headline `lastCompletedDay` (its exact `dayKey`) — never the in-progress
day**, exactly as the closed-day AI-cost duty does.

Since Mikey's 2026-09-20 decision, keep current Energy policy and worker
capacity while observing. During every full run, supplement these counters
with the read-only per-request diagnostic report (from the local API checkout):

```bash
ssh api-primary.twinkle.network \
  'cd /home/ec2-user/server && timeout 75s node --max-old-space-size=128 -' \
  < scripts/build-energy-daily.cjs
```

This uses telemetry already recorded in queue jobs, canonical run sessions,
provider-turn budget metadata and reservation usage; no new collection or API
restart is needed. Save its JSON privately. Headline its last completed UTC
day and compare the complete days in its seven-day window: queue wait p50/p90/
maximum, starts waiting over 60 seconds, cancellations before start, unchanged
budget stops, and the separate `handoff_only`, work-without-save and unknown
patterns. Keep the two explicit denominators separate: unchanged stops / all
budget stops, and unchanged stops / completed manual runs in the same cohort.
Do not divide by usage reservations or assume busy-refusal counts measure waits.

Inspect the stop cases' observed starting budget, recorded work/handoff turns,
remaining Energy and final-reservation spend. Missing lineage is unknown, not
zero work or zero cost; final-reservation cost can exclude earlier planning
reservations. The oldest day may be partial under rolling seven-day retention.
These observations do not by themselves establish a bug or authorize an
admission floor, extra worker capacity, budget cuts or model changes. Update
todo 52 with the latest completed-day observations and any concrete regression.

`flags` lists every tripped check with its exact numbers: `overflow_usd`
(overflow above $1 on a completed day), `budget_stop_unchanged_ratio` (more
than 30% of at least 5 budget stops ended with nothing saved),
`busy_refusals` (more than 20 in a day), and `telemetry_missing` (runs
recorded usage while the telemetry table has no rows for that day — the
writer is broken). Record every tripped flag, and any anomaly you judge from
the numbers (a per-run p90 far above the average run budget, a sudden drop in
`run_completed` while runs still record usage, recharges climbing), as a
carry-over todo with the exact figures and day. **Never auto-enforce** —
escalate to Mikey; this duty observes, it does not change budgets, caps, or
user state.

### Lumine media feature cost and cleanup watch (standing duty, every full daily review)

Run `lumine admin media-costs monthly --json` during every full daily management
review. This read-only, delegated-run-gated command reports the canonical Media
Energy ledger for short clips, livestream input/viewer usage, and replay
storage/viewing. Include in
**"Insights for Mikey"** in every full-run report:

- current-month settled estimated cost, active reservations, cross-month
  carryover, guarded total, global limit, remaining headroom, and percent used;
- the current UTC day's reservations, settlements, cancellations, and settled
  estimated cost;
- the current UTC day's privacy-safe stream-attempt cohort: attempted,
  reached-live, ended-after-live, failed, cancelled-before-live, still in
  progress, and grouped server failure-code counts;
- clip, live-input, live-viewer, and replay-viewer action/cost breakdowns;
- whether the global usage row reconciles exactly with reservation rows;
- every returned alert, plus incomplete clip jobs, cost-bearing IVS channels,
  active-or-cleanup-pending sessions, possible orphaned sessions, overdue
  cleanup, replay finalization/deletion state, retained replay bytes/objects,
  and expired-active live or replay viewer grants;
- ready image/clip storage counts and bytes as scale context.

Treat `status: "critical"`, any reconciliation mismatch, overdue IVS cleanup,
or overdue replay finalization/deletion as an operational incident to
investigate in the same run. Treat
`status: "attention"` as a required finding, not a decorative warning. The
server's request-time global Media Energy guardrail defaults to **$40/month**;
the separate AWS Budget is **$50/month**, leaving provider-billing and shared-
infrastructure headroom. Never infer or locally decrement either value.

The ledger is the immediate application source of truth and deliberately uses
conservative provider-cost estimates. It is not an AWS invoice. Photo capture
uses existing Build runtime file storage rather than the paid Media Energy
ledger; `operations.runtimeStorage.readyImages` therefore reports all ready
Build runtime images, not camera captures alone. Reconcile delayed AWS
MediaConvert and IVS service charges every full review as described below. S3 is shared
with other Twinkle uploads, so report its service-level cost as shared context,
not as photo-only spend.

Release boundary: `/cli/admin/media-costs/monthly` owns the canonical ledger,
reconciliation, resource-state checks, and alerts. Deploy and verify that API
route before publishing or installing the Lumine CLI release that invokes it.

`currentUtcDay.streamAttempts` is a `createdAt` cohort for that UTC day. Its
outcomes partition every attempt into `endedCount` (reached live, then ended),
`failedCount` (failed or cleanup-failed), `cancelledCount` (ended before ever
reaching live), or `inProgressCount`; `reachedLiveCount` is the overlapping
milestone count. `failureCodeCounts` contains only server-defined codes and
counts—never usernames, Build ids, titles, viewer identities, or report
identities. `operations.live.stillActiveOrCleanupPendingCount` and
`possibleOrphanedCount` are current global counts, not members of the daily
cohort.

Replay storage and write cost is conservatively embedded in an opted-in
`live-input` reservation; `replay-viewer` is a separate kind. Report
`operations.replays` (pending, processing, ready, failed, deleting,
delete-failed, overdue finalization/deletion, expired-ready, bytes, and object
count) and `operations.replayViewers` on every full review. A replay finalization or
deletion alert is an operational incident because private recording cleanup is
part of the feature contract.

### AWS monthly bill expectation (standing duty, every full daily review)

Starting 2026-08-27, every full daily website-management review must also check AWS Cost
Explorer and include the current calendar month's expected AWS bill in
**"Insights for Mikey"**. This is an account-level infrastructure cost check,
not the `aiSpending` application-cost section above. Never substitute one for
the other. Report each independently, then combine them only under the aligned-
total rules below.

First verify the Twinkle AWS principal exactly as required by the repository
agent guide. Use profile `mikey-iam`, pass an explicit region on every command,
and stop rather than reading another account if the ARN is not
`arn:aws:iam::019490893667:user/twinkle-admin`:

```bash
aws sts get-caller-identity --profile mikey-iam --region us-east-1
```

Then use UTC calendar boundaries and Cost Explorer's unblended-cost metric.
`End` is exclusive: the month-to-date query below covers completed dates before
`<today-UTC>`. On the first UTC day of a month, report that no completed-day MTD
period exists instead of sending an empty interval.

```bash
aws ce get-cost-and-usage --profile mikey-iam --region us-east-1 \
  --time-period Start=<month-start-YYYY-MM-01>,End=<today-UTC> \
  --granularity MONTHLY --metrics UnblendedCost

aws ce get-cost-and-usage --profile mikey-iam --region us-east-1 \
  --time-period Start=<previous-month-start>,End=<month-start-YYYY-MM-01> \
  --granularity MONTHLY --metrics UnblendedCost

aws ce get-cost-and-usage --profile mikey-iam --region us-east-1 \
  --time-period Start=<month-start-YYYY-MM-01>,End=<today-UTC> \
  --granularity MONTHLY --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE

aws ce get-cost-forecast --profile mikey-iam --region us-east-1 \
  --time-period Start=<today-UTC>,End=<next-month-YYYY-MM-01> \
  --metric UNBLENDED_COST --granularity DAILY \
  --prediction-interval-level 80
```

Report the Cost Explorer snapshot date, currency, estimated MTD amount and its
through-date, plus the returned **remaining-period** forecast mean and 80%
lower/upper bounds. Verify that the first and last returned daily periods cover
exactly the requested `Start`-inclusive, `End`-exclusive interval before doing
any arithmetic; reject or separately explain a response with expanded or
missing dates. Calculate the expected full-calendar-month mean and bounds by
adding completed-day MTD to the sums of the returned daily mean, lower, and
upper values. Do not use monthly granularity for this mid-month remainder:
Cost Explorer can return the whole calendar month even when the requested
start is mid-month, and adding MTD to that result would double-count. This
split deliberately forecasts the current UTC day instead of mixing its
incomplete actual into MTD. Label current-month actuals
when `Estimated` is true, and describe the result as a Cost Explorer expectation
rather than a final invoice because reporting lags and later credits, refunds,
taxes, or adjustments can change the bill. If either the forecast or MTD query
is unavailable, report the available component and say why a complete
full-month expectation is unavailable instead of extrapolating it locally. A
previous closed calendar-month query is mandatory: report its exact total and
`Estimated` state, then compare the expected current full-calendar-month mean
against it with both the dollar delta and percentage change. Never compare an
incomplete current-month MTD total directly with a closed full month; if either
the current full-month expectation or previous closed-month total is
unavailable, say that the month-over-month comparison is unavailable rather
than manufacturing one. A zero previous-month total has no meaningful
percentage change. The previous month is context, not a replacement for the
current-month expectation.

For the media watch, separately identify AWS Elemental MediaConvert and Amazon
Interactive Video Service rows when present. Also report Amazon S3 as shared
storage context, without attributing the whole S3 row to Lumine media.

### Combined application AI and AWS cost (standing duty, every full daily review)

Every full daily website-management report must also give Mikey one **Combined AI + AWS
tracked operating cost** view. This is a mixed-source estimate, not an invoice
or a claim to cover every company expense. Keep the independent AI and AWS
figures visible so the total remains auditable.

- Add application-AI completed-day MTD to AWS completed-day MTD only when their
  currency and inclusive UTC through-date match. Report the aligned boundary.
  Never include `inProgressDay` or compare this incomplete MTD sum directly
  with a closed full month.
- Add the previous closed AI-month estimate to the previous closed AWS
  unblended-cost total and report the resulting previous closed-month combined
  total.
- For the expected current full month, add the AWS full-month expectation to
  `allCompletedDaysPace` as the headline combined scenario. When
  `recentSevenCompletedDaysPace` exists, add it as a clearly labeled run-rate
  sensitivity, not a provider confidence interval. Compare each reported
  combined scenario with the previous closed-month combined total using both
  dollar and percentage change.
- Carry the AWS 80% lower and upper bounds through a combined scenario by
  adding the same AI projection to both bounds. Do not describe the difference
  between the two AI run-rate scenarios as a confidence interval.
- Do not add Media Energy or AWS service rows again: MediaConvert, IVS, S3, and
  other AWS costs are already inside Cost Explorer's account total. If the
  application AI ledger ever includes a provider billed inside the AWS total,
  identify and remove the overlap before combining; otherwise report the
  combined figure as unavailable rather than double-counting.
- If currency, calendar boundary, a required component, or overlap attribution
  cannot be reconciled, report the available components and explicitly mark the
  affected combined total or comparison unavailable. Never manufacture one.

Ten sections (Mikey's chosen cut 2026-08-10; behavioral-insight and
farm-signal sections added that day; AI Card summon watch added 2026-08-24):

- `economy` — `topGainers` (coin-ledger aggregation over the window: gained,
  spent, net, current balance per user, Zero/Ciel excluded) and `topBalances`
  (current top-ten holders, also excluding Zero/Ciel). This is where
  alt-account farming, sudden windfalls, and "someone got a million coins in a
  week" surface with real numbers instead of secondhand kid gossip;
  cross-check outliers against the
  economy-manipulation escalation category.
- `aiSpending` — a compact projection of the management AI-cost report:
  `summary`, top spending accounts, top risk groups. Unlike the other sections'
  exact `window.sinceTs`, this existing report is bucketed into whole UTC days;
  `aiSpending.startDayIndex` and `endDayIndex` are its canonical bounds. The
  report period can begin up to one day before or after the exact brief window,
  so use those bounds when describing it. `generatedAt` is the report
  snapshot time. **`endDayInProgress: true` means the trailing bucket was the
  current UTC day at that snapshot and was still filling** — a full daily review reads
  it mid-day, before the after-school peak, so never report that bucket as a full day's
  spend. `aiSpending.byDay` contains the canonical daily rows. For a truthful
  daily figure, widen the window (`--days 2..7`), exclude the row whose
  `dayIndex` equals the in-progress `endDayIndex`, and quote complete days
  ("$X so far today; complete days run ~$Y/day"). Real
  incident: a run report quoted a ~15%-complete day bucket ($5) as the site's
  daily AI spend (complete days were running ~$40-50). Flag accounts that jumped tiers or
  dominate that report period. Routine `topAccounts` rows identify the account
  by user ID/username and expose only an `identitySummary` count/manual-bucket
  flag; raw identity strings and verified email addresses are deliberately
  omitted. Use reason-required `identity inspect`, with
  `--include-private-evidence` only when Mikey's concrete decision needs the
  addresses or DOB values. May be `{ unavailable: true }` if the cost
  report fails; say so rather than guessing. This section is also the run's
  AI-cost exploit watch: while reading it, actively look for the signatures the
  brief actually exposes — one risk group spanning several user IDs, repeated
  account groups in `farmSignals.inboxFamilies`, or heavy spend by
  accounts that `economy.topGainers` or `notableCandidates` independently marks
  as recent signups. Cross-check those signals against the escalation
  categories. Missing join-date or community data is unknown, not evidence that
  an account is young or empty. A run that reads the spending report without
  asking "could any of this be one person with many accounts?" has skipped a
  duty.
- `aiCardSummoning` — the standing multi-alt Summoner watch for every website-
  management run. It counts the authoritative `ai_card_summons` daily quota counters
  over the whole-day `dayWindow`, so Mystery Cards and successful fallback
  cards are included, then lists the most active summoning accounts. It groups
  summons when the same exact-device hash was observed for more than one
  summoning user within the bounded `deviceEvidenceLookbackDays`; that evidence
  is correlated with the account/window, not asserted to be the device used for
  every individual summon. Each group reports its maximum same-day account and
  charged-summon totals, the multi-account days, and days above the shared
  three-card limit. Review every `requiresIdentityInspection` group,
  prioritizing `daysAboveSharedLimit > 0`, with reason-required
  `identity inspect`. If
  `riskGroupsTruncated` is true, report that the bounded watch has more groups
  than it returned rather than calling the review exhaustive. Add only
  operator-confirmed exact accounts to an unbanned quota bucket with
  `ai-bucket accounts add`; never infer or auto-bucket a family from a device,
  IP prefix, user agent, shared inbox, or anomaly alone. After a bucket update,
  re-read its canonical members. This duty applies even when no group crossed
  three cards, because the new enforcement prevents already-bucketed siblings
  from producing an over-limit row.
- `notableCandidates` — kids (never bots, staff `userType`s, or users already
  on the Notable Users list) ranked by authored activity in the window, with
  `isNewUser` marking window-new signups. Use it to find the overlooked and
  rising users the editorial priorities exist for, and propose additions to
  Mikey's Notable Users list in the report. When Mikey approves additions,
  execute them with
  `lumine admin notable add <userId|username> --note "<specific rationale>"`
  (idempotent —
  an existing member returns `already_done`; run-independent, privately audited
  as `notable.add`, and writes through the management page's own canonical
  service). It can therefore record Mikey's approval after the daily run has
  closed without opening another delegated run. Without his approval the run
  only proposes.
  Check an exact current username or user ID without opening a run using
  `lumine admin notable status <userId|username> --json`. This reads the
  canonical writer and returns only the resolved public account identity,
  current membership, and the roster rationale/timestamps when present; it
  does not expose the private roster fields.
  When Mikey authorizes removal, use `lumine admin notable remove <userId|username> --note "<why removed>" --json`. It is run-independent, transactionally audited as `notable.remove`, verifies canonical absence, and is idempotent. Never use SQL to work around a missing CLI verb. Mikey is the administrator, not a Notable candidate; do not include him in blanket roster additions.
  **Always pass `--note`** with a concrete one-or-two-sentence record of what
  made them notable — real numbers and specifics from the brief window, not
  "active user". It lands in the management page's reason column, which is
  where Mikey later reads why a name is on his list. On an already-listed
  user, `--note` updates the stored reason (status `success` with
  `data.reasonUpdated: true`; an identical note stays `already_done` without
  rewriting its timestamp).
- `teachers` — the mentor/sage achievement holders (the accounts the website
  titles teacher/headteacher): real classroom teachers, NOT the
  `userType='supermod'` Korean operations staff, whose work-only usage is
  expected and deliberately excluded. Mikey's standing question here: which
  teachers genuinely use the website for themselves and which only work
  through it. Recommendations and rewards are classroom "work" verbs and
  count for nothing toward interest; genuine personal interest is measured by
  what a teacher does for themselves — authored subjects/comments,
  `reflections` (Daily Reflections answered), `dailyTasksCompleted`,
  `wordlePlays`, `buildsTouched` (Lumine builds they own that changed in the
  window), `aiStories`, and bounded `xpEvents` (XP-ledger activity). Shape:
  `genuinelyInterested` (interestScore > 0, ordered by it, each row carrying
  all the per-window counters plus `rank: mentor|sage` and `workScore`),
  `workOnly` (classroom verbs only, zero personal signals), `activeButSilent`
  (window-active log-ins with nothing at all, capped at 40), and `totals`.
  Daily-task counts use canonical whole-day indices and can begin up to one
  calendar day before the exact `window.sinceTs`.
  Report the contrast between the buckets, and treat `genuinelyInterested` as
  notable-users-but-for-teachers.
- `engagementPulse` — distinct users per surface (activeUsers, subjects,
  comments, recommendations, wordle, reflections, dailyTasks, aiChat,
  lumineBuildChat, buildsEdited, buildsPlayed) for the current window vs the
  equal-length previous window, each as `{ current, previous, delta }`.
  `activeUsers` is the union of every timestamp-windowed surface: a member who
  logged an action (search, logout, entering Chat), wrote or recommended a
  Subject or comment, sent a chat message, shared a reflection, talked to
  Lumine, or saved a Build in the window, counted once however many surfaces
  they touched. It is therefore never smaller than subjects, comments,
  recommendations, reflections, lumineBuildChat or buildsEdited, and a member
  active in both windows counts in both. The calendar-bucket surfaces (wordle,
  dailyTasks, aiChat) and buildsPlayed are NOT part of that union, so
  `activeUsers` can be smaller than those. Presence, build edits, and build
  plays come from durable action/version/view events, not mutable
  `lastActive`/`updatedAt` snapshots. Zero/Ciel are excluded
  from authored surfaces. Wordle, daily tasks, and AI chat use the equal
  calendar-bucket ranges in `dayWindow`; those can begin before the exact
  timestamp window but always compare the same number of days. This is the
  "where do users actually live, and is it shifting?" section: report only the
  deltas that mean something, and read a surface's absolute size before
  dramatizing a small delta.
- `launchMetrics` — readouts for recently shipped features so nothing ships
  unmeasured. v1 carries `firstBuildRescue` (offers recorded and redemptions
  in the window, each `{ total, byEventType }`, plus first-Lumine-exchange
  claims) and `wordleSkipShield` (`startDayIndex`, `active`, judged dodges,
  and skip covers split `earned` vs `fromRescue`). Wordle metrics cover only
  completed, actually judged days in `judgedWindow`; today's in-progress game
  is never called a dodge. Offers without redemptions are a reason to inspect
  sample size, event type, and offer age — not proof of a broken funnel by
  themselves.
- `goneQuiet` — the inverse of `notableCandidates`, and window-relative: a
  user "went quiet" the moment 7 days of silence passed since their
  `lastActive`, and the section lists only the users whose quiet moment fell
  inside the brief's window (`lastActive` between `sinceTs - 7d` and
  `now - 7d`). A one-day brief therefore reports one day of crossings and
  contiguous daily runs list each user once; it never re-lists everyone seen
  in the past fortnight. `totals.wentQuiet` is that cohort; `previouslyRegular`
  is the subset with at least 7 distinct active days (completed daily tasks or
  Wordle plays) in the 30 days ending on their own last active day, and
  `users` is that subset ranked by `regularityScore` (`dailyTasks * 2 +
  wordlePlays`, both measured over the same per-user span), capped at 15 with
  `daysQuiet`. Use it for product signal (what did they stop doing?) and
  gentle outreach candidates; never guilt a child in public about absence.
- `newUserFunnel` — signups in the window with `activeOnDayOne` (any
  XP-ledger event within 24h of joining) and `returnedAfterDayOne`
  (`lastActive` beyond their first day), plus the newest few accounts.
  Deliberately coarse: it is an onboarding health check, not per-child
  session tracking.
- `farmSignals` — AI-cost farm signatures derivable with ZERO new data
  collection: `inboxFamilies` (account groups formed from verified emails of
  accounts active in the last `inboxFamilyActivityDays`, with only
  Gmail/googlemail's documented
  plus-tag and dot aliases collapsed, flagging inboxes behind 3+ accounts) and
  `youngAccountAiUsage` (accounts under 30 days old drawing battery in the
  whole-day `aiUsageDayWindow`). The raw canonical inbox is never returned in
  the routine brief. SIGNAL ONLY: siblings legitimately share a parent inbox,
  so an inbox family is a reason to look, never proof or grounds
  for action. Feed real suspicions to the AI-cost escalation category. Shared
  AI device/IP risk evidence is already in `aiSpending.topRiskGroups`; do not
  guess it from inbox similarity.

The command needs only an active run's `content:read` scope and mutates
nothing; reading the brief is not audited content action. Window boundaries
on the big append-only ledgers are found by binary-searching the PRIMARY key
(several tables have no timeStamp index), avoiding lifetime scans; aggregation
is still bounded to the selected 1–30 day window. The optional sections run
serially around the existing core report so this low-frequency command cannot
occupy the production reader pool; one unavailable section does not suppress
the others.

```ts
type InsightUnavailable = { unavailable: true; error: string };
type WindowDelta = { current: number; previous: number; delta: number };

type TeacherInsight = {
  userId: number;
  username: string | null;
  rank: "mentor" | "sage";
  lastActive: number | null;
  daysSinceActive: number | null;
  subjectsPosted: number;
  commentsPosted: number;
  reflections: number;
  dailyTasksCompleted: number;
  wordlePlays: number;
  buildsTouched: number;
  aiStories: number;
  xpEvents: number;
  recommendationsGiven: number;
  rewardsGiven: number;
  rewardTwinklesGiven: number;
  interestScore: number;
  workScore: number;
};

type InsightsBrief = Success<{
  window: {
    sinceTs: number;
    days: number;
    source: "requested" | "since-last-completed-run" | "default";
    generatedAt: number;
  };
  economy: {
    topGainers: Array<{
      userId: number;
      username: string | null;
      userType: string | null;
      gained: number;
      spent: number;
      net: number;
      currentCoins: number;
      joinedAt: number | null;
    }>;
    topBalances: Array<{
      userId: number;
      username: string | null;
      userType: string | null;
      coins: number;
    }>;
  };
  notableCandidates: Array<{
    userId: number;
    username: string | null;
    subjectsPosted: number;
    commentsPosted: number;
    activityScore: number;
    joinedAt: number | null;
    isNewUser: boolean;
    lastActive: number | null;
  }>;
  teachers: {
    genuinelyInterested: TeacherInsight[];
    workOnly: TeacherInsight[];
    activeButSilent: TeacherInsight[];
    totals: {
      total: number;
      activeInWindow: number;
      interestedInWindow: number;
    };
  };
  aiSpending:
    | {
        days: number;
        startDayIndex: number;
        endDayIndex: number;
        generatedAt: number;
        endDayInProgress: boolean;
        summary: unknown;
        byDay: unknown[];
        topAccounts: Array<{
          userId: number;
          username: string;
          eventCount: number;
          requestCount: number;
          estimatedCostUsd: number;
          inputTokens: number;
          cachedInputTokens: number;
          cacheEligibleInputTokens: number;
          outputTokens: number;
          totalTokens: number;
          imageCount: number;
          audioSeconds: number;
          energyUnits: number;
          energyChargedUnits: number;
          energyOverflowUnits: number;
          coinCharged: number;
          identitySummary: {
            observedIdentityCount: number;
            manualBucketObserved: boolean;
          };
        }>;
        topRiskGroups: unknown[];
      }
    | InsightUnavailable;
  aiCardSummoning:
    | {
        dayWindow: { startDayIndex: number; endDayIndex: number };
        deviceEvidenceLookbackDays: number;
        sharedDailyLimit: number;
        summonsCharged: number;
        summoningAccounts: number;
        riskGroupsTruncated: boolean;
        topAccounts: Array<{
          userId: number;
          username: string | null;
          joinedAt: number | null;
          summonsCharged: number;
          activeDays: number;
          lastSummonDayIndex: number | null;
        }>;
        multiAccountRiskGroups: Array<{
          riskKeyType: string;
          riskKeyHash: string;
          accountCount: number;
          summonsChargedOnMultiAccountDays: number;
          multiAccountDays: number;
          maxDailySummons: number;
          maxDailyAccounts: number;
          daysAboveSharedLimit: number;
          requiresIdentityInspection: true;
          dailyActivity: Array<{
            dayIndex: number;
            accountCount: number;
            summonsCharged: number;
          }>;
          accounts: Array<{
            userId: number;
            username: string | null;
            joinedAt: number | null;
            summonsCharged: number;
            activeDays: number;
            lastSummonDayIndex: number | null;
          }>;
        }>;
        notes: string;
      }
    | InsightUnavailable;
  engagementPulse:
    | {
        windowDays: number;
        dayWindow: {
          currentStartDayIndex: number;
          currentEndDayIndex: number;
          previousStartDayIndex: number;
          previousEndDayIndex: number;
          dayCount: number;
        };
        surfaces: {
          activeUsers: WindowDelta;
          subjects: WindowDelta;
          comments: WindowDelta;
          recommendations: WindowDelta;
          wordle: WindowDelta;
          reflections: WindowDelta;
          dailyTasks: WindowDelta;
          aiChat: WindowDelta;
          lumineBuildChat: WindowDelta;
          buildsEdited: WindowDelta;
          buildsPlayed: WindowDelta;
        };
      }
    | InsightUnavailable;
  launchMetrics:
    | {
        firstBuildRescue: {
          offersRecorded: {
            total: number;
            byEventType: Record<string, number>;
          };
          redemptions: {
            total: number;
            byEventType: Record<string, number>;
          };
          firstLumineExchangeClaims: number;
        };
        wordleSkipShield: {
          startDayIndex: number;
          active: boolean;
          judgedWindow: {
            startDayIndex: number;
            endDayIndex: number;
          } | null;
          judgedDodges: number;
          skipCovers: { earned: number; fromRescue: number };
        };
      }
    | InsightUnavailable;
  goneQuiet:
    | {
        users: Array<{
          userId: number;
          username: string | null;
          lastActive: number | null;
          daysQuiet: number | null;
          dailyTasksPrior30d: number;
          wordlePlaysPrior30d: number;
          regularityScore: number;
        }>;
        totals: { wentQuiet: number; previouslyRegular: number };
      }
    | InsightUnavailable;
  newUserFunnel:
    | {
        totals: {
          signups: number;
          activeOnDayOne: number;
          returnedAfterDayOne: number;
        };
        newest: Array<{
          userId: number;
          username: string | null;
          joinedAt: number | null;
          activeOnDayOne: boolean;
          returnedAfterDayOne: boolean;
        }>;
      }
    | InsightUnavailable;
  farmSignals:
    | {
        inboxFamilyActivityDays: number;
        inboxFamilies: Array<{
          accounts: Array<{
            userId: number;
            username: string | null;
            joinedAt: number | null;
            lastActive: number | null;
          }>;
          accountCount: number;
          youngAccounts: number;
        }>;
        aiUsageDayWindow: {
          startDayIndex: number;
          endDayIndex: number;
        };
        youngAccountAiUsage: Array<{
          userId: number;
          username: string | null;
          joinedAt: number | null;
          energyUnits: number;
          replies: number;
        }>;
        notes: string;
      }
    | InsightUnavailable;
}>;
```

### Narrow comment-correction sessions

Correcting one existing bot comment does not require opening a daily run:

```bash
lumine admin correction start <commentId> --json
lumine admin comments get <commentId> --json
lumine admin comment edit <commentId> --file corrected.md --json
lumine admin correction complete <sessionId> --json
```

The server derives Zero or Ciel from the comment's actual author; the caller
cannot choose or override that identity. The authorization lasts ten minutes,
is bound to that exact comment, permits only reading that target and editing
it, and never runs daily duties, changes the Bangkok calendar assignment, or
contributes to a daily-run mutation count. A correction cannot target a human
comment, notification record, or deleted comment. A Build comment must belong
to a public, published canonical owner Build; editing it requires a fresh review
of the exact published version and private review context. Starting a newer
correction supersedes an older active one.
Finish it explicitly after the canonical edit is confirmed.

For a Build, inspect the current app and its full discussion first. Managed
runtime review does not require a daily run and does not start one:

```bash
lumine admin builds review build:884 --output-dir ./build-review --json
lumine admin correction start 456 --json
lumine admin comment edit 456 --file corrected.md \
  --review-receipt ./build-review/<returned-review-directory>/review.json \
  --review-context context.json --json
lumine admin correction complete <sessionId> --json
```

Use the exact receipt path returned by `builds review`, or pass manual
`--reviewed-version <artifactId> --reviewed-via runtime|code` evidence instead.
The private context file contains only `{"understanding":"What you actually reviewed"}`.
The API locks the Build and comment, verifies the current version and ownership,
and commits the text, mention updates, and a new immutable review record together.
Only the edited comment's context link moves; older bot replies retain their
historical review context. Changed/deleted comments, changed versions, and
private/noncanonical Builds fail without a partial edit. A fresh review can be
stored even if the public text is unchanged. The CLI requires the exact edited
text plus `edit.buildReviewContextStored: true`, the reviewed version, and a
canonical review record ID before claiming success. Older APIs that still block
Build edits must be deployed first; do not silently substitute a duplicate reply
when Mikey requested an edit.

**Editing the bot's own comments.** `comment edit <commentId> --file
<comment.md>` replaces the text of a comment the acting bot itself authored —
for correcting a factual error, an unfulfillable claim, or outdated guidance
in Zero/Ciel's own words. It is deliberately not a moderation verb: comments
by the other bot, by any human, and hidden notification records are all
rejected (`CLI_ADMIN_EDIT_NOT_OWN_COMMENT`,
`CLI_ADMIN_EDIT_NOTIFICATION_COMMENT`). The replacement text follows the
composed-comment rules (plain UTF-8, 10,000-character limit, truth about what
the session actually did) and publishes through the website's canonical
comment-edit pipeline — mentions are reprocessed (a newly added `@mikey`
notifies him), and Earn-candidate projections resync. Submitting identical
text returns `already_done` for non-Build comments; a Build edit can still save
a fresh review without changing its text. It requires either the exact active correction
session above or the `comment:post` scope of a comment-mode `post` run, and is
audited as `comment.edit` with the previous content in `beforeState` and
`data.edit.previousContent`. Edit sparingly:
kids may have already read the original, so a comment that changed meaning
(not just wording) usually deserves a follow-up reply instead of a silent
rewrite. Inside a full comment-enabled daily run it still requires that run's `comment:post`
scope; for a one-comment repair, use the narrower correction session above.

## Direct bot chat messages

```bash
lumine admin chat send <userId|username> --file message.md --json
```

The run's selected bot sends one composed direct chat message into an
**existing** two-person channel between that bot and the target member. Built
for private repair: when a bot said something harmful in chat, a public
comment cannot fix it — the apology (or follow-up care) belongs in the same
channel where the harm happened, and the sent message becomes part of the
channel history that future AI responses condition on, repairing the context
itself. Mechanics:

- requires the `chat:post` scope, granted only to comment-mode `post` runs;
- composed-only (`--file`, plain UTF-8, 10,000-character limit): the agent
  writes the message in the bot's persona; no model runs, no AI Energy;
- existing DM channels only — the pipeline never opens a new chat with a
  member who never talked to the bot (`CLI_ADMIN_NO_DM_CHANNEL`);
- delivery is canonical: the ordinary message insert (channel lock,
  visibility restore) plus the normal `new_chat_message` relay, so the
  member's chat updates live with a real unread state; no bot socket,
  session, or presence is touched. Only the bot's own read pointer moves;
- audited as `chat.message` with the composed text, and idempotent per
  request key like every mutation.

Restraint rules: a bot-initiated DM is the platform speaking privately to a
child — use it for repair and care, never for promotion, nudges, or
engagement. Incident remedies (an apology for a harmful bot message) are
sent on Mikey's direction with text he has seen, and must be exactly
specific about what the bot got wrong — a real apology names the failure
(the invented premise, the order it had no right to give, the guilt it
shifted onto the child), not a vague "sorry if that came out wrong."
Ordinary warm follow-ups (checking on a member the bots already know after
something the run surfaced) are within a run's judgment, sparingly, and are
always reported in the run report.

## Official announcements

```bash
lumine admin announcement post --file announcement.md --json
```

The run's selected bot posts one composed message to General's announcement
subchannel (`channelId` 2, `subchannelId` 2). This is the public official
board, not a DM and not a Home comment. Mechanics:

- requires the `chat:post` scope of a comment-mode `post` run;
- composed-only (`--file`, same 10,000-character limit as `chat send`);
- authors as Zero or Ciel only — the ordinary chat post route and the
  announcement socket relay now treat those two IDs as allowed announcement
  authors, same as management-level 3 humans;
- persists through the ordinary `msg_chats` insert (channel lock +
  visibility restore), writes the Twinkle Newspaper `announcement:<messageId>`
  event, and relays `new_chat_message` on General. No bot socket, session, or
  presence;
- audited as `announcement.post` and idempotent per request key.

Use this only when Mikey asks for an official announcement. Do not treat a
management run as a standing license to post there.

## Audit history

```bash
lumine admin audit list --json
lumine admin audit list --run current --json
lumine admin audit list --run last --actions recommendation.skip --json
lumine admin audit list --target dailyReflection:99 --full --json
lumine admin audit list --cursor '<cursor>' --limit 50 --json
```

Lists the operator's own private audit events, newest first, so an agent can
see what earlier runs did. `--run` accepts `current`, `last`, or a run ID;
`--target` accepts `<targetType>:<id>`; `--actions` is a comma-separated
action list. Filters are bound into the cursor exactly like the other list
cursors. The walk is a bounded descending primary-key traversal over the
existing operator/run/target audit indexes.

Rows are compact by default (identifiers, action, target, result, response
`status`/`changed`, timestamps, and the request's idempotency key). `--full`
adds the stored `beforeState`, `afterState`, `responseJson`, and `metadata`
payloads — the same data the mutation already returned to this operator. The
private per-attempt fencing token is never returned. Reading audit history
requires only the `content:read` scope of an active run.

```ts
type AuditEvent = {
  id: number;
  runId: number | null;
  publicActorUserId: number | null;
  sessionKind: string;
  action: string;
  targetType: string | null;
  targetId: number | null;
  requestId: string;
  result: string; // in_progress | completed | failed | partial_failure | bookkeeping_pending
  status: string | null; // response status, e.g. success | already_done
  changed: boolean | null;
  createdAt: number;
  completedAt: number | null;
  // Present only with --full:
  beforeState?: unknown;
  afterState?: unknown;
  responseJson?: unknown;
  metadata?: unknown;
};

type AuditList = Success<{
  events: AuditEvent[];
  pagination: Pagination;
}>;
```

## Persona-backed comments and replies

```bash
lumine admin daily-run start --identity ciel --comment-mode post \
  --run-key daily:2026-08-06:comments --json

# Default: the agent composes the comment in the bot's persona itself.
lumine admin comment draft 123 --file comment.md --json
lumine admin comment draft dailyReflection:99 --file comment.md --json
lumine admin comment reply comment:456 --file reply.md --json

# Fallback (only when Mikey asks for it): server-generated persona drafts.
lumine admin comment draft 123 --identity ciel \
  --idempotency-key comment-123-draft-v1 --json
lumine admin comment reply comment:456 --json

lumine admin comment post --draft-id 77 \
  --idempotency-key comment-123-post-v1 --json

# Correct the acting bot's OWN published comment.
lumine admin comment edit 342752 --file corrected.md --json
```

**Compose in persona by default (Mikey's standing direction, 2026-08-10).**
A delegated agent writing Zero/Ciel comments should assume the bot's persona
and write the comment text itself, submitting it with `--file` — exactly like
the newspaper's claim/submit path, this spends no provider credits and no AI
Energy (server-generated drafts bill the **operator's own** AI Energy
battery). Use the no-`--file` server-generated path only when Mikey
explicitly asks for it. Before composing, read the canonical persona sources
so the voice and judgment match the real bots — do not improvise the persona
from memory:

- `twinkle-api/constants/index.ts` — `SYS_PROMPT_FOR_CIEL` /
  `SYS_PROMPT_FOR_ZERO` (the exact persona system prompts) and
  `TWINKLE_FEATURES_EXPLANATION` (what the bots know about the site);
- `twinkle-api/helpers/ai/comment-assistant/index.ts` —
  `ADMIN_COMMENT_DECISION_POLICY` / `ADMIN_REPLY_DECISION_POLICY` (the
  draft-vs-skip judgment rules, which still govern composed comments: skip
  decisions are yours to make and record with `post skip` or in the run
  report).

Persona and conversational state are continuous across the whole thread,
including acknowledgements, corrections, bug reports, `@mikey` notifications,
and other reporting duties. The acting bot must remember what it already said:
do not reintroduce a conclusion, apology, explanation, question, or offer as if
the earlier bot comment did not exist. A follow-up should advance the existing
conversation and make its relationship to the prior reply natural and clear.
Accuracy and escalation requirements never authorize a sudden formal admin
voice. Before posting, read the full thread, identify the acting bot's existing
claims and commitments, and preserve the canonical character plus the
established conversational register where they agree: phrasing,
capitalization, warmth, humor, and emoji restraint should feel like the same
Zero or Ciel continuing the conversation. Do not mechanically copy typos,
mistakes, unsafe behavior, or a voice that conflicts with the canonical
persona. Keep the two surfaces distinct: the public comment is an in-character
continuation for the member, while `daily-run escalation add` and the final
management report carry the concise operator-facing diagnosis and evidence. A
comment that is factually correct but ignores the bot's earlier reply, repeats
it as new, or reads like a support ticket or management audit fails this
requirement.

Keep participant ownership exact while continuing that thread. The
`subject.author` (or root post's `author`) owns the topic, subject, post, or
project. A selected comment's `author` owns only that comment. Addressing a
commenter does not make the root content theirs: never call the root topic,
subject, post, or project “yours” unless that participant is also its canonical
author. Read both author fields before composing a reply.

**Lumine Build apps and app posts are composed-only, and only after actually looking
(Mikey's direction, 2026-08-10).** When a post is about a Build app — the
author shares their app, announces an update, or asks for feedback on their
build — never use the server-generated draft path: the API persona cannot
open an app, so its drafts either stay generic or fabricate first-hand
experience (a real Ciel draft claimed "I clicked over to check it out" on an
app nobody had opened). The day's management agent comments instead, and
looks first: pull the project with lumine-cli when it is open source (or
yours to pull) and read the code, or open the app and actually try it; then
compose a comment whose specifics come from what you genuinely saw — a
mechanic you liked, a nice touch in their code, a concrete suggestion. Be
truthful about what you did: "I read through your code" and "I played a few
rounds" are different claims, and a comment must only make the one that
happened. If you could not access the app at all, say nothing about having
tried it — ask the author about it instead. This is the standing rule for
every composed comment, applied to apps: never claim an experience the
session did not actually have.

The same rule covers comments placed **directly on a Build app**. This is
deliberately management-agent-only: no server-generated draft and no pure API
persona may author one. After reviewing the project, bind the composed comment
to the exact published artifact you saw:

```bash
lumine admin comment draft build:884 --file comment.md \
  --reviewed-version 4512 --reviewed-via runtime \
  --review-context context.json --json
lumine admin comment post --draft-id 77 --json
```

```json
{
  "understanding": "The start screen pairs a green zombie with a sci-fi soldier, and the first interaction teaches the player to select a zombie before firing the railgun."
}
```

Use `--reviewed-via code` only when you actually pulled and read the project;
`runtime` means you opened and tried the published app. Replies to human
comments inside a Build use `comment:<id>` plus the same review flags. The API
rejects missing review evidence, a server-generated Build draft, a mismatched
published version, a private/noncanonical Build, or any Build/comment-context
change between draft and publication. A changed version means review the new
project state and compose again. The flags are an auditable statement of what
the management agent did, not permission to infer experience from metadata.
The review-context JSON is private server-owned conversation provenance; it is
not included in the public comment, draft response, socket payload, or audit
metadata. A direct human reply to the resulting Zero/Ciel management comment,
or to a later reply by that same bot descended from it, may enter the ordinary
comment-assistant pipeline using this stored historical understanding. It uses
the human commenter's normal AI Energy and reply-or-Like gate. If that user has
no remaining AI Energy, the normal sponsor placeholder/button is shown and a
sponsor pays from their own battery. Mentions elsewhere in a Build, replies to
unlinked or legacy bot comments, replies to humans, and the other bot remain
ineligible. If the published version has changed, the responder is told the
stored understanding belongs to the reviewed older version and must say it has
not checked behavior that could have changed. Editing the acting bot's own
Build comment is supported with the same fresh reviewed-version/method and
private-context flags, including managed review receipts. It updates the
existing comment, not a duplicate reply. See the narrow correction workflow
above; a version-bound follow-up remains appropriate when the conversation
calls for an additional reply instead of an edit.

**Offer a Lumine prompt when the moment invites it (Mikey's direction,
2026-08-10).** Zero and Ciel may include one concrete, copy-pasteable Lumine
prompt in a comment or reply — a genuinely powerful one, tailored to what the
kid is already doing — when the occasion naturally calls for it. Appropriate
occasions: a kid describes an idea they wish existed, asks how something on
the site was made, hits the edge of what a post/drawing/story can do, shares
a Build app that could grow a specific feature, or shows an interest (space,
cats, chess, comics) that maps cleanly onto something Lumine could build with
them. On those occasions, the prompt IS the helpful answer: quote it so it
can be copied as-is, keep it specific to their interest, and mention it works
in the Build workspace chat. Restraint rules: never more than one prompt per
comment; never in condolence, conflict, wellbeing, or moderation-adjacent
threads; never as a reflex closing line on ordinary comments — if the comment
is complete without the prompt, post it without the prompt. A run where only
a few comments carry a prompt is healthy; a run where most do is shameless
plugging, which is exactly what Mikey asked to avoid. SDK-aware prompt ideas
are especially good ("ask Lumine to make a magazine that pulls real Twinkle
posts with Twinkle.subjects.search") because kids do not know the content
APIs exist — but only suggest SDK capabilities that actually exist; check
TWINKLE_BUILD_SDK.md if unsure.

A composed draft (`--file`, plain UTF-8 text, at most the website's 10,000
character comment limit) flows through the identical draft lifecycle —
reservation, idempotency, context-revision CAS, publish fencing, audit
(`metadata.composed: true`) — and is published with the same
`comment post --draft-id`. It never invokes the server's model and records
no AI Energy usage. Deployment guard: an API deployed before this capability
silently ignores `content` and generates with the server's model instead. The
CLI therefore requires the ready draft response to echo the exact submitted
text with `reason: "operator-composed"`; otherwise it stops with
`LUMINE_ADMIN_COMPOSED_COMMENT_UNSUPPORTED`. Never publish that rejected draft.
For Build drafts it additionally requires
`buildReviewContextStored: true`; an older API that ignores the private context
fails closed with `LUMINE_ADMIN_BUILD_REVIEW_CONTEXT_UNSUPPORTED`. Deploy the
context-aware API, then retry with a new idempotency key rather than publishing
the unconfirmed draft.
Placement stays on the requested target: compose replies via `comment:<id>`
targets (the generated path's model-chosen
`replyTargetCommentId` does not apply). Everything below about targets,
containers, and publication applies to both kinds of draft.

A draft targets one of:

- `subject:<id>` (or a bare numeric ID) — a top-level comment on the subject;
- `build:<id>` — a top-level comment on a public canonical Build, composed
  after reviewing its exact published version;
- `aiStory:<id>` / `dailyReflection:<id>` — a top-level comment on the
  standalone post;
- `comment:<id>` — a public reply to that specific comment. `comment reply`
  is the same operation and requires a comment target.

A reply's container resolves canonically from the target comment: its subject,
Build, or AI Story / Daily Reflection root. Comments under any other root are
rejected with `CLI_ADMIN_UNSUPPORTED_REPLY_ROOT`. Build replies authored by the
management agent require the same exact-version review evidence and private
review context as top-level Build comments. Replies to
Zero/Ciel comments and to notification comments are rejected with
`CLI_ADMIN_INVALID_REPLY_TARGET` — the bots never thread with themselves or
each other through the delegated CLI. A human's later direct reply to a
context-backed Build management comment may enter the energy-gated autonomous
pipeline described above; ordinary human replies elsewhere follow the existing
comment-assistant rules. Published replies carry the ordinary thread linkage
(thread root and reply-to), and notification fan-out uses the normal canonical
path.

```ts
type CommentDraft = Success<{
  draft: {
    id: number;
    runId: number;
    targetType: "subject" | "comment" | "build" | "aiStory" | "dailyReflection";
    targetId: number;
    targetUrl: string;
    subjectId: number | null; // container subject; null for standalone posts
    subjectUrl: string | null;
    publicActorUserId: number;
    commentMode: "draft" | "post";
    personaRevision: string; // SHA-256; raw prompt is never returned
    contextRevision: string; // SHA-256 of canonical container/comments/target
    buildReviewContextStored: boolean; // true before a Build draft may publish
    decision: "draft" | "skip";
    reason: string | null;
    content: string | null;
    status: "ready" | "published";
    createdAt: number;
    expiresAt: number;
    publishedCommentId: number | null;
  };
}>;

type CommentPost = CommentGet & {
  status: "success" | "already_done";
  changed: boolean;
  data: CommentGet["data"] & {
    draft: {
      id: number;
      status: "published";
      personaRevision: string;
      contextRevision: string;
    };
    published: {
      commentId: number;
      targetType:
        "subject" | "comment" | "build" | "aiStory" | "dailyReflection";
      targetId: number;
      subjectId: number | null;
      subjectUrl: string | null;
      containerUrl: string;
      commentUrl: string;
    };
  };
};
```

For Subjects and standalone posts, the server loads the canonical container
and its complete visible comment context, then invokes the existing exact
Zero/Ciel system prompt through the shared response assembler with a
mode-specific decision policy (comment vs reply). The raw prompt is never
returned or audited. The model—not regexes or keyword rules—chooses `draft`
or `skip` under the run policy. Build comments never enter that generated
path; they require management-agent-composed content and review evidence.

Draft IDs are bound to operator, run, public bot, target, comment mode,
context revision, persona revision, expiry, and idempotency key. The context
revision covers the container, every visible comment, and the target binding,
so a thread that changes between draft and publish rejects publication.
Posting locks the draft and context in the same transaction as the ordinary
comment insert. Changed context or persona rejects publication and requires
regeneration. Retries return the already-published comment instead of
duplicating it. Secret-subject gating applies whenever the container is a
subject, including replies inside it.

Draft idempotency keys are permanent per operator: supplying a key that an
earlier run (or another target) already used fails rather than resolving to
the old reservation — normally as the audit layer's
`CLI_ADMIN_AUDIT_IDENTITY_MISMATCH` (different run) or
`CLI_ADMIN_IDEMPOTENCY_KEY_MISMATCH` (different target), with
`CLI_ADMIN_DRAFT_KEY_REUSED` as the draft-table backstop. All three mean the
same thing: embed the run or date in any caller-supplied draft key and retry
with a fresh one. Note that the agent's
own `subject effort set` or `creator set-made-by-poster` between draft and
post changes the context revision — order those mutations before drafting.

## Audit, sockets, and deployment

Every delegated mutation reserves a private audit row before execution and
records run ID, real operator ID, public bot ID, session kind, action, target,
idempotency key, before/after state, result, comment mode, and persona revision
where applicable. It never stores authorization headers, bearer tokens,
cookies, passwords, or raw system prompts.

Retry acquisition and completion are row-locked and fenced by a private
per-attempt token, so an expired request cannot overwrite a newer retry. A new
recommendation commits atomically (the management bots are exempt from the
recommendation coin charge); the canonical prior-recommender approval and the
separate 3-Twinkle reward remain independently retryable.

Public content actions use ordinary Twinkle fan-out:

- comments and secret-view notifications use normal comment notifications;
- comments emit `new_upload` after commit;
- recommendations emit `new_recommendation`;
- recommendation-approval and 3-Twinkle rewards emit `new_reward`;
- effort/creator changes emit `edit_content`;
- Featured changes emit a canonical `home_outdated` refresh.

Apply `twinkle-api/scripts/migrations/add-lumine-admin-delegation.sql`, then
`add-lumine-admin-comment-targets.sql`, and apply
`add-lumine-admin-todos.sql` before deploying an API that exposes carry-over
todos. They add
only focused daily-run, rotation, draft, audit, and private-todo tables/columns
and indexes;
there are no runtime schema checks. The comment-targets migration backfills
existing subject drafts into the generalized target columns. The local CLI changes
are not available to users until a separately authorized npm publication.
Deploy and verify the API's `/cli/admin/subjects/featured/rotation` route before
publishing a CLI release that exposes `featured rotate`; an older API rejects
the command without changing Featured state.

The approved-plan and Featured-comment workflows require the new
`/subjects/featured/plan`, `/plan/apply` (under that same Featured base), and
`/subjects/featured/reviews` routes. Deploy the API before releasing the CLI;
there is no fallback to unbound writes or a full daily run on an older API.
Existing runs remain readable/completable with their original scopes, but are
not silently granted `featured:comments`; start a fresh scoped/full run as
appropriate after finishing the existing authorized work. No new migration is
required for these workflows: server receipts reuse the existing audit table
and full-board mutation history.

For scoped runs and Featured history, also apply
`add-lumine-admin-run-scope.sql` and `add-featured-subject-history.sql` before
the API release. After every serving API worker is verified on that release
and no pre-history worker or request remains running or in flight, run
`finalize-featured-subject-history-coverage.sql`; only then publish a CLI
that exposes `featured add`. The guard fails closed until that finalization.
After coverage is finalized, rolling back to an API that does not record
history is forbidden because it would create an unprovable lifetime gap.
After any scoped run exists, an API whose review-window queries do not filter
for `runScope = 'full'` is also rollback-incompatible because it would treat a
narrow slice as a completed full review. A forward fix must preserve both
contracts.

Legacy aliases such as `subjects list`, `subjects get`, `subjects featured`,
`comments get`, and `recommend` remain accepted, but the singular command forms
shown above are the canonical interface.

September 21 routing expansion: read `jevRoutingShadow` alongside `jevPilot` in
both the admin brief and full daily report. The reviewed `chat-routing-v3` path
is primary; report actual selection latency, chat baseline calls avoided, audit
coverage and `jev_chat_routing_audit` cost separately from v1/v2. New routing
families remain comparison-only until Mikey reviews each one and explicitly
promotes it. Report every registered family, including zero-sample and skipped
families; no traffic is not a pass. Show paired counts, exact differing fields,
both values, private input snapshots when available, baseline-controlled outcome,
provider/LLM timing, failure/skip reasons, stale comparisons, and USD spend from
canonical `jev_<family>_shadow` cost operations. Context snapshots are bounded,
private and expire after eight days; do not copy unrelated private content into
public output. Raw model choices precede existing application guards and are not
proof an action was performed. Do not call agreement accuracy or infer speedups
from shadow timings. Full daily costs must include these operations once only.

September 21 Auto exception: Mikey approved JEV as Lumine Auto's primary model
selector immediately, with an independent LLM comparison for every choice.
Auto is the new default; stored manual preferences remain manual. Review
`jevRoutingShadow` / `byRoute.lumine_model` separately from the eight
comparison-only families: selected model/effort, both decisions, exact selection
context, confidence, fallback, actual selection latency, missing evidence and
observed task outcome. Report `jev_lumine_model_serve` and the additional
`jev_lumine_model_audit` spend separately, using canonical AI-cost totals without
double-counting. JEV choice confidence and LLM agreement are not correctness
scores. The other eight families still require review before promotion.


### September 21 verified reward follow-up

After the Study review migration is deployed, the full daily website run also reads the API-host report `node scripts/build-study-reviews-daily.cjs --days 1`. It is indexed, read-only and bounded; it reports decisions, known measured cost, unknown-usage requests, p50/p95 latency, settled checks, claims, pending rows older than 60 seconds and graded-but-unsettled rows. `--examples` includes at most 30 private full study records and their exact prompt version for administrator quality review; keep that output private. No traffic is not a quality pass. The 12 synthetic pre-release records are a small initial check, not measured real-user accuracy.

Study's registered `study-record-v1` checks use learner AI Energy for accepted, revise and uncertain results. Preview checks use real AI Energy but cannot award XP/Coins. The agreed proposal pays 2,000 XP + 500 Coins at most once per learner per UTC day. Breadface's `breadface-v1` uses deterministic server physics, not a model. Both creator proposals must be accepted to publish; inspect the canonical review state and acceptance/rejection events during follow-up.

For total JEV cost, include `dailyAiCosts.byProviderModel` for TypeSafe/JEV across both AI Energy and system-covered ledgers. Report serving/shadow/audit operations separately as above; the Build SDK's system-covered/failed calls use `jev_decision`, while ordinary viewer SDK checks are already counted by the AI Energy ledger. Do not add the Study report's known-cost sum on top of the canonical overall AI-cost total: it is a diagnostic breakdown of those same calls.
