# Provider migration snapshots and rollback

Open Claudia takes a verified, byte-for-byte snapshot before any provider-aware state, session, or job schema is activated. At startup it then migrates and verifies all three stores before loading the vault, scheduler, router, or channel adapters. Direct and agent modes use the same modular `bot.js` runtime, so there is only one active state writer during this process.

Snapshots live under the configured Open Claudia data directory:

```text
migration-backups/provider-parity-v1.1/snapshots/<timestamp-id>/
  originals/
  manifest.json
  manifest.sha256
```

The allowlist contains `state.json`, `sessions.json`, `jobs.json`, legacy `crons.json`, an earlier `crons.json.migrated`, and relevant `.bak` siblings. It deliberately excludes `.env`, vault, authentication, identity, transcript, and derived SQLite files. The manifest records relative names, byte sizes, SHA-256 hashes, and advisory schema guesses; it never records source contents or absolute installation paths. Malformed and provider-ambiguous records are preserved as `unresolved` rather than assigned to Claude, Codex, or a project by guesswork. Removed-provider selections, settings, session pointers/history, and jobs are retained only in non-selectable or disabled archives; they are never reassigned to Claude or Codex and never fired.

Snapshot directories are private operator data. They can contain conversation titles, prompts, native session identifiers, and scheduled-job text. Keep them mode `0700` with files mode `0600`; do not share, publish, or commit them.

## Rollback procedure

Rollback is an offline operation. Do not restore files while any Open Claudia process can write them.

1. Stop every Open Claudia bot and scheduler process. Confirm no old or new runtime remains active.
2. Identify the snapshot referenced by `migration-backups/provider-parity-v1.1/journal.json`. Do not select a partial orphan directory merely because its timestamp is newer.
3. Validate the complete snapshot with `validateMigrationSnapshot(snapshotDir)`. A manifest mismatch, missing file, corrupt copy, symlink, or checksum failure blocks rollback before any live target changes.
4. Call `rollbackMigrationSnapshot` from the installed Open Claudia package with the configured state/session/job/cron paths, as in the example below. The helper stages exact bytes, restores the complete allowlisted set, restores original absence where applicable, and verifies every target against the manifest.
5. Validate the snapshot again and verify the restored live files with the same source list. Keep the snapshot; rollback never deletes it.
6. If the restored files contain Cursor Agent state, restart the matching older Cursor-capable Open Claudia version. The current runtime intentionally cannot execute that provider. Otherwise restart the upgraded version only after its migration journal and schema expectations have been reviewed.

Rollback restores the snapshot's exact pre-upgrade bytes. Post-upgrade changes are not merged into those older files and will be discarded by the restore. If those changes must be audited, make a separate private copy of the current files while every writer is stopped; never combine records from the two schema generations by hand.

Example from a Node maintenance script run inside the installed package:

```js
const migration = require("./core/migration-backup");
const config = require("./core/config");
const configDir = config.CONFIG_DIR;
const snapshotDir = process.env.OPEN_CLAUDIA_MIGRATION_SNAPSHOT;
const sources = migration.defaultMigrationSources({
  configDir,
  stateFile: config.STATE_FILE,
  sessionsFile: config.SESSIONS_FILE,
  jobsFile: config.JOBS_FILE,
  cronsFile: config.CRONS_FILE,
});

migration.validateMigrationSnapshot(snapshotDir);
migration.rollbackMigrationSnapshot(snapshotDir, { configDir, sources });
migration.validateMigrationSnapshot(snapshotDir, { sources, verifySources: true });
```

Never mark a migration globally activated after migrating only one store. The journal tracks `state`, `sessions`, and `jobs` separately and reaches `activated` only after all three component migrations have completed and verified their writes. Each component certificate records the live file's logical path, byte size, schema version, and SHA-256; readiness recomputes those values, so a legacy, corrupt, or subsequently replaced file cannot satisfy an old activation flag.

Startup beyond the migration preflight is gated on that globally activated journal. A missing component, corrupt manifest, stale live-file proof, or partial migration blocks normal runtime initialization. After the gate, an old Cursor selection becomes an explicit “provider selection required” state; the user must choose Claude Code or OpenAI Codex with `/backend`. Old callbacks show a removal notice and do not mutate the provider, model, or session.

Each component write is two-phase. Open Claudia records a `pendingComponents` entry against the verified pre-migration snapshot before the atomic live-file rename, validates the new schema after the rename, and only then marks that component activated. If the process stops between those steps, startup keeps the original snapshot pinned; the component migration must validate and finish the pending write instead of creating a mixed-schema replacement snapshot.

Provider-blind conversation-history records are retained as non-selectable `legacy` entries unless exactly one provider adapter can prove ownership with a read-only probe. New history is keyed by canonical user, project, and provider. Chat buttons contain only short-lived opaque server-side tokens; provider-native session IDs are never embedded in callback payloads.

Provider and project switches restore only that tuple's active conversation, while model, effort, budget, permission mode, and worktree choices remain scoped to the provider. `/new` clears the selected project's active pointer for the current provider only. `/end` closes the project selection and settles transient queued work without deleting provider pointers or conversation history.

Compaction captures one immutable canonical-user/project/provider/session tuple. Summary and seed runs cannot persist ordinary transcripts, usage, or session pointers, and queued user turns remain paused until both stages settle. Claude seeds use its native one-turn limit in read-only plan mode. Codex seeds use a read-only sandbox plus a bounded JSON acknowledgement contract; they deliberately use a persistent compaction invocation instead of Codex's ephemeral utility flag because the resulting seed must remain resumable. No unsupported Codex `maxTurns` or complete tool-suppression claim is made.

The pointer, provider-tagged history lineage, and tuple-scoped usage checkpoint are committed through a recoverable prepared/committed marker. Restart rolls a prepared operation back to the old pointer and history, while a committed marker is finalized without replaying history. Legacy usage-log records expose their existing `backend` as `provider` when read, without rewriting the original JSONL file.

Scheduled work uses a versioned `{schemaVersion, jobs, archived}` document. Every runnable occurrence stores its canonical user, channel, project, provider, provider-settings snapshot, native session lineage, origin run, and stable occurrence ID. A fresh foreground turn can create a job before its native session is known; terminal session persistence binds that job by origin run ID. Compaction descendants are followed only when recorded lineage proves the relationship, and an older scheduled branch cannot replace an unrelated live session pointer.

Wakeups and crons wait for terminal provider completion, transcript/usage/session persistence, and required delivery. A one-shot receives a durable success marker before deletion. Busy occurrences retain their occurrence ID across restart and wait until the chat is idle, without consuming execution attempts or expiring under the missed-fire grace window. Genuine execution failures still have bounded retries; exhaustion and wakeups never admitted within the missed-fire grace window remain visible as disabled dead letters.

A watch persists its trigger (or expiry) before delivery and does not poll its check again while that occurrence is pending. Repeat watches resume polling only after successful delivery advances the occurrence. An in-process guard prevents overlapping callbacks from starting the same occurrence, even when a run lasts longer than its retry delay. Advisory notifications are marked before sending and attempted at most once per occurrence across retries/restarts; a crash between marking and sending can omit this optional notice. An interrupted provider run may still retry with the same occurrence ID if no durable success marker exists; this is not an exactly-once guarantee for external side effects. Previously dead-lettered jobs are not automatically revived because their queued actions may now be stale.

Provider fallback is empty by default. When `PROVIDER_FALLBACKS` is explicitly configured, a fallback starts a fresh provider session with an archived provider-neutral brief and never receives another provider's native session ID. Legacy `crons.json` remains byte-for-byte untouched after its ambiguous records are copied into the disabled archive.
