# Architecture

lvrged-factory is deliberately small in code and large in knowledge. The
extension owns the state machine; the skills own the how-to; the provider
tables own the interfaces. Nothing else.

## Design rules

1. **The core abstraction is the deployment, not the provider.** A deployment
   is `{id, provider, gpu, runtime, model, endpoint, ssh, status, cost_hr}` —
   the stable handle everything else hangs off. Jobs reference deployments;
   spend rolls up from jobs; the agent never SSHes around for the happy path.
2. **Models are deployable artifacts.** `models.json` holds manifests
   (source, VRAM, GPU class, runtime) with an explicit `verified` flag. The
   extension doesn't know how to deploy H3 — the agent discovers that path via
   the lvrged-factory-model-deployment skill and records the receipt.
3. **Workflows are first-class.** A workflow is executable infrastructure —
   a ComfyUI graph, an HTTP contract, or a provider API call. `lvrged_factory_job action=run`
   targets `workflow_id`, not "the model".
4. **The ledger is the source of truth.** `jobs.jsonl` is append-only. Every
   generation records model, workflow + version, deployment, GPU, duration,
   cost, artifacts, error. Cost answers ("last 100 videos?") are queries, not
   recollections.
5. **The extension is dumb but strict.** Tools validate state, enforce the
   spend policy, and execute provider command templates — they do not reason
   about CUDA errors or pick runtimes. The agent does that, guided by skills.
   Deterministic mechanics live in code (the provision tool's capacity
   fallback is a fixed rotation with error-string classification, not
   reasoning); judgment stays with the agent.
6. **Providers are capability tables.** A provider is `{cli, authCheck,
   provision, stop, start, destroy, list, notes}`. RunPod is the one
   first-class adapter (runpodctl is an agent-friendly CLI). Any
   other provider is added by the user as a table entry + a setup-skill
   recipe + a cheat sheet — not a release. The extension stays dumb; the
   agent learns the provider from the skills.
7. **The golden path is scripted.** `scripts/h3-pod-up.sh` → `h3-install.sh`
   → `h3-run.py` is the zero-decision H3 lane, shared by the pi tools and the
   Claude Code (bash-only) skill set. Skills reference the scripts instead of
   re-deriving the commands.

## Layer map

```
extension/
  index.ts       entry: registers tools + commands, startup nudges
  registry.ts    .pi/lvrged-factory state: providers, machines, deployments,
                 models, workflows, jobs ledger, policy + rollups
  providers.ts   capability tables + execFile runner + id parsing
  tools.ts       lvrged_factory_* tool implementations (state machine + gates)
  commands.ts    /lvrged-factory dashboard commands

skills/          the knowledge (loaded on demand by the agent)
  lvrged-factory-setup      detect/install/auth provider CLIs (runpodctl, vastai)
  lvrged-factory-gpu-onboarding first-run flow: H3 economics card, provider lanes, first deploy
  lvrged-factory-gpu-ops        lifecycle, keep-alive economics, failure protocol
  comfyui        runtime install, weights, HTTP API, health checks
  lvrged-factory-model-deployment  the "make model X runnable" protocol
  lvrged-factory-video-models   H3 / Wan / Hunyuan / Seedance manifests + H3 Turbo stack
  lvrged-factory-provider-adapters  RunPod cheat sheet + add-a-provider recipe
  lvrged-factory-runpod-api     RunPod REST API surface (logs endpoint, API create, CLI-vs-API rule)

claude-skills/   the same skills adapted for Claude Code (bash-only, no
                 lvrged_factory_* tools), sharing the same .pi/lvrged-factory JSON registry

docs/
  h3-economics/  H3 production economics: Turbo stack, benchmarks, lanes
  pricing/       provider price snapshots (2026-08-12) + live-verify notes
  providers/     deep per-provider playbooks
```

## The lifecycle

```
discover → provision → install → configure → health-check → register → run → retrieve artifacts → shutdown
   lvrged_factory_setup   lvrged_factory_pod action=provision   (skills)   (skills)    test gen    lvrged_factory_pod action=register   lvrged_factory_job action=run   sync off   lvrged_factory_pod action=destroy
```

Every stage has a registry artifact: `providers.json` → `machines.json` →
`deployments.json` (READY only after health-check) → `jobs.jsonl` (completed
only with artifact URIs).

## Economics decision (agent-side)

For a queue of N jobs the agent compares `keep-alive` (idle hours × cost_hr)
vs `restart` ((N−1) × startup cost). Policy `idle_shutdown_after_min` is the
default kill switch; the agent may override only by proposing it to the user.
The pricing snapshots exist to make this decision fast — the agent verifies
live rates before committing.

## Trust boundary

The spend policy is enforced in code (`spendGate` in tools.ts): anything above
`confirm_above_usd` requires a user confirmation, and per-job/daily/monthly
ceilings block silent spend. This is the one place the extension is strict by
design — GPU CLIs move real money with no undo.
