---
name: gpu-ops
description: >
  The GPU workload lifecycle for Claude Code: provision → install → configure →
  health-check → register → run → retrieve artifacts → shutdown, plus
  keep-alive-vs-destroy economics, idle policies, and failure debugging.
  Use for any end-to-end GPU job. The registry is plain JSON under .pi/gpu/ —
  you maintain it with file writes, there are no helper tools.
---

# GPU Operations (Claude Code)

Lifecycle:

```
discover → provision → install → configure → health-check → register → run → retrieve artifacts → shutdown
```

## The state machine (plain files, you are the tool)

- `.pi/gpu/providers.json` — installed/authenticated providers
- `.pi/gpu/machines.json` — raw instances
- `.pi/gpu/deployments.json` — the stable handle: `{id, provider, instance_id, gpu, runtime, model, endpoint, ssh, status, cost_hr}`
- `.pi/gpu/models.json` + `.pi/gpu/workflows.json` — manifests
- `.pi/gpu/jobs.jsonl` — append-only ledger, one JSON object per line
- `.pi/gpu/policy.json` — spend ceilings; read it before every money action

**Discipline rules (same as the pi version):**

1. A deployment is only `"status": "ready"` after a test generation succeeded.
2. Every job gets a line in jobs.jsonl — append, never rewrite history:
   `{"id":"J-000123","deployment_id":"dep_h3_01","provider":"vast","model":"minimax-h3","workflow":"h3-text-to-video","workflow_version":1,"inputs":{...},"status":"completed","started_at":"...","completed_at":"...","duration_s":192.4,"cost_usd":0.014,"gpu":"RTX 5090","artifacts":[{"name":"out.mp4","uri":"s3://..."}],"ts":"..."}`
3. Artifacts are synced off the instance **before** destroy — outputs on the
   pod die with the pod.
4. Destroyed deployments stay in deployments.json with status `destroyed` —
   history is never deleted.

## Keep-alive vs destroy — the math

```
total_runtime = N × avg_job_time
startup_cost  = (startup_minutes / 60) × cost_hr
restart_cost  = (N − 1) × startup_cost
```

Keep alive when the idle gap between jobs costs less than restarting. ComfyUI
+ big video model on a $0.70/hr 4090: ~10–15 min startup (~$0.15) → keep alive
unless the gap exceeds ~10 min. On an H100 at $3/hr: destroy aggressively.
Prefer spot/interruptible for retryable queues.

## Failure debugging protocol

1. Record the failure in jobs.jsonl (status failed + error) — honesty first.
2. Read logs (ComfyUI: container stdout / workflow `/history` entry; native:
   process stdout).
3. Classify:
   - **CUDA OOM** → lower res / `--lowvram` / fp8 weights / bigger GPU class. Never blind-retry.
   - **CUDA driver/runtime mismatch** → reinstall the torch wheel matching the image's CUDA, don't rebuild the world.
   - **Custom node import fails** → pin the node repo to the last-known-good commit.
   - **Weights fail / hash mismatch** → re-download with checksum; check disk.
4. Retry once. Fail again → mark the deployment error in deployments.json and
   report — every retry costs money.

## Spend hygiene

- Read the ledger (`jq -s 'map(select(.status=="completed")) | map(.cost_usd // 0) | add' .pi/gpu/jobs.jsonl`) before proposing a batch.
- Respect policy.json ceilings as hard stops — never loop a retry past the
  per-job ceiling without checking in with the user.
- Reuse a live deployment before provisioning a new one.
