---
name: gpu-ops
description: >
  The GPU workload lifecycle: provision → install → configure → health-check →
  register → run → retrieve artifacts → shutdown, plus keep-alive-vs-destroy
  economics, idle policies, and failure debugging. Use for any end-to-end GPU
  job: batch video generation, model deployment, queue draining.
---

# GPU Operations

The lifecycle is:

```
discover → provision → install → configure → health-check → register → run → retrieve artifacts → shutdown
```

## The state machine

- **Deployments are the handle.** After provisioning, always `gpu_register` the
  instance once it is reachable — from then on, every job goes through
  `gpu_run deployment_id=...` and you never SSH for the happy path again.
- **Machines cost money while alive.** A deployment in status `ready` with no
  running job is burning `cost_hr`. The policy idle timeout
  (`policy.idle_shutdown_after_min`, default 30) is the line: past it, propose
  `gpu_destroy` unless the queue is about to use it again.

## Keep-alive vs destroy — the actual math

For a queue of N jobs:

```
total_runtime = N × avg_job_time
startup_cost  = (startup_minutes / 60) × cost_hr        # pull image, load weights
keep_cost     = idle_gap_hours × cost_hr                # time between jobs
restart_cost  = (N − 1) × startup_cost                  # if you destroy between jobs
```

Keep alive when `idle_gap_hours < (N − 1) × startup_minutes / 60`, i.e. when
restarting costs more than idling. For ComfyUI with a big video model on a
$0.70/hr 4090: startup is ~10–15 min (~$0.15), so keep alive unless the gap
between jobs is over ~10 minutes. On an H100 at $3/hr the threshold is ~4
minutes — destroy aggressively, and prefer spot/interruptible for queues.

## Retrieving artifacts — before you destroy

Outputs on the pod die with the pod. For every completed job:

1. Sync outputs off the instance (rsync/scp, or the provider's volume/object
   storage) to a durable URI (R2/S3/local path).
2. Record the URIs in `gpu_job_finish artifacts=[{name,uri}]`.
3. Only then consider `gpu_destroy`.

## Failure debugging protocol

When a job fails, follow this order instead of guessing:

1. `gpu_job_finish status=failed error=...` so the ledger stays honest.
2. Read the logs: ComfyUI → `journalctl`/`/logs` in the container, or the
   workflow's `/history` entry; native server → stdout of the process.
3. Classify:
   - **CUDA out of memory** → lower resolution/`--lowvram`, switch to fp8
     weights, or move to a bigger GPU class. Do NOT just retry.
   - **CUDA driver/runtime mismatch** → reinstall the torch wheel matching the
     image's CUDA (see comfyui skill), don't rebuild the world.
   - **Custom node fails to import** → pin the node repo to a known-good commit
     (the version that worked last time), check Python package conflicts.
   - **Weights fail to load / hash mismatch** → re-download with checksum
     verification; check disk space on the instance.
4. Retry once after the fix. If it fails again, downgrade the deployment status
   (`gpu_register status=error` or destroy) and report to the user — don't loop
   silently, every retry costs money.

## Spend hygiene

- Check the ledger (`gpu_spend`) before proposing any new batch.
- Never let a retry loop exceed the per-job ceiling — after one failed retry,
  come back to the user.
- Prefer `gpu_ensure` over fresh provisioning: reusing a live deployment is
  always cheaper than a new one.
