---
name: gpu-onboarding
description: >
  First-run onboarding for video-factory: when a user installs the package and
  has no registry yet, walk them through the H3 production economics, pick a
  provider and GPU for their actual workload, and make the first deployment.
  Use when the registry is empty, when the user asks what to deploy / which
  provider to use, or when /gpu onboard is invoked.
---

# GPU Onboarding

Goal: a new user goes from "I installed this" to "my first video is queued"
without reading a manual, and makes an informed provider choice instead of
defaulting to the first name they recognize.

## When to run

The extension nudges on first session start (no registry). Trigger the flow
whenever the user hasn't chosen a lane yet. **Do not** run the full flow for an
existing user with deployments — they already have opinions.

## Step 1 — The one-paragraph context

Say, in plain words:

> H3 generates up to 15 seconds at 2K with audio in one pass, so a "minute" of
> video is four clips stitched together. The big cost lever is the workflow:
> Turbo LoRA 4–8 steps + SageAttention + INT8 runs 4–10x faster than the
> default 20-step setup, which is what most public "slow H3" numbers come from.

## Step 2 — The decision card

Present (docs/h3-economics.md has the detail):

```
H3 PRODUCTION — planning numbers (2026-08-12, verify live)
  1 finished minute = 4 × 15s clips
  anchors: RTX 6000 480p/15s: 56s @ 4-step Turbo · 5090 15s/0.4MP: ~3 min
  per finished minute: Salad 5090 ~$0.06 · Vast 5090 ~$0.14 · RunPod 5090 ~$0.32
  $15 buys: Higgsfield 1 min · Vast 5090 ~100 min · Salad 5090 ~250 min
```

## Step 3 — Four questions (then stop talking)

1. **Batch size** — how many finished minutes do you need per batch? (Drives
   keep-alive vs destroy math; 30+ minutes → keep the instance alive.)
2. **Quality floor** — 0.4 MP quick pass, or full 2K? (Turbo 4-step is the
   fast lane; 8-step for quality; 2K needs the big cards.)
3. **Budget ceiling** — set the policy with gpu_policy, don't just quote it.
4. **Provider preference** — explain the three lanes, don't ask them to pick
   blindly:
   - **Vast 5090 (~$0.69/hr)** — best balance; the default recommendation.
   - **Salad 5090 (~$0.29/hr batch)** — ~2x cheaper, distributed fleet,
     availability caveats; the aggressive lane for retryable batches.
   - **RunPod PRO 6000 ($1.69/hr)** — the measured 56s/4-step anchor card;
     most predictable operations; choose for polish over price.

## Step 4 — First provision

- `gpu_ensure` with the chosen model (default: `minimax-h3`, but if the batch
  is large and cheap is the goal, offer `wan-2.2-5b` on a 24GB card as the
  draft lane and H3 for the final pass — that split is often the cheapest
  production setup).
- Follow the plan: provider search → provision → comfyui install → weights →
  **benchmark the exact workflow** (docs/h3-economics.md anchors) →
  health-check with a test generation → gpu_register → gpu_run.
- Record the measured per-clip time + cost in the job ledger and in the model
  manifest notes (`gpu_model_add notes=...`) — the next user/agent gets real
  numbers instead of this research.

## Step 5 — The one-line close

State the standing plan: *"Everything is recorded in .pi/gpu — tomorrow the
agent will know what's deployed, what it costs per minute, and what's queued."*

## Onboarding variants

- **No provider CLIs at all:** run the gpu-setup flow first (install CLIs,
  auth), then come back to this.
- **User already has a provider:** skip Step 3's lane explanation, ask only
  budget + quality, then deploy on what they have.
- **User only wants the API path** (no infra): point at the hosted options in
  docs/h3-economics.md (Higgsfield / H3 API / cheap hosted) — the package is
  still useful for the job ledger and cost comparison.
