# H3 Economics — the planning model

**Compiled 2026-08-12 from MiniMax's H3 announcement, community benchmarks
(Reddit r/StableDiffusion, r/comfyui), and live provider pricing.** This is the
planning baseline for the first workload: one finished minute = **4 × 15-second
H3 generations stitched together** — H3's native maximum is 15 seconds at 2K
with stereo audio, so "a minute of H3" is never one diffusion pass.

## The single most important finding

**The GPU price is secondary to the workflow.** A 5090 at $0.69/hr running the
optimized stack is radically cheaper per finished minute than a cheaper GPU
running the vanilla 20-step workflow. Anyone quoting "H3 takes 10 minutes for
15 seconds" is almost certainly describing launch-week default settings.

### The optimization stack

| Layer | What it does | Status |
| --- | --- | --- |
| **Base H3, 20 steps** | The launch-week default — where the scary 5–15 min/clip benchmarks came from | Baseline, don't run this |
| **INT8 / FP8 / NVFP4** | Quantization — cuts memory, can improve throughput | Shipped |
| **SageAttention / ck-attention** | Faster attention kernels | Shipped |
| **Spectrum / EasyCache / FirstBlockCache** | Skips or reuses computation across steps | Shipped |
| **Turbo LoRA** | The big one: ~20 steps → **4–8 steps** | Shipped; community standard is 6–8 |
| **LightX2V 4-step** | Newer 4-step workflow | In testing |

Production configuration today: **Turbo LoRA (4–8 steps) + SageAttention2 +
INT8 models (+ a cache).**

### Benchmark anchors (measured, not modeled)

| Setup | Clip | Time |
| --- | --- | --- |
| RTX 6000, 480p, 15s, SageAttention2 | 4-step Turbo | **56 s** |
| RTX 6000, 480p, 15s, SageAttention2 | 8-step Turbo | **1 m 47 s** |
| RTX 5090, 0.4 MP, 15s | 4-step Turbo + INT8 | **~3 min** (was ~12 min on 20-step) |
| RTX 5090, 1 MP, 5s | 4-step Turbo + Sage | **~58 s** |
| 5070 Ti, 5s, 0.4MP, 20-step (default) | sampling only | 121 s — plus ~240 s model load |

Rule: **never trust an H3 benchmark without** checkpoint, quantization,
resolution, duration, step count, sampler, attention implementation, cache,
**CUDA/torch build (cu128 vs cu130)**, and whether model loading is included.
A 10x spread across "H3 benchmarks" is normal and meaningless. The cu128 vs
cu130 split alone is ~2x on the INT8 convrot path — a number without the CUDA
build attached is uninterpretable.

## Provider prices (planning baseline)

| Provider | GPU | $/hr | Notes |
| --- | ---: | --- | --- |
| RunPod | RTX PRO 4500 32GB | 1.15 | No public H3 benchmark yet — do not assume 5090-like speed (10.5k CUDA cores / 896 GB/s vs 5090's ~21.8k / 1.79 TB/s) |
| RunPod | RTX 5090 32GB | 1.58 | Polished API/CLI |
| RunPod | RTX 4090 24GB | 1.10 | Needs quantized/low-res workflow |
| RunPod | RTX PRO 6000 96GB | 1.69 | The measured 56 s/4-step anchor card |
| Vast.ai | RTX 5090 | ~0.69 | **Best balance of price/perf/ops** |
| Vast.ai | RTX 4090 | ~0.34 | |
| Vast.ai | RTX 3090 | ~0.22 | Slow but extremely cheap |

Snapshot date differs per source; **always verify live before committing spend**
(lvrged-factory-gpu-ops rule). RunPod and Vast are the first-class adapters —
other providers are added per the lvrged-factory-provider-adapters recipe.

## Cost per finished minute (4 × 15s)

| Provider / GPU | Config | Time/15s | Time/minute | GPU cost/minute |
| --- | --- | ---: | ---: | ---: |
| RunPod PRO 6000 $1.69 | 4-step Turbo + Sage | 56 s measured | ~3.7 min | ~$0.105 |
| Vast 5090 $0.69 | 4-step Turbo, 0.4 MP | ~3 min | ~12 min | ~$0.14 |
| Vast 4090 $0.34 | optimized/quantized | ~4 min* | ~16 min | ~$0.09 |
| Vast 3090 $0.22 | optimized/quantized | ~9.7 min* | ~39 min | ~$0.14 |
| RunPod PRO 4500 $1.15 | 4–8 step Turbo, INT8, Sage | 4–6 min* | 16–24 min | ~$0.19–0.29 |

\* modeled from the 5090/4090 performance relationship, not measured. The
PRO 4500 row is the reason to prefer the 5090: same price class, roughly double
the hardware, and the 4500 has no public H3 inference benchmark at all.

## The Higgsfield comparison

| $15 budget | Finished 1-min equivalents |
| --- | ---: |
| Higgsfield (~$10–15/min) | 1–1.5 |
| H3 on Vast 5090 | ~100–110 |
| H3 on RunPod PRO 6000 | ~140 |

The gap is real but it's not apples-to-apples: Higgsfield is a finished
creative service (sequencing, continuity, audio, editing, multi-model access:
Seedance 2.0, Kling 3.0, Veo 3.1), not raw diffusion compute. Its Explainer
product: 20s narrated video ≈ 72 credits ≈ $3.60 → ~$10/minute. Market
reference points: official H3 API ~$0.13/s; a cheap hosted H3 ~$0.32 per 15s
768p clip (~$1.30/min).

**Planning number: treat ~50–80 finished one-minute equivalents per $15 on a
$0.72–1.15/hr 32GB card as conservative, ~100–110 on a Vast 5090, and RunPod
PRO 6000 (~140) as the predictable upper bound.**

## The lane (v2: RunPod-only first-class)

**RunPod RTX PRO 6000, secure cloud, cu130 image.** The 5090 rows above are
kept for reference but the community verdict is blunt: H3's INT8 peaks
~31.7GB VRAM — right at a 5090's ceiling — and the unquantized pipeline is
~110GB, so the 96GB card keeps everything resident while 32GB cards swap.
Expect to pay **secure $2.09/hr**: community $1.69 exists on paper but the
session-verified reality is create-rejections across every DC. Budget rule:
keep a $5–10 balance buffer — hitting $0 kills pods and deletes disks.

Other providers plug in via the adapter recipe
(lvrged-factory-provider-adapters) — they are deliberately not first-class.

## Benchmark protocol v1 (the standardized test)

One fixed test, run on every GPU/provider/workflow combo, one row per combo.
**Never change the fixed settings between runs.**

### The fixed test

- **Model:** MiniMax H3, Turbo LoRA v4 step600 EMA
  (`minimax_h3_turbo_v4_step600_ema.safetensors`), pruned INT8 convrot base.
- **Settings:** 8 steps, euler sampler (MiniMaxH3TurboSampler), **beta**
  scheduler, LoRA strength 1.0.
- **Attention:** SageAttention2 ON — *verified in the boot log*, not "the
  module imports". Spectrum/caches OFF.
- **Clip:** 15 seconds (362 frames), 480p (864×480, 0.4MP).
- **Prompt (fixed):** "A barista pours latte art in a sunlit cafe, steam
  rising, camera slowly pushes in, ambient cafe sounds" · **Seed: 42.**

### Metrics per run (from log timestamps)

| Metric | How |
| --- | --- |
| T1 provision | provider create call → SSH/pod ready |
| T2 install+load | pod ready → workflow loaded, models in VRAM |
| T3 sampling | generation start → last frame done |
| T4 total wall | T1+T2+T3 (first run) / T3 only (warm) |
| $/finished min | (warm T3 hours × billed rate) × 4 |
| Quality | blind A/B vs the 20-step base, same seed, 1–5 rubric |
| Reliability | jobs completed / attempted (run the clip 5×) |

Rubric: 5 = indistinguishable from 20-step base · 4 = minor softness ·
3 = artifacts but publishable for social · 2 = plastic skin / mushy motion /
audio glitches · 1 = unusable.
Score = (Quality/5) × (1 / $-per-finished-min) × Reliability.

### Rules

1. Run each config **5 times; report the median**, not the best.
2. Cold start counts once; warm runs are what you budget from.
3. A provider preempt/kill is a **reliability failure**, not a retry.
4. Log the exact checkpoint filename, ComfyUI version, **and torch/CUDA
   build** in Notes — numbers are meaningless without them.
5. Re-run the whole table when the Turbo LoRA version changes.

### Results table

| Provider | GPU | CUDA | $/hr billed | Warm 15s clip | $/fin min | Cold total | Quality | Reliab. | Notes |
| --- | --- | --- | ---: | ---: | ---: | ---: | --- | --- | --- |
| RunPod | PRO 6000 (secure) | cu128 | 2.09 | — | — | 215.1s¹ | — | 1/1 | ¹8-step 864×480/15s, **stock attention** (sage stub), includes first VRAM load — the reference that forced the cu130 redeploy |
| RunPod | PRO 6000 (secure) | cu130 | 2.09 | *pending* | | | | | runpod/comfyui:cuda13.0, Sage built from source |

Community reference points (2026-08-12, verify before trusting): RTX 6000
480p/15s Sage2 — 56s @4-step / 107s @8-step (reddit); RTX 5090 4-step
~19–48s (resolution unstated); RTX 4070 12GB 4-step 88s; RTX 3060 12GB
8-step 5s@864×480 4.5min (larryvrh HF discussion). The spread is exactly why
this protocol exists.

## Sources

- MiniMax H3 announcement: https://minimaxi.com/blog/minimax-h3
- RTX 6000 4/8-step benchmark: reddit.com/r/StableDiffusion/comments/1vlo9se
- 5090 4-step Turbo + INT8 reports: reddit.com/r/comfyui/comments/1vm32qs
- Turbo LoRA: reddit.com/r/StableDiffusion/comments/1vgxf4x
- 5070 Ti default-workflow timings: reddit.com/r/StableDiffusion/comments/1vg1qve
- Model-load overhead discussion: reddit.com/r/StableDiffusion/comments/1vhbik7
- H3 API vs RunPod/Vast: reddit.com/r/StableDiffusion/comments/1vfmrvh
- RunPod pricing: runpod.io/pricing · Vast: vast.ai
- NVIDIA RTX PRO 4500 specs: nvidia.com/en-gb/products/workstations/professional-desktop-gpus/rtx-pro-4500
