# H3 Economics — the planning model

**Compiled 2026-08-12 from MiniMax's H3 announcement, community benchmarks
(Reddit r/StableDiffusion, r/comfyui), and live provider pricing.** This is the
planning baseline for the first workload: one finished minute = **4 × 15-second
H3 generations stitched together** — H3's native maximum is 15 seconds at 2K
with stereo audio, so "a minute of H3" is never one diffusion pass.

## The single most important finding

**The GPU price is secondary to the workflow.** A 5090 at $0.69/hr running the
optimized stack is radically cheaper per finished minute than a cheaper GPU
running the vanilla 20-step workflow. Anyone quoting "H3 takes 10 minutes for
15 seconds" is almost certainly describing launch-week default settings.

### The optimization stack

| Layer | What it does | Status |
| --- | --- | --- |
| **Base H3, 20 steps** | The launch-week default — where the scary 5–15 min/clip benchmarks came from | Baseline, don't run this |
| **INT8 / FP8 / NVFP4** | Quantization — cuts memory, can improve throughput | Shipped |
| **SageAttention / ck-attention** | Faster attention kernels | Shipped |
| **Spectrum / EasyCache / FirstBlockCache** | Skips or reuses computation across steps | Shipped |
| **Turbo LoRA** | The big one: ~20 steps → **4–8 steps** | Shipped; community standard is 6–8 |
| **LightX2V 4-step** | Newer 4-step workflow | In testing |

Production configuration today: **Turbo LoRA (4–8 steps) + SageAttention2 +
INT8 models (+ a cache).**

### Benchmark anchors (measured, not modeled)

| Setup | Clip | Time |
| --- | --- | --- |
| RTX 6000, 480p, 15s, SageAttention2 | 4-step Turbo | **56 s** |
| RTX 6000, 480p, 15s, SageAttention2 | 8-step Turbo | **1 m 47 s** |
| RTX 5090, 0.4 MP, 15s | 4-step Turbo + INT8 | **~3 min** (was ~12 min on 20-step) |
| RTX 5090, 1 MP, 5s | 4-step Turbo + Sage | **~58 s** |
| 5070 Ti, 5s, 0.4MP, 20-step (default) | sampling only | 121 s — plus ~240 s model load |

Rule: **never trust an H3 benchmark without** checkpoint, quantization,
resolution, duration, step count, sampler, attention implementation, cache,
and whether model loading is included. A 10x spread across "H3 benchmarks" is
normal and meaningless.

## Provider prices (planning baseline)

| Provider | GPU | $/hr | Notes |
| --- | ---: | --- | --- |
| RunPod | RTX PRO 4500 32GB | 1.15 | No public H3 benchmark yet — do not assume 5090-like speed (10.5k CUDA cores / 896 GB/s vs 5090's ~21.8k / 1.79 TB/s) |
| RunPod | RTX 5090 32GB | 1.58 | Polished API/CLI |
| RunPod | RTX 4090 24GB | 1.10 | Needs quantized/low-res workflow |
| RunPod | RTX PRO 6000 96GB | 1.69 | The measured 56 s/4-step anchor card |
| Vast.ai | RTX 5090 | ~0.69 | **Best balance of price/perf/ops** |
| Vast.ai | RTX 4090 | ~0.34 | |
| Vast.ai | RTX 3090 | ~0.22 | Slow but extremely cheap |
| SaladCloud | RTX 5090 | ~0.294 batch | **Economic outlier** — distributed consumer GPUs |
| SaladCloud | RTX 4090 | ~0.204 batch | |
| SaladCloud | RTX 3090 | ~0.124 batch | |
| Modal | RTX PRO 6000 | ~3.03 | Serverless, zero idle |

Snapshot date differs per source; **always verify live before committing spend**
(gpu-ops rule).

## Cost per finished minute (4 × 15s)

| Provider / GPU | Config | Time/15s | Time/minute | GPU cost/minute |
| --- | --- | ---: | ---: | ---: |
| RunPod PRO 6000 $1.69 | 4-step Turbo + Sage | 56 s measured | ~3.7 min | ~$0.105 |
| Vast 5090 $0.69 | 4-step Turbo, 0.4 MP | ~3 min | ~12 min | ~$0.14 |
| Salad 5090 $0.294 | same | ~3 min | ~12 min | **~$0.059** |
| Vast 4090 $0.34 | optimized/quantized | ~4 min* | ~16 min | ~$0.09 |
| Salad 4090 $0.204 | same | ~4 min* | ~16 min | ~$0.056 |
| Vast 3090 $0.22 | optimized/quantized | ~9.7 min* | ~39 min | ~$0.14 |
| RunPod PRO 4500 $1.15 | 4–8 step Turbo, INT8, Sage | 4–6 min* | 16–24 min | ~$0.19–0.29 |

\* modeled from the 5090/4090 performance relationship, not measured. The
PRO 4500 row is the reason to prefer the 5090: same price class, roughly double
the hardware, and the 4500 has no public H3 inference benchmark at all.

## The Higgsfield comparison

| $15 budget | Finished 1-min equivalents |
| --- | ---: |
| Higgsfield (~$10–15/min) | 1–1.5 |
| H3 on Vast 5090 | ~100–110 |
| H3 on Salad 5090 | ~250 |
| H3 on RunPod PRO 6000 | ~140 |

The gap is real but it's not apples-to-apples: Higgsfield is a finished
creative service (sequencing, continuity, audio, editing, multi-model access:
Seedance 2.0, Kling 3.0, Veo 3.1), not raw diffusion compute. Its Explainer
product: 20s narrated video ≈ 72 credits ≈ $3.60 → ~$10/minute. Market
reference points: official H3 API ~$0.13/s; a cheap hosted H3 ~$0.32 per 15s
768p clip (~$1.30/min).

**Planning number: treat ~50–80 finished one-minute equivalents per $15 on a
$0.72–1.15/hr 32GB card as conservative, ~100–110 on a Vast 5090, and Salad
5090 (~250) as the upper-bound model** — Salad's batch pricing on a distributed
fleet means availability/startup can differ from dedicated clouds.

## Provider priority for the agent

1. **Vast first** — live offers/pricing API, agent-friendly CLI, best
   price/perf balance for H3 (5090 at ~$0.69).
2. **RunPod second** — most polished API/CLI, predictable, PRO 6000 anchor
   card at $1.69.
3. **Salad third** — the pure-economics outlier (5090 at ~$0.294 batch);
   container-based REST API, so the agent's "give me any 24–32GB GPU → run
   this container → return output → terminate" pattern fits naturally. Treat
   it as the aggressive lane with an availability caveat.

The decision flow: discover offers → score by $/generated-second (not $/hr) →
deploy cheapest that fits the quality bar (≥0.4 MP, Turbo 4–8 step, Blackwell
preferred) → if startup is slow or availability is thin, fall back to Vast →
if the output benchmark fails, switch configuration or provider.

## Sources

- MiniMax H3 announcement: https://minimaxi.com/blog/minimax-h3
- RTX 6000 4/8-step benchmark: reddit.com/r/StableDiffusion/comments/1vlo9se
- 5090 4-step Turbo + INT8 reports: reddit.com/r/comfyui/comments/1vm32qs
- Turbo LoRA: reddit.com/r/StableDiffusion/comments/1vgxf4x
- 5070 Ti default-workflow timings: reddit.com/r/StableDiffusion/comments/1vg1qve
- Model-load overhead discussion: reddit.com/r/StableDiffusion/comments/1vhbik7
- H3 API vs RunPod/Vast: reddit.com/r/StableDiffusion/comments/1vfmrvh
- RunPod pricing: runpod.io/pricing · Vast: vast.ai · Salad: salad.com/pricing · Modal: modal.com/pricing
- NVIDIA RTX PRO 4500 specs: nvidia.com/en-gb/products/workstations/professional-desktop-gpus/rtx-pro-4500
