---
name: video-models
description: >
  Manifests for the first-class video generation models: MiniMax H3, Wan 2.2,
  HunyuanVideo, Seedance — VRAM requirements, sources, runtime notes, and the
  GPU class that fits each. Use when picking a model for a batch, planning a
  deployment, or estimating what GPU to provision.
---

# Video Models

These are the seed entries in `models.json` (all `verified: false` — treat the
numbers as starting hypotheses and confirm from the model card before
provisioning; the model-deployment skill is the verification protocol).

## MiniMax H3 (Hailuo H3) — flagship text/image-to-video

- **Source:** https://github.com/minimax-ai/Hailuo-H3 (weights via HF/GitHub releases)
- **Capability:** up to **15 seconds at 2K with native stereo audio** in one generation — a "one-minute video" is 4 × 15s clips stitched.
- **VRAM:** 24GB+ with quantized weights (INT8/FP8); 48GB+ for fp16 comfort. Community runs it on 24GB cards with offload.
- **Runtime:** ComfyUI with community nodes is the fastest path; the repo's native inference code is the reference path. Check both at deploy time — H3 support moves fast.
- **THE workflow rule:** run **Turbo LoRA 4–8 steps + SageAttention2 + INT8 models** (+ a cache: Spectrum/EasyCache/FirstBlockCache). The vanilla 20-step workflow is 4–10x slower and is what most public "H3 is slow" benchmarks measured. A new 4-step LightX2V workflow is in testing.
- **Benchmark anchors (measured):** RTX 6000 480p/15s → **56 s @ 4-step**, 1m47s @ 8-step (SageAttention2); RTX 5090 15s/0.4MP → **~3 min** with 4-step Turbo + INT8 (was ~12 min); RTX 5090 5s/1MP → ~58 s (4-step + Sage). Full model + sources: docs/h3-economics.md.
- **Benchmark hygiene:** never trust an H3 number without checkpoint, quantization, resolution, duration, steps, sampler, attention impl, cache, and whether model load is included — a 10x spread between "H3 benchmarks" is normal.
- **Workflows:** `h3-text-to-video`, `h3-image-to-video` (registry ids).
- **Use for:** the flagship quality pass. With the Turbo stack it's affordable; without it, it's a money pit — always deploy the optimized workflow first.

## The H3 provider lanes (planning, 2026-08-12)

| Lane | Best GPU | $/hr | Per finished minute | When to pick it |
| --- | --- | ---: | ---: | --- |
| **Vast** | RTX 5090 | ~0.69 | ~$0.14 | Default — best price/perf/ops balance |
| **Salad** | RTX 5090 | ~0.294 (batch) | ~$0.06 | Pure economics, retryable batches; availability caveats |
| **RunPod** | RTX PRO 6000 | 1.69 | ~$0.105 | Measured 56s/4-step anchor card; most predictable ops |
| RunPod | RTX PRO 4500 32GB | 1.15 | ~$0.19–0.29 | Only if you must — no public H3 benchmark, ~half a 5090's cores/bandwidth |

Vast 5090 is the recommendation for H3. Salad is the upper-bound economics
model (treat ~250 finished minutes/$15 as optimistic). Always verify live
prices; always benchmark the exact workflow before committing a batch.

## Wan 2.2 (Alibaba) — the workhorse

- **Sources:** HF `Wan-AI/Wan2.2-T2V-14B`, `Wan-AI/Wan2.2-T2V-5B` (+ I2V variants)
- **VRAM:** 14B ≈ 40GB full / ~26GB fp8; 5B ≈ 16GB — fits a 24GB card.
- **Runtime:** ComfyUI native nodes (`ComfyUI-WanVideoWrapper`). Mature.
- **Use for:** the default batch model — cheapest quality/price ratio, and the
  5B runs on the cheapest cards (RTX 4090/5090, L4).

## HunyuanVideo (Tencent) — 13B, needs the big cards

- **Source:** HF `tencent/HunyuanVideo`
- **VRAM:** ≈60GB full / ~40GB fp8 → 80GB class, or 4090 with aggressive
  offload (slow). Requires `ComfyUI-HunyuanVideoWrapper`.
- **Use for:** when Wan quality isn't enough and H3 is overkill.

## Seedance (ByteDance) — check the access path

- **Source:** HF `ByteDance-Seed/Seedance-1.0-lite` (verify open-weight status
  at deploy time); the pro tier is API-only (Volcano Engine).
- **VRAM:** 40GB+ class.
- **Use for:** a second opinion pass; API path has no GPU management at all.

## Model → GPU quick table

| Model | Min VRAM | Comfortable card | Hourly ballpark (Vast/RunPod) |
| --- | --- | --- | --- |
| Wan 2.2 5B | 16GB | RTX 4090/5090, L4 | $0.25–0.70 |
| Wan 2.2 14B (fp8) | ~26GB | RTX 5090, L40S | $0.31–0.99 |
| HunyuanVideo (fp8) | ~40GB | L40S, A100 | $0.43–1.49 |
| MiniMax H3 (INT8/FP8, Turbo) | 24GB | RTX 5090 (Vast ~$0.69), PRO 6000 | $0.29–1.69 |
| Seedance lite | 40GB | L40S, A100 | $0.43–1.49 |

Pricing is a snapshot (docs/pricing/) — always query live rates for a real
batch (gpu-ops: keep-alive math decides which card is actually cheapest for a
given queue).
