---
name: model-deployment
description: >
  The protocol for turning "make model X runnable" into a working deployment —
  for Claude Code: find the model, determine requirements, pick a runtime and
  GPU class, compare providers live, provision, install, health-check, and
  register in .pi/gpu/deployments.json. Use whenever the user names a model
  (H3, Wan, HunyuanVideo, Seedance, anything) and no ready deployment exists.
---

# Model Deployment (Claude Code)

A model is a deployable artifact whose shape you discover — never hardcode a
deployment path for a model you haven't verified.

## The discovery protocol

1. **Find the model.** Hugging Face (`curl "https://huggingface.co/api/models?search=<name>"`), GitHub, the provider's template library. Record the source URL in `.pi/gpu/models.json`.
2. **Requirements.** From the model card: parameter count, VRAM (fp16 and fp8), torch/CUDA versions, custom nodes or codebase.
3. **Runtime.** Usually ComfyUI (default) → native repo code → provider template → provider-native API.
4. **GPU class.** `vram_fp8` if running quantized, else `vram_min` + headroom for activations. 24GB cards handle 5B–13B class video models (H3 quantized, Wan 5B); 48GB handles fp8 of most 14B; 80GB is the safe class above.
5. **Compare providers live** — never trust the snapshot:
   - Vast: `vastai search offers 'gpu_name=RTX 5090' --order 'dph_total asc' --limit 10 --raw`
   - RunPod: template + pool choice matters (`runpodctl get templates`)
   - Salad: batch pricing, container-based, availability caveats
   - The deciding question is $/generated-second for the actual batch, not $/hr (see docs/h3-economics.md for the H3 model).
6. **Provision** (provider-adapters skill) → **install** (comfyui skill) → **health-check**: endpoint answers AND a test generation produces an output. A deployment that has never produced an output is not a deployment.
7. **Register:** write the deployment to `.pi/gpu/deployments.json` with status ready, endpoint, ssh, real cost_hr; set `"verified": true` in the model manifest and record the working workflow in `.pi/gpu/workflows.json`.

## When it breaks

| Symptom | First move |
| --- | --- |
| CUDA OOM | fp8 / lower res / bigger GPU class — never blind retry |
| torch CUDA mismatch | reinstall torch wheel matching the image's CUDA |
| custom node import error | pin node to last-known-good commit |
| output is garbage | VAE/text-encoder mismatch — check the model card's file set |
| /prompt 400s | validate graph against /object_info |

## The verification rule

Every manifest field marked `"verified": false` is a hypothesis. Confirm from
the source before spending on it.
