---
name: model-deployment
description: >
  The protocol for turning "make model X runnable" into a working deployment:
  find the model, determine its requirements, pick a runtime and GPU class,
  compare providers, provision, install, health-check, and register. Use
  whenever the user names a model (H3, Wan, HunyuanVideo, Seedance, anything)
  and no ready deployment exists.
---

# Model Deployment

A model is a deployable artifact with a shape you discover, not a name you
hardcode. The extension holds registries (`models.json`, `workflows.json`) but
the *path* from model name to running service is found by investigation:

## The discovery protocol

1. **Find the model.** Hugging Face (`huggingface.co/api/models?search=<name>`),
   GitHub, the provider's template library (RunPod has one-click templates for
   popular models). Record the source URL in `gpu_model_add`.
2. **Determine requirements.** From the model card: parameter count, VRAM
   (full precision and fp8 if published), required torch/CUDA versions, whether
   it needs a specific custom node or codebase.
3. **Pick the runtime.** Usually: ComfyUI (default, most models have community
   workflows) → native repo code (model ships its own `generate.py`) → provider
   template (someone already packaged it) → provider-native API (API-only
   models like Seedance via Volcano Engine).
4. **Pick the GPU class.** `vram_gb_fp8` if you'll run fp8, else `vram_gb_min`;
   add headroom for activations (a model with 40GB of weights wants a 48GB+
   card, not a 40GB one). 24GB cards (RTX 4090/5090, L4) handle 5B–13B class
   video models; 48GB (L40S) handles fp8 of most 14B models; 80GB (H100/A100)
   is the safe class for anything above.
5. **Compare providers.** Query live prices, don't trust the snapshot:
   - Vast: `vastai search offers 'gpu_name=RTX 5090' --order 'dph_total asc' --limit 10 --raw` — cheapest market rate, but storage is billed separately and offers vanish.
   - RunPod: community pool is cheaper than secure; template selection matters (`runpodctl get templates` or the console).
   - Lambda/Modal: check their pricing docs; Modal has no idle cost (serverless), Lambda has no egress fees.
   The deciding question is usually: *does the batch fit in the idle window at
   the hourly rate?* (see gpu-ops economics).
6. **Provision** via `gpu_provision` (or manually per the provider-adapters
   skill), **install** the runtime + weights per the comfyui skill,
   **health-check** (endpoint answers, a tiny test generation runs — always
   run a test generation before registering READY; a deployment that has never
   produced an output is not a deployment),
7. **Register** with `gpu_register endpoint=... ssh=...` → status ready, and
   **record the receipt**: `gpu_model_add verified=true` + `gpu_workflow_add`
   for the working graph. The next request for this model is then a one-liner.

## When it breaks

Follow gpu-ops' failure protocol (logs → classify → fix → retry once). The most
common deployment failures and their first fix:

| Symptom | First move |
| --- | --- |
| CUDA out of memory | fp8 weights / lower res / bigger GPU class — never blind retry |
| torch CUDA mismatch | reinstall torch wheel matching the image's CUDA version |
| custom node import error | pin node to last-known-good commit |
| model loads but output is garbage | VAE/text-encoder mismatch — check the model card's exact file set |
| endpoint up but /prompt 400s | validate graph against `/object_info` |

## The deploy-time verification rule

Every field in a model manifest that says `verified: false` is a hypothesis.
Verify from the source (model card, repo README) before spending on it — the
manifest exists to make hypotheses explicit, not to make them true.
