# Modal Adapter

**Interface:** `modal` CLI. **Kind:** serverless — there are no machines. **Zero idle cost, cold starts.**

## Install & auth

```bash
pip install modal
modal token new          # browser flow
modal profile current    # verify
```

## Deploy (there is no "provision")

A Modal deployment is a Python app. The video-factory pattern:

```python
# app.py
import modal
app = modal.App("video-factory")
image = modal.Image.deploy_sandbox(  # or Image.from_registry("runpod/comfyui:...")
    "python:3.11"
).pip_install("torch", "comfyui-reqs...")

@app.function(image=image, gpu="H100", timeout=600, allow_concurrent_inputs=4)
def run_workflow(prompt: str, ...):
    # start ComfyUI in-process/on port 8188, queue the workflow, return artifact URI
    ...
```

```bash
modal deploy             # from the app dir → live URL
modal app logs video-factory   # tail logs
```

GPU strings: `"T4"`, `"L4"`, `"A10G"`, `"L40S"`, `"A100"`, `"A100-80GB"`,
`"H100"`, `"H200"`, `"B200"`, `"B300"`, with `:N` for count (`"H100:8"`).

## Destroy (there is no destroy)

`modal app stop video-factory` stops the app; billing stops automatically when
no containers run. Scale-to-zero is the whole point.

## Gotchas

- **Cold starts:** image pull + model load on every scale-up. For a dense
  queue this repeats constantly; for sparse batches it's free. The gpu-ops
  keep-alive math decides which regime you're in.
- Secrets (HF tokens for gated weights): `modal secret create hf --env HF_TOKEN=...`
- Network egress from Modal containers is billed; big weight downloads count.
- `gpu="H100"` auto-upgrades to H200 when available — use `"H100!"` to force.

## Reference

- Docs: https://modal.com/docs/guide/gpu (GPU types, pricing semantics)
- Pricing snapshot: docs/pricing/modal.md
