---
name: deploy-model
description: Package and ship a model — versioned artifact, serving API, model card, and a monitoring plan for production.
keywords: deploy, serving, mlops, model card, monitoring, versioning, registry
---

# Deploy Model

> Pairs with `superpowers:requesting-code-review` at Gate 5 — once the model is packaged and the monitoring plan is written, open the PR with the ticket link for peer review before any production rollout.

---

## When to use this skill

- Gate 5 of the ML workflow, after the evaluation has been APPROVED
- Any time a trained model artifact needs to be promoted from experiment to production
- When a new model version replaces an existing one in a registry or serving endpoint

---

## Steps

### 1. Version the artifact

Register the model in the artifact store or model registry (MLflow Model Registry, Weights & Biases Artifacts, or equivalent). The registered entry must include:

- Model version number
- Training data version (DVC hash, dataset URI, or snapshot timestamp)
- Code commit SHA that produced the artifact
- Config file used for training (hyperparameters, seed)

Without these four identifiers, the artifact cannot be reproduced or audited. Do not promote an unversioned artifact to production.

### 2. Package the full inference pipeline

Save the complete end-to-end inference pipeline — preprocessing transforms plus the model — as a single serializable object (sklearn `Pipeline`, MLflow `pyfunc` model, ONNX graph, TorchScript, TF SavedModel, etc.). The packaging must guarantee **train/serve parity**: the identical preprocessing that was applied during training must run at inference time, with no manual re-implementation in the serving layer.

Pin the serving environment to the same dependency versions used during training. Document any framework-specific serialization caveats (e.g., custom PyTorch modules, sklearn custom transformers).

### 3. Expose a serving API

Build or configure an inference endpoint with the following properties:

- **Input validation:** reject malformed or out-of-range inputs with a clear error before they reach the model
- **Batch and timeout handling:** define the maximum batch size and inference timeout; return a graceful error on timeout
- **Health check endpoint:** a lightweight `/health` or equivalent that confirms the model is loaded and responsive — required for orchestrators and load balancers
- **Versioned endpoint path:** include the model version in the route so multiple versions can coexist during a canary or A/B rollout

Example minimal stack: FastAPI + uvicorn with a Pydantic input schema for validation.

### 4. Define the monitoring plan

Specify what will be observed in production and what actions each signal triggers. The plan must cover:

- **Input drift:** statistical test (e.g., KS test, PSI) on feature distributions vs the training baseline; alert threshold
- **Prediction distribution drift:** shift in output scores or class probabilities; alert threshold
- **Latency:** p50 / p95 / p99 response times; SLA breach alert
- **Error rate:** HTTP 5xx and model exception rate; alert threshold
- **Retraining trigger:** the specific condition (drift score, metric degradation, elapsed time, data volume) that initiates a retraining pipeline run

Document the monitoring plan in the model card and link it to the observability tooling (Prometheus, Grafana, Arize, Evidently, etc.).

### 5. Finalize the model card

Update the model card drafted at Gate 4 with:

- Intended use and out-of-scope uses
- Training data description (source, version, date range, class balance)
- Evaluation metrics (held-out results, baseline comparison, per-slice breakdown)
- Known limitations and failure modes identified in error analysis
- Monitoring and retraining plan summary
- Owner name and contact; approval date

The model card is the primary audit document for the deployed model.

Output language: auto-detect from the ticket/task input — see `custom/rules/output-language.md` (Vietnamese input → Vietnamese output; otherwise English).

### 6. Create the PR

Invoke `superpowers:requesting-code-review`. Open a PR that includes: serving code, model card, monitoring plan, and any pipeline changes. Link the ticket. Reference the eval report and the registered artifact version in the PR description.

---

## Completion Checklist

- [ ] Artifact registered with data version, code commit, and config
- [ ] Full inference pipeline packaged — preprocessing + model, train/serve parity guaranteed
- [ ] Serving environment pinned to training-time dependency versions
- [ ] Inference API includes input validation, batch/timeout handling, and health check
- [ ] Monitoring plan defined: input drift, prediction drift, latency, error rate, retraining trigger
- [ ] Model card finalized: intended use, data, metrics, limitations, monitoring summary, owner
- [ ] PR created with ticket link, eval report reference, and artifact version noted
