---
name: prompt-eval
description: Use when a prompt needs its quality measured — after drafting or revising a prompt, before certifying a library entry, or when comparing prompt variants. Runs the golden-set eval loop with rubric grading and reports deltas. Triggers on /prompt-eval, "eval this prompt", "does this variant score better", "run the golden set", "certify this prompt".
---

# Prompt Eval — the golden-set loop

A prompt's quality claim is its eval result, nothing else. This skill runs the loop; conventions for golden sets, rubrics, and thresholds are `.claude/rules/evals.md`.

## Procedure

1. **Locate the family's golden set** at `evals/{family}/` — the owning feature doc names the family. A prompt with no golden set gets one *first* (eval-first: the set ships with the prompt, and `data-generator` drafts edge-case items for human review).
2. **Run the set** against the prompt via `eval-runner` — every item, not a sample. Batch variants in one run so scores are comparable.
3. **Grade with the family's rubric.** Mechanical checks (format, required fields, refusal correctness) grade in code; judgment dimensions grade against the rubric's anchors. Report per-dimension scores and the delta against the incumbent prompt's certified score.
4. **Gate:** meeting the family's threshold on iteration is a *candidate*; certification for the library additionally requires the `judge` ship gate (frontier, one dispatch — the only frontier call in the loop).
5. **Record:** the winning variant's scores go into its `prompt-library` entry; the losing variants' deltas belong in the changelog entry's refinement context.

## Rules

- **Iteration stays cheap.** The draft → eval → revise loop runs on economy (`eval-runner`) and standard (`prompt-iterator`) tiers; the frontier judge appears once, at the gate — never inside the loop.
- **Never weaken a rubric to pass a prompt.** A rubric change is a design change to the family's eval suite: it goes through the owning feature doc, re-runs the whole family, and re-certifies every entry.
- **A failing eval is a result, not an obstacle.** Report it honestly; if the family genuinely needs a stronger tier, that is the `escalate` skill's evidence.
