---
type: Template
title: Eval Rubric Template
description: The scaffold for an evaluation rubric instance, including weighted dimensions, partial credit, and the sprint contract. Copy into /planning/evals/ and fill.
tags: [template, eval, rubric, scoring, planning]
timestamp: 2026-06-28
---

# Eval Rubric Template

**Situating context:** This template was authored as the canonical scaffold the
[eval-rubric skill](/skills/eval-rubric.skill) instantiates, so every rubric uses the same
weighted-scoring math and carries its sprint contract. Copy the block into
`/planning/evals/<slug>-eval-rubric.md` and fill it; the copy is operational and does not live in the
OKF bundle.

> Usage: replace every `<…>`. Weights must sum to 1.0. The threshold is pre-committed — set it before
> any output exists. See [output-eval](/okf/core/concepts/output-eval.md) for the gate model.

---

```markdown
---
type: Rubric
title: <initiative title>
eval_type: capability | regression   # capability = push frontier; regression = guard against backslide
initiative_id: <id>
product_slug: <product_slug>
pm_slug: <pm_slug>
parent_KR: <reference OR health:<id> (D59)>
pass_threshold: <0.00–1.00>          # pre-committed, before output exists
timestamp: <YYYY-MM-DD>
---

# Eval Rubric: <initiative title>

## Gate assignment
<!-- State which gate each criterion serves. Acceptance Gate is PM-only and NOT rubric-graded. -->
- Quality Gate (deterministic): <compile / lint / tests that must pass>
- Review Gate (rubric-scored): the weighted dimensions below
- Acceptance Gate (PM, strategic): NOT graded here — PM decides "is this the right thing?"

## Weighted dimensions (Review Gate)
<!-- Outcome-only: grade what was produced, not the path/tool-calls. Partial credit required. -->

| Dimension              | Weight | Score (0–1) | Weighted = W×S |
|------------------------|-------:|------------:|---------------:|
| <dimension 1>          |  <0.x> |       <0.x> |          <0.x> |
| <dimension 2>          |  <0.x> |       <0.x> |          <0.x> |
| <dimension 3>          |  <0.x> |       <0.x> |          <0.x> |
| **Total**              | **1.0**|             |   **<sum>**    |

**Result:** total `<sum>` vs threshold `<pass_threshold>` → PASS / FAIL.

## Sprint contract (pre-run)
initiative_id: <id>
run_id: <id>
done_criteria:
  1. <testable criterion>
  2. <testable criterion>
test_method:
  1. <how criterion 1 is verified, and at which gate>
  2. <how criterion 2 is verified, and at which gate>
PM_approved: true | false   <!-- must be true before work starts; P0 = temp file + PM sign-off -->
```

---

## Partial credit / weighted scoring — how to use

The rubric is **not** all-or-nothing. Score each dimension on a continuous 0–1 scale (partial credit
required), multiply by its weight, and sum:

```
total = Σ (weight_i × score_i),   with Σ weight_i = 1.0
PASS iff total ≥ pass_threshold
```

Rules:

- **Weights sum to 1.0.** If they do not, the rubric is malformed.
- **Partial credit is mandatory.** A dimension that can only score 0 or 1 hides quality gradients and
  defeats weighted scoring — split it or define intermediate anchors.
- **Anchor the scale — richer anchors lift agreement.** For each dimension, define what the
  score levels look like so scoring is repeatable and calibratable against
  [eval-calibration.md](/planning/evals/eval-calibration.md). The **minimum** is `0 / 0.5 / 1`;
  **prefer a 5-level anchored scale** (`0 / 0.25 / 0.5 / 0.75 / 1`, or a 1–5 scale mapped onto it) with
  a concrete one-line descriptor per level. Anchored multi-level scales are the single **largest
  evidenced lever on judge agreement** — Cohen's κ rises materially versus a bare numeric or a
  three-point scale ([llm-judge-calibration](/planning/research/llm-judge-calibration.md)): the more
  precisely each level is described, the less two evaluators (or the evaluator vs the PM) diverge.
  Spend the extra levels on the dimensions where the quality gradient actually matters, not every one.
- **Threshold is pre-committed.** Never tune the threshold after seeing the output.
- **Evidence per dimension.** A Review-Gate verdict must carry, for each scored dimension, a short
  verbatim quote or precise pointer (file+line, query output, CI line) from the evaluated artifact;
  a dimension without evidence is marked `unverified` and scored conservatively (≤ 0.5). Rubric
  authors should write dimensions so that evidence is *citable* — observable outcomes, not
  impressions ([llm-judge-calibration](/planning/research/llm-judge-calibration.md)).
- **Outcome-only.** Score the artifact produced, not the tool-call sequence or reasoning path.
- **Coherence criteria are semantic.** A "rule holds everywhere" criterion must enumerate its
  live-doc set — every living mirror of a live surface (skills, templates, playbooks, concepts,
  protocol docs, tool schemas + their reference/data-model mirrors, UI copy, intake instruments) —
  and be tested by a semantic pass (reading for meaning); a grep sweep is evidence,
  never the definition of done (health-anchors lesson —
  [eval-calibration](/planning/evals/eval-calibration.md)).

## Authoring checklist

- [ ] `eval_type` set.
- [ ] Weights present and sum to 1.0.
- [ ] Each dimension supports partial credit with anchored 0/0.5/1.
- [ ] `pass_threshold` pre-committed.
- [ ] Gate assignment stated; Acceptance Gate marked PM-only, not graded.
- [ ] Dimensions written so a verdict can cite verbatim evidence per dimension (observable
      outcomes, not impressions).
- [ ] Any rule-coherence criterion enumerates its live-doc set and tests by semantic pass.
- [ ] Sprint contract present with `PM_approved`.
