---
type: Concept
title: Eval-Driven Product Management
description: PMOS's core method — write the eval before the build. The rubric is the spec; the PM manages by defining what "good" means, not by inspecting work after the fact.
tags: [evaluation, method, product-management, eval-driven]
timestamp: 2026-06-28
---

# Eval-Driven Product Management

**Situating context:** This concept names PMOS's central operating method — the thing that makes it a
distinct way to do product management rather than just a folder of skills. It ties together the
[SDLC loop](/okf/core/concepts/sdlc-loop.md), the [eval gate](/okf/core/concepts/output-eval.md), and
[pre-run contracts](/okf/core/concepts/sdlc-loop.md). It informs why the eval-rubric skill exists and
why contracts come before builds.

## The core move: write the eval first

Eval-driven PM is test-driven development lifted to the product altitude: **define how you will judge
the output before any output exists.** The eval rubric is written, and the pre-run contract agreed,
*before* the agent builds anything.

This inverts the default. The default is: build, then ask "is this good?" Eval-driven PM is: decide
what "good" means, then build to it. The rubric **is** the spec; building is the act of satisfying a
spec that was made measurable in advance.

## Why before, not after

Writing the eval after the build corrupts the judgment in three ways:

1. **The output anchors the standard.** Once you have seen what the agent produced, your idea of
   "good enough" silently bends to match it. A rubric written in ignorance of the output is the only
   honest one.
2. **There is nothing to fail against.** Without a prior standard, evaluation degenerates into "does
   this look fine?" — which is exactly the [evaluator leniency](/okf/core/concepts/output-eval.md)
   failure, made structural.
3. **Drift has no tripwire.** A pre-committed standard is what lets you notice
   [silent drift](/okf/core/concepts/watermelon-flag.md): the output meets the rubric or it does
   not, and you defined the rubric back when you still remembered the intent.

## The PM manages by defining "good," not by inspecting work

In eval-driven PM, the PM's primary leverage is *upstream*: authoring the rubric, agreeing the
contract, setting the threshold. By the time work comes back, most of the management has already
happened — it lives in the standard the work is measured against.

This is why the PM's end-of-run role is the [Acceptance Gate](/okf/core/concepts/output-eval.md), not
line-by-line inspection. The Quality and Review gates check the work against the pre-written rubric
mechanically; the PM's scarce attention is spent on the one question the rubric cannot answer — *is
this the right thing?* — and on whether the rubric itself was right.

## How it shapes the loop

- **Spec stage** authors the rubric alongside the PRD — the
  [eval-rubric skill](/skills/eval-rubric.skill) and [prd skill](/skills/prd.skill) are
  paired for this reason.
- **Contract stage** turns the rubric into agreed, testable done-criteria the generator and evaluator
  both sign off on before work starts.
- **Evaluate stage** runs the output against the pre-committed rubric — outcome-only, weighted,
  partial-credit.
- **Learn stage** feeds rubric calibration back into
  [eval-calibration.md](/planning/evals/eval-calibration.md): when the PM and evaluator disagree, the
  rubric or the evaluator prompt — not the memory of what happened — gets fixed.

## Calibration is part of the method

An eval-driven system is only as trustworthy as its evals, and PMOS's evals start
**uncalibrated** (evaluator leniency). Eval-driven PM therefore includes a calibration discipline:
treat early P1 runs as calibration cycles, log PM overrides, and patch the evaluator until PM and
evaluator agree. The method is not "trust the eval" — it is "make the eval trustworthy, then manage
through it."
