---
type: Concept
title: Output Evaluation — The Three-Tier Gate
description: How PMOS evaluates agent output through three escalating gates — Quality, Review, Acceptance — and why the PM owns only the last one.
tags: [evaluation, gates, quality, acceptance, governance]
timestamp: 2026-06-28
---

# Output Evaluation — The Three-Tier Gate

**Situating context:** Every turn of the [SDLC loop](/okf/core/concepts/sdlc-loop.md) ends by
judging what an agent produced. This concept defines the three-tier gate that judgment runs through,
so the [eval-rubric skill](/skills/eval-rubric.skill), AGENTS.md, and the PM all use the same
vocabulary. It informs where automation stops and human judgment begins.

## The three tiers

Output passes through three escalating gates. Each answers a different question, with a different
nature and a different owner.

### 1. Quality Gate — *does it work?*

- **Nature:** deterministic, automated.
- **Checks:** compile, lint, unit tests — toolchain-enforced pass/fail.
- **Owner:** the toolchain. No judgment involved; either the tests pass or they do not.
- A failure here never reaches a human; it bounces back to the agent.

### 2. Review Gate — *does it satisfy the rubric?*

- **Nature:** probabilistic, adversarial.
- **Checks:** does the output meet the rubric criteria for this initiative? Applied by a separate
  **evaluator agent** or by the PM acting as critic.
- **Owner:** evaluator agent (or PM-as-critic).
- This gate is **uncalibrated** until 5–10 P1 cycles are validated — the evaluator defaults to
  generous scoring (**evaluator leniency**). Treat its verdict as advisory until calibrated against
  [eval-calibration.md](/planning/evals/eval-calibration.md).

### 3. Acceptance Gate — *is it the right thing?*

- **Nature:** subjective, strategic.
- **Checks:** strategic fit. Is this what the product actually needs right now? Does it move the
  parent KR — or, for health-anchored work, honestly serve its declared budget (D59)? **This is
  NOT a re-check of the rubric.**
- **Owner:** the **human PM only**. Non-delegable.
- The PM's role at the end of every agent run **is** the Acceptance Gate.

## Why three, and why the order

The gates escalate from cheap-and-certain to expensive-and-strategic:

- Quality is **cheap and certain** — run it first, fail fast, never spend human attention on broken
  output.
- Review is **mid-cost and probabilistic** — it filters rubric misses before they reach the PM, but
  its judgment is fallible.
- Acceptance is **expensive and irreducibly human** — it is the only gate that asks whether the work
  *should* exist, not whether it was done correctly.

Running them out of order wastes the scarce resource (PM attention) on problems the cheap gates
would have caught.

## The Acceptance Gate is not the Review Gate

The most important distinction in this concept: **passing the rubric is not the same as being
right.** An output can satisfy every rubric criterion and still be the wrong thing to build — that
is exactly the [silent drift](/okf/core/concepts/watermelon-flag.md) failure. The Acceptance Gate
exists *because* rubric satisfaction is not strategic correctness. Collapsing the two — letting a
green Review Gate auto-accept — removes the only structural defense against drift.

## Grading discipline

- **Outcome-only:** gates check what the agent *produced*, not the path it took. Do not grade tool
  call sequences.
- **Weighted, with partial credit:** the Review Gate scores weighted dimensions and totals against a
  threshold; it is not all-or-nothing. (See the
  [eval-rubric skill](/skills/eval-rubric.skill) and
  [template](/okf/core/templates/eval-rubric-template.md).)

## Relationship to autonomy

As PMOS moves L2 → L3, more of the Quality and Review gates run unattended. The Acceptance Gate
stays human at every autonomy level up to the declared L3 ceiling. Keeping that gate human is what
makes higher autonomy safe rather than reckless.
