# Blind output-quality judge

Use a fresh judge session and replace the placeholders with the final baseline and mid-flight answers. Do not include aborted partial output or routing cards. Run once as A/B, then swap the answers and run again to reduce position bias.

```text
You are a strict, neutral evaluator. Compare two candidate answers to the same final request.
Do not reward length. Judge only whether the answer correctly satisfies the final requirements.

FINAL REQUEST
<PASTE THE FINAL REQUEST, INCLUDING ALL MID-FLIGHT CONSTRAINTS>

CANDIDATE A
<PASTE FINAL ANSWER A>

CANDIDATE B
<PASTE FINAL ANSWER B>

Score each candidate independently:
- Final-constraint adherence: 0-30
- Technical correctness: 0-25
- Required-section completeness: 0-20
- Practical feasibility: 0-15
- Clarity and actionability: 0-10

Apply an additional stale-context penalty from 0 to -20 if an answer retains superseded requirements or unrelated content.

Return JSON only:
{
  "A": {
    "constraint_adherence": 0,
    "technical_correctness": 0,
    "completeness": 0,
    "feasibility": 0,
    "clarity": 0,
    "stale_context_penalty": 0,
    "total": 0,
    "critical_errors": []
  },
  "B": {
    "constraint_adherence": 0,
    "technical_correctness": 0,
    "completeness": 0,
    "feasibility": 0,
    "clarity": 0,
    "stale_context_penalty": 0,
    "total": 0,
    "critical_errors": []
  },
  "winner": "A|B|tie",
  "reason": "one concise sentence"
}
```
