# interview-eval output templates

Two per-session artifacts (`report.md`, `scorecard.yaml`) and one rollup (`candidate-rollup.md`). Fill every `‹…›`. Keep citations to real `U<n>` turns with short verbatim quotes.

---

## `report.md` (per session)

```markdown
# Steering evaluation — ‹session title›

`‹session-id›` · ‹session-type› · ‹N› human turns · ‹wall-clock› · cli ‹version›

> **Steering level: ‹Exemplary | Strong | Competent | Developing | Weak›** (‹normalized›/100) ‹One-sentence verdict — what kind of operator this session shows, in plain language.›

**Headline:** ‹2–3 sentences. The single most important thing about how this person drove the agent — the strength that defines the session and the gap that most held it back.›

## Scores

| Dim | Dimension | Band | Wt | Note |
| --- | --- | :-: | :-: | --- |
| A | Problem Framing & Specification | ‹0–5› | 15 | ‹≤10 words› |
| B | Decomposition & Sequencing | ‹0–5› | 12 | ‹≤10 words› |
| C | Context Curation | ‹0–5› | 10 | ‹≤10 words› |
| D | Verification & Skepticism | ‹0–5› | 18 | ‹≤10 words› |
| E | Debugging Collaboration | ‹0–5 / N/A› | 10 | ‹≤10 words› |
| F | Course Correction & Steering | ‹0–5› | 10 | ‹≤10 words› |
| G | Technical & Architectural Judgment | ‹0–5› | 15 | ‹≤10 words› |
| H | Workflow & Tooling Leverage | ‹0–5› | 5 | ‹≤10 words› |
| I | Outcome & Efficiency | ‹0–5› | 5 | ‹≤10 words› |
|  | **Weighted total** |  |  | **‹normalized›/100** |

‹If any gate fired:› **Gate: ‹GATE-NAME›** — ‹what triggered it, with the turn cite›. Level capped at ‹cap›.

## Per-dimension findings

For each dimension, 2–4 sentences of rationale + the turn citations that justify the band.

**A — Problem Framing (‹band›).** ‹rationale› _Evidence:_ ‹U1: "…"›, ‹U4: "…"›.

**B — Decomposition (‹band›).** ‹rationale› _Evidence:_ ‹…›.

‹…repeat C–I…›

## Standout moments

- ✅ ‹The 1–3 best steering moves, each with a turn cite and why it was good.›

## Weaknesses

- ⚠ ‹The 1–3 most costly gaps, each with a turn cite and the consequence.›

## What a stronger operator would have done differently

‹3–5 concrete, specific alternatives — the actual better prompt or move at a named turn. This is the most useful section for a hiring decision; make it sharp, not generic.›

## Caveats

‹Harness noise excluded from scoring (e.g. "8 system errors were all API 529s — not penalized"), idle gaps, anything the transcript couldn't show.›
```

---

## `scorecard.yaml` (per session, machine-readable)

```yaml
session_id: ‹id›
title: ‹ai-title›
session_type: ‹greenfield-build | feature-add | debugging | refactor | ops | research | review›
turns: ‹N›
duration_human: ‹3h42m›

dimensions:
  framing: { band: ‹0-5›, weight: 15, rationale: "‹…›", cites: ["U1", "U4"] }
  decomposition:
    { band: ‹0-5›, weight: 12, rationale: "‹…›", cites: ["U2", "U4"] }
  context: { band: ‹0-5›, weight: 10, rationale: "‹…›", cites: ["U9"] }
  verification:
    { band: ‹0-5›, weight: 18, rationale: "‹…›", cites: ["U7", "U11"] }
  debugging:
    {
      band: ‹0-5|null›,
      weight: 10,
      na: ‹false|true›,
      rationale: "‹…›",
      cites: ["U9"],
    }
  correction: { band: ‹0-5›, weight: 10, rationale: "‹…›", cites: ["U13"] }
  judgment: { band: ‹0-5›, weight: 15, rationale: "‹…›", cites: ["U23"] }
  tooling: { band: ‹0-5›, weight: 5, rationale: "‹…›", cites: ["U17"] }
  outcome: { band: ‹0-5›, weight: 5, rationale: "‹…›", cites: ["U26"] }

gates_triggered: [] # e.g. [BLIND-SHIP]
green_flags: [] # e.g. [empirical-debugging, unprompted-operability]

weighted_total: ‹float› # Σ(band/5 × weight) over scored dims
available_weight: ‹int› # Σ weight of scored (non-N/A) dims
normalized: ‹float› # weighted_total / available_weight × 100, after gate caps
level: ‹Exemplary | Strong | Competent | Developing | Weak›

evidence_excluded_from_scoring:
  - "‹e.g. 8 API 529 errors — Anthropic-side overload, not candidate fault›"
```

---

## `candidate-rollup.md` (only when >1 session)

```markdown
# Candidate steering rollup — ‹candidate / label›

‹K› sessions · ‹total turns› · ‹combined wall-clock›

> **Overall steering level: ‹level›** ‹One-paragraph calibrated read of the candidate as an AI-pairing operator across all sessions.›

## Sessions

| Session | Type   | Level   | Norm. | Headline    |
| ------- | ------ | ------- | :---: | ----------- |
| ‹title› | ‹type› | ‹level› |  ‹n›  | ‹≤12 words› |

## Dimension trend across sessions

| Dim       | S1  | S2  | …   | Read                                        |
| --------- | :-: | :-: | --- | ------------------------------------------- |
| A Framing | ‹b› | ‹b› |     | ‹consistent? improving? context-dependent?› |
| …         |     |     |     |                                             |

## Consistency

‹Does the candidate steer the same under pressure (debugging) as greenfield? Where does quality drop — and is the drop in a dimension that matters for the role?›

## Strongest evidence (top 3 across all sessions)

1. ‹session · U‹n› · what it shows›

## Weakest evidence (top 3 across all sessions)

1. ‹session · U‹n› · what it shows›

## Interview follow-ups to probe the gaps

- ‹Concrete question that would confirm/deny the weakest dimension, e.g. "Walk me through how you'd have verified the worker actually drained the queue — what would you have checked?"›
```
