---
name: interview-eval
description: "Use when assessing how a human steered a Claude Code session — prompting, decomposition, verification, debugging, judgment — from `.jsonl` transcripts. For interview/hiring of engineers using AI coding agents; writes a per-session report + scorecard plus a rollup. Judges the human, not the code."
---

# interview-eval

Judge **how a human drove a Claude Code session** — not how good Claude's output was. The unit of evaluation is the person's steering: how they framed the problem, decomposed it, fed context, verified results, debugged, corrected course, and revealed engineering judgment through their prompts. One or many session transcripts in → per-session reports + scorecards out, plus a candidate-level rollup when there are several.

This is the **counterpart** to the sibling `evaluator` skill: `evaluator` scores a _repo's artifacts_; `interview-eval` scores the _conversation that produced them_. They share no rubric.

## The one rule that governs every score

**You are scoring the human, not the model.** Claude's code being clean does not earn the candidate points; Claude's code being buggy does not lose them points — _unless_ the human's prompt caused it or the human failed to catch it. Reward the prompt that set Claude up to succeed, the verification that caught Claude's mistake, the redirect that pulled Claude off a bad path. A candidate who wrote three sharp prompts and let a capable agent run is steering _well_, even if the turn count is low. A candidate who typed forty vague nudges and rubber-stamped broken output is steering _badly_, even if something shipped. Turn count, prompt length, and verbosity are **not** quality signals — precision, judgment, and verification are.

## Input

- **One or more session transcripts** — Claude Code `.jsonl` files (e.g. `0fe33297-….jsonl`). Accept a list of files, a glob, or a directory (evaluate every `*.jsonl` inside).
- **An output directory** — `out_dir` (default: `./interview-eval-out/`). One subfolder per session, plus a rollup at the root when >1 session.
- **Optional context** the user may volunteer: the candidate's seniority/role, the task brief they were given, time limits, or whether sessions are independent tasks vs. one continued task. Fold it in; never invent it.

If the user just points at this folder, treat every `*.jsonl` as one candidate's sessions unless told otherwise.

## Process

### 0. Preflight (cheap)

Only dependency is `python3` (stdlib only). No installs, no network. If `python3` is missing, say so and stop.

### 1. Extract an evidence pack per session (deterministic)

Run the extractor — it never scores, it only surfaces signal:

```bash
python3 <skill_dir>/extract_session.py SESSION.jsonl --out <out_dir>/<session-slug>/ --json
```

It emits a markdown evidence pack on stdout and writes `<sessionId>.session.json`. The pack gives you:

- **Turn-by-turn**: each human prompt (IDE tags stripped) followed by a summary of what the agent did (tool counts + sample targets). This is your primary evidence — quote `U<n>` turn numbers in every finding.
- **Metrics**: prompt count, wall-clock & idle gaps, tool histogram, permission-mode transitions, plan-mode / subagent / AskUserQuestion / slash-command usage, interrupts, system/API errors.
- **Auto red-flag signals**: secrets pasted into prompts, destructive commands the agent ran (`rm -rf ~`, force-push, `DROP TABLE`, …). These are **leads to verify**, not verdicts.

**Read the full evidence pack before scoring.** For long sessions, also open the raw `.jsonl` around any turn you intend to quote, to confirm context. Do not score from metrics alone — a session can have great metrics and terrible steering, or vice-versa.

**Separate noise from signal.** `system_errors` that are `status: 529` / `ECONNRESET` are Anthropic-side overload, not the candidate's fault — never penalize them. `interrupts` may be a sharp early redirect (good) or thrash (bad) — read the surrounding turns to tell which.

### 2. Classify the session type

Pick the dominant type (a session can be a blend — note it): **greenfield-build · feature-add · debugging · refactor · ops/infra · research/Q&A · review/audit**. The type sets which dimensions are _primary_, _secondary_, or _N/A_ — see `rubric.md` → "Session-type matrix". A pure debugging session, for example, makes **E** the main event and may mark **B** N/A.

### 3. Score the 9 dimensions

Full band anchors (0–5, with concrete "looks like" descriptions and green/red flags) live in **[`rubric.md`](rubric.md)** — read it before every evaluation; do not score from memory. Summary:

| # | Dimension | Weight | The senior question it answers |
| --- | --- | :-: | --- |
| **A** | Problem Framing & Specification | 15 | Did they tell the agent _what_ and _why_, at the right altitude, with constraints + acceptance criteria? |
| **B** | Decomposition & Sequencing | 12 | Did they break work into coherent, dependency-ordered, verifiable increments? |
| **C** | Context Curation | 10 | Did they ground the agent in real artifacts (files, logs, outputs) instead of making it guess? |
| **D** | Verification & Skepticism | 18 | Did they treat output as a draft to check — run it, read it, catch errors — rather than rubber-stamp? |
| **E** | Debugging Collaboration | 10 | When it broke, did they give high-signal diagnostics (repro, exact error, expected-vs-actual, hypothesis)? |
| **F** | Course Correction & Steering Control | 10 | Did they catch drift early and redirect surgically, instead of letting bad paths run or thrashing? |
| **G** | Technical & Architectural Judgment | 15 | Do the prompts reveal engineering maturity — operability, failure modes, trade-offs, right-sizing? |
| **H** | Workflow & Tooling Leverage | 5 | Did they use the harness well (plan mode, subagents, git, slash/skills) and pick the right autonomy? |
| **I** | Outcome & Efficiency | 5 | Did the session reach a correct, working end state with economy proportional to scope? |

Each dimension gets an integer **0–5 band**, a one-paragraph **rationale**, and **≥2 quoted turn citations** (`U7: "i see the rows but they are empty"`). No citation → no score; write `INSUFFICIENT_EVIDENCE` and exclude from the total. Mark genuinely inapplicable dimensions `N/A` and renormalize over the remaining weight.

### 4. Apply gates (hard caps), then green flags

**Gates** cap the overall band regardless of dimension scores — see `rubric.md` → "Gates". They fire only on _confirmed_ unsafe steering (pasted live secrets, approved a destructive/irreversible op without checking, knowingly shipped broken output). Auto-detected red flags are candidates for a gate — confirm against the transcript before firing one, and record why.

**Green flags** are noted as evidence of strength but never push a dimension above its band's ceiling.

### 5. Compute totals and level

```
weighted = Σ(band_d / 5 × weight_d)  over scored dimensions
normalized = weighted / Σ(weight_d of scored dimensions) × 100
```

Map `normalized` to a **steering level** (after gate caps):

| Level | Normalized | Reads as |
| --- | :-: | --- |
| **Exemplary** | 85–100 | Operates the agent as a true force-multiplier; senior+ signal |
| **Strong** | 70–84 | Confident, effective, low-waste steering; clear senior signal |
| **Competent** | 55–69 | Gets results; real gaps in verification or framing; mid-level |
| **Developing** | 40–54 | Works, but reactive/under-specified/under-verified; junior steering |
| **Weak** | 0–39 | Ineffective or unsafe steering |

The level is a **calibration aid, not a verdict** — the human reviewer decides hire/no-hire. Present the level with the evidence that earned it, and call out the 1–2 dimensions that most moved it.

### 6. Write outputs (see [`TEMPLATE.md`](TEMPLATE.md))

Per session, write to `<out_dir>/<session-slug>/`:

- **`report.md`** — narrative: summary line, level, per-dimension table with citations, standout moments, weaknesses, gates/flags, and a "what a stronger operator would have done differently" section. Shape = TEMPLATE.md.
- **`scorecard.yaml`** — machine-readable: per-dimension band + rationale + citations, gates, green flags, totals, level.

When >1 session, also write `<out_dir>/candidate-rollup.md`:

- Per-session level + headline.
- Per-dimension trend across sessions (is verification consistently weak? is framing consistently strong?).
- **Consistency** read — does the candidate steer the same way under pressure (debugging) as when greenfield?
- A single calibrated **overall steering level** with the 3 strongest and 3 weakest pieces of evidence across all sessions, and concrete interview follow-up questions to probe the gaps.

## Scoring discipline (read before every run)

- **Evidence or it didn't happen.** Every band cites real turns. "Seems thorough" is not a finding; `U7–U11 debugged the empty-rows bug by pasting the live API response (U9) and isolating restart behavior (U11)` is.
- **Absolute bands, not curve.** A lone session is scored against the anchors in `rubric.md`, not against other candidates. With multiple sessions you may _report_ relative trends, but each session's band is absolute.
- **Don't confuse activity with skill.** A 40-turn session is not 40-turns-good. Ask what each turn _accomplished_. Penalize wasted loops; reward turns that each move a verifiable step.
- **Don't penalize the harness.** API 529s, tool-permission prompts, and model mistakes the candidate then caught are not the candidate's failures. Idle gaps may be the candidate thinking, testing out-of-band, or away — note, don't assume.
- **Charity with a spine.** Read terse prompts in the best plausible light (typos, brevity ≠ low skill), but do not invent verification or judgment the transcript doesn't show. If you can't see it, it isn't scored there.
- **Surface the model's contribution honestly.** If Claude proposed the good architecture and the candidate merely said "yes", that is weak **G** with strong-agent luck — say so. If the candidate's prompt _specified_ the operational concern, that is strong **G**.

## What good and weak steering look like (calibration vignettes)

- **Strong A+G+D:** Opens with a data model, the four API operations, and an explicit constraint ("use an in-memory DB"). Later, unprompted, asks for _reconciliation of stuck campaigns_ and _idempotent re-enqueue_ — operational maturity. Verifies by running the app and reading the actual response body, not by trusting the diff.
- **Strong E:** On a bug, pastes the literal API response (`U9`), then narrows it — "I see the rows but they are empty" → "restarting the server still shows old rows" — feeding the agent expected-vs-actual until root cause. Converges in a few turns.
- **Weak D:** Runs the whole session in an auto-accept mode and never once runs, reads, or questions the output; "make it work" → "still not working" → "ok ship it." No evidence the human ever confirmed correctness.
- **Weak F:** Lets the agent build for many turns down a path the candidate didn't want, then restarts from scratch instead of redirecting; or re-sends the identical prompt three times when it fails instead of adding information.
- **Gate fires:** Pastes a live API key into a prompt, or tells the agent to `git push --force` to a shared branch and approves it without looking — cap at **Developing** and flag, no matter how good the rest is.

## Notes

- The skill is read-only over the transcripts; it only writes into `out_dir`. Never mutate the source `.jsonl`.
- If a transcript is corrupt/empty or has zero human turns, write a stub report saying so rather than fabricating a score.
- Keep reports honest and specific enough that the candidate, if shown the report, would recognize the session. That is the bar for evidence quality.
