# hr — AI-Steering Evaluation harness

This repo evaluates **how a person steered a Claude Code session** — for interview and hiring assessment. The harness reads session transcripts (`.jsonl`) and produces a calibrated, evidence-anchored read on the candidate's ability to drive an AI coding agent.

The work is the `interview-eval` skill (its `rubric.md` holds the band anchors, gates, and session-type matrix; `extract_session.py` is the deterministic extractor; `TEMPLATE.md` holds the output shapes). This `CLAUDE.md` is the **policy** over that skill: the pipeline, the discipline that keeps scores honest, and the standing prohibitions.

## Two evaluation modes

Pick the mode from what the user hands over:

- **Session steering** (`.jsonl` transcripts) — score how the candidate drove the agent. That is the `interview-eval` skill and the pipeline below; it is the default and the bulk of this file.
- **Code-project quality** (one or more repos / paths) — score the project the candidate produced. For a *set* of projects, follow the **`code-eval-batcher`** rule (it fans the `evaluator` skill across them and ranks the result); for a single project, invoke the `evaluator` skill directly. The session-steering discipline below does not apply to this mode — the skill and the rule are self-contained.

---

## The one rule above all others

**Score the human, not the model.** Claude writing clean code earns the candidate nothing; Claude writing buggy code costs them nothing — _unless_ the candidate's prompt caused it or they failed to catch it. Reward the prompt that set the agent up, the verification that caught its mistake, the redirect that pulled it off a bad path. Turn count, prompt length, and verbosity are **not** quality — precision, judgment, and verification are. Re-read this before every scoring pass.

---

## Step 0 — first run

If the `interview-eval` skill isn't resolvable (no `.claude/skills/interview-eval/`), this project isn't set up — tell the user to run `npx --yes @nurix/etna --name=hr` (or point at an existing copy of the skill) before continuing. Confirm `python3` is on PATH; the extractor is stdlib-only and needs nothing else.

## Step 1 — intake interview

Before scoring anything, establish scope with a short `AskUserQuestion` round (skip a question if the user already answered it):

1. **What am I evaluating?** → one candidate (one or many sessions) · several candidates · a cohort.
2. **What context governs expectations?** → the candidate's target role/seniority, the task brief they were given, and any time box. This adjusts _expectations_ (a time-boxed take-home weights efficiency up, polish down) — it never changes the rubric's absolute band anchors.
3. **Where do the transcripts live, and where should output go?** → input `.jsonl` paths / a folder, and an `out_dir` (default `./interview-eval-out/`).

Keep it to those. Carry the answers forward; don't re-interview per session.

## Step 2 — run the pipeline (delegate to `interview-eval`)

For each candidate, invoke the `interview-eval` skill. It owns the mechanics; this policy drives it and holds the discipline below. The flow it runs:

| Step | Mechanics | Output |
| --- | --- | --- |
| **Extract** | `python3 extract_session.py SESSION.jsonl --out <out_dir>/<slug>/ --json` — read the FULL evidence pack before scoring; open raw `.jsonl` around any quoted turn | evidence pack in `<out_dir>/<slug>/` |
| **Classify** | dominant session type (greenfield / feature / debugging / refactor / ops / research / review) | which dimensions are primary / secondary / N/A |
| **Score** | 9 dimensions × 0–5 absolute bands, ≥2 quoted turn cites each; gates; green flags | `<out_dir>/<slug>/{report.md, scorecard.yaml}` |
| **Roll up** | if >1 session: dimension trends + consistency + one calibrated level | `<out_dir>/candidate-rollup.md` |

After the rollup, **always produce a one-page candidate brief** (verdict + level, the standout evidence, the single biggest risk, and 3–4 interview follow-ups that would confirm or kill that risk). That brief is the deliverable a hiring manager reads first.

## Step 3 — calibration discipline (non-negotiable)

- **Evidence or it didn't happen.** Every band cites real `U<n>` turns with short verbatim quotes. "Seems thorough" is not a finding.
- **Absolute bands, not a curve.** Score each session against `rubric.md`'s anchors. With multiple candidates, relative ranking may be _reported_ — but each band is assigned absolutely.
- **Exclude harness noise.** API 529s / connection drops are Anthropic-side — never penalize them. Tool-permission prompts and model slips the human then caught are not candidate failures. Distinguish a sharp early interrupt (good) from thrash (bad) by reading the surrounding turns.
- **Gates require human direction.** A gate fires only on confirmed _unsafe steering_ the human directed or ratified (pasted a live secret, approved a destructive op unchecked, knowingly shipped broken output) — not on a model mistake. Verify the auto-detected red-flag leads against the transcript before firing one.
- **Charity with a spine.** Read terse prompts and typos in the best plausible light, but never invent verification or judgment the transcript doesn't show.

## Standing prohibitions

- **The hire decision stays with the human.** The output is a level and a brief; frame the level as a calibrated lens, never a verdict.
- **Source transcripts are never mutated.** The skill is read-only over `.jsonl`; writes land only in `out_dir`.
- **Claude's code quality is never scored as the candidate's.** If the agent proposed the good design and the human only assented, say so — that's weaker judgment than originating it.
- **A score is never fabricated from thin evidence.** A one-turn fragment or an empty/corrupt transcript gets a stub that says so and defers to the rollup — not an invented band.

Every output lands under `out_dir` and is plain markdown/YAML — reviewable, quotable, and reversible.

## Delegation

- **Delegate to keep bulk tool output out of the main context.** A sub-agent's raw output dies with it; the same output read inline is re-read and re-billed on every later turn.
- **Batch code-project scoring is the canonical fan-out here** — the `code-eval-batcher` rule dispatches the `evaluator` skill across projects in parallel, one project per sub-agent. That per-project scoring is exactly the independent, repetitive work this section exists to push out of the main context; nothing here narrows it.
- **Session-steering scoring stays inline.** Reading a full transcript for evidence and banding it against `rubric.md` is judgment work, not mechanical fan-out — never delegate a scoring pass by reflex; a second opinion is for a high-stakes call or an explicit request.
- **A delegated result is a report, not verifiable ground truth** — a returned scorecard or rollup still needs its own validation (`validate_scorecard.py`) before it's trusted.
- **Tiers**: mechanical fan-out (batch per-project dispatch) runs on the cheapest capable tier; judgment-heavy delegation runs on the standard tier; the strongest tier fires only on explicit user request.
- **An explicit user directive about tier, cost, or delegation overrides this section.** No dispatch mechanism ⇒ inline.
