# interview-eval rubric — band anchors

Score each dimension **0–5** against the anchors below. Bands are **absolute** (not curved against other candidates). Every band needs ≥2 quoted turn citations. Mark a dimension `N/A` only when the session type makes it genuinely inapplicable (see the matrix at the end), and renormalize over remaining weight.

Read the whole file before scoring. The general meaning of the 0–5 scale:

| Band | Meaning |
| :-: | --- |
| **5** | Exemplary. The behavior a senior+ operator exhibits deliberately and consistently. |
| **4** | Strong. Mostly present and effective, minor lapses. |
| **3** | Competent. Present but inconsistent or shallow; gets by. |
| **2** | Developing. Mostly absent or reactive; visible cost to the session. |
| **1** | Weak. Counterproductive or near-absent. |
| **0** | None / actively harmful. |

When between two bands, pick the lower unless a green flag justifies the higher. Half-points are not used.

---

## A — Problem Framing & Specification (weight 15)

_Did the human tell the agent what to build and why, at the right altitude, with the constraints and acceptance criteria it needed?_ The opening prompt matters most, but framing recurs every time a new sub-task starts.

**Look for:** goal + intent ("why"), not just mechanics; explicit constraints (stack, storage, perf, scope); data shapes / interfaces / acceptance criteria; stated non-goals; the _right altitude_ — outcome-with-guardrails, neither a vague one-liner nor line-by-line dictation that wastes the agent.

- **5** — Specs are rich and structured: goal, constraints, data/interface shapes, and acceptance criteria are all present at the right altitude; non-goals stated; the agent has everything it needs and freedom in _how_. Re-framing on each new sub-task is equally crisp.
- **4** — Clear, well-scoped specs with constraints and shapes; occasionally leaves a gap the agent has to guess, but the guess space is small.
- **3** — Gets the core ask across but under-specifies constraints or acceptance; the agent fills meaningful gaps, sometimes wrongly, requiring a follow-up.
- **2** — Vague or mechanics-only ("add a function that does X") with little context; the agent frequently guesses; or over-dictates exact code, wasting the agent and inviting transcription errors.
- **1** — One-line asks with no context ("build the thing", "make the UI"); intent unclear.
- **0** — Unintelligible or contradictory framing.

**Green flags:** stating non-goals; giving a concrete example input/output; "for a demo, so don't…" right-sizing. **Red flags:** contradictory requirements within one prompt; dictating implementation while omitting the goal.

---

## B — Decomposition & Sequencing (weight 12)

_Did the human break the work into coherent, dependency-ordered, individually verifiable increments?_

**Look for:** a sensible build order (foundations before features, data before UI, happy-path before edge-cases); each turn advancing a bounded step you could check; dependencies respected; no big-bang dump that forces the agent to do everything blind; no chaotic re-ordering or contradictory re-scoping.

- **5** — Clean incremental arc: each step is bounded, builds on the last, and is independently verifiable; dependencies are respected; scope changes are deliberate and explained.
- **4** — Mostly well-sequenced with bounded steps; one or two steps bundled or slightly out of order.
- **3** — Reasonable overall shape but some steps too large to verify, or some avoidable back-tracking.
- **2** — Either everything-at-once mega-prompts, or scattershot ordering that creates rework; the agent often builds on unstable ground.
- **1** — Incoherent sequencing; frequent contradictory re-scoping; constant rework.
- **0** — No discernible decomposition.

**Green flags:** explicitly descoping then re-adding ("remove credentials for now"); building a foundation before layering on workers/UI. **Red flags:** asking for the UI before the data model exists; reversing a decision every few turns.

---

## C — Context Curation (weight 10)

_Did the human ground the agent in the real artifacts it needed instead of making it guess?_

**Look for:** pasting actual error text, logs, API responses, or stack traces; pointing at specific files (IDE-open context, `@file`, paths); referencing prior decisions to keep the agent coherent; managing the context window (not dumping irrelevant noise, not starving it); using CLAUDE.md / plan files / memory where appropriate.

- **5** — Consistently feeds precise, relevant context: real outputs and errors verbatim, exact file references, prior-decision callbacks; nothing the agent needs is left to guess, nothing irrelevant pollutes the window.
- **4** — Usually grounds the agent in concrete artifacts; occasional reliance on the agent to re-derive context it could have supplied.
- **3** — Provides context when prompted by failure but not proactively; mixes some vague description with concrete artifacts.
- **2** — Mostly asks the agent to operate blind ("it's not working") without supplying the evidence; or floods irrelevant context.
- **1** — Almost never grounds the agent; near-zero artifacts.
- **0** — No usable context at all.

**Green flags:** pasting a literal JSON response / log line to ground a bug; opening the relevant file in the IDE before asking about it. **Red flags:** describing an error in prose when the exact text was one paste away.

---

## D — Verification & Skepticism (weight 18 — the strongest senior signal)

_Did the human treat the agent's output as a draft to be checked — running it, reading it, probing it — rather than rubber-stamping?_ This is where seniors separate from juniors. Weight it accordingly.

**Look for:** running the app/tests and reporting results; reading diffs and pushing back on questionable choices; inspecting actual behavior/output; catching subtle or superficial work the agent slipped in; confirming a fix actually fixed it before moving on; not declaring "done" on faith.

> Permission mode is **not** the signal. `acceptEdits` / bypass mode is a legitimate pragmatic choice — judge whether the human _verified behavior_, by any means (running, testing, reading output, inspecting the UI), not whether they clicked approve on each edit.

- **5** — Systematic verification throughout: runs the thing, inspects real output, catches agent errors the agent missed, confirms each fix before advancing, and pushes back on shaky choices. Skepticism is the default stance.
- **4** — Verifies most consequential output and catches real issues; a few steps taken on faith.
- **3** — Verifies reactively (notices when something is visibly broken) but doesn't proactively probe; misses non-obvious issues.
- **2** — Little verification; mostly trusts the agent; "done" declared without evidence it works; bugs surface late.
- **1** — Effectively no verification; accepts and moves on regardless of correctness.
- **0** — Blindly ships output known or shown to be broken.

**Green flags:** "restarting the server still shows old rows" (confirming a fix _didn't_ take); reading the actual response body rather than the diff; spotting that the agent stubbed/faked something. **Red flags:** "looks good" with no evidence; shipping immediately after a failed run; never once running the code in a build session.

**Gate interaction:** if the human is _shown_ the output is broken and ships/commits it anyway, cap **D ≤ 2** and consider the BLIND-SHIP gate.

---

## E — Debugging Collaboration (weight 10)

_When something broke, did the human drive toward root cause with high-signal diagnostics, or just nudge vaguely?_

**Look for:** precise reproduction; exact error/output pasted; expected-vs-actual framing; stated hypotheses; isolating variables one at a time; converging instead of flailing. (In a pure debugging session this is the main event; in a clean build it may be minor or `N/A`.)

- **5** — Drives debugging like an engineer: repro + verbatim evidence + hypothesis, isolates one variable at a time, narrows steadily to root cause, confirms the fix. Tight convergence.
- **4** — Good diagnostics and mostly systematic; occasional vague nudge but recovers.
- **3** — Provides some evidence and gets there, but with avoidable thrash or symptom-level (not cause-level) reasoning.
- **2** — Mostly "still broken / fix it" with thin evidence; relies on the agent to guess; loops without narrowing.
- **1** — Repeats the same vague complaint; no evidence; no convergence.
- **0** — Actively misleads or abandons; lets the agent flail indefinitely.

**Green flags:** a chain like "fetching 2 campaigns but not executing" → "I see rows but they're empty" → "restart still shows old rows" — each turn adds a discriminating fact. **Red flags:** re-sending "it doesn't work" unchanged; never pasting the error.

---

## F — Course Correction & Steering Control (weight 10)

_Did the human catch the agent drifting and redirect surgically — and know when to refine vs. undo vs. restart?_

**Look for:** early detection of off-track work; minimal, targeted redirects that preserve good work; tightening constraints constructively ("run 10 at a time, spend 1s per task"); well-timed interrupts; choosing the cheapest correction (refine > undo > restart). Avoid: letting a wrong path run for many turns; chaotic yanking; re-sending an identical failing prompt; nuking good work to fix a small issue.

- **5** — Catches drift fast and corrects surgically; tightens scope constructively; interrupts are well-timed; always picks the cheapest effective correction.
- **4** — Generally redirects well and early; an occasional late catch or slightly heavy-handed correction.
- **3** — Corrects eventually but sometimes lets bad paths run, or over-corrects (restart when a tweak would do).
- **2** — Frequently late to catch drift; corrections are blunt (restart-heavy) or repetitive.
- **1** — Lets the agent run far off-track repeatedly; re-sends failing prompts unchanged.
- **0** — No steering control; the session runs the human, not the reverse.

**Green flags:** a precise interrupt that saves a wrong path; "kill the docker process and only run redis" (decisive, scoped). **Red flags:** identical prompt sent 3× hoping for a different result (distinguish from harness retries / auth hiccups, which are not the candidate's fault).

---

## G — Technical & Architectural Judgment (weight 15)

_Do the prompts themselves reveal engineering maturity?_ This reads the **content** of the human's instructions for evidence of senior thinking — independent of whether the agent executed well.

**Look for:** anticipating operational concerns unprompted (idempotency, reconciliation of stuck state, observability/logging, backpressure, failure/retry, race conditions, data consistency); sound trade-offs; right-sizing to context (descoping auth for a demo; not gold-plating); awareness of the system beyond the happy path; sensible tech choices for the stated constraints.

> Attribute carefully: if _Claude_ proposed the mature design and the human only assented, that is weaker **G** than if the human's prompt _specified_ it. Note who originated the judgment.

- **5** — Prompts repeatedly surface senior-level concerns the agent wouldn't have without being told — operability, failure modes, consistency — and make sound, explicit trade-offs right-sized to the context.
- **4** — Clear engineering maturity in most prompts; raises real operational concerns; a few missed opportunities.
- **3** — Solid happy-path thinking; some good instincts (e.g. observability) but misses obvious operational concerns until forced.
- **2** — Mostly happy-path; accepts naive designs; little evidence of system-level thinking.
- **1** — Prompts reveal weak fundamentals; misses obvious risks even after they bite.
- **0** — Actively poor judgment driving the design.

**Green flags:** asking for reconciliation of campaigns "not updated for some time"; a self-perpetuating queue (enqueue-next-batch) pattern; adding logging at the right seams; descoping credentials for a demo. **Red flags:** designing UI-first with no data model; ignoring concurrency in an obviously concurrent system; no failure handling anywhere.

---

## H — Workflow & Tooling Leverage (weight 5)

_Did the human use the agent and harness well, and pick the right level of autonomy?_ Credit when applicable; **do not penalize** a session that legitimately didn't need advanced features.

**Look for:** plan mode for genuinely ambiguous/large work; subagents for parallel or large-scope fan-out; slash commands / skills; sane git hygiene (branch, review, meaningful commits); MCP/CI where useful; matching autonomy to risk (high autonomy on safe mechanical work, tighter control on risky/irreversible work); economy of interaction.

- **5** — Deliberately reaches for the right tool/mode at the right moment (plan mode to de-risk ambiguity, subagents to parallelize, clean git flow), and tunes autonomy to risk. Visibly fluent with the harness.
- **4** — Good, idiomatic use of the harness; misses an opportunity or two for leverage.
- **3** — Competent single-mode use; gets the job done without reaching for available leverage even where it would clearly help.
- **2** — Clumsy or mismatched tooling (e.g. reckless full-auto on risky ops, or heavy micromanagement of safe work).
- **1** — Fights the tool; ignores obvious leverage; risky autonomy with no guardrails.
- **0** — Tooling actively works against the goal.

**Green flags:** plan mode before a big ambiguous build; subagents for an audit sweep; a clean commit/push flow with a written README. **Red flags:** bypass-all-permissions on destructive ops; no version control on substantial work. _(Absence of advanced features in a small, well-handled task is fine — score around 3, not down.)_

---

## I — Outcome & Efficiency (weight 5)

_Did the session reach a correct, working end state, with economy proportional to scope?_ A capstone — partly downstream of the other dimensions, so weighted light to avoid double-counting.

**Look for:** a working, correct result the human verified; effort (turns/time/tokens) proportional to the scope; little wasted looping; a clean stopping point (committed/pushed/documented).

- **5** — Clear goal achieved and verified; efficient path; minimal waste; clean close (e.g. committed + README).
- **4** — Goal achieved; minor inefficiency or a loose end.
- **3** — Mostly working outcome; noticeable wasted effort or unverified gaps.
- **2** — Partial or unverified outcome; significant waste relative to scope.
- **1** — Little working outcome; heavy waste.
- **0** — No working outcome.

**Note:** scope-adjust. Achieving a small task crisply is a 5; half-finishing a huge one with good reason may still be a 3–4. Don't reward raw volume.

---

## Gates (hard caps on the overall level)

Gates fire only on **confirmed** behavior in the transcript (use auto red-flags as leads, then verify). When a gate fires, cap the overall level as stated, record the gate, and quote the turn. Multiple gates stack downward.

| Gate | Fires when | Cap |
| --- | --- | --- |
| **SECRETS-LEAKED** | The human pastes a live secret (API key, token, private key, real password) into a prompt, or directs the agent to commit one, and doesn't immediately remediate. | ≤ **Developing** |
| **DESTRUCTIVE-UNCHECKED** | The human approves an irreversible/destructive op (`rm -rf` on a real path, `git push --force` to a shared branch, `DROP TABLE`/prod data loss, force-pushing over others' work) without verifying intent/scope. | ≤ **Developing** |
| **BLIND-SHIP** | The human is _shown_ output is broken (failed run, wrong result) and ships/commits/declares-done anyway, with no remediation. | **D ≤ 2** and overall ≤ **Competent** |
| **PROMPT-INJECTION-NAIVE** | The human pipes untrusted external content into the agent as instructions and acts on it without scrutiny in a context where that's dangerous. | ≤ **Competent** |

A gate is about _unsafe steering_, not mere mistakes. A bug the candidate then caught is **not** a gate — it's evidence _for_ dimension D. Distinguish "the harness/model did X" from "the human directed X": gates require human direction or human ratification.

---

## Green-flag ledger (evidence of strength; never inflates beyond a band ceiling)

- **Empirical debugging** — grounds bugs in verbatim outputs/logs (→ C, E).
- **Unprompted operability** — idempotency, reconciliation, observability, backpressure, retries asked for without being bitten first (→ G).
- **De-risking with plan mode** — uses plan mode before an ambiguous/large build (→ H, A).
- **Wise descoping** — cuts non-essential scope explicitly for the context ("demo, drop auth") (→ A, G).
- **Catching the agent** — spots stubbed/faked/superficial work or a regression the agent introduced (→ D).
- **Decisive surgical redirect** — a single precise interrupt/constraint that saves a path (→ F).

---

## Session-type matrix — which dimensions are primary / secondary / N-A

Classify the dominant type; blends are common — note them. "Primary" = weigh heavily and expect strong evidence; "Secondary" = score if present; "N/A" = exclude and renormalize unless the session actually exercises it.

| Type | Primary | Secondary | Often N/A |
| --- | --- | --- | --- |
| **greenfield-build** | A, B, G, D | C, F, I | — |
| **feature-add** | A, D, G | B, C, F, I | — |
| **debugging** | E, D, C | F, G | B (little to decompose) |
| **refactor / cleanup** | D, G, B | C, F, I | E |
| **ops / infra** | A, D, G, H | C, F, I | B, E |
| **research / Q&A** | A, C, F | G, I | B, D (no code to verify), E, H |
| **review / audit** | C, G, D, H | A, F | B, E, I |

If the user states a different intent (e.g. "this was a time-boxed take-home"), let that adjust expectations (efficiency matters more, polish less) but not the band anchors themselves.
