---
name: evaluator
description: "Use when scoring a single software project for production-readiness + product quality. Takes a repo + output dir; writes a per-project workspace (README, `scorecard.yaml`, flat `raw/`) scoring 8 sub-criteria (4 Code + 4 Product) to /100 with evidence and gates. One project, no cohort."
---

# Evaluator

Single-project evaluator. One project in → one per-project workspace out. The caller supplies the output directory; the skill writes `<out_dir>/<project>/{README.md, scorecard.yaml, raw/<sub-criterion>-<artifact>.ext}` and returns that path.

This skill scores **exactly one project**. It does not rank, normalize across projects, weight sub-criteria differently, or know anything about a batcher, a cohort, or other projects. Every band it assigns is **absolute** — computed from per-signal thresholds in each recipe — so the scorecard for a project is identical whether the project is scored alone or as one of fifty. Cross-project weighting and comparison, if wanted, is a separate concern layered on top by whoever calls this skill; the skill itself stays unaware of it.

## Preflight (run once per host, idempotent)

This skill depends on external tools (`gitleaks`, `lizard`, `scc`, `semgrep`, `hadolint`, `ruff`, `biome`, `knip`, `deptry`, `shellcheck`, `golangci-lint`, `gh`, `jq`, `pip-audit`, `govulncheck`, `trivy`). Before scoring, run [`preflight.md`](preflight.md). It checks each tool, installs any that are missing (via `brew` / `npm -g` / `pipx` / `go install`), and fails loud if an install fails. **Do not silently fall back to grep-or-LLM-only substitutes** — that poisons the gate checks (VULN / SECRETS / AUTH / OBSERVABILITY) and makes the scorecard dishonest. If preflight exits non-zero, return a clear failure message naming the missing tools.

## Hard rule: every slow command gets a timeout

Every command the skill runs that can hang or burst must be wrapped in a bounded wall-clock guard. Unbounded retries are how a single project eats 20+ minutes. Apply one of:

- `timeout 60 <cmd>` for any `gh api`, `semgrep`, `npm install`, `pip install`, `go build`, `go vet`, or `curl` — even when the recipe doesn't show it. If the recipe shows a naked command, wrap it yourself.
- `--maxdepth N` and `-not -path '*/node_modules/*' -not -path '*/.venv/*' -not -path '*/.git/*'` on every `find`.
- `git log -n 500` (or similar ceiling) instead of walking full history. Never `git log --all` without a limit.
- `semgrep`: always pass `--timeout 30` (per-rule), `--disable-version-check`, `--metrics=off`. Avoid `--config p/code-quality` — that registry entry 404s in retry-storm; use per-language packs (`p/python`, `p/typescript`, `p/javascript`, `p/go`) or `p/default`.
- Dep-vuln scans: `pip-audit --timeout 120`, `govulncheck` in a 180 s `timeout`, `trivy image --timeout 180s`. `npm audit --json` doesn't take a flag; wrap the whole call in `timeout 120`.

If a single mechanical command still exceeds ~2 min, something is wrong with the project (giant `node_modules/`, generated bundle, vendored deps). Skip it with `_NOT_RUN: true, reason: "timeout on <cmd>"` rather than retrying — a sparse signal is tolerable; a hang is not.

## Input

A repo reference + an output directory:

- **Local clone path** — e.g. `/path/to/project/`. Used by filesystem reads (Semgrep, grep, file sampling, git analytics). This is the primary input.
- **GitHub ref** (optional) — e.g. `<org>/<repo>`. Used by `gh api` calls for PR / code-review history when the project is hosted on GitHub. Absence just means those signals are skipped, not a failure.
- **`out_dir`** — caller-supplied base directory. The skill writes into `<out_dir>/<project>/` and returns that path. The skill creates the folder; `raw/` is flat — one file per signal named `<sub-criterion>-<artifact>.ext`, no subdirectories.

`<project>` is a stable slug for the project (its directory name is fine). Optional baseline stats (kLoC, commit count, contributor count) may be supplied to skip a pre-pass; otherwise derive them from the local clone.

## Output — per-project workspace

Every invocation writes a folder at `<out_dir>/<project>/` with two parts plus raw dumps:

```
<out_dir>/<project>/
├── README.md            # Human-readable narrative. Shape: § "README shape" below.
├── scorecard.yaml       # Structured scorecard (below). Machine-readable.
└── raw/                 # Flat directory — every file named `<sub-criterion>-<artifact>.ext`
    ├── code-quality-*            # llm-sample.jsonl, lint-{ruff,biome,eslint,golangci}.json, semgrep-quality.json, cr-comments.jsonl
    ├── testing-*                 # tests-inventory.jsonl, llm-sample.jsonl, semgrep-resilience.json, gha-runs.json, auth-scan-{http,ws,grpc,trpc}.json
    ├── deployment-readiness-*    # artifacts.json, hadolint.json, localhost-refs.txt, observability-markers.json, llm-read.md
    ├── maintainability-*         # gitleaks.json, lizard.csv, knip.json|deptry.json|gomod-diff.txt, file-sizes.txt, tech-debt.txt, dep-vuln-{pip-audit,npm-audit,govulncheck,trivy}.json, llm-read.md
    ├── problem-substantiveness-* # llm-read.md, proxies.json
    ├── ux-mechanics-*            # path-detection.json + path-specific files (ux-heuristics, routes, cli-probes, shellcheck, openapi-quality, api-surrogate), llm-read.md
    ├── value-articulation-*      # llm-read.md, proxies.json
    └── mechanical-ambition-*     # scc.json, stubs.json, git-analytics.json, tier-a-build.log, llm-read.md
```

`raw/` is a **flat directory** — no subfolders. Every file is named `<sub-criterion>-<artifact>.ext` (e.g. `code-quality-semgrep-quality.json`, `maintainability-deptry.json`). A signal that couldn't be extracted (tool absent, language not present, repo shape mismatch) produces **no file** — absence is the signal, and `scorecard.yaml` flags the gap. Per-recipe artifact lists are authoritative in each recipe's "Raw dumps" section.

### Scorecard shape

```yaml
project: my-project

# Run metadata (also rendered into README.md)
commit_sha_pinned: 7423b5c... # HEAD at extraction — audit only
rubric_version: 1             # bump when sub-criteria / gates / banding change (see "Rubric version" below)
scored_at: 2026-06-15T23:10:00Z # ISO 8601 UTC timestamp of this run
agent_model: claude-opus-4-8  # model that executed this pass (or null)
approx_tokens: 92000          # rough total tokens consumed (or null)

# 1.x — Production-Grade Code (12.5 pts each)
code_quality: { score: 4, evidence: "...", data: { ... } }
testing: { score: 2, evidence: "...", data: { ... } }
deployment_readiness: { score: 3, evidence: "...", data: { ... } }
maintainability: { score: 3, evidence: "...", data: { ... } }

# 2.x — Product Thinking (12.5 pts each)
problem_substantiveness: { score: 4, evidence: "...", data: { ... } }
ux_mechanics: { score: 4, evidence: "...", data: { ... } }
value_articulation: { score: 3, evidence: "...", data: { ... } }
mechanical_ambition: { score: 3, evidence: "...", data: { ... } }

gates_triggered: [UNTESTED, OBSERVABILITY]
short_circuited: false # true when ≥3 gates stack; see "Dead-project short-circuit"
scored_sub_criteria: 8 # items with a non-null score
total_sub_criteria: 8  # 4 Code + 4 Product

total: 66.25     # Σ (score × 2.5) over scored items; max 100
available: 100.0 # max points from currently-scored items (lower if any item is null)
normalized: 66.25 # total ÷ available × 100
```

**Each item scored 0–5** → points = `score × 2.5` (every item is weighted 12.5 pts, so the step is 2.5). `total` is the sum over scored items; `available` is the max points those scored items could have reached (12.5 × count); `normalized = total ÷ available × 100`. With all 8 items scored, `available == 100.0` and `total == normalized`. If a sub-criterion is genuinely unscoreable (`score: null`), it drops out of both `total` and `available` so the percentage stays honest.

The `total` here is an **equal-weight** convenience score for reading one project on its own. The skill does not apply per-sub-criterion weights — it has no opinion on relative importance. A caller that wants weighted, cross-project-comparable scores reads the raw 0–5 `score` fields and applies its own weighting; that logic lives entirely outside this skill.

### Workspace-writing order

Write the outputs in this order — each depends on the prior:

1. **`raw/<sub-criterion>-<artifact>.ext`** — run each recipe (see table below), dump raw tool output verbatim to the flat `raw/` directory with per-recipe filename prefixes. This is the extraction pass.
2. **`scorecard.yaml`** — fold each recipe's aggregates into the sub-criterion block, apply gates, compute `total` / `available` / `normalized`. Pin `commit_sha_pinned` (audit metadata) and `rubric_version`.
3. **`README.md`** — render the README shape below with values drawn from `scorecard.yaml` and pointers to `raw/` artifacts. The raw-dump section is a pointer index, not a re-embedded dump.

### README shape

`README.md` is a self-contained human narrative — no external template. Render these sections from `scorecard.yaml`:

1. **Title + verdict** — `# <project> — <normalized>/100` and a 2–3 sentence overall read.
2. **Run metadata** — `rubric_version`, `scored_at`, `agent_model`, `commit_sha_pinned`, gates triggered, short-circuit flag.
3. **Scorecard table** — one row per sub-criterion: `| Sub-criterion | Band (0–5) | Points / 12.5 | One-line evidence |`, split into a Code block and a Product block, with the `total / available / normalized` footer.
4. **Per-sub-criterion detail** — for each item, the band, the full `evidence` string, and the key numbers from `data`.
5. **Raw artifact index** — a pointer list of the `raw/<sub-criterion>-*` files that were produced (and a note on any that were skipped and why).

### Rubric version

`rubric_version` is a simple integer marker recorded on each scorecard. Bump it whenever a recipe adds/drops a signal, a gate threshold moves, a banding table changes, or the scorecard shape changes. It exists so a reader can tell which rubric produced a scorecard and so a caller can decide whether a cached workspace is stale. It is **not** a cohort cache key and carries no cross-project semantics — one project, one rubric version, decided at extraction time.

## Scoring model

- **100 points total, 50 per half:**
  - **1.x Production-Grade Code** — 4 sub-criteria × 12.5 pts each = **50 pts**
  - **2.x Product Thinking** — 4 sub-criteria × 12.5 pts each = **50 pts**
- **Each item scored 0–5** → points = `score × 2.5`.
- **Scorecard reports `total`, `available`, `normalized`.**

### Bands (absolute)

| Band  | Meaning                                                          |
| ----- | --------------------------------------------------------------- |
| **5** | Excellent — strong across all recipe criteria, no material gaps |
| **4** | Strong — strong on most criteria, one minor gap                 |
| **3** | Solid — meets criteria in basic form, material gaps present     |
| **2** | Weak — thin execution, multiple criteria missing or shallow     |
| **1** | Token effort — minimal execution, most criteria absent          |
| **0** | Absent                                                          |

Bands are **absolute, not relative to any other project**. Per-signal numeric thresholds live in each recipe's _Banding_ section. The rubric is single-project-invokable by design — there is no cohort pass, no quintile math, no re-banding step.

### Gates (cap specific items; stack additively)

| Gate | Trigger | Effect |
| --- | --- | --- |
| `DEMO` | (`has_container_artifact == false` AND `has_any_deploy_artifact == false` AND `readme_has_run_section == false` AND `llm_has_run_section == false`) OR (README claims a deploy artifact — Dockerfile / compose / k8s / target — that does not exist in the tree) | `deployment_readiness ≤ 1` |
| `BROKEN` | Tier-A build check fails (`ruff` / `tsc --noEmit --skipLibCheck` / `go vet` / `py_compile`) — NOT on `DEPS_NOT_INSTALLED` errors | `mechanical_ambition ≤ 2` |
| `UNTESTED` (hard) | `test_file_count == 0` | `testing ≤ 1` |
| `TESTS_TRIVIAL` (soft) | Tests exist but `assertion_density < 1.0` per test file OR `trivial_body_ratio > 0.8` | `testing ≤ 2` |
| `SECRETS` | `gitleaks_findings ≥ 1` — any secret in git history, even rotated out of HEAD | `maintainability ≤ 2` |
| `VULN` | Dep-scan reports ≥ 1 CVE at CRITICAL OR ≥ 3 at HIGH across any dep (direct or transitive). Tools: `pip-audit` / `npm audit --json` / `govulncheck` / `trivy image`. Skipped if no recognized package manifest. | `maintainability ≤ 2` |
| `AUTH` | Network-routed handlers exist AND ≥ 1 mutating endpoint is not behind an auth middleware/dependency/decorator. Covers HTTP (FastAPI/Express/Flask/Gin/Django POST/PUT/PATCH/DELETE), WebSocket (ws/Socket.io/SignalR event handlers), gRPC (grpc/grpcio unary/streaming), tRPC (@trpc/server mutations). Skipped if no routed handlers detected. | `testing ≤ 2` |
| `OBSERVABILITY` | Structured-logging library absent (no structlog / winston / pino / log-slog / zap AND no JSON log format detected in code) **AND** metrics/tracing library absent (no OpenTelemetry, Prometheus client, Datadog agent, statsd client, or custom `/metrics` endpoint). Both must be absent for the gate to fire. | `deployment_readiness ≤ 3` |

Gates cap the **band** of the named sub-criterion. When multiple gates cap the same item, the lowest cap wins. Gates are recorded in `gates_triggered`; a reader (or a downstream caller) can see exactly which caps applied.

### Dead-project short-circuit

If **3 or more gates** fire on one project, skip the interpretive LLM passes for all sub-criteria, emit a fixed floor scorecard, and set `short_circuited: true`:

- **Non-gate-capped items**: `score: 1, evidence: "Floor — 3+ gates triggered; interpretive pass skipped"`.
- **Gate-capped items**: take the cap value (e.g. `deployment_readiness: 1` when DEMO fires, `testing: 1` when UNTESTED fires, `maintainability: 2` when SECRETS or VULN fires, `deployment_readiness: 3` when OBSERVABILITY fires). If multiple gates cap the same item, use the lowest cap.
- **Totals**: `total`, `available`, `normalized` computed normally over the floored items.

Rationale: a project that's undeployable + untested + broken is clearly dead. Running the expensive interpretive pass produces generous LLM hallucinations on dead code. Short-circuit saves compute and keeps the score honest.

### CodeRabbit — calibrating modifier with null-data semantics

If the project's GitHub history includes CodeRabbit review comments, they act as a **calibrating modifier on 1.1 (Code Quality) and 1.2 (Testing)**, never a primary signal.

- When a project has CR data (`cr_pr_count > 0` **and** non-empty actionables): `cr_*_per_kloc` values feed the recipe's banding table as calibrating modifiers.
- When a project has no CR data (no CR history, or walkthrough-only with `cr_actionable_total == 0`): every `cr_*_per_kloc` is emitted as `null`. The scorer treats `null` as a no-op; 1.1 / 1.2 bands are decided by the remaining signals.

Design principle: `null ≠ 0`. `null` means "no CR data, skip the modifier"; `0` means "CR reviewed the PR and found nothing" — a real signal pulling the band up. Projects not hosted on GitHub, or with no PRs, simply emit `null` and the modifier is a no-op.

## Sub-criterion extraction recipes

One recipe file per sub-criterion. All 8 are AI-scored from repo evidence.

| # | Item | Half | Contender | Recipe |
| --- | --- | --- | --- | --- |
| 1.1 | Code Quality | Code | LLM interpretive (primary) + lint (`ruff` / `biome` / `golangci-lint`) + `lizard` (per-language CCN: Py 10 / JS-TS 10 / Java 12 / Go 15) + `semgrep` quality packs + CodeRabbit (calibrating modifier) | [criteria/code-quality.md](criteria/code-quality.md) |
| 1.2 | Testing & Resilience | Code | Filesystem + GHA + Semgrep + LLM (load-bearing) + multi-protocol AUTH route-scan (HTTP/WS/gRPC/tRPC); CodeRabbit (calibrating modifier) | [criteria/testing.md](criteria/testing.md) |
| 1.3 | Deployment-Readiness | Code | Filesystem + Hadolint + observability markers + LLM | [criteria/deployment-readiness.md](criteria/deployment-readiness.md) |
| 1.4 | Maintainability | Code | gitleaks + lizard (per-language) + knip/deptry + dep-vuln scan (pip-audit/npm audit/govulncheck/trivy) + LLM | [criteria/maintainability.md](criteria/maintainability.md) |
| 2.1 | Problem substantiveness | Product | LLM-primary README + sample data + target-user read. Detects toy-problem patterns. | [criteria/problem-substantiveness.md](criteria/problem-substantiveness.md) |
| 2.2 | UX mechanics | Product | 3-path: web-UI (UX heuristics + routes) / CLI (clig.dev + shellcheck) / API-only (OpenAPI quality + versioning + idempotency + rate-limiting + input validation — **no cap, backends score up to 5**) + shared LLM UX read | [criteria/ux-mechanics.md](criteria/ux-mechanics.md) |
| 2.3 | Value articulation | Product | README parse for quantified impact claims + scope coverage ratio | [criteria/value-articulation.md](criteria/value-articulation.md) |
| 2.4 | Mechanical ambition | Product | scc + Semgrep stubs + git analytics + Tier-A build + LLM README-vs-shipped + COCOMO | [criteria/mechanical-ambition.md](criteria/mechanical-ambition.md) |

## Workflow

The skill scores one project end-to-end:

1. **Preflight** — run [`preflight.md`](preflight.md). Halt with a clear error if it fails.
2. **Mechanical extraction** — per recipe, run the shell commands under each recipe's "Extraction" section. Dump every tool output **verbatim** to the flat `<out_dir>/<project>/raw/` directory using the `<sub-criterion>-<artifact>.ext` naming convention. A signal that can't be extracted produces **no file** — absence is the signal.
3. **LLM interpretive passes** — per recipe, run the LLM reads (sample files, score axes, README parse, shipped-vs-promised diff). Write `raw/<slug>-llm-sample.jsonl` or `raw/<slug>-llm-read.md` per sub-criterion. These are the load-bearing interpretive judgments.
4. **Apply gates** — evaluate each gate (DEMO / BROKEN / UNTESTED / TESTS_TRIVIAL / SECRETS / VULN / AUTH / OBSERVABILITY) using the triggers above. If **≥3 gates** fire, emit the dead-project short-circuit floor scorecard (`short_circuited: true`) and skip remaining interpretive passes.
5. **Emit `scorecard.yaml`** — apply each recipe's absolute-threshold band table directly. Record the run-metadata block (`commit_sha_pinned`, `rubric_version`, `scored_at` via `date -u +"%Y-%m-%dT%H:%M:%SZ"`, `agent_model`, `approx_tokens`). Compute `total` / `available` / `normalized`.
6. **Render `README.md`** — fill the README shape above with values from `scorecard.yaml` and pointer references to `raw/` artifacts.
7. **Return the workspace path** — `<out_dir>/<project>/`. Write the files and leave them uncommitted; whoever invoked the skill decides whether to commit.

### Scorecard schema validator

[`validate_scorecard.py`](validate_scorecard.py) — a Python script that asserts a `scorecard.yaml` has every required top-level field (including `total` / `available` / `normalized`), all 8 sub-criterion blocks, and `score` + `evidence` + `data` on each. Exits 0 on valid; exits 1 on first failure with a specific error on stderr. Run it after emitting a scorecard to catch shape errors early. **Extending:** if a scorecard-shape bump introduces a new required field, append it to `REQUIRED_TOP` in the validator.

## Not in scope

- **Cross-project comparison, weighting, or ranking.** The skill emits raw 0–5 bands and an equal-weight `total`. Anything that compares projects, applies non-uniform weights, or normalizes a set against itself is a layer above this skill and must not leak into it.
- **Runtime verification.** The skill reads the repo statically. Actually running `docker compose up` / `npm run dev` and clicking through the product is a separate manual pass a reviewer may add on top; the skill does not drive it.
