# Evaluator — Preflight

> **Run before scoring a project.** This skill depends on external tools that must exist on the host. If a tool is missing, **install it** — do not silently fall back to grep / LLM-only. A silent fallback poisons the gate checks (VULN / SECRETS / AUTH / OBSERVABILITY) and makes the scorecard dishonest.

The evaluator is allowed to install its own tools via `brew` / `pipx` / `npm -g`. Installing is cheap (~1 minute total) and idempotent. Ignoring a missing tool is not.

---

## Required tools

| Tool | Purpose | Recipes that use it | Install (macOS, primary) |
| --- | --- | --- | --- |
| `gh` | GitHub API (CR scrape, GHA runs, PR enumeration) | 1.1, 1.2, 2.4 | `brew install gh` |
| `jq` | JSON parsing in shell pipelines | all | `brew install jq` |
| `gitleaks` | committed-secret scan — fires `SECRETS` gate | 1.4 | `brew install gitleaks` |
| `lizard` | polyglot complexity / god-function detection (per-language CCN: Py 10 / JS-TS 10 / Java 12 / Go 15) | 1.1, 1.4 | `brew install lizard` |
| `scc` | LoC + cocomo estimate (replaces tokei/cloc) | 1.4, 2.4a | `brew install scc` |
| `semgrep` | polyglot static analysis (quality + resilience + multi-protocol auth) | 1.1, 1.2, 2.4a | `brew install semgrep` |
| `hadolint` | Dockerfile lint | 1.3 | `brew install hadolint` |
| `ruff` | Python lint + formatter | 1.1 (Python repos), 2.4a (tier-A) | `brew install ruff` |
| `shellcheck` | shell-script lint | 2.2a (CLI path) | `brew install shellcheck` |
| `tree` | folder-tree read for LLM context | 1.4, 2.2a, 2.4a | `brew install tree` |
| `golangci-lint` | Go aggregate linter | 1.1 (Go repos) | `brew install golangci-lint` |
| `biome` | TS/JS lint + format (preferred over eslint for speed) | 1.1 (TS/JS repos) | `npm i -g @biomejs/biome` |
| `knip` | TS/JS dead-code + unused-deps | 1.4 | `npm i -g knip` |
| `pipx` | isolated install for Python CLIs below | — | `brew install pipx && pipx ensurepath` |
| `deptry` | Python unused-deps | 1.4 | `pipx install deptry` |
| `pip-audit` | Python dep-vuln scan — feeds `VULN` gate | 1.4 (Python repos) | `pipx install pip-audit` |
| `govulncheck` | Go dep-vuln scan — feeds `VULN` gate | 1.4 (Go repos) | `go install golang.org/x/vuln/cmd/govulncheck@latest` (conditional — skipped if go not on PATH) |
| `trivy` | Container + JS dep-vuln scan — feeds `VULN` gate | 1.4 | `brew install trivy` |

Host prerequisites that must already exist:

- `brew` (macOS) or `apt` (Linux)
- `node` + `npm` (≥ 18)
- `python3` (≥ 3.10) — `pipx` layers on top

---

## Invocation

Run once per host: `bash .claude/skills/evaluator/preflight.sh`. Idempotent — re-running no-ops when every tool is already present. Canonical install commands live in `preflight.sh` (the executable contract); edit there, not here. Host prerequisites (`brew`, `node`/`npm`, `python3`) must already exist — the script halts loud if any are missing.

---

## Required behaviour for scoring agents

1. **Run preflight first.** Before extracting any signal, execute the preflight script.
2. **If preflight exits non-zero, STOP.** Do not proceed to scoring with missing tools. Return a clear failure message to the caller naming which tools failed to install.
3. **No silent fallback.** A signal that requires a tool must either run the tool or emit `{signal}_NOT_RUN: true` with `reason: "preflight failed"` in the matching sub-criterion's `data` block in `scorecard.yaml`. This is a loud failure mode, not a graceful one.
4. **Re-running preflight is free.** Every invocation should run it — the first run installs; subsequent runs no-op.

---

## When a tool does not apply

Some tools apply only when a repo contains the target language (e.g. `ruff` on Python, `golangci-lint` on Go, `knip` on TS/JS). In that case:

- The tool is **not** required to be installed for repos where the language is absent.
- The signal is emitted as `{signal}: null, reason: "language_not_present"` — NOT as `_NOT_RUN`.
- Preflight still installs all tools idempotently; language filtering happens at the recipe level, not preflight.

This separates "tool missing from host" (a failure) from "tool not applicable here" (expected).

---

## Rationale

Silent fallback to grep / LLM equivalents corrupts the scorecard in two ways:

1. **Wrong bands.** The absolute banding tables in each recipe are calibrated against real tool output (e.g. `lizard` CCN, `semgrep` finding density). A grep-estimated substitute lands the project in the wrong band.
2. **Under-reporting bias.** `gitleaks` catches entropic tokens `grep` cannot; `pip-audit` / `trivy` catch CVEs no heuristic will. Missing them turns a SECRETS- or VULN-gate hit into silence — a dishonest pass.

Installing tools is cheap and one-time. Skipping them silently produces a scorecard that looks complete but isn't.
