# Code Quality — Extraction Plan

- **Sub-criterion:** 1.1 (AI-only)
- **Weight:** 12.5 pts (4 Code AI × 12.5 = 50 pts half)
- **Contenders:** LLM interpretive read (primary, load-bearing) + lint (`ruff` / `biome` / `golangci-lint`) + `lizard` complexity (**per-language CCN**: Python 10 / JS-TS 10 / Java 12 / Go 15) + `semgrep` quality packs + CodeRabbit scrape (**calibrating modifier**).
- **Source quote:** _"Readable, modular, well-structured. Good naming, clean APIs, no code smells."_

Bands are **absolute** — the _Banding_ table in this recipe declares per-signal numeric thresholds. `lizard_god_per_kloc` is computed against **per-language CCN thresholds**: Python / JS / TS functions are "god" at CCN > 10; Java at CCN > 12; Go at CCN > 15 — a per-language lookup keyed on file extension. `_NOT_RUN: true` is supported for tool-missing cases but not for language-not-present (emit `null` aggregates instead). 1.1 has no gates of its own; secret/vuln signals are owned by `maintainability.md`.

> **Why CR is a modifier, not primary.** Null-data semantics and the plan-upgrade escape hatch are spelled out once in [`SKILL.md → CodeRabbit — calibrating modifier with null-data semantics`](../SKILL.md#coderabbit--calibrating-modifier-with-null-data-semantics). Summary: CR joins lint / complexity / file-size / semgrep as a **calibrating modifier**; contributes when a project has non-empty CR data, no-op (`null`) when empty.

> **Source-quote is semantic.** "Readable", "good naming", "clean APIs", "no code smells" are all reader-experience judgments. No single mechanical tool captures them — so the LLM read is the load-bearing signal, with lint / complexity / file-size / semgrep / CR acting as calibrating modifiers.

---

## Preflight

Before running this recipe, execute [`../../preflight.md`](../../preflight.md). The recipe hard-requires:

- `ruff` (Python repos)
- `biome` (TS/JS repos) — preferred over eslint for speed; eslint is a per-project fallback if biome rejects the repo's config shape
- `golangci-lint` (Go repos)
- `lizard` (polyglot complexity)
- `semgrep` (polyglot static analysis)
- `gh`, `jq` (CR scrape + utilities)

If any of the above is missing **and preflight cannot install it**, the recipe must emit `{signal}_NOT_RUN: true` + `reason` and flag the shortfall loudly in `§6.1` of the analysis file. No silent grep-estimation substitutes.

Tools are called only for languages present in the repo (see each signal below).

---

## Extraction recipe

### Signal 1 — LLM interpretive read (**primary**, load-bearing)

Sample **6 source files** that give the reader a representative slice:

1. **Largest source file** by LoC (one per primary language)
2. **Highest-churn file** from `git log --numstat --all --no-merges | awk ...` (top by lines touched)
3. **Entrypoint / handler**: matches `*/main.{py,ts,js,go}`, `*/app.{py,ts}`, `*/handlers/*`, `*/routes/*`, `*/api/*`
4. **3 random** across distinct directories, excluding `tests/`, `__tests__/`, generated, vendored, lockfiles. Prefer files > 50 LoC.

Reviewer prompt (apply identically across all six files):

> Score 0–5 on each axis:
>
> - **Naming** — intent-revealing identifiers, consistent vocabulary, no cryptic abbreviations.
> - **API shape** — function signatures clear, arguments typed / labeled, side-effects explicit, return types obvious.
> - **Abstraction** — one responsibility per module, boundaries intentional, no god-objects.
> - **Modularity** — low coupling, reasonable import graph, shared concerns extracted.
> - **Code smells** — no duplicated blocks, no magic numbers, nesting ≤ 3, no dead branches.

Emit **per file** `{path, naming, api_shape, abstraction, modularity, smells, one_sentence}`.

Emit **aggregate**:

- `llm_quality_score` = mean of per-file axis means, clamped to 0–5, to one decimal.
- `llm_notes` — 2–3 sentences summarising the strongest signal that drove the score (good or bad).
- `llm_files_sampled` — list of 6 paths.

### Signal 2 — Lint cleanliness

Run the lint tool for every language present in the repo. Skip the tool for languages absent.

```bash
# Detect primary languages (reuse the detection from 1.2 / 2.4 recipes)
LANGS=$(find <repo> -type f \
  \( -name '*.py' -o -name '*.ts' -o -name '*.tsx' \
     -o -name '*.js' -o -name '*.jsx' -o -name '*.go' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' \
  | sed 's/.*\.//' | sort -u)

# Python
if echo "$LANGS" | grep -q '^py$'; then
  ruff check <repo> --output-format json > /tmp/ruff.json || true
fi

# TS / JS — prefer biome. Use `biome lint` (not `biome ci` or `biome check`):
# `ci` / `check` bundle lint + format + organizeImports as diagnostics, which
# inflates `lint_findings_per_kloc` with stylistic `category: "format"` noise
# that the scorer doesn't want (one observed project showed 11.4/kLoC with
# format, 5.1/kLoC without — crossing the band-4 threshold of 10). `biome lint`
# is pure lint and matches the intent of the signal.
if echo "$LANGS" | grep -qE '^(ts|tsx|js|jsx)$'; then
  cd <repo>
  biome lint . --reporter=json > /tmp/biome.json 2>/dev/null \
    || npx --yes eslint . --format json > /tmp/eslint.json 2>/dev/null \
    || true
fi

# Go
if echo "$LANGS" | grep -q '^go$'; then
  cd <repo> && golangci-lint run --out-format json ./... > /tmp/golangci.json 2>/dev/null || true
fi
```

Parse (same shape across tools):

- `lint_findings_total` — sum across tools
- `lint_findings_per_kloc` — `total ÷ (scc_code_loc ÷ 1000)`; reuse `scc_code_loc` from 2.4 recipe
- `lint_by_severity` — `{error: N, warning: N, info: N}` mapped from each tool's severity field
- `lint_by_tool` — `{ruff: N, biome: N, eslint: N, golangci: N}` for transparency
- `lint_has_config_in_repo` — bool: does the repo itself ship a `ruff.toml` / `biome.json` / `.eslintrc*` / `.golangci.yml`? (rewards repos that committed to a standard).

### Signal 3 — Complexity density (`lizard`; shared with 1.4)

`lizard` is run once by 1.4 Maintainability — **read its output**, do not re-run. 1.4's Signal 2 applies per-language CCN thresholds (Python 10 / JS-TS 10 / Java 12 / Go 15 — see [maintainability.md](maintainability.md#signal-2--polyglot-complexity-lizard-per-language-ccn)); 1.1 reuses those aggregates verbatim.

Reuse:

- `lizard_mean_ccn` — mean cyclomatic complexity across functions
- `lizard_god_functions` — count with CCN > per-language threshold (Py/JS/TS 10, Java 12, Go 15, other 15)
- `lizard_god_by_language` — map `{py, js, ts, go, java, other}` → god count per language
- `lizard_god_per_kloc` — `god_functions ÷ (scc_code_loc ÷ 1000)`
- `lizard_god_ratio` — `god_functions ÷ total_functions`
- `lizard_thresholds_applied` — `{py: 10, js: 10, ts: 10, java: 12, go: 15, default: 15}` (audit constant)

If 1.4 has not been run yet for this repo, run lizard now (same command, same CSV location) and share the output.

### Signal 4 — File-size distribution (filesystem)

```bash
find <repo> -type f \
  \( -name '*.py' -o -name '*.ts' -o -name '*.tsx' \
     -o -name '*.js' -o -name '*.jsx' -o -name '*.go' \
     -o -name '*.java' -o -name '*.kt' -o -name '*.rs' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' \
  -not -path '*/dist/*' -not -path '*/build/*' \
  -exec wc -l {} + 2>/dev/null | sort -rn | head -20
```

Capture:

- `files_over_500_loc` — count of source files ≥ 500 LoC
- `largest_file_loc` — largest source file
- `mean_source_file_loc` — mean over all source files

Shared with 1.4; avoid duplicate computation.

### Signal 5 — Semgrep quality rule-set

```bash
# NOTE: `--config p/code-quality` is deprecated on Semgrep Registry and
# returns HTTP 404 with retry-storm. Stick to the per-language packs that
# resolve cleanly. If you need an umbrella, use `p/default` (lower signal).
timeout 180 semgrep --config p/python --config p/typescript \
        --config p/javascript --config p/go \
        --json --timeout 30 --disable-version-check --metrics=off \
        <repo> > /tmp/semgrep-quality.json
```

Filter out findings that belong to 1.2 (resilience) or 1.4 (SECURITY). The `p/security-audit` and resilience rules carry distinct `check_id` prefixes — drop findings whose `check_id` matches `(bare-except|broad-except|swallow|missing-(validation|input-validation)|injection|deserial|unsafe|hardcoded-password|hardcoded-api-key|ssrf|sqli|xxe|eval-)`.

Capture:

- `semgrep_quality_findings` — post-filter count
- `semgrep_quality_per_kloc` — `findings ÷ (scc_code_loc ÷ 1000)`
- `semgrep_quality_by_severity` — `{ERROR, WARNING, INFO}`

### Signal 6 — CodeRabbit scrape (**calibrating modifier**)

Scrape CR comments from both endpoints CR posts to — inline review comments and issue-style walkthrough comments. Both are needed: walkthrough-only repos look empty on the pulls endpoint alone. CR's label text is authoritative; emoji is a fallback (CR has rotated emoji twice recently).

```bash
# Cap each endpoint at 90s so a CR-heavy repo can't stall the run. The
# `--paginate-limit` flag was removed in gh 2.89+, so we rely on `--paginate`
# (fetches all pages) wrapped in `timeout` for the actual cap. Typical comment
# counts fit easily inside the 90s window.
(
  timeout 90 gh api "repos/<org>/<repo>/pulls/comments?per_page=100" \
    --paginate \
    --jq '[.[] | select(.user.login == "coderabbitai[bot]")]'
  timeout 90 gh api "repos/<org>/<repo>/issues/comments?per_page=100" \
    --paginate \
    --jq '[.[] | select(.user.login == "coderabbitai[bot]")]'
) > /tmp/cr.json
```

Parse each CR comment; prefer label text, fall back to emoji:

| Label text (authoritative) | Emoji (may drift) | Category      |
| -------------------------- | ----------------- | ------------- |
| `Refactor suggestion`      | `🛠️`              | `refactor`    |
| `Potential issue` / `Bug`  | `🐛`              | `bug`         |
| `Warning`                  | `⚠️`              | `warning`     |
| `Nitpick`                  | `📝` / `🧹`       | `nitpick`     |
| `Security`                 | `🔒`              | `security`    |
| `Performance`              | `⚡`              | `performance` |

**Walkthrough detection (positive-match, not silence-inferred).** CR posts one issue-level walkthrough per PR it reviews, regardless of plan tier. Detect walkthroughs positively via any of these markers in the body (robust across CR template drift):

- HTML comment: `<!-- walkthrough_start -->` (current stable anchor)
- Section heading: `## Walkthrough`
- Collapsible summary: `<summary>📝 Walkthrough</summary>`

Any one match ⇒ count the comment as a walkthrough and record `cr_walkthrough_count`. This is separate from `cr_pr_count` (unique PRs with ≥1 CR comment of any kind).

**Actionable-count capture.** When CR has actionable inline feedback, the walkthrough summary carries a line of the form `Actionable comments posted: N` (optionally bold-wrapped). Match with a tolerant regex:

```regex
\*{0,2}Actionable comments posted:\s*(\d+)\*{0,2}
```

Sum captured `N` across walkthroughs → `cr_actionable_total`. If no walkthrough contains this string **and** at least one walkthrough contains a Free-plan disclaimer (`Summarized by CodeRabbit Free` / `is on the Free plan`), set `cr_free_plan: true` — this distinguishes "zero actionables because Free-plan walkthrough-only" from "zero actionables because the regex is blind to a new template." Unmatched inline comments bucket to `uncategorized`; surface the count so loss is visible.

Capture:

- `cr_pr_count` — unique PRs with ≥1 CR comment
- `cr_walkthrough_count` — CR walkthrough comments matched via positive marker
- `cr_free_plan` — bool: body contains Free-plan disclaimer; explains zero-actionable outcomes
- `cr_actionable_total` — Σ captured `N` across walkthroughs (regex above)
- `cr_actionable_per_kloc` — `cr_actionable_total ÷ (scc_code_loc ÷ 1000)` (lower-is-better; `null` when `cr_pr_count == 0`)
- `cr_by_category` — map of category → count (`uncategorized` bucket when neither label nor emoji matches)
- `cr_severe_count` — `by_category.bug + by_category.security + by_category.warning`
- `cr_severe_per_kloc` — `cr_severe_count ÷ (scc_code_loc ÷ 1000)` (lower-is-better; `null` when `cr_pr_count == 0`)

**Banding rule:** CR is one of several calibrating modifiers (see _Banding_ below). When `cr_pr_count > 0`, `cr_actionable_per_kloc` and `cr_severe_per_kloc` join lint/lizard/semgrep as density signals that cap the band. When `cr_pr_count == 0` or all walkthroughs are actionable-free (the Free-plan default for many projects), both fields are `null` and the modifier is a no-op — the LLM + remaining density signals decide the band. Scorer cites the top 3 CR comments in the evidence string when `cr_actionable_total > 0`.

---

## Raw dumps (flat — files under `raw/` named `code-quality-*`)

Each signal writes its raw output verbatim to the flat `raw/` directory at the workspace root. Absent files = "not extracted" (tool missing / language not present). Aggregates land in `scorecard.yaml` under `data.`.

| File | Source signal | Shape / note |
| --- | --- | --- |
| `raw/code-quality-llm-sample.jsonl` | Signal 1 | 6 lines, one per sampled file: `{path, naming, api_shape, abstraction, modularity, smells, one_sentence}` |
| `raw/code-quality-lint-ruff.json` | Signal 2 | Raw `ruff check --output-format json` (Python only) |
| `raw/code-quality-lint-biome.json` | Signal 2 | Raw `biome ci --reporter json` (TS/JS preferred) |
| `raw/code-quality-lint-eslint.json` | Signal 2 | Raw `eslint --format json` (TS/JS fallback when biome rejects config) |
| `raw/code-quality-lint-golangci.json` | Signal 2 | Raw `golangci-lint run --out-format json ./...` (Go only) |
| `raw/code-quality-semgrep-quality.json` | Signal 5 | Post-filter semgrep output (resilience + security rules stripped) |
| `raw/code-quality-cr-comments.jsonl` | Signal 6 | One line per CR comment: `{pr, file, line, category, label, body_excerpt}`. Empty file when `cr_pr_count == 0` — still emit it (absence of file would be ambiguous with "scrape not run"). |

`raw/maintainability-lizard.csv` and `raw/maintainability-file-sizes.txt` are shared with 1.4 — read, don't duplicate.

---

## Banding → 0–5

**Primary axis:** `llm_quality_score` (higher-is-better).

**Calibrating modifiers** (lower-is-better for all):

- `lint_findings_per_kloc`
- `lizard_god_per_kloc`
- `files_over_500_loc`
- `semgrep_quality_per_kloc`
- `cr_actionable_per_kloc` _(participates only when `cr_pr_count > 0`; null = no-op)_
- `cr_severe_per_kloc` _(participates only when `cr_pr_count > 0`; null = no-op)_

Absolute, per-signal thresholds. CR caps apply only when `cr_pr_count > 0` (null = no-op).

| Band | `llm_quality_score` | Density caps |
| --- | --- | --- |
| **5** | ≥ 4.5 | AND `lint_findings_per_kloc` ≤ 5 AND `lizard_god_per_kloc` ≤ 2 AND `files_over_500_loc` ≤ 1 AND (when non-null) `cr_severe_per_kloc` ≤ 0.5 |
| **4** | 3.5 – 4.5 | AND `lint_findings_per_kloc` ≤ 10 AND `lizard_god_per_kloc` ≤ 5 AND (when non-null) `cr_actionable_per_kloc` ≤ 5 |
| **3** | 2.5 – 3.5 | — |
| **2** | 1.5 – 2.5 | OR `lint_findings_per_kloc` > 20 OR `lizard_god_per_kloc` > 10 OR (when non-null) `cr_severe_per_kloc` ≥ 3 |
| **1** | < 1.5 | OR `files_over_500_loc` > 5 OR (when non-null) `cr_severe_per_kloc` ≥ 5 |
| **0** | no source code (empty repo) | — |

---

## Output (what the scorer emits)

```yaml
sub_criterion: code_quality
score: 4
evidence: "LLM sample (6 files) scored 4.2 — naming consistent, API signatures typed, one god-function in src/orchestrator.py flagged by lizard (CCN 18). ruff 3.2 findings/kLoC; semgrep 1.1 quality findings/kLoC; CodeRabbit 12 PRs, 4 actionables (refactor×3, nitpick×1) — 0.3/kLoC, no severe. No code smells beyond the one god-function."
data:
  llm_quality_score: 4.2
  llm_files_sampled:
    [
      src/orchestrator.py,
      src/api/routes.py,
      src/db.py,
      src/worker.py,
      internal/retry.py,
      scripts/seed.py,
    ]
  llm_notes: "..."
  lint_findings_total: 42
  lint_findings_per_kloc: 3.2
  lint_by_severity: { error: 0, warning: 12, info: 30 }
  lint_by_tool: { ruff: 42, biome: 0, eslint: 0, golangci: 0 }
  lint_has_config_in_repo: true
  lizard_mean_ccn: 4.1
  lizard_god_per_kloc: 1.8
  files_over_500_loc: 1
  largest_file_loc: 580
  mean_source_file_loc: 112
  semgrep_quality_findings: 14
  semgrep_quality_per_kloc: 1.1
  semgrep_quality_by_severity: { ERROR: 0, WARNING: 4, INFO: 10 }
  cr_pr_count: 12
  cr_actionable_total: 4
  cr_actionable_per_kloc: 0.3
  cr_severe_count: 0
  cr_severe_per_kloc: 0.0
  cr_by_category:
    {
      refactor: 3,
      nitpick: 1,
      warning: 0,
      bug: 0,
      security: 0,
      performance: 0,
      uncategorized: 0,
    }
  signals_used: [llm, ruff, lizard, filesystem, semgrep, coderabbit]
  scoring_method: llm_primary_with_mechanical_modifiers
```

When `cr_pr_count == 0` or all walkthroughs are actionable-free (the Free-plan default for most projects), emit:

```yaml
cr_pr_count: 12
cr_actionable_total: 0
cr_actionable_per_kloc: null
cr_severe_count: 0
cr_severe_per_kloc: null
cr_by_category: {}
signals_used: [llm, ruff, lizard, filesystem, semgrep, coderabbit_empty]
```

`null` values signal the scorer to skip the CR modifier for this project — banding leans on the remaining density signals.

If any required tool failed preflight:

```yaml
data:
  lint_findings_total: null
  lint_NOT_RUN: true
  lint_reason: "preflight could not install ruff — host lacks compatible Python"
  scoring_method: llm_primary_with_mechanical_modifiers_degraded
```

---

## Caveats

- **LLM sample size = 6 files.** Evidence must name every file sampled; the reader should be able to audit each axis score.
- **Lint tool coverage varies.** A TS repo without a `biome.json` or `.eslintrc` will emit different default rules than one with either — `lint_has_config_in_repo` surfaces the difference.
- **Mixed-language repos.** Sum findings across tools; normalise by `scc_code_loc` (not per-language LoC) to avoid double-normalising.
- **`lizard` CCN thresholds are language-idiomatic** — handled by per-language lookup (Py/JS/TS 10, Java 12, Go 15, other 15). See [maintainability.md → Signal 2](maintainability.md#signal-2--polyglot-complexity-lizard-per-language-ccn).
- **CodeRabbit Free plan** sparsity → `cr_*_per_kloc: null`, modifier no-op. Full rationale + plan-upgrade escape hatch in [`SKILL.md → CodeRabbit — calibrating modifier with null-data semantics`](../SKILL.md#coderabbit--calibrating-modifier-with-null-data-semantics).
- **`AI_AUTHORED_SUSPECTED` signal** (>50% added lines in commits > 500 LoC) remains useful — LLM reviewer should weigh LLM-generated code patterns in `llm_notes` (repetitive scaffolding, over-engineered edge-case branches, inconsistent naming within a single file, etc.).
- **No silent fallback.** If a tool is absent despite preflight, the signal is emitted as `_NOT_RUN` with `reason`. Grep-based substitutes break the absolute band thresholds and are banned.
