# Testing & Resilience — Extraction Plan

- **Sub-criterion:** 1.2 (AI-only)
- **Weight:** 12.5 pts (4 Code AI × 12.5 = 50 pts half)
- **Contenders:** Semgrep + filesystem + GHA + LLM (joint load-bearing) + CodeRabbit resilience reuse (**calibrating modifier**) + **multi-protocol route-auth scan** (HTTP + WebSocket + gRPC + tRPC — feeds the `AUTH` gate).
- **Source quote:** _"Unit/integration tests, error handling, input validation. It should survive more than the happy path."_

Bands are **absolute** — thresholds in the _Banding_ section below are fixed numeric cutoffs. Gates this recipe owns: **`UNTESTED`** (hard, `testing ≤ 1` when `test_file_count == 0`), **`TESTS_TRIVIAL`** (soft, `testing ≤ 2` when `assertion_density < 1.0` OR `trivial_body_ratio > 0.8`), and **`AUTH`** (fires `testing ≤ 2` when network-routed handlers exist and ≥ 1 mutating endpoint isn't behind an auth middleware/dependency/decorator, detected across HTTP [FastAPI/Express/Flask/Gin/Django POST/PUT/PATCH/DELETE], WebSocket [ws/Socket.io/SignalR], gRPC [grpc/grpcio unary+streaming], tRPC [@trpc/server mutations]). Raw dumps: `raw/testing-auth-scan-{http,ws,grpc,trpc}.json`. Data fields: `routed_mutating_endpoints_total`, `routed_mutating_endpoints_unguarded`, `auth_by_protocol: {http: N, ws: N, grpc: N, trpc: N}`. AUTH gate is skipped if no routed handlers are detected.

> Four load-bearing signals + one calibrating modifier (CR), integrated by the scorer. No weighted-math internal split — the LLM emits a single 0–5 band and cites the most load-bearing signal.

---

## Preflight

Before running this recipe, execute [`../../preflight.md`](../../preflight.md). Required:

- `semgrep` (polyglot static analysis — resilience rule-set)
- `gh`, `jq` (CI workflow introspection + CR evidence scrape)

If `semgrep` cannot be installed, emit `semgrep_resilience_findings: null, semgrep_NOT_RUN: true, reason: "..."` and flag loudly in `§6.1`. Do not substitute grep for it — the patterns Semgrep catches (`bare-except`, `empty-catch`, `unhandled-promise`) are structural, not lexical.

---

## Extraction recipe

### Signal 1 — Filesystem (tests exist, framework configured, tests meaningful)

```bash
# Test file count
find <repo> -type f \
  \( -name 'test_*.py' -o -name '*_test.py' \
     -o -name '*.test.ts' -o -name '*.test.js' \
     -o -name '*.spec.ts' -o -name '*.spec.js' \
     -o -name '*_test.go' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' \
  | wc -l

# Source file count (for ratio) — language-aware; see "Per-language thresholds" below
find <repo> -type f \
  \( -name '*.py' -o -name '*.ts' -o -name '*.tsx' \
     -o -name '*.js' -o -name '*.jsx' -o -name '*.go' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' | wc -l

# Framework configs
ls <repo>/{pytest.ini,jest.config.*,vitest.config.*} 2>/dev/null
grep -lE 'pytest|jest|vitest|\[tool\.pytest\]' \
  <repo>/{pyproject.toml,package.json} 2>/dev/null

# Test meaningfulness — assertion density + trivial-body ratio
TEST_FILES=$(find <repo> -type f \
  \( -name 'test_*.py' -o -name '*_test.py' \
     -o -name '*.test.ts' -o -name '*.test.js' \
     -o -name '*.spec.ts' -o -name '*.spec.js' \
     -o -name '*_test.go' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*')

# Assertions across all test files
echo "$TEST_FILES" | xargs grep -cE \
  'assert |expect\(|\.to(Equal|Be|Throw|Have|Contain|Match)|require\.|assertEqual|assertTrue' \
  2>/dev/null | awk -F: '{sum+=$2} END {print sum}'

# Trivial bodies (assert True / pass / expect(true))
echo "$TEST_FILES" | xargs grep -cE \
  '^\s*(pass|assert True|expect\(true\)|return)\s*$' 2>/dev/null \
  | awk -F: '{sum+=$2} END {print sum}'
```

Capture: `test_file_count`, `src_file_count`, `test_ratio`, `framework_configured`, `assertion_density` (assertions per test file), `trivial_body_ratio` (trivial bodies ÷ total test bodies).

**Primary language detection** (for thresholds — see Aggregation):

```bash
find <repo> -type f \
  \( -name '*.py' -o -name '*.ts' -o -name '*.tsx' \
     -o -name '*.js' -o -name '*.jsx' -o -name '*.go' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' \
  | sed 's/.*\.//' | sort | uniq -c | sort -rn | head -1
```

Capture: `primary_language` (py / go / ts / js / tsx / jsx).

### Signal 2 — GitHub Actions (did CI test and pass?)

```bash
# Recent runs
gh api "repos/<org>/<repo>/actions/runs?per_page=30" \
  --jq '[.workflow_runs[] | {name, conclusion, event, head_branch}]'

# Workflow files → find test invocations
gh api repos/<org>/<repo>/contents/.github/workflows --jq '.[].path' | while read wf; do
  gh api "repos/<org>/<repo>/contents/$wf" --jq '.content' | base64 -d \
    | grep -E 'pytest|jest|vitest|go test|cargo test|npm test|yarn test'
done
```

Capture: `ci_has_test_job` (bool), `latest_test_conclusion` (SUCCESS/FAILURE/NONE), `ci_test_invocation`.

### Signal 3 — CodeRabbit reuse (**calibrating modifier**)

Reuse the CR scrape from 1.1 (see [code-quality.md](code-quality.md#signal-6--coderabbit-scrape-calibrating-modifier) Signal 6). Filter to resilience-relevant categories `[bug, warning]` — these are CR's signal for missing error handling, swallowed exceptions, or validation gaps on the PRs it reviewed.

```yaml
categories_in_scope: [bug, warning]
```

Capture:

- `cr_resilience_flags = by_category.bug + by_category.warning`
- `cr_resilience_per_kloc = cr_resilience_flags ÷ (scc_code_loc ÷ 1000)` (lower-is-better; `null` when `cr_pr_count == 0` or `cr_resilience_flags == 0` on a zero-actionable walkthrough)
- `cr_semgrep_overlap` — count of `file:line` flagged by both CR and Semgrep Signal 4. Prefer CR's context in evidence when they overlap; surface the overlap count for transparency.

**Banding rule:** CR is a calibrating modifier — it pulls the band down when `cr_resilience_per_kloc` is high and (when non-null) participates in the aggregation table below alongside `semgrep_resilience_per_kloc`. When CR data is empty, `cr_resilience_per_kloc` is `null` and the modifier is a no-op. Full rationale + plan-upgrade escape hatch in [`SKILL.md → CodeRabbit — calibrating modifier with null-data semantics`](../SKILL.md#coderabbit--calibrating-modifier-with-null-data-semantics).

**Semgrep remains the broadest resilience signal** (full-tree scan, not just reviewed diffs). Signal 4 below.

### Signal 4 — Semgrep scan

```bash
semgrep --config p/python --config p/javascript --config p/typescript \
        --config p/go \
        --json <repo> > /tmp/semgrep.json
```

**Filter to resilience-relevant rules only.** Tight pattern (catches the actual phenomena, drops noise):

```bash
jq '[.results[] | select(
  .check_id | test(
    "(swallow|bare-except|broad-except|empty-catch|unhandled-(rejection|promise|exception)|missing-(validation|null-check|type-check|input-validation)|no-(validation|check)|unchecked-(await|promise))"
  )
)]' /tmp/semgrep.json
```

**What the pattern deliberately excludes, and why:**

- `error|exception` — matched every rule name containing those words (most security rules do), inflating findings ~10×.
- `injection|deserial|unsafe` — security territory; belongs in 1.4 Maintainability / SECURITY gate to avoid double-counting.
- `await|promise` (bare) — scoped to `unchecked-await` / `unhandled-promise` so async-correctness rules aren't indiscriminately matched.

`p/security-audit` config is excluded — it produces security-skewed noise. Language packs alone carry the resilience rules we want.

Capture:

- `semgrep_total_findings`
- `semgrep_resilience_findings` (after filter)
- `semgrep_resilience_by_severity` (ERROR / WARNING / INFO)
- `semgrep_resilience_per_kloc`

**Note:** Semgrep ERROR severity does **not** act as a hard score cap — signal only.

### Signal 5 — LLM sample (3 handler-or-resilience-dense files)

Tiered file selection — walk tiers until 3 files are picked:

**Tier 1 — request handlers (preferred):**

- Matches `*route*`, `*handler*`, `*controller*`, `*api*`, `*server*`, `*endpoint*`, `*view*`
- Under `src/`, `app/`, `internal/`, `pkg/` or root

**Tier 2 — resilience-dense files (for non-API repos: CLIs, agents, batch jobs):** Rank source files by `(try|catch|raise|throw|except|panic|recover) count ÷ LOC`; pick top 3. These are where resilience _matters_, regardless of repo shape.

```bash
for f in $(find <repo>/src <repo>/app <repo>/internal \
  -type f \( -name '*.py' -o -name '*.ts' -o -name '*.go' -o -name '*.js' \) 2>/dev/null); do
  loc=$(wc -l < "$f")
  hits=$(grep -cE '\b(try|catch|raise|throw|except|panic|recover)\b' "$f" 2>/dev/null)
  [ "$loc" -gt 20 ] && echo "$(awk -v h="$hits" -v l="$loc" 'BEGIN{print h/l}') $f"
done | sort -rn | head -3
```

**Tier 3 — fallback** (if Tier 1 & 2 yield nothing): 3 largest files under `src/` / `app/`.

Score each on:

- Error handling shape (bare except / granular try / explicit raise?)
- Input validation at boundaries (Pydantic / Zod / class-validator / manual / none?)
- Resilience patterns (retry decorators, timeout wrappers, circuit breakers?)

Capture: `llm_resilience_score` (0–5), `llm_notes` (one sentence), `llm_files_sampled` (list), `llm_sample_tier` (1/2/3).

### Signal 6 — Multi-protocol AUTH route-scan (AUTH gate trigger)

Enterprise backends ship mutating endpoints. A POST/PUT/PATCH/DELETE (HTTP) — or a WebSocket message handler, gRPC unary/streaming method, or tRPC mutation — that doesn't run behind an auth middleware / dependency / decorator is an enterprise-grade resilience failure: anyone can mutate server state. This signal scans for routed mutating handlers across four protocols and checks each one for an auth guard. **Skipped (gate does not fire) if no routed handlers are detected** — a CLI or batch job isn't expected to carry auth.

```bash
# HTTP — FastAPI / Flask / Django / Express / Gin / Echo
# Mutating verbs: POST / PUT / PATCH / DELETE. GET is informational, not gated.
semgrep --config p/python --config p/javascript --config p/typescript --config p/go \
        --timeout 30 --metrics=off --disable-version-check --json \
        --pattern-either \
          --pattern '@app.post(...)' --pattern '@app.put(...)' \
          --pattern '@app.patch(...)' --pattern '@app.delete(...)' \
          --pattern '@router.post(...)' --pattern '@router.put(...)' \
          --pattern '@router.patch(...)' --pattern '@router.delete(...)' \
          --pattern 'app.post(...)' --pattern 'app.put(...)' \
          --pattern 'app.patch(...)' --pattern 'app.delete(...)' \
          --pattern 'router.post(...)' --pattern 'router.put(...)' \
          --pattern 'router.patch(...)' --pattern 'router.delete(...)' \
          --pattern 'r.POST(...)' --pattern 'r.PUT(...)' \
          --pattern 'r.PATCH(...)' --pattern 'r.DELETE(...)' \
        <repo> > /tmp/auth-http.json 2>/dev/null || true

# WebSocket — ws / Socket.io / SignalR / Django Channels / FastAPI WebSocket
# Mutating WS handlers are identified by event registration patterns.
semgrep --config=auto --timeout 30 --metrics=off --disable-version-check --json \
        --pattern-either \
          --pattern 'socket.on("$EVENT", $HANDLER)' \
          --pattern 'io.on("connection", ...)' \
          --pattern '@websocket.event' \
          --pattern 'WebSocketHandler' \
          --pattern '@sio.on(...)' \
        <repo> > /tmp/auth-ws.json 2>/dev/null || true

# gRPC — grpc / grpcio unary + streaming server methods
# Service implementations typically subclass *Servicer; methods are RPC endpoints.
semgrep --config=auto --timeout 30 --metrics=off --disable-version-check --json \
        --pattern-either \
          --pattern 'class $S($B.Servicer): ...' \
          --pattern 'def $METHOD(self, request, context): ...' \
          --pattern 'grpc.UnaryUnaryClientInterceptor' \
          --pattern 'grpc.ServerInterceptor' \
        <repo> > /tmp/auth-grpc.json 2>/dev/null || true

# tRPC — @trpc/server mutations (not queries — queries are reads, not mutations)
semgrep --config=auto --timeout 30 --metrics=off --disable-version-check --json \
        --pattern-either \
          --pattern '$R.mutation(...)' \
          --pattern 'router({ ...mutation: ... })' \
          --pattern 't.procedure.mutation(...)' \
          --pattern 't.procedure.input(...).mutation(...)' \
        <repo> > /tmp/auth-trpc.json 2>/dev/null || true
```

Parse each tool's output for the number of mutating handlers and, critically, whether each one is adjacent to or wrapped by an auth guard. A handler is "guarded" when any of these are present in the same scope:

- **Python/FastAPI**: `Depends(...)` call on the function signature that resolves to an auth dep, `@require_auth`/`@login_required` decorator, or a router prefix/dependency declaring auth
- **Python/Django**: `@login_required` / `@permission_required` / DRF `permission_classes`
- **Python/Flask**: `@login_required` / `@jwt_required` / flask-login / flask-security decorator
- **JS/TS/Express**: middleware ordering — auth middleware registered on the app or router before the route, or a `passport.authenticate(...)` / `requireAuth` / `verifyJWT` call in the handler chain
- **JS/TS/tRPC**: `t.procedure.use(authMiddleware).mutation(...)` or a `protectedProcedure` base
- **Go/Gin/Echo/chi**: auth middleware chain (`r.Use(AuthMiddleware)` / `router.Group("/x", AuthMiddleware)`)
- **WebSocket**: handshake-time auth in `connect` handler OR per-message token validation
- **gRPC**: `ServerInterceptor` registered at `grpc.Server(interceptors=[...])` OR explicit `context.authenticate(...)` in the method

The scan runs with a 30 s per-rule cap + 120 s wall-clock ceiling per protocol; tools that time out emit `<protocol>_NOT_RUN: true, reason: "semgrep timeout"` with no effect on the gate.

Aggregate per protocol, then across all protocols:

```bash
# Count mutating handlers and unguarded ones, per protocol.
# The guard detection is scorer-side — LLM reads the semgrep output + surrounding
# source for each mutating handler and emits {path, line, protocol, guarded: bool, guard_mechanism: string|null}.
```

Capture:

- `routed_mutating_endpoints_total` — int; sum across HTTP + WS + gRPC + tRPC
- `routed_mutating_endpoints_unguarded` — int; unguarded count, also summed
- `auth_by_protocol` — `{http: {total: N, unguarded: N}, ws: {total: N, unguarded: N}, grpc: {total: N, unguarded: N}, trpc: {total: N, unguarded: N}}`
- `auth_unguarded_examples` — list of `{path, line, protocol}` for the first 5 unguarded handlers (for evidence citation)
- `auth_NOT_RUN` — bool; true when no routed handlers were detected at all (CLI / batch / library repo)
- `auth_reason` — string; populated when `_NOT_RUN` or when a protocol scan timed out

**AUTH gate trigger:** `routed_mutating_endpoints_total ≥ 1` AND `routed_mutating_endpoints_unguarded ≥ 1`. Caps `testing ≤ 2`. Skipped (gate does not fire) when `auth_NOT_RUN == true`.

---

## Raw dumps (flat — files under `raw/` named `testing-*`)

Each signal writes its raw output to the flat `raw/` directory at the workspace root. Absent files = not extracted. Aggregates go to `scorecard.yaml` under `data.`.

| File | Source signal | Shape / note |
| --- | --- | --- |
| `raw/testing-tests-inventory.jsonl` | Signal 1 | One line per test file: `{path, assertions, trivial_body, loc}` |
| `raw/testing-llm-sample.jsonl` | Signal 5 | 3 lines, one per resilience-dense sampled file: `{path, tier, error_handling_shape, validation_shape, resilience_patterns, one_sentence}` |
| `raw/testing-semgrep-resilience.json` | Signal 4 | Post-filter semgrep output (regex narrowed to swallow / bare-except / unhandled-promise / missing-validation); security rules stripped |
| `raw/testing-gha-runs.json` | Signal 2 | Recent workflow runs (conclusion, event, branch) + workflow files with test invocations |
| `raw/testing-auth-scan-http.json` | Signal 6 | Semgrep output for HTTP mutating handlers (FastAPI/Express/Flask/Gin/Django POST/PUT/PATCH/DELETE). Absent when no HTTP server detected. |
| `raw/testing-auth-scan-ws.json` | Signal 6 | Semgrep output for WebSocket event handlers (ws/Socket.io/SignalR/Channels). Absent when no WS detected. |
| `raw/testing-auth-scan-grpc.json` | Signal 6 | Semgrep output for gRPC service implementations. Absent when no `grpc`/`grpcio` imports detected. |
| `raw/testing-auth-scan-trpc.json` | Signal 6 | Semgrep output for tRPC mutation procedures. Absent when no `@trpc/server` imports detected. |

CodeRabbit comments are dumped once by 1.1 under `raw/code-quality-cr-comments.jsonl`; this recipe reuses that file (see Signal 3) rather than re-scraping.

---

## Aggregation → 0–5 band

Scorer integrates the 5 signals. Rough directionality:

| Signal | Pulls score UP | Pulls score DOWN |
| --- | --- | --- |
| `test_ratio` (per-language — see table) | Above language threshold | 0 (triggers `UNTESTED` gate) |
| `assertion_density` + `trivial_body_ratio` | ≥ 2 assertions/file, < 10% trivial | < 1 assertion/file OR > 80% trivial → triggers `TESTS_TRIVIAL` gate |
| `ci_has_test_job` + `latest_test_conclusion` | Both + SUCCESS | Missing or FAILURE |
| `semgrep_resilience_per_kloc` | < 1 | > 5 |
| `llm_resilience_score` | 4–5 | 0–2 |
| `cr_resilience_per_kloc` _(modifier; participates only when non-null)_ | 0 on a reviewed repo (clean on reviewer context) | > 1 |

### Per-language test ratio thresholds

Go colocates `foo_test.go` with `foo.go`; Python typically uses a separate `tests/` tree; TS varies. Fixed thresholds distort across stacks — use language-aware floors:

| Primary language        | Top-band ratio | Floor (pulls band down) |
| ----------------------- | -------------- | ----------------------- |
| Go (`go`)               | ≥ 0.50         | < 0.10                  |
| Python (`py`)           | ≥ 0.30         | < 0.05                  |
| TypeScript / JavaScript | ≥ 0.20         | < 0.03                  |
| Mixed / unknown         | ≥ 0.25         | < 0.05                  |

**Bands (absolute):**

- **5** — All banded signals top-tier (CI green, ratio above language threshold, Semgrep < 1/kLoC, LLM ≥ 4, assertion density healthy)
- **4** — Strong but one weaker signal
- **3** — Mixed: tests exist but resilience signals middling
- **2** — Tests present but resilience poor OR vice versa; or `TESTS_TRIVIAL` fires
- **1** — Near-zero tests (`UNTESTED` gate enforces this), many resilience hits
- **0** — Suppressed (gates raise to 1)

---

## Gates

| Gate | Trigger | Cap |
| --- | --- | --- |
| `UNTESTED` (hard) | `test_file_count == 0` | `testing` ≤ 1 |
| `TESTS_TRIVIAL` (soft) | `test_file_count > 0` AND (`assertion_density < 1.0` per test file OR `trivial_body_ratio > 0.8`) | `testing` ≤ 2 |
| `AUTH` | `routed_mutating_endpoints_total ≥ 1` AND `routed_mutating_endpoints_unguarded ≥ 1`. Covers HTTP (POST/PUT/PATCH/DELETE), WebSocket event handlers, gRPC unary+streaming, tRPC mutations. Skipped when `auth_NOT_RUN == true` (no routed handlers detected — CLI/batch/library). | `testing` ≤ 2 |

All three can fire independently and stack. `AUTH` is the "enterprise-grade mutation without authentication" case — a POST endpoint that lets any unauthenticated caller mutate server state is a testing failure regardless of how clean the happy-path tests are. `TESTS_TRIVIAL` is the "tests exist but are meaningless" case (a single `def test(): pass` file passes `UNTESTED` but should not reach band 3+).

---

## Fallback — signals missing

- `test_file_count == 0` → `UNTESTED` gate fires regardless; cap at 1.
- No CI workflows → drop signal 2, rely on 1, 4, 5.
- `cr_pr_count == 0` OR CR empty (Free-plan walkthrough-only) → `cr_resilience_per_kloc: null`, Signal 3 modifier is a no-op; rely on 1, 2, 4, 5 for the band.
- Semgrep doesn't support a repo's primary language → drop signal 4, rely on 1, 2, 5.
- Semgrep not installed (preflight failed) → emit `semgrep_NOT_RUN: true`, rely on 1, 2, 5, flag in §6.1.
- LLM tier 1 & 2 yield nothing → use tier 3, annotate `llm_sample_tier: 3` (confidence is lower).

Annotate evidence with `"Signals used: [...]"` so the audit trail is explicit.

---

## Output (what the scorer emits)

```yaml
sub_criterion: testing
score: 3
evidence: "Tests present (test_ratio 0.40); 2.1 avg assertions/test; CI test job passes; Semgrep 3.1 resilience hits/kLoC (swallowed exceptions dominate); LLM (tier 2, resilience-dense) sees inconsistent try/except in sampled handlers."
data:
  test_file_count: 23
  src_file_count: 57
  test_ratio: 0.40
  primary_language: py
  framework_configured: true
  assertion_density: 2.1
  trivial_body_ratio: 0.04
  ci_has_test_job: true
  latest_test_conclusion: SUCCESS
  cr_resilience_flags: 8
  cr_resilience_per_kloc: 0.8
  cr_semgrep_overlap: 2
  semgrep_total_findings: 142
  semgrep_resilience_findings: 31
  semgrep_resilience_by_severity: { ERROR: 4, WARNING: 20, INFO: 7 }
  semgrep_resilience_per_kloc: 3.1
  llm_resilience_score: 3
  llm_sample_tier: 2
  llm_files_sampled: [src/orchestrator.py, src/retry.py, src/handlers/inbox.py]
  signals_used: [filesystem, gha, semgrep, llm, coderabbit]
  gates_triggered: []
```

When CR data is empty (`cr_pr_count == 0` or all walkthroughs actionable-free):

```yaml
cr_resilience_flags: 0
cr_resilience_per_kloc: null
cr_semgrep_overlap: 0
signals_used: [filesystem, gha, semgrep, llm, coderabbit_empty]
```

`null` signals the scorer to skip the CR modifier for this project.

---

## Caveats

- Semgrep pre-built packs skew toward security. A tight regex keeps resilience focus; worth 3–5 custom rules if scoring reveals gaps.
- CodeRabbit reuse is a calibrating modifier, not a primary axis; `null` on Free-plan walkthrough-only repos means no-op. Full rationale + plan-upgrade escape hatch in [`SKILL.md → CodeRabbit — calibrating modifier with null-data semantics`](../SKILL.md#coderabbit--calibrating-modifier-with-null-data-semantics).
- `AI_AUTHORED_SUSPECTED` flag (from 1.1) remains informational — LLM reviewer should note AI-generated resilience patterns (e.g. copy-pasted try/except blocks with identical logging).
- LLM sample size is 3. Diverse codebases may score unevenly. Evidence must name the files.
- `trivial_body_ratio` heuristic may false-positive on pytest `parametrize` stubs where the real body lives in fixtures. Acceptable — over-triggering a soft gate is safer than under-triggering.
- Per-language `test_ratio` thresholds are **absolute**. Each language's floor/top band is fixed in the table above.
- AUTH scan's guard-detection is LLM-adjudicated — Semgrep surfaces the mutating handler; the scorer reads surrounding context to decide whether a guard is present. This is a deliberate trade-off: pattern-matching every auth idiom across 4+ protocols and 10+ frameworks is combinatorial, and an LLM with the source in front of it is more accurate than a brittle megapattern. When the LLM's confidence on guard-detection is low (unusual framework, ambiguous middleware chain), emit `auth_reason: "guard detection uncertain"` and let the gate fire conservatively — false-positive AUTH is cheaper than false-negative.
