# Rules audit

Use when the repo's written rules (`AGENTS.md`, contributing guides,
style docs) are not being followed, or when the same correction keeps
coming up in review and is not written down. One agent checking a diff
against twenty rules skips some; one checker per rule does not.

Two modes. **Check** asks "does this change follow our rules?".
**Mine** asks "which rules are we missing?".

Check on someone's PR runs as delegates; check before merging your own
work, and mine always, run in a workstream. Violations and candidate
rules are findings: [findings](findings.md) says where they live.

## Check: a change against the rules

1. **List the rules.** A scout (a delegate call) extracts every
   checkable rule from the rule files, one per line with its source,
   into every checker's brief (delegate mode) or the umbrella's note
   (workstream mode):

   ```text
   RULE: r7 AGENTS.md:52 "Throw typed error classes, not bare Error"
   ```

   Skip rules no diff can show (team habits, meeting norms). Record the
   count.
2. **Pin the target**: the diff, branch, or PR, as in
   [review-panel](review-panel.md) step 1.
3. **One checker per rule**, as a [delegate call](tasks-or-calls.md#delegate-call)
   The checker gets one rule and the diff, and records each violation
   as a finding ([findings § Record](findings.md#record); title
   `<severity>: r7 <what>`), or reports `r7: no violations` with what it
   looked at. One rule per checker keeps it from skimming; batch under
   the cap ([orchestrator-loop § Concurrency](orchestrator-loop.md#concurrency)).
4. **A skeptic pass** over the flags, also a delegate call: one fresh agent reads each flagged
   line and the rule, and drops false positives (the rule does not
   apply here, the code already complies, the rule has a stated
   exception). In workstream mode, a dropped flag is
   `close --as rejected --why '<reason>'`; a confirmed one is accepted.
5. **Report or fix** confirmed violations, as in review-panel step 5.

Done when every rule from step 1 has a checker result, and every flag
is decided.

## Mine: rules you keep stating but never wrote

1. **Collect corrections.** Sources, any that exist: review comments
   on merged PRs, `REJECT` notes from [adversarial-review](adversarial-review.md)
   tasks, and past agent sessions (with a session archive such as
   [museum](https://github.com/mu-crew/museum), search for user turns
   that correct the agent: "no, use", "don't", "we always"). One reader
   (a delegate call) per source batch returns each correction as one
   line:

   ```text
   CORRECTION: k12 <source ref> "use the shared retry helper instead of a hand-rolled loop"
   ```

2. **Cluster** the corrections in one delegate call: the same lesson in
   different words is one cluster. Keep clusters with at least two
   corrections from different occasions; a one-off is not a rule.
3. **One `OPEN/triage` task per candidate rule**, the rule as the
   title and its cluster of corrections in the note.
4. **Refute each candidate** with a [delegate call](tasks-or-calls.md#delegate-call): would it have
   prevented the real mistakes in its cluster? Does it contradict an
   existing rule? Is it already enforced by a linter or test (then it
   needs no prose)? Dropped candidates close `--as rejected`.
5. **Propose, don't commit.** The candidates left in triage are the
   proposal: write them as a diff to the rule file in the umbrella's
   note, worded per [brief](brief.md) (positive, specific, with the
   reason). A rule change is a human decision: the human accepts or
   rejects each candidate task.

Done when every cluster has a verdict and the proposal lists each kept
rule with the corrections behind it.

## Traps

- **A rule a tool can check belongs in the tool.** If a linter, type,
  or test can catch it, propose that instead of prose.
- **Mined rules overfit.** Two corrections in one week about one file
  may be about that file, not a rule. The refute step asks.
