# Assertion judge - a fresh reader decides what the screen shows

A visual or semantic assertion ("the error message tells the user what to do",
"no text is clipped", "the selected plan is highlighted") cannot be decided by
an exact comparison. It is decided by a judge: a fresh subagent that sees the
assertion and the current evidence, and nothing about how the run got there.

Consumers: `/multi-agent:test` (`commands/sim-test.md` Step 5),
`/multi-agent:bug-bash` (charters and verification). Script:
`$HOME/.claude/scripts/explore-findings.mjs judge-packet` / `judge-verdict`.

## 1. Why the judge is separate

The agent that drove the app knows what it meant to do, what it already
believes, and what it was told about when a step counts as done. Every one of
those lowers the bar for "the screen shows it". A judge that reads only the
statement and the evidence answers the question that was asked. The actor's
transcript, its step summaries, earlier findings and its instructions never
reach the judge.

What does reach it is the app's own vocabulary, when the assertion depends on
it ("plans are called tiers"): a sentence that would still be true for a
different tester on the same app.

## 2. When to use it

| Assertion | Decided by |
|---|---|
| exact text, count, total, route, element present | an expect step (`expect_text`, `expect_element`, `expect_url`) or a UI-tree read; no judge |
| layout, contrast, clipping, overlap, an image's content | judge, with a screenshot |
| copy meaning ("explains the error"), state ("the right tab is selected") | judge, with the screenshot and the UI tree |

If an exact check can decide it, it is not a judge question.

## 3. Flow

1. Capture the evidence now, for this assertion: a screenshot file and, when
   useful, the UI tree saved to a file. Evidence from an earlier step is not
   current.
2. Build the packet:

   ```bash
   node "$HOME/.claude/scripts/explore-findings.mjs" judge-packet \
     --assertion "<one sentence>" \
     --evidence screenshot=<png> --evidence ui-tree=<json> \
     [--vocabulary "<app terms>"] > <dir>/judge-packet.json
   ```

   The script refuses any field that describes the run.
3. Dispatch one fresh subagent (Agent tool, no conversation context carried
   over; Copilot CLI: a new task). Its whole prompt is the packet: the
   `instructions` field, the assertion, the evidence paths to open, and the
   vocabulary when present.
4. Save its JSON answer to `<dir>/judge-verdict.json` and check it:

   ```bash
   node "$HOME/.claude/scripts/explore-findings.mjs" judge-verdict \
     --packet <dir>/judge-packet.json --verdict <dir>/judge-verdict.json
   ```

## 4. Verdicts

| Judge answer | Result |
|---|---|
| `passed`, citing evidence from the packet | passed |
| `failed`, citing evidence from the packet | failed, `ASSERTION_FAILED` |
| `inconclusive` | failed, `ASSERTION_INCONCLUSIVE` |
| a verdict citing a path not in the packet, no reason, or no citation | failed, `ASSERTION_INCONCLUSIVE` |

Inconclusive is never a pass. It means the evidence did not show enough:
capture a better screenshot (scroll, wait for the content) and judge again once,
or record the step as failed with that code. The checked result goes on the
finding as `judge: {verdict, code, reason}`.

The actor never overrides a judge verdict. When it disagrees, it captures new
evidence and asks a new judge; it does not argue with the old one.
