# Exploratory testing - findings, budgets, stop rules

<!-- toc -->
- [1. The run file](#1-the-run-file)
- [2. A finding](#2-a-finding)
- [3. Budget and stop rules](#3-budget-and-stop-rules)
- [4. Product failure or blocked](#4-product-failure-or-blocked)
- [5. Driving the app: locators and expect steps](#5-driving-the-app-locators-and-expect-steps)
- [6. Report](#6-report)
<!-- /toc -->

An exploratory run drives a live app toward a goal and reports what looks
wrong. What it produces is a set of claims, and this contract keeps them
checkable: every finding carries what was expected, what was seen and how to
get there again; every run has a budget and stops on its own; a step that could
not run is reported apart from a product that ran and was wrong.

Consumers: `/multi-agent:test` (the UI Bug Hunter, `commands/sim-test.md`),
`/multi-agent:bug-bash` (one run per charter). Script:
`$HOME/.claude/scripts/explore-findings.mjs`. Run file:
`$HOME/.claude/schemas/explore-run.schema.json`. Assertions that need judgement go
to a fresh judge: `features/assertion-judge.md`.

## 1. The run file

Write the run as it goes to `<log dir>/run.json`: `goal`, `platform`, `budget`,
`usage`, `steps[]`, `findings[]`. One step is one charter-sized move ("change
the quantity on the first cart row and read the total"), not one tap.

```json
{"index": 3, "title": "Change quantity", "status": "failed", "code": "ASSERTION_FAILED",
 "detail": "total stayed 12.00", "trace": "<path of the toolkit's trace.md, if any>"}
```

`status` is `passed | failed | blocked | exhausted`. `code` is the toolkit's
code when the step did not pass; it decides how the step counts (section 4).

## 2. A finding

| Field | Rule |
|---|---|
| `kind` | `issue` (functional defect, fails the run) or `warning` (cosmetic, does not) |
| `severity` | integer 1-5: 1 trivial, 2 low, 3 medium, 4 high, 5 critical |
| `title`, `screen` | short, and where it was seen |
| `expected` | what should have happened, from the spec when there is one (`spec` quotes it) |
| `observed` | what happened instead |
| `repro` | the steps that reach it from launch, at least one |
| `evidence` | at least one `{kind, path}`: `screenshot`, `ui-tree`, `video`, `trace`, `log`, `text` |
| `step` | the step that reported it |

Report a finding the moment its evidence is on disk, not at the end. A finding
that misses a field is listed under "failed the contract" and cannot fail the
run: `explore-findings.mjs report` holds it out of the verdict.

Secrets never appear in a finding. A credential is named, not written:
`secret_ref:<name>`. The script refuses `password: <value>`, `token=<value>`
and the like in any finding text.

## 3. Budget and stop rules

| Budget | Default | Range |
|---|---|---|
| steps | 8 | 1-12 |
| minutes (wall clock) | 10 | 3-15 |
| tokens | 150000 | 20000-400000 |

Values outside the range are clamped and named in `clamped`. Before each new
step run:

```bash
node "$HOME/.claude/scripts/explore-findings.mjs" check --run "<run.json>"
```

`{"stop": true}` ends the run with that `reason`:

- `step-limit`, `time`, `tokens`: the budget is spent.
- `stuck`: the last three steps were blocked and none of them reported a
  finding. Three blocked steps in a row means the environment, not the app, is
  in the way, and more steps would spend the budget learning nothing.

A failed step does not stop the run; only the rules above do.

## 4. Product failure or blocked

| Step | Counts as |
|---|---|
| `failed` + `ASSERTION_FAILED` / `ASSERTION_INCONCLUSIVE` | product failure |
| `CREDENTIAL_MISSING` | blocked: credentials |
| `ENVIRONMENT_UNAVAILABLE` | blocked: environment |
| `SEED_DATA_MISSING` | blocked: seed data |
| `LOCATOR_NOT_FOUND`, `LOCATOR_AMBIGUOUS`, `REPLAY_DIVERGED` | blocked: automation |
| `BUDGET_EXCEEDED` | exhausted (neither) |
| `failed` with no code | blocked: unknown (never silently a product failure) |

A blocked step is a fact about the test setup. It goes in its own report
section, grouped by cause, so "the login screen is broken" and "the run had no
test account" are never the same line.

Verdict, from `explore-findings.mjs report`:

- **failed**: at least one valid `issue`.
- **blocked**: no step ran, or every step was blocked, and no `issue`.
- **passed**: steps ran, no `issue`, not every step blocked. Warnings allowed.

## 5. Driving the app: locators and expect steps

With `@mmerterden/multi-agent-toolkit-mcp` ≥ v3.19.0, prefer these over
coordinates; on an older toolkit, fall back to the UI tree plus frame-centre
taps as before and say so in the report:

- **Locate, then act.** `ui_locate` resolves an element by accessibility
  identifier, by role plus label, or by visible text; `ios_tap_element` /
  `android_tap_element` tap what it resolves. Identifier first, role plus label
  second, text last (text changes with locale). `LOCATOR_NOT_FOUND` and
  `LOCATOR_AMBIGUOUS` are automation blocks, not product findings: narrow the
  locator and retry once before recording the step as blocked.
- **Assert with expect steps.** In `agent_run_steps`, `expect_text`,
  `expect_element` and `expect_url` decide a deterministic check and return
  `verdict: passed | failed | blocked` with a code. Use them for anything
  exact: a total, a label, a route. Only a visual or semantic claim ("the
  layout is not clipped", "the error explains what to do") goes to the judge.
- **Secrets by reference.** The type tools take `secret_ref` instead of the
  value; the value never enters the transcript, the run file or a finding.
- **Traces.** A failed `agent_run_steps` writes a `trace.md`; put its path in
  the step's `trace` and in the finding's `evidence` as `kind: "trace"`.
- **Replay.** `save_as` stores a step sequence that reached a finding, and
  `replay` runs it again. `REPLAY_DIVERGED` means the app no longer follows the
  recorded path: an automation block, and a hint that the screen changed.
- **Web sessions.** Pass `session` to keep parallel browser sessions apart
  (one per bug-bash charter).

Tool inputs are the tool's own schema; the names above are the contract, the
argument shapes are read from the server.

## 6. Report

`explore-findings.mjs report --run <run.json>` prints the verdict, why the run
ended, the budget used, issues before warnings with the most severe first,
blocked steps by cause, and the findings that failed the contract. The UI Bug
Hunter keeps its own screen table and summary around this block.
