# Visual evidence  -  before/after screenshots and the UI flow video

A UI fix that reads as three changed files in a diff is not reviewable. The
reviewer cannot see what was wrong, and the tester cannot see what to look for.
This contract makes the pipeline carry the picture: the state the reporter saw,
the state the fix produces, and where possible a recording of the flow running.

The artefacts live as **Jira attachments** and are referenced from two places -
the Phase 7 Jira comment, where the picture belongs next to the work summary and
the test scenarios, and the PR body, which tells the reviewer they exist.

Consumers: Phase 3 (capture), Phase 5 (video, preferred host), Phase 6 (blocker),
`channels/jira.md` and `channels/pr.md` (render). Gate: `smoke-visual-evidence.sh`.

## 1. When it is required

Decided mechanically from `taskType` plus the changed-file list, never from the
model's reading of the task. A file counts as UI by stack:

| Stack | UI file test |
|---|---|
| iOS | `.swift` declaring `: View`, `: ViewController`, or under a `Views/` path |
| Android | `.kt` with `@Composable`, or declaring `: Activity` / `: Fragment` |
| Web | `.tsx` / `.jsx` / `.vue` / `.svelte` |

| Case | Before | After | Video |
|---|---|---|---|
| `taskType` is `bugfix` AND a UI file changed | required | required | when a tier allows |
| Design work: `taskType` is `component`, or Figma was referenced, or a UI file was **added** | not applicable | required | when a tier allows |
| Anything else | - | - | - |

`state.visualEvidence.required` records the verdict and which rule produced it.
Nothing about this section is optional-by-omission: when it is required and an
artefact is absent, the absence is written down with its reason (section 5).

## 2. Before  -  the reporter's screenshot, or nothing

**The pipeline does not build the old state to photograph it.** Reproducing a
pre-fix screen costs a second build and a second launch on every UI bug, to
recreate evidence the reporter has usually already attached.

So: Phase 0 already fetches the issue. Any image attachment on it (`.png`,
`.jpg`, `.jpeg`, `.gif`, `.heic`) is downloaded to `$WORKTREE/.pipeline/evidence/`
and recorded as `state.visualEvidence.before[]`. More than one is kept - a bug
reported on two platforms has two.

No image on the ticket means no "before". Record the gap
(`before_missing: "ticket carries no image attachment"`) and continue. The
section still renders, saying so.

## 3. After  -  Phase 3, not Phase 5

Phase 5 is the natural home: the simulator is already up. It is also **dropped by
every `autopilot` and `--local` entry** (`phase-5-test.md` TLDR), so a capture
that lives only there produces nothing for unattended runs - which are exactly
the runs where nobody watched the screen.

The capture therefore happens in **Phase 3, after the build+test gate passes**,
which is in every mode's phase set. Phase 5 may add richer evidence on top; it is
never the only source.

```bash
bash $HOME/.claude/scripts/capture-evidence.sh after \
  --task "$TASK_ID" --platform "$PLATFORM" --label "<screen-slug>"
```

What the capture is worth checking against - the per-component catalog, the
measure-never-the-token rule, the presence gate - is
`$HOME/.claude/multi-agent-refs/features/design-conformance.md`. This contract
owns getting the picture; that one owns reading it.

The script boots the device if needed, launches the built app, cleans the status
bar (`ios_status_bar preset:clean` / the Android equivalent) so the frame carries
no clock, carrier or battery noise, captures at device resolution, downscales to
at most 1242px wide, and writes
`$WORKTREE/.pipeline/evidence/<TASK_ID>-<label>-after.png`.

A clean status bar is not cosmetic: without it two captures of the same screen
differ by the clock, which makes every "after" look like a change.

## 4. Video  -  the recording rides on a test run

The flow video is not recorded on its own. It wraps something that drives the
screen, and what drives it is chosen by the user at intake, because running a UI
suite costs minutes and that is their time to spend.

### 4.1 Probe first

`probe-evidence-capability.sh` measures, before the question is asked, what this
machine and this repo can actually do: the UI test target, the tests matching
this change, the device, the recorder CLI, and whether the toolkit MCP is
registered. Detection is delegated to `run-ui-tests.sh detect`, which is also
what runs the tests, so there is one implementation of the answer.

Two rules the probe keeps, and the reason for each:

- **Absence carries a reason.** `no booted simulator, but one is available to
  boot` and `no iOS simulator available on this machine` close the same menu row
  and ask the user for completely different things.
- **Unmeasurable is `null`, never `false`.** With no `adb` on the PATH the probe
  cannot see whether a device is attached; reporting that as "no device" is a
  negative nobody looked for, and it reads exactly like one somebody checked.

### 4.2 Then the question

Phase 0 Step 7.7 asks the depth, with the options built from the probe. A closed
option stays on the list carrying its reason; when every option but the first is
closed, nothing is asked and `testDepthSource` records `forced`. The full menu
rules are in `phases/phase-0-init.md`.

### 4.3 Then the recording

| `testDepth` | Tier | What drives the screen |
|---|---|---|
| `unit+ui` | 1 | The repo's own UI test, selected by the changed files. Most faithful, and it is a test that already runs in CI |
| `unit+mcp` | 2 | `mcp__multi-agent-toolkit__agent_run_steps` drives the flow. No test target needed |
| `unit` | 3 | Nothing. No recording, and the reason is the user's own answer |

The recorder itself is `capture-evidence.sh video start|stop`, which writes
`<task-id>[-<label>]-flow.mp4` into the evidence directory. The label is optional
because one recording per task is the common case, and both renderers cite the
unlabelled form. It is shell rather than MCP for three reasons: the sibling still capture already shells out,
a host with no toolkit MCP registered still produces evidence, and Phase 3 and
Phase 5 are exactly where an MCP call is contested.

**iOS UI test targets are not named, they are detected.** A path containing
`UITests` is not the signal: in the reference app 477 files sit under such a path
and exactly 2 drive the UI, the rest being snapshot tests that render a view and
compare pixels without ever launching the app. The signal is `XCUIApplication`,
the only API that drives another process's interface, which is precisely the
precondition a flow recording has. On Android the signal is the `androidTest`
source set, which is instrumentation by definition.

**A real app has many candidates** - 17 UI test directories on the reference iOS
app, 8 instrumentation modules on the Android one. Detection reports the whole
set and lets the changed files choose; taking the first off a `find` is a guess
wearing a measurement's clothes.

### 4.4 Re-check before recording

The probe runs at intake and the recording happens in Phase 3. A simulator booted
then can be gone by the time the build goes green, so Phase 3 re-measures the
device row alone and downgrades the tier if it has to, recording the transition
(`tier 1 -> 3: simulator no longer booted`). A tier taken from a stale
measurement is a promise the run cannot keep.

### 4.5 Duration, and the recording that shows nothing

The cap is `visualEvidence.maxVideoSeconds` (default 60), read through
`capture-evidence.sh limits` so one value serves the phase doc and the recorder.
Android's `screenrecord` has its own ceiling of 180 seconds that no setting can
lift, so the script clamps to it and says when it did: one preference honoured on
one platform and silently halved on the other is worse than a stated limit.

Both recorders encode on change. A flow over a screen that never moved therefore
produces a valid two-frame mp4 a fraction of a second long - not a broken file,
but not evidence of a flow either. `video stop` says so on stderr when the result
is under a second, and the caller records that as a gap reason. The duration is
never asserted against wall clock anywhere, because doing so fails a correct
capture of a static screen.

`visualEvidence.enabled` turns the whole feature off - capture, upload, both
render sections and the Phase 6 blocker with it.

## 5. Size, and what happens when it does not fit

Jira's attachment ceiling is an instance setting, so it is a preference:
`visualEvidence.maxAttachmentMb` (default `10`).

Order of attempts, each one recorded:

1. Upload as captured.
2. On `413` or a local size overrun: re-encode. Video drops to 720p and a lower
   bitrate; PNG converts to JPEG at quality 80. **Reducing quality is allowed;
   dropping the artefact is not.**
3. Still over: replace the video with a four-frame contact sheet (start, two
   midpoints, end) as a single PNG, and say in the caption that the recording
   exceeded the limit.

Never silently attach nothing.

## 5b. Where the artefacts live  -  the host

Resolved in Phase 6 Step 2.9, recorded as `state.visualEvidence.host`:

| Order | Host | Condition | Stills | Video |
|---|---|---|---|---|
| 1 | `jira` | `jiraId` present | Jira attachments | Jira attachment |
| 2 | `github-public` | no Jira, GitHub remote, public repo | `evidence/<task-id>` orphan branch, embedded in the PR body | none |
| 3 | `github-private` | no Jira, GitHub remote, private repo | same branch, blob permalink in the PR body | none |
| 4 | `none` | anything else, or `githubHost: off` | not published, gap recorded | none |

**GitHub cannot be given a file.** There is no API that attaches an image to an
issue or a pull request; the web uploader posts to an endpoint that needs a
browser session, so no token can drive it. The PR body can only point at
something already hosted, and an orphan branch is the one mechanism a script has
that neither touches the PR diff nor publishes a release.

**A private repo cannot show an inline image.** GitHub renders markdown images
through its own proxy, which carries no credentials for a private repo, so an
embedded `raw.githubusercontent.com` URL renders broken for every reader
including the author. That is worse than a link, because a broken image looks
like missing evidence. The private variant therefore links rather than embeds.

**Video is Jira-only, by decision.** On a GitHub-hosted run none is recorded at
all. An mp4 behind a blob link is a download rather than something a reviewer
opens mid-review, and paying minutes of UI-test time for a recording nobody
watches is worse than saying plainly that there is none. The gap line carries
that reason.

## 6. Rendering

### Jira comment (`channels/jira.md`)

The comment's section order is fixed and nothing may be inserted between the
three sections. The evidence goes **inside** two of them:

- `summary`  -  after the summary sentences, the before/after pair as a Jira wiki
  thumbnail row: `!<TASK>-<label>-before.png|thumbnail! !<TASK>-<label>-after.png|thumbnail!`
  preceded by one line naming which is which in `outputLanguage`
  (`Düzeltme öncesi / Düzeltme sonrası`).
- `test_scenarios`  -  the flow video attached under the scenario it demonstrates:
  `!<TASK>-flow.mp4!`, one line stating the tier used.

Missing artefacts print their reason on the same line, never an empty frame.

### PR body (`channels/pr.md`)

A `visuals` section whose job is to tell the reviewer the evidence exists and
where it lives - the images themselves stay on the ticket:

```markdown
## Görsel Kanıt

- Düzeltme öncesi: <before-filename> (PROJ-123 ekinde)
- Düzeltme sonrası: <after-filename> (PROJ-123 ekinde)
- Akış videosu: <flow-filename>, tier <N> (PROJ-123 ekinde)
```

## 7. Blocker

When section 1 says required and `state.visualEvidence` carries neither an
artefact nor a recorded reason for its absence, **Phase 6 Step 3 blocks** - the
same shape as the `risk` section blocker. A recorded reason is enough to pass:
the gate is against silence, not against an honest "no image on the ticket".

## 8. State

```json
"visualEvidence": {
  "required": true,
  "requiredBy": "bugfix + ui-file-changed",
  "platform": "ios",
  "before": [{"file": "...", "source": "ticket", "jiraFilename": "...", "url": "..."}],
  "after":  [{"file": "...", "capturedAt": "phase-3", "jiraFilename": "...", "url": "..."}],
  "videoTier": 2,
  "video": {"file": "...", "seconds": 41, "jiraFilename": "...", "url": "..."},
  "gaps": [{"what": "before", "reason": "ticket carries no image attachment"}]
}
```
