# xqa-test-plan — manual test plan

Manual verification checklist for the `xqa-test-plan` skill. Run when adding/changing skill behavior.

## Activation

- [ ] Skill activates on `/xqa-test-plan`
- [ ] Skill activates on implied intent ("what should I QA?")

## Detect state

- [ ] Detect state correctly computes slug for current branch
- [ ] Detect state auto-prunes stale sibling dirs but never current
- [ ] Auto-prune iterates branches forward (applies `branchToSlug`); never tries to invert slug
- [ ] Auto-prune leaves directory when slug-match is uncertain

## Generate flow

- [ ] Handles 0 booted simulators (error), 1 (auto), >1 (prompt)
- [ ] Uses AskUserQuestion (or platform fallback) for simulator selection when >1 booted
- [ ] Passes `--intent` correctly
- [ ] Detects base ref from open PR (`gh pr view`)
- [ ] Falls back to upstream tracking when no PR exists
- [ ] Omits `--base` + emits broader-diff warning when both detection methods fail
- [ ] Never fabricates `origin/HEAD` fallback claim

## Approval

- [ ] Approval loop calls `xqa plan edit` per file
- [ ] Emits Run-go gate after plan approval for all scenarios (never skips; plan approval ≠ run approval)
- [ ] Run-go gate question references expected app state (first scenario's precondition label)
- [ ] Emits flat `- [ ]` checklist for approval (not numbered scenario+steps)

## Run

- [ ] Preflights simulator before dispatching
- [ ] Setup coordination asks per-group before dispatching when scenarios have divergent setups
- [ ] Batches scenarios with identical setup without re-asking
- [ ] Ordering puts non-destructive setups before no-wallet/delete-wallet setups
- [ ] Cancellation handles abort/skip/skip all/rerun

## Gates

- [ ] Does not emit free-text "Reply X / Y / Z" prompts at any decision gate
- [ ] Uses AskUserQuestion for plan approval gate
- [ ] Uses AskUserQuestion for run-go gate
- [ ] Uses AskUserQuestion for sim-state question (multi-profile case)
- [ ] Uses AskUserQuestion for transition-ready prompts between groups
- [ ] Uses AskUserQuestion for destructive-delete seed-backup confirmation
- [ ] Uses AskUserQuestion for existing-plan Rerun/Regenerate/Extend choice
- [ ] Uses AskUserQuestion for post-interruption Resume/Report/Abort choice
- [ ] Still accepts free-form replies (AskUserQuestion's "Other" path)
- [ ] Detects `AskUserQuestion` availability at conversation start
- [ ] When available: uses AskUserQuestion at all 11 gates
- [ ] When unavailable: uses structured platform-fallback format at all 11 gates
- [ ] Fallback format never degenerates into loose "Reply X / Y / describe" prompts

## Report

- [ ] Renders `passed` scenarios with green check (no strikethrough)
- [ ] Renders `failed` scenarios with strikethrough + inline finding + screenshot link
- [ ] Renders `not_run` scenarios with outcome verb (errored/timed_out/aborted) or "(no run record)"
- [ ] Opens with correct sentence based on bucket distribution
- [ ] Footer summary line shows counts when >1 scenario
- [ ] Uses correct `xqa plan report` flags (`--findings`, `--specs`) with default paths from `.xqa/test-plan/<slug>/`
- [ ] Omits `--runs` by default (`scenario-runs.json` resolved from same dir as `findings.json`)

## Rerun / Regenerate / Extend

- [ ] Rerun doesn't regenerate specs
- [ ] Regenerate wipes specs before re-invoking xqa plan
- [ ] Regenerate preserves `runs/`
- [ ] Extend appends scenario-N+1 without re-writing existing scenarios

## PR detection

- [ ] Probe runs in parallel with sim probe and slug computation (single Bash batch)
- [ ] Branch with no open PR: skips PR-integration silently, proceeds local-only
- [ ] Branch with open PR: fetches PR body, computes diagnostic signals
- [ ] Signal A: coverage count identifies items with changed-surface token overlap
- [ ] Signal B: items containing vitest/eslint/pnpm test/ci passes are flagged
- [ ] Signal C: items containing TODO/TBD/??? are flagged
- [ ] Diagnostic emitted before gate (N_covered/M, K CI steps, P placeholders)
- [ ] Same PR body always produces same diagnostic (deterministic)
- [ ] AskUserQuestion gate offers all three options: Use as-is / Enrich / Regenerate
- [ ] "Use as-is" skips planner, converts PR checklist to approval checklist
- [ ] "Enrich" runs planner then merges with PR checklist, dedupes by intent
- [ ] "Regenerate" ignores PR body, runs normal Generate flow
- [ ] PR with no test plan section proceeds with normal Generate flow (no gate shown)

## Update PR

- [ ] Update PR gate fires after report renders when open PR exists (single 3-way gate, not two sequential y/n gates)
- [ ] `Write plan + tACK`: executes write-back then self-tACK posting sequentially, no additional confirmation
- [ ] `Write plan only`: executes write-back only, no self-tACK
- [ ] `Skip`: PR body unchanged, no tACK

## Write-back

- [ ] Posts every approved checklist item verbatim as `- [ ]` bullets (not scenario titles)
- [ ] Single-scenario plan has no subheader; multi-scenario plan has one `### <Scenario name>` per scenario
- [ ] No `[x]` boxes at write-back time — all unchecked
- [ ] Replaces existing `## Test plan` section without corrupting other sections
- [ ] Appends `## Test plan` section when none existed
- [ ] Preserves all content below the original test plan section verbatim
- [ ] Uses `gh pr edit <number> --body-file` (not --body string interpolation)
- [ ] Surfaces `gh pr edit` errors verbatim, no silent retry
- [ ] `## Test plan` bullets stay verbatim after run completes (no scenario-title rewrite, no `<!-- failed -->` injection, no `[x]` flip)
- [ ] Post-run status goes in optional footnote below checklist, not inline in bullets

## Self-tACK

- [ ] No prior tACK: renders proposed comment for transparency, posts immediately (no separate gate — Update PR was the confirmation)
- [ ] Prior tACK identical: reports "no update needed" without gating
- [ ] Prior tACK differs: renders diff before gating update (self-tACK update confirmation gate fires)
- [ ] Update uses `gh api --method PATCH /repos/{owner}/{repo}/issues/comments/<id>` with `--field body=@file`
- [ ] Comment body: first line `tACK`, blank line, all items `[x]`, no commentary
- [ ] Validation: item count matches, all boxes `[x]`, no paraphrase, no invented items
- [ ] Post-run PR body fetch and tACK lookup run in parallel (Probe A + Probe B)
- [ ] tACK body is subset of PR `## Test plan` — every tACK line matches a test plan line verbatim (modulo `[ ]` → `[x]`)
- [ ] tACK item count equals PR `## Test plan` item count (no missing, no extra)
- [ ] Scenario subheaders (`### <name>`) preserved when present in test plan
