---
name: uat
description: >
  Full UAT pipeline for a generated SmartStack client app: generate a
  deterministic test plan (`.plantest.yml`) from the LIVE surface (SQL nav/RBAC
  + componentRegistry + controllers), provision a UAT tenant + one user per
  role through the admin API, execute the plan on BOTH axes — backend (every
  endpoint × role: expected status, duration, response size) and frontend (a
  real Chromium driven like a user: login page, sidebar clicks, row clicks,
  real form submits on UAT-marked data, per-page timings/weight/console) — and
  render a single-file HTML report with every metric. Use when the user asks to
  prepare or RUN UAT, verify role permissions end-to-end, or get a UAT report.
argument-hint: "[run|plan|provision|api|ui|report] [Application[/Module[/Section]]] [--dry-run]"
phase: uat
cli: cli/uat-run
allowed-tools: [Read, Glob, Grep, Bash]
---

# uat

User-acceptance testing of a generated SmartStack app, end to end: **what each
role can reach (and must NOT reach), whether the pages chain without errors at
acceptable speed, and whether the real user write journey works** — clicked in
a real browser, not simulated by HTTP alone.

## Commands (sub-CLIs under `cli/`)

| Command | CLI | Role |
|---------|-----|------|
| `/uat run` | `cli/uat-run` | **All-in-one orchestrator**: app health (boots via `ss dev up` if down) → plan (drift-gated) → provision → API axis → UI axis → HTML report, under ONE run id. Health runs first so the booted app materializes the DB connection config the plan discovery reads. |
| `/uat plan` | `cli/uat-plan` | Generate the deterministic `.plantest.yml` + `.signature` (M1, unchanged) |
| `/uat provision` | `cli/uat-provision` | UAT tenant (created + bootstrapped: every application activated via `tenants/{id}/applications/bulk` — a fresh tenant sees 0 roles otherwise) + 1 user per role via the app's admin API, roles joined by the plan's `role_catalog` ids (idempotent, login-verified) → gitignored `uat-users.json` |
| `/uat api` | `cli/uat-api` | Execute `endpoints[]` × roles: login per role (JWT), assert expected status, measure duration + size → `api-results.json` |
| `/uat ui` | `cli/uat-ui` | Execute `routes[]` × roles in a real Chromium (Playwright): real clicks, write journeys, perf + weight + console/network capture → `ui-results.json` |
| `/uat report` | `cli/uat-report` | Render the self-contained `report.html` from a run's artifacts |

Every CLI: `npx --prefer-offline tsx skills/uat/cli/<cli>/index.ts --spec '<JSON>' [--dry_run]`,
JSON envelope on stdout (`lib/output.ts`). **Findings (assertion failures) exit 0
with `allPassed:false` — they are the product; only infrastructure errors
(missing plan/credentials, unreachable app, Playwright absent, login failures)
exit 1.**

## Prerequisites

- Generated SmartStack app with a reachable SQL database (plan discovery) —
  `ss dev up` once so `appsettings.Local.json` exists (DB connection + the
  `Security.InitialAdmin` password that provisioning logs in with).
- **UI axis only**: Playwright. The skills runtime ships the npm package;
  the browser binary is a one-time download:
  `npx playwright install chromium` (run from the deployed skills directory).
  `uat-ui` fails with exactly these instructions when missing.

## The two axes

**Backend (`uat-api`)** — per (endpoint × role): logs in via `POST
/api/auth/login`, calls with `Authorization: Bearer` + `X-Tenant-Slug`,
asserts the plan's expected status and records duration (ms) + response size
(bytes). Covers BOTH strata: the integration controllers (`/api/{module}/{section}`
via `[NavRoute]` — the app-less client form is requalified into the app-rooted
scope, so a 3-level client app keeps its whole integration surface in the plan)
and the screen-driven controllers (`/api/screens/{plural}`, correlated back to
their owning section).

**The expected permission is the DECLARED one** — each endpoint's own
`[RequirePermission]` expression, resolved against the live permission set
(`permission_source: declared`); the legacy verb→action guess is gone (it
planned a `POST …/approve` gated `.Approve` as `.create` — verdicts computed
for the wrong permission). Two findings the plan carries LOUDLY instead of
certifying them silently:
- `ungated: true` — no `[RequirePermission]`, no `[AllowAnonymous]`: every
  authenticated role passes. The plan still expects the deployed behaviour,
  but the report's **Security surface** section names the endpoint as a
  defect (fix = audit-dev-api DEV-API-033) — those green rows are never a
  certification.
- `permission_source: declared-unseeded` — the declared permission exists in
  NO live permission row: every role is denied (lost-seed drift; the
  build-time twin is audit-dev-core DEV-CORE-011, which also compares grants
  BA ↔ seed in both directions — the UAT itself validates the DEPLOYED app,
  spec-conformity belongs to that rule).

DENIED write verbs are always probed:
the role must get exactly 401/403 (the gate answers before model binding — empty
body, nothing persists). ALLOWED write verbs are **NOT** probed by default
(`includeWriteProbes: false`): an empty-body POST would EXECUTE for real against
any action that binds the body with `EmptyBodyBehavior.Allow` (a parameterless
custom/bulk action). Enable `includeWriteProbes` to also check write authorization
— then a real 2xx is surfaced as `real_write_executed` (never a silent pass). The
REAL 2xx write is exercised by the UI axis. Parameterized routes (`{id}`) are
reported as honest skips, never silently dropped.

**Frontend (`uat-ui`)** — per role, a fresh browser context that lives the
plan like a user:

1. **Login through the login page** (`login-email`/`#email` …) — also asserts
   the login journey itself per role.
2. **Navigation by REAL clicks**: sidebar links for `menu_click` routes (with
   collapse/expand handling, goto fallback recorded as a degradation), first
   row click for `click_row_in_parent` (`/:id` pages), placeholder-id goto only
   to prove DENIED verdicts. `anonymous` browses logged out and must be
   redirected to `/login` everywhere.
3. **Assertion per page**: access verdict vs the plan. A denial on a generated
   route is detected THREE ways (it has no single signal): the scaffolded
   `permission-denied` testid / "Accès refusé" copy (fr/en/it/de), a redirect
   AWAY from the requested route (the platform redirects module/section denials
   to `/applications`), or the `/login` redirect for unauthenticated. Also: zero
   non-whitelisted console errors and zero failed (4xx/5xx) requests on allowed
   pages; INDETERMINATE on empty parent lists (never a false denial).
4. **Write journeys** (`writes:true`, default): on each ALLOWED create page —
   fill the real form (testid-first field discovery, EntityLookup pick, typed
   values), submit, expect the POST to succeed; then edit and **delete ONLY the
   rows carrying this run's `UAT-…` marker**. The EDIT fiche is **read-first**
   by default: the driver opens every section via its `section-edit-*` toggle
   BEFORE looking for an input (a page without toggles is direct-edit, no-op),
   and a Save left `disabled` (form not dirty) is an INDETERMINATE fact —
   never a 4s click-timeout painted as an error. A validation-rejected
   synthetic payload is INDETERMINATE, not a failure; a 403 on a permitted
   write is an RBAC finding.
5. **Metrics per page**: nav→ready timings (warn ≥ 3 s, slow ≥ 8 s — plan
   thresholds), transferred bytes + request count, console errors, failed
   requests, screenshot (`all` | `failures` | `none`).

## Coverage & limitations (what the UI axis does NOT exercise yet)

The UI axis is a **route-and-CRUD walker**: it navigates every route, asserts
the access verdict, and runs the create → edit → delete journey on list pages.
It does **not** yet drive **in-page subsets** — tabs, toolbar/row action buttons,
and **custom-action payload dialogs** (`open_modal`) are modeled in the plan
(`route.subsets`) but not clicked. **Kanban** cards and **dashboard** widgets are
observed as pages (access + perf + console) but not interacted with (the row
walker targets table rows). Custom / bulk **action payloads** (`payloadParameters`,
bulk `ids`) are not modeled, so those endpoints are not exercised with a real body
on either axis. These are tracked as follow-up scope, not silent gaps.

## Artifacts (under `projectPath`)

```
.application-test/uat/{application}/
├── {name}.plantest.yml + .signature      # the plan (committable)
├── uat-users.json                        # credentials — enforced-gitignored
└── runs/{runId}/                         # runs — enforced-gitignored
    ├── api-results.json                  # status/duration/size per endpoint×role
    ├── ui-results.json                   # verdict/perf/weight/errors per route×role
    ├── screenshots/{role}/{ROUTE_ID}.png
    └── report.html                       # single-file, no CDN
```

Both result files additionally carry, per row, the plan's resolved `permission`
(+ `controller` on the API axis) and — on failed API calls only — a bounded
`bodyExcerpt` of the server response (ProblemDetails title/detail extracted when
JSON). All three fields are optional: artifacts from older runs still parse.

The report reads failures-first: a **PASS/FAIL verdict banner** with run
duration, inline SVG **charts** (pass/fail donut per axis, latency/readiness
distribution — no CDN, theme tokens), **coverage cards** (executed vs skipped,
broken down by skip reason), then a **Failures first** section grouping every
failed check by cause (RBAC mismatch, 5xx, network, console, failed requests…)
with the permission and the response excerpt inline. Next the **RBAC matrices**
— role × endpoint and role × route heat-tables, grouped by module (collapsed
unless the group has a failure, expand/collapse-all tools, hover a cell for the
expected/actual detail) — then per-role/per-module rollups and the full backend
(method, endpoint, role, permission, assert mode, expected/actual, duration
banded, size, notes + response excerpt) and frontend (route, URL, role,
permission, nav, verdicts, nav/TTFP/ready timings, weight, requests, issues,
screenshot) tables, filterable by role/status/**search** and sortable
(aria-sort). **Failure screenshots are embedded as data URIs** (≤ 300 KB each,
≤ 8 MB total — beyond the caps the report falls back to the relative link with
an explicit note; passing rows always link relatively). Dark-scheme and print
styles ship inline. Open it in any browser — no server, no network — and the
failure evidence travels WITH the file.

## Key spec fields

`uat-run` (orchestrator — the usual entry point):

| Field | Default | Role |
|-------|---------|------|
| `projectPath`, `path` | required | App root + scope `Application[/Module[/Section]]` |
| `refreshPlan` | `auto` | `auto` = generate if missing, regenerate on signature drift; `always`; `never` |
| `bootIfDown` | `true` | Boot via `ss dev up --json` when a side is down |
| `skipProvision/skipApi/skipUi/skipReport` | `false` | Phase toggles |
| `headless` / `screenshots` / `writes` | `true` / `failures` / `true` | UI axis |
| `includeWriteProbes` | `false` | API-axis write-gate probes on ALLOWED writes (see "The two axes" — off by default because an empty-body probe can really execute) |
| `roles`, `runId`, `apiUrl`, `frontendUrl`, `adminEmail/adminPassword`, `tenantName/tenantSlug/emailDomain`, `title` | resolved | Overrides |

Sub-CLIs accept the same vocabulary (`path` or `planFile`, `usersFile`,
`runId`, `outDir`, timeouts) — see each `types.ts` for the full Zod contract.

## Safety model

- Provisioning writes ONLY through the app's own admin API (never SQL), into a
  dedicated UAT tenant; it is idempotent and login-verifies every credential
  (rotating via `change-password` when the backend flags
  `mustChangePassword`). The tenant is BOOTSTRAPPED on every run: all
  applications are activated on it (roles are a global catalogue filtered per
  tenant by its active applications — a fresh B2C tenant sees 0 roles and
  grants 0 access until then). Roles are joined plan → app by the
  `role_catalog` GUIDs the plan carries, never by display name (the API
  localizes names per Accept-Language); a plan without `role_catalog`
  (pre-1.1.0) is refused with "regenerate the plan".
- UI deletes touch ONLY rows created by the run (the `UAT-` marker); the
  plan's destructive-subset policy (`log_only`) still applies to plan subsets.
- `uat-users.json` and `runs/` are appended to the app's `.gitignore`
  automatically (idempotent block).
- Drift: `uat-run` recomputes the source signature before running and
  regenerates a stale plan (policy `refreshPlan`); the report displays the
  signature it ran against.
