# MCP pipeline operations, generate → validate → render → feedback

Operator workflows against the a2ui MCP server. Tool schemas + wrappers:
[mcp-tool-reference](mcp-tool-reference.md).

## Full pipeline, one command

```bash
node scripts/mcp-pipeline.cjs "dashboard with 4 stat cards and a revenue chart"
```

Pipes `intent → generate_ui → validate_schema → a2ui-to-html →
check_anti_patterns` and prints the scores. Fastest whole-stack confirmation
after a change.

## Step-by-step (inspect intermediates)

```bash
node scripts/mcp-call.cjs generate_ui '{"intent":"dashboard with 4 stat cards","mode":"instant"}'
node scripts/mcp-call.cjs validate_schema '{"messages":"<paste-messages-json>"}'
echo '<paste-messages-json>' | node scripts/a2ui-to-html.cjs
node scripts/mcp-call.cjs check_anti_patterns '{"html":"<paste-rendered-html>"}'
```

`validate_schema` is fast and deterministic, batch it over
`packages/gen-ui/engine/corpus/chunks/*.json` templates to surface corpus drift without
re-running the generator.

## Compose from the chunk corpus

When the intent matches a known page-shape (auth flow, dashboard layout, error
shell), prefer `compose_from_chunks` over `generate_ui`:

```bash
node scripts/mcp-call.cjs compose_from_chunks '{"intent":"sign-in card with email + password + OAuth"}'
```

Tier 1 returns a matched chunk's HTML when the retrieval score ≥ 8; otherwise
the LLM picks a `{page, slot_bindings}` plan from a pre-filtered ~30-entry
catalog and the chunk composer materializes it, validating slot-name +
chunk-kind contracts. Returns `{state_id, html, plan, candidates}`.

## Multi-turn refinement (state_id chain)

```bash
# Turn 1 → returns state_id "abc123"
node scripts/mcp-call.cjs compose_from_chunks '{"intent":"sign-in card with email + password"}'
# Turn 2 → updateComponents messages + new state_id
node scripts/mcp-call.cjs refine_composition '{"state_id":"abc123","intent":"add OAuth row with Google + GitHub"}'
```

Two-pass synthesis (locator → modifier), validator-driven retry
(`maxAttempts=2`). Each refinement chains through `parent_state_id`; walk
history with `get_state`. The state cache is in-memory and bounded, after an
MCP server restart, re-run `compose_from_chunks` for a fresh `state_id`.

## Reporting issues

When the engine breaks expectations, fire `report_issue` with the most recent
`state_id` (reporter: `llm` for agent self-fire, `user` for a human request).
The record lands in the engine-internal telemetry store, scratch data for
diagnosis. **[corrected 2026-08-23, ADR-0008 amendment]** that store is
`qa/findings/issues/<issue_id>.json`
(`packages/gen-ui/engine/compose/strategies/zettel/issue-reporter.js:10,26`,
`DEFAULT_STORAGE_ROOT`), not the original `.brain/audit-history/issues/`
path, which was retired. Anything worth durable tracking (a recurring
pattern, a fix proposal) belongs in a GitHub issue or the PR description of
the fixing change.

## Validation checks

`validate_schema` runs a weighted checklist (target aggregate ≥ 80). Common
failures when the corpus drifts: `hasRootComponent` (missing `id: "root"`),
`cardContentModel` (section without Column wrapper, or heading inside section
instead of header), `headingHierarchy` (skipped levels), `flatAdjacency`
(nested components instead of sibling id references). Full list + weights:
`packages/gen-ui/mcp/TOOLS.md` (the `gen-ui` section).

## Feedback loop

Score runs with `submit_feedback` keyed on the `executionId` from
`generate_ui`. The analyzer
(`packages/gen-ui/a2ui/retrieval/feedback/feedback-analyzer.js`) aggregates
`corpus/feedback/*.jsonl` into per-intent trends, promotion candidates
(≥95 score + ≥4 rating across 3+ runs → `npm run feedback:promote`), and the
gap registry (`packages/gen-ui/engine/corpus/gaps/registry.json`).
`npm run feedback:report` surfaces the current state.

### Human signal (gh#668)

Every score in the store except a rating is self-graded, the validator marking
its own homework. Human thumbs are the only outside signal, and they arrive two
ways, both through the SAME function
(`packages/gen-ui/a2ui/retrieval/feedback/submit-feedback.js`) into the same JSONL:

- `submit_feedback` (MCP), and
- `POST /api/feedback` from a rendered surface, today the gen-UI gallery's
  `<agent-feedback-bar-ui>` row (thumbs-up = rating 5, thumbs-down = 2).

`get_training_gaps` ranks weak domains on `blendedScore`
(`retrieval/feedback/human-signal.js`): `0.7 * humanScore + 0.3 * selfScore`
where a domain has both signals, `humanScore` alone where it has no self-grade,
`selfScore` alone where it has no thumbs. **A missing signal is not a zero**, weighting an absent self-grade as 0 would rank a domain humans unanimously
approved below an unrated one the validator liked. So a domain the pipeline
scores 99 and humans thumb down ranks weak, not strong; `selfScore` averages
execution scores plus any `score` a rating carries inline, and `engine` /
`strategy` are triage context on the log, not blend inputs.

Verify the whole loop in a browser with `npm run probe:feedback-loop`. It is
scratch-by-default (temp store + temp screenshot dir); `PORT=…` if 3456 is
taken, `PROBE_REAL_STORE=1` to write the actual corpus feedback log.
