# Changelog

All notable changes to BALDART will be documented in this file.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [7.9.1] - 2026-07-24

### Changed

- **Prose diet manuale del trittico delicato** (chiude il programma dieta sul payload
  pesante): `/new` SKILL.md 3.13.0 (70.4→64.5KB, −8.5% — consolidata la regola
  anti-standalone-tracker-Edit ripetuta in 3 sezioni, compressi numeri misurati e
  razionali storici), `new2` 1.13.0 (41.6→40.6KB), `prd` 1.28.0 (37.5→35.9KB — worked
  example worktree → tre failure mode nominati, HARD RULES e costante invariate).
  Metodo /prose-diet integrale: ancore guard censite PRIMA (17/17 verificate
  byte-identiche dopo), verificatore adversariale sul diff (1 perdita trovata —
  l'eccezione phase-boundary del tracker-write — ripristinata riconciliata), golden
  parity 286 obbligazioni + new-parity + skill-renders + codex-parity verdi.

Codex parity: N/A (varianti native Codex non toccate; base bundle Claude-native).

## [7.9.0] - 2026-07-24

Wave "prose diet fleet" — la spina dorsale evoluta per la generazione Claude 5:
22 superfici messe a dieta in un giro orchestrato multi-agente.

### Added

- **Skill maintainer `/prose-diet` v1.0.0** (repo-root `.claude/skills/`, NON distribuita)
  — il metodo codificato: pipeline misura → dieta (6 transizioni Claude 5 + regole
  GPT-5.6) → **verificatore adversariale separato che refuta sul diff** → restore →
  bookkeeping → guard; invarianti intoccabili (gate deterministici, ancore guard, H2,
  slot, costanti); ordinamento peso×frequenza per il programma a loop.
- **`framework/agents/prd-agent-templates.md`** — nuovo modulo on-demand
  (`SECTION=technical-summary|prd-template|plan-template`): i template inline
  dell'agente `prd` esternalizzati (progressive disclosure).

### Changed

- **Prose diet Wave A — 14 agenti** (orchestrata: dieter opus effort-low → verificatore
  adversariale → repair; 34 subagent, 0 errori): prd 1.2.0 (−15.7%, template
  esternalizzati), ui-expert 2.5.0, security-reviewer 1.2.0, plan-auditor 2.3.0,
  doc-reviewer 2.2.0, visual-fidelity-verifier 1.1.0, api-perf-cost-auditor 1.1.0,
  ui-quality-critic 1.2.0, hybrid-ml-architect 1.1.0 (−16.7%), markup-fidelity-verifier
  1.1.0, deep-human-insight 1.1.0, email-deliverability-architect 2.1.0,
  onboarding-architect-lead 1.2.0, legal-counsel-gdpr 1.1.0. Totale ~260KB→241KB.
  6 superfici con perdite normative trovate dal gate e ripristinate dal repair.
  Dettaglio per-agente in `framework/.claude/agents/CHANGELOG.md`.
- **Prose diet Wave B — 8 skill** (solo corpi SKILL.md; references/scripts intoccati):
  worktree-manager 1.11.0, e2e-review 1.8.0 (−8.8%), bug 2.11.0 (ancore lifecycle
  byte-identiche, verificate), ui-design 2.7.0, design-system-init 1.6.0 (−14.4%),
  simplify 1.5.0, ui-implement 1.4.0, prd-add 1.4.0. Totale ~288KB→269KB. 3 skill con
  perdite trovate e ripristinate. Entry per-skill nei rispettivi `CHANGELOG.md`.
- Esclusioni deliberate dal giro automatizzato (restano al giro manuale): `/new`/`new2`
  (golden parity + workflow-paired), `prd` SKILL.md (costanti/validator), skill-creator
  (appena aggiornata), agenti già a 3 strati minori.

Tutti i guard verdi (golden parity 286 obbligazioni, new-parity 29 policy, codex-parity,
agent-runtimes, capabilities, skill-renders). Codex parity: adapter-generated — stessi
`.md` SSOT, transpile TOML invariato.

## [7.8.1] - 2026-07-24

### Changed

- **Prose diet `prd-card-writer` v1.11.0** (context engineering gen-5, secondo giro della
  wave 7.8.0) — l'agente più pesante della fleet (74.4KB, ri-pagato a ogni spawn) scende
  del 10% (1083→1017 righe): rimossi esempi worked-out, incident story e razionali
  misurati. Il gate adversariale sul diff ha trovato 3 perdite normative + 4 puntatori
  minori, tutti ripristinati prima della release — 47 MUST/NEVER/FORBIDDEN, soglie, enum
  e criteri Rule A/A.1/B/C integri; H2 stabili.
  Prossimo target documentato: esternalizzazione dei template inline di `prd.md`
  (progressive disclosure) come intervento dedicato.

Codex parity: adapter-generated (stesso `.md` SSOT, transpile TOML invariato).

## [7.8.0] - 2026-07-24

Wave "Claude 5 context engineering + field guide Fable" — distillazione operativa dei due
articoli Anthropic archiviati in `docs/research/anthropic-context-engineering-claude-5.md`
e `docs/research/anthropic-field-guide-fable-unknowns.md`.

### Added

- **Blind Spot Pass in `/prd` Discovery** (skill `prd` 1.27.0) — Step 2.0: sweep nativo
  degli unknown unknowns all'ingresso del loop (gemello early del Codex Completeness
  Cross-Check di uscita, invariato), alimentato dal nuovo output block
  `## Unknowns & Blind Spots` del `PROFILE=discovery` (`agents/analysis-profiles.md`,
  profile-specific → sopravvive a `OUTPUT=terse`); selezione delle dimensioni
  **architecture-impact first** (checklist order come tie-break); Reference-Implementation
  Check nel processing delle risposte.
- **`links.reference_impl`** (`card-schema.md`, `prd-card-writer` 1.10.0, `/new` 3.12.0) —
  campo card opzionale: path a codice esistente che già implementa il comportamento
  voluto; l'owner lo legge (read-only) e ne reimplementa le SEMANTICHE nello stack del
  progetto. Gemello backend/logic di `links.design_src`. Additivo: il validator card
  deriva i campi dalla matrice, nessun HALT su card legacy. NON è una chiave
  `baldart.config.yml` → la schema-change propagation rule non si applica.
- **Deviation Log in `/new`** (`coder` 2.5.0, `/new` 3.12.0) — policy conservative-option
  nel briefing + campo `deviations:` (planned/hit/did/evidence) nel Completion Report;
  esplicitamente NON un canale di deferral (status per-AC binari invariati); i reviewer
  lo consumano come priority review target via `completion_report_path`. Due nuove RULE
  in `scripts/check-new-parity.mjs` tripwirano deviation log e reference-impl chain.
- **Disciplina context-engineering per skill** (skill-creator 1.2.0) —
  `references/skill-structure.md` § "Context-engineering discipline (Claude 5
  generation)": rules→judgment, examples→interface design, progressive disclosure,
  no repetition, rich references — REQUIRED per ogni skill nuova o revisionata.
- Libreria ricerca: i due articoli Anthropic distillati + "Applicazione decisa v7.8.0"
  con i prossimi target di dieta prioritizzati (audit meccanico densità MUST).

### Changed

- **Prose diet dei root-file primitives** — `AGENTS.md` skeleton 1.20.0 (299→257 righe:
  compressi i MUST codebase-architect / lifecycle / bat-beacon / bug-routing, homogeneity
  guard, markdown-lint gate — ZERO perdita normativa, esempi worked-out e war-story
  rimossi, heading H2 e ancore guard preservati) e `CLAUDE.md` stub 1.8.0 (il paragrafo
  Plan Mode cita la regola tier-ortogonale invece di ri-affermarla).

Codex parity: portable/adapter-generated — primitives e moduli sono cross-tool; la
variante Codex nativa di `/new` riceve deviation policy + ramo `reference_impl` nel
prompt del worker e nella prosa del phase-contract (canale v1 = final-summary, schemi
envelope invariati); guard verdi (`check-new-parity` 31 policy, golden parity 286+
obbligazioni, codex-parity, agent-runtimes, capabilities, skill-renders).

## [7.7.0] - 2026-07-24

### Added

- **Centralizzazione automatica della ricerca cross-project** — la conoscenza
  generalizzabile prodotta da `/research` nei consumer fluisce automaticamente
  nella nuova **libreria centrale** `framework/research-library/` e torna a
  TUTTI i consumer via `baldart update` (payload subtree). Flusso a consenso
  design-time, zero interazione runtime (intent esplicito del maintainer):
  - **Skill `research` v1.3.0**: classificazione end-of-run GENERALIZABLE vs
    PROJECT-SPECIFIC (criteri in `references/centralization.md`; dubbio =
    locale), coda durabile `.baldart/research-candidates.json`, trasporto in
    background, retry self-healing al kickoff di ogni run successiva (niente
    cron/routine), righe di trasparenza obbligatorie nel return finale.
    Cascata reuse pre-flight: libreria locale PRIMA, poi INDEX centrale
    (`[central library]`, silent-skip se assente).
  - **CLI `baldart push --research-only --auto`** (`src/utils/research-sync.js`):
    non-interattivo, scoped alla SOLA `framework/research-library/` (mai
    skill/agenti, mai VERSION/tag), autofix path silenzioso, contamination
    scan **fail-closed** (finding `requires-decision`/`block` → report
    TRATTENUTO in coda con motivo; override manuale = push interattivo),
    trasporto fail-soft (fallimento rete → coda ritenuta, retry alla
    prossima run).
  - Doppio gate indipendente: giudizio semantico (skill, bias conservativo) +
    scanner deterministico (CLI). Il push completo del framework resta
    manuale e interattivo.
  - Codex parity: portable (skill markdown + stesso CLI zero-dep; push
    foreground come ultima azione invece del background task).
  - Nessuna nuova chiave `baldart.config.yml` (ride su `paths.research_dir`)
    → schema-change propagation rule non applicabile.

## [7.6.1] - 2026-07-24

### Changed

- **Output styles `Terse` + `Terse Ultra`: regole ADHD-friendly distillate** da
  [ayghri/i-have-adhd](https://github.com/ayghri/i-have-adhd) (MIT). Import come skill
  REFUTATO (categoria già coperta dagli output styles; una skill di formato always-on è
  un antipattern nel nostro modello). Delle 10 regole, 3 erano genuinamente nuove e sono
  state integrate in entrambi gli style: (1) cap liste a 5 voci con ranking + riga di
  taglio, (2) state marker a inizio risposta nei processi multi-step, (3) chiusura con UN
  solo next step concreto <2 minuti quando resta azione lato utente (unico "closer"
  ammesso — eccezione esplicita alla regola no-postamble). Nessuna nuova config key.
  Codex parity: N/A (output styles Claude-only, gap già documentato v4.50.0).

## [7.6.0] - 2026-07-24

### Added

- **Parent-ownership gate (centralize-first) nella pipeline PRD/new** (#122): nuova regola
  `design-system-protocol.md § Parent-Ownership` — un nuovo elemento UI che rende dentro il
  subtree di un'istanza di una primitive registrata (per `mockup_analysis.layout`) o si ripete
  per-item nel suo scope di iterazione ha quella primitive come OWNER: il default è estendere
  il contratto dell'owner (route `/ds-edit` / binding `reuse-variant`), il render ad-hoc
  sibling richiede `component_bindings[].ownership_exception` sulla card. Domanda di ownership
  nello step 4 della Component Reconciliation di `/prd` (variante Claude + Codex-native +
  Reconciliation-lite), obligation both-sides in `check-prd-native-contract.mjs`, vocabolario
  advisory `DS_PARENT_BYPASS` nei finder di `codexreview` e nei reviewer. Skill prd 1.26.0,
  codexreview 1.3.1; agenti plan-auditor/code-reviewer 2.2.1, ui-expert 2.4.1. Deferral
  dichiarato: nessun lint AST (stack-specifico, precedente v5.10.0). Codex parity: portable.
- **Stray-spec adoption in `/ds-edit` + advisory `DS_SPEC_STRAY` in `ds-gate`** (#123): gli
  spec componente co-located fuori da `${paths.design_system}/components/` producevano tre
  esiti incoerenti (ds-edit crash "missing" al serializer, ds-gate cieco con ok:true,
  workaround manuale). `/ds-edit` P1.1: fallback deterministico su miss del path canonico
  (glob sotto `${paths.components_root}`, STOP su match multipli, adopt route con
  prose/agentic portati esplicitamente nel bundle step 6, stray rimosso dopo la rigenerazione
  canonica + riga INDEX); `ds-reuse-gate` emette l'advisory diff-scoped `DS_SPEC_STRAY`
  (mai blocking, fail-open) con remediation `/ds-edit <Name>`; +2 test. Skill ds-edit 1.3.0.
  Gate natività Codex: CONFERMA (gpt-5.6-terra, effort high, 2 passaggi). Codex parity:
  portable.

## [7.5.1] - 2026-07-24

### Fixed
- **Runtime-ignore backfill: newly-ignored paths are now untracked (#118).** `.gitignore` rules added by the v6.24.0 runtime-artifacts block never untracked paths already in the index, so consumers that committed `.baldart/bat-beacon/` before the rule kept permanent `git status` noise (and a cross-terminal race on `last-retro.json`). New `untrackIgnoredRuntimePaths()` in `src/utils/runtime-ignores.js`: `git rm --cached` + an atomic deletion-only commit built from an isolated temp index (CAS `update-ref`, index rollback on mid-flight error, hooks/signing bypass documented as deliberate plumbing), wired into `update`'s post-update reconcile and the `doctor` fix action (now confirmation-gated — it creates a commit). `checkRuntimeIgnores` reports tracked-but-ignored paths. Hardened through a two-pass Codex nativity gate. Codex parity: portable (CLI Node, identical on both runtimes).

## [7.5.0] - 2026-07-24

### Fixed
- **`/new` workflows: persist-result agent-scribe removed (#117).** The 3 dynamic workflows spawned a final agent told to `Write` the terminal result JSON verbatim to `/tmp` — on some host runtimes the safety classifier denies that write as a forged complete/PASS attestation (retries read as bad-faith tunneling). The scribe is eliminated; the truncation-safe disk source is the Workflow task harness's own output-file (`tasks/<id>.output` terminal `<result>`, also via TaskOutput), and the "return truncated" prose (`review-cycle.md` #102-2, `final-review.md` F.1.5) now reads that. Skill `new` 3.11.0. Codex parity: N/A (workflows are Claude-only; the Codex Model B `result_persist_path` direct-write path is untouched — verified by the Fase 3.6 gate).
- **`agent-discovery-info` hook: capability filter silently dead in consumers (#119).** The hook resolved `capability-gate.js` at `.framework/src/…`, a path the subtree payload never ships — the fail-open left the gated-off filter empty and capability-excluded agents were reported as MISSING on Codex session start. Resolution cascade added: `.framework/src` → local `node_modules/baldart` → global npm root. Codex parity: portable (same dual-runtime script).

### Added
- **Worktree stale advisory + registry port reclamation (#116 F1).** `setup-worktree.sh` (worktree-manager 1.10.0) now prunes, under the existing lock, registry entries whose `path` no longer exists (their ports are reclaimed), emits the additive manifest field `stale_worktrees:` (merged-ancestor + tip≠trunk + clean + no-listener discriminant) surfaced as a non-blocking WARN by `/new` `setup.md` §6a, `/nw`, and `new2.js` (`staleWorktrees`), and makes the port-saturation error enumerate the holding entries. Destruction deliberately stays in the human `/cw` path (adversarial gate verbale in the issue). Codex parity: portable (shared bash).

## [7.4.0] - 2026-07-24

### Fixed
- **`/new` final-review no longer auto-applies a manual-confirmation finding (#111).** `new-final-review.js` gains a deterministic **O5(c)** reconciliation (twin of the existing O5(b) build-claim demotion): a `VERIFIED` finding whose OWN author signals a manual owner-confirmation intent — the `NEEDS_MANUAL_CONFIRMATION` token in title/evidence, or an owner-decision conditional (`confirm … otherwise …`) in `minimal_fix_direction` — is demoted to `NEEDS_MANUAL_CONFIRMATION` **before** the Fix phase, so it surfaces to the human residual gate instead of being auto-applied against a mismatched `classification` field. Discriminator verified (no false demotion of real actionable findings). Claude-only workflow.
- **Backlog-card baseline gate catches trailing prose/markdown the lenient reader dropped (#114, gate-not-fired, Codex-observed).** `validate-card-baseline.js` gains `scanTrailingGarbage()` — a runtime-agnostic, Node-core, dep-free blocker on an indent-0 line that a well-formed machine-generated card can never contain (skipping comments, `---`/`...` markers, `- ` items, `[…]`/`{…}` flow collections, and valid `key:` lines). This is the bottleneck fix for MODE apply-findings appending a raw `## Applied by quality audit …` block that read as a YAML comment to the lenient reader while breaking strict PyYAML (6 cards shipped unparseable with a green gate). The producer is fixed at both reaches: `prd-card-writer` body (Claude) writes the audit trail as a `notes:` block scalar; the **Codex runtime brief** (`prd-agent-brief.mjs`) carries the format rule inline, since the agent body is not in the Codex TOML. Per-file test added. **Codex nativity gate** (`gpt-5.6-terra`, high) CONFIRMED after applying its one constructive refutation (skip col-0 flow collections to remove a false positive on valid YAML). Codex parity: adapter-generated + inline brief + runtime-agnostic backstop.
- **`/e2e-review` honours the authored `test_plan.e2e_required: false` signal (#116 F2).** New pre-flight **Branch A**: a card that set `e2e_required: false` AND has no mockup auto-skips with `reason: e2e_not_required` — an enumerated PRD-time decision the skill consumes without discretion, fixing the case where such a card was forced through because its diff merely touched a `*.tsx` with no markup change. Never skips when a mockup is present. The old 3-AND pure-backend heuristic is unchanged (Branch B).
- **`/new` implement briefing declares an invariant-stub exception (#116 F3).** `implement.md` File Permissions now sanctions a **minimal invariant stub** to a FORBIDDEN domain-override doc when a project pre-commit hook blocks the card's in-scope commit without it AND a repo non-negotiable already assigns the coder that invariant — resolving the fork between violating the FORBIDDEN list and bypassing a hook with `--no-verify` (which stays forbidden). Declared via `[INVARIANT-STUB]` in the completion report for the Final doc-reviewer to curate.

### Notes
- #116 F1 (port-starvation worktree reaper) deferred to the next sweep as a 2-slot adversarial-gate item (it deletes worktrees + touches the cross-consumer port-cap interaction). Plan anchored in the issue.
- #113 closed as not-actionable: the `mandatory-graphify` nag is a Graphify-native pre-tool hook installed by `graphify <platform> install`, not a BALDART artifact — its heuristics live upstream, not in this repo.

## [7.3.1] - 2026-07-23

- bat-beacon 2.2.1: send-time truncation of over-limit legacy outbox bodies (observed live GraphQL 'Body is too long'). Codex parity: portable.

## [7.3.0] - 2026-07-23

### Added

- bat-beacon 2.2.0 channel hardening: `beacon-scan` gains classes `ownership-underscope-respawn` + `harness-security-false-positive` and an `unmodeled-signal-present` safety net (never silent on unmodeled `[...]` prefixes when count==0); transcript uploads verified (parts+manifest re-listed, one retry, `transcript_upload` receipt field); `sent/` retention (30d, `--retention-days`) + verified-upload snapshot deletion; `tools.enabled` parsed from the config (always-unknown bug); stub beacons flagged (`<!-- stub: analysis-missing -->` + mandatory Expected/Observed/Improvement block); inline excerpt capped at 4KB. (#109)
- `/new` inline ownership expansion: on a single-card batch, a hard boundary on a file demonstrably required by an in-scope AC expands the map inline (`[SCOPE-EXPANDED]`) instead of partial + re-spawn; `prd-card-writer` (v1.9.0) derives mount/wiring `files_likely_touched` by following the data flow. (#108)

### Fixed

- Review workflows: a degraded/dead finder can no longer read as a clean pass — top-level `status: complete|degraded`, `mergeReady: false` + `degradedReasons[]` (engine fallback/TIMEOUT/none, dead reviewer, zero agents on non-empty scope, missing required qa), loud args-degeneration guard (the 17ms no-op), and a prose contract: degraded ⇒ re-run inline. (#105)
- Review workflows: degenerate placeholder findings (identical/'test'-like fields) are discarded at fan-in before classification — never VERIFIED, never in residual. (#107)
- Review workflows: diff base pinned to an immutable sha recorded in the receipt (was a moving branch name — concurrent trunk merges polluted diff-scoped gates); fix pass constrained to in-scope fixes only. (#106)
- Consumer worktree residue: `.baldart/bat-beacon/` ignored wholesale (orphan files outside the three old entries); Phase 8 telemetry (tracker archive + skill-runs.jsonl) now lands in a surgical `chore(metrics)` satellite commit instead of permanent untracked files. (#109)

Codex parity: the two review workflows are Claude-only by design (the Codex-native /new has its own review controller, untouched); beacon scripts zero-dep portable; prd-card-writer transpiled automatically. Guards green (286 + 27 + renders + 36 agents).

## [7.2.0] - 2026-07-23

### Added

- `/new` user-constraint discipline: operational constraints stated at batch start ("handle the merge later", "don't push") are registered machine-readably in the tracker's `## User Constraints` at kickoff and re-read FROM DISK by the new blocking Phase 6.0 merge gate (`merge: deferred` → skip merge, branch left ready, `[MERGE-DEFERRED]` in the report); new2 gains the deterministic deferred-merge skip. Codex deferred-merge needs a coordinated engine change — declared follow-up in production-policy.md. new 3.8.0. (#99)
- `/bug` record-only entry mode (`/bug record <desc>`): scaffold → closed-vocabulary fill → exit-0 validation → record-only commit, for bugs fixed outside the skill; consumer pre-commit gate advisory documented. bug 2.10.0. (#101)

### Fixed

- `/new` Codex-native E2E lane: authentic non-zero Playwright telemetry now binds (`completed` OR `failed` with integer exit code — terminality no longer conflated with exit-zero); command identity canonicalization strips the standard capture wrapper (`set -o pipefail` + trailing `2>&1 | tee <abs-temp-path>`, temp roots only, traversal segments rejected) then requires exact equality; harness-start failures get one in-budget retry before reporting a gap (stale-mutex recovery). Codex-native /new 2.2.3. (#95 #96 #97)
- Shared git-config false attribution (#91) recurrence closed as duplicate — fix shipped in v7.1.2. (#98)
- ds-gate 7-vs-40 sha mismatch closed as recurrence of #93-F3 — `shaEq` shipped in v7.1.2 handles the mixed fleet without manual realignment. (#104)

Codex parity: #95–#97 gate-verified (traversal refutation fixed in second pass); #99 Claude-enforced now with the Codex engine change declared as follow-up; #101 portable prose.

## [7.1.2] - 2026-07-23

### Fixed

- `/new` Codex-native engine: cleanup verification scoped to run-owned resource identities with exact offending IDs in `CLEANUP_INCOMPLETE` evidence (concurrent runs' resources can no longer fail a completed delivery); worker broker attributes shared common-dir `.git/config` deltas to the leaf worker ONLY with writer attribution from telemetry (otherwise a non-aborting `concurrent_shared_mutation` journal event with key-level evidence; downgrade scoped to `config` only — shared hooks/info stay attributable); union review budget proven to cover the lane's worst-case legal sequence up front, with gate-ordered levers (grounding reclaim first, companion downgrade as a declared last-resort exception recorded in the plan). Codex-native /new 2.2.2. (#90 #91 #94)
- `/new` team-mode: the per-card commit sequence decides outcome from REAL HEAD advancement (never exit 0), resets inherited staging between cards, retries bounded with a loud `[COMMIT-BLOCKED]`, and guards against commit collapse (file-set ⊆ ownership map); the card-review fix pass runs under `agentSafe` and reconciles from on-disk state when the fixer dies. (#92)
- `/new` + satellites, 7-finding wave: mid-implementation agent death now inspects the worktree and resumes via the CAP-HANDOFF machine (never a clean re-spawn over partial work; Codex twin in recovery.md); `beacon-scan` matches closed-set codes (`[AGENT-CRASH]`, …) tolerant of markdown list/emphasis prefixes and language-independent; sha comparison normalized (`shaEq`, 7-vs-40 prefix match) killing the structural `DS_COMPONENT_STALE` false positive; `checksFailed` is merge-blocking only when corroborated by a gateTable FAIL row and the fixer lints diff-scoped; markdownlint recipes assert linted-file count > 0 before trusting exit 0 (5 locations); the e2e-review skip clause filters `components_primitives` by UI-source extension. Versions: new 3.7.1, bat-beacon 2.1.0, e2e-review 1.6.1, prd 1.24.1, qa-sentinel 2.1.5, AGENTS primitive 1.19.2. (#93)

Codex parity: #90/#91/#94 are Codex-native mechanics, gate-verified with two constructive refutations fixed (config-only downgrade, gate-ordered levers + declared exception); #92/#93 prose + workflow fixes with Codex twins where applicable. Guards green (286 native obligations, 27 parity policies, renders).

## [7.1.1] - 2026-07-23

### Fixed

- `/prd` graph-integrity wave: `prd-card-writer` (v1.8.0) now self-validates against the whole-pack gate (`validate-prd-pack.mjs`, exit-0 return) with the composite constraint documented (parallel_group == containing group level; cap-split renumbers downstream levels); FEAT-id collision closed at three points — global allocation scan (origin/trunk + all worktree heads), pre-commit re-check (HALT + renumber), post-merge re-check (renumber + reseal) — on BOTH runtimes; staged JSON artifacts formatted BEFORE `prd-final-gate seal`; task-spine fallback generalized from "Codex" to "host without task spine" via `plan_surface: file-backed`; state-template YAML blocks fenced (markdownlint-clean). prd 1.24.0. (#85 #86 #87)
- `return-contract-protocol.md`: a memory-maintenance note can NEVER be an agent's final message — it appends to the report, it never replaces it. (#86)
- `/new` Codex-native planner: typed `BASE_SNAPSHOT_MISSING_SELECTOR` when the frozen inventory lacks the selector (was a misleading `FRAMEWORK_DEFECT EPIC_EMPTY`); `EPIC_EMPTY` reserved for a present tracker with zero children; `EPIC_NO_ELIGIBLE_CHILDREN` for trackerless related matches; EPIC_* map to `NOT_READY`. (#88)
- `/new` Codex-native planner honors `execution_strategy.recommended_mode: sequential` (`EPIC_SEQUENTIAL_PARTITION_REQUIRED` + deterministic dependency-ordered `partition_proposal`) and attaches the safe partition to `BUDGET_PLAN_UNFIT`; the immutable worker cap and fail-closed behavior are unchanged. Codex-native /new 2.2.1. (#89)

Codex parity: #85–#87 replicated in the prd Codex-native phases (gate-refuted first pass fixed); #88/#89 are Codex-runtime mechanics confirmed native by the Codex gate (gpt-5.6-terra, effort high); golden-parity guards green (286 + 669 obligations).

## [7.1.0] - 2026-07-23

### Added

- `/email` init scaffolder — `email-init-scaffold.mjs` (zero-dep): discovery JSON → schema-valid registry entries (`TODO-curate` markers) + `--ds-bootstrap` create-if-missing email DS declination; wired into the init flow. Codex parity: portable. (#80)
- `/ui-design` `study` lane — read-only exploratory lane (discovery + compressed Design Read + 2–3 decisive directions with token-anchored trade-offs + choice gate); approval routes into the full pipeline reusing done work; Step 0 exception so decision requests on existing surfaces reach the lane. Codex parity: portable (textual choice fallback declared). (#84)

### Fixed

- serialize-spec data-loss fail-safe: HEAD-only bundles no longer clobber curated spec prose (existing body re-read and preserved, `prose_preserved` counter); `source_sha` always quoted (leading-zero YAML corruption); lint-clean fill output (MD012/MD036/MD056); `--verify-usage` grouped by cause + projection-consumer downgrade to info; ds-reuse-gate emits `new-component` info instead of `DS_COMPONENT_STALE` for sources absent from the baseline ref. (#82)
- allocate-id.sh: lockdir heartbeat (live-but-slow holder no longer steal-eligible), stale registry entry prune inside `reserve`, progress logs on stderr; AGENTS.md primitive 1.19.1 scopes the terminal retrospective + `/bat-beacon` to the session principal only (delegated/companion tasks never run it). (#83)
- baldart-update: troubleshooting recipe for rebasing across the v6.37.0 framework relocation. (#81)

Codex parity: portable (all fixes are markdown prose + zero-dep scripts); #80/#84 passed the Codex nativity gate (gpt-5.6-terra, effort high) with 3 constructive refutations fixed.

## [Unreleased]

## [7.0.2] - 2026-07-23

### Fixed

- **Migrazione di consumer legacy** — normalizza anche il bulk symlink
  `agents/` preesistente, sostituendo il target assoluto storico con quello
  relativo portabile prima della riconciliazione dei runtime.

## [7.0.1] - 2026-07-23

### Fixed

- **Consumer clone portability** — il bulk symlink `agents/` ora usa un
  target relativo a `.framework/framework/agents`; una checkout clonata o
  spostata non conserva più il path assoluto dell'installazione originaria.

## [7.0.0] - 2026-07-22

### Breaking

- **Distribuzione interamente privata** — repository, CLI, payload, issue e
  transcript bat-beacon passano a GitHub Releases private. Il package npm resta
  solo come release ponte con `baldart private-bootstrap`; nessuna release v7+
  ordinaria viene pubblicata su npm. `git subtree` è sostituito da bundle
  payload versionati e verificati.

### Added

- **Access layer unico** — `BALDART_GH_TOKEN` → `GH_TOKEN` → credential store
  `gh`, resolver `state.framework_repo`, profili `payload-read`, `beacon-write`,
  `contribute`, tassonomia distinta per auth/permission/not-found/rate-limit/
  offline/integrity e comandi `baldart auth setup|status [--json]`.
- **Asset privati allowlisted** — CLI e framework separati, manifest file-level,
  `SHA256SUMS`, workflow tag-driven e installazione `.framework` atomica dopo
  verifica completa. State ledger v2 registra distribuzione, asset e checksum,
  mai credenziali.
- **Guard privacy e migrazione zero-loss** — consumer pubblici rifiutati;
  visibilità non verificabile richiede `--ack-private-storage`; update/migrate
  confrontano il payload installato col release asset esatto e bloccano drift
  non preservato.
- **Bat-beacon Issues-only** (skill v2.0.0) — issue idempotente tramite
  `beacon-id`, transcript completo gzip+base64 in commenti da 48.000 caratteri,
  manifest SHA-256, retry dei chunk mancanti e receipt in outbox fino al
  completamento. `batcave-transcript-digest` ricostruisce e verifica i commenti.
- **Contributi via PR privata** (`baldart-push` v2.0.0) — diff contro asset
  pristine, scan contaminazione/segreti, clone temporaneo, mapping nel payload,
  branch `codex/...` e PR privata; nessun file consumer.
- **Automazioni private** — GitHub App cross-repo preferita, secret token come
  fallback, outbox bat-beacon preservata come artifact; cron e Claude Code
  Cloud diagnosticano esplicitamente l'assenza di accesso.
- Guida `framework/docs/PRIVATE-DISTRIBUTION.md`, bootstrap, recovery e release
  protocol aggiornati. Claude/Codex parity: payload/runtime condiviso e adapter
  specifici, senza regressioni ai path installati.

## [6.37.0] - 2026-07-22

### Added

- **`/prd` Step 3.pre — Surface Coverage Checkpoint** (skill prd v1.23.0, variante Codex v0.12.0) — prima di qualsiasi lavoro mockup (incluso Step 3.0) l'utente valida il perimetro completo delle superfici toccate dalla feature: schermate nuove, schermate esistenti modificate (con `isa_ref` al touchpoint ISA che le implica) e superfici individuate ma escluse dallo scope UI **con motivo obbligatorio** — l'esclusione silenziosa non è più uno stato legale. Sposta la validazione di copertura dal punto più costoso (iterazione post-mockup/post-build) al più economico (una tabella prima di disegnare); chiude la classe di fallimento FEAT-0068 (perimetro ISA errato non visto fino all'implementazione). Elastico sul Tier Gate 2.1: STOP pieno su `lane: full` / ≥2 superfici / esclusione HIGH, inline `auto-validated` sul caso compact banale. `screens_in_scope[]` ha ora un unico punto di costruzione (righe taggate `origin`/`kind`) consumato da Branch B handoff, Hybrid routing e card; nuovo blocco `surface_coverage` nello state template; correzioni isa-origin sincronizzano la Integration Surface Map nello stesso edit. Codex parity: portable + gate bounded nativo **G22** (interaction-policy, phase-contract, playbook ui-design) con ancore nel golden-parity guard `check-prd-native-contract.mjs`. Nessuna nuova chiave `baldart.config.yml`.

## [6.36.0] - 2026-07-22

### Added

- **Codex Multi-agent v2 opt-in nativo** — install, update e migrate dei consumer con Codex abilitato aggiungono `multi_agent_v2 = true` alla `.codex/config.toml` di progetto. Merge additivo: configurazione esistente preservata e un valore esplicito `true`/`false` non viene sovrascritto. L'adapter usa i tool v2 nativi (`spawn_agent`, mailbox, follow-up, interrupt, tree names); nessuna dipendenza dai job CSV rimossi in Codex 0.145.0. Codex parity: adapter-generated; Claude invariato.
- **Fleet completa e skill portabili su v2** — 35/35 agenti dichiarano sandbox Codex; 18 skill orchestratrici usano ruolo `agent_type`, identità `task_name` valida (`lowercase_digits_underscores`), fork contestuale esplicito e lifecycle v2 senza degradazioni inline/sequenziali. `/prd` e `/prd-add` usano host-v2 come backend primario con fallback session-wide; `/new` rileva v2 ma resta esplicitamente `subprocess-required` finché i worker host non possono entrare nella transazione Model B (receipt, accounting, rollback, worktree/process-group cleanup). Claude invariato.
- **Guard CI Multi-agent v2** — `scripts/check-multi-agent-v2.mjs` controlla tutti i 45 entrypoint skill, non solo i tre bundle render-mode, e blocca degradazioni obsolete, spawn ambigui, `task_name` illegali e binding senza ruolo/task/fork. Il merge opt-in preserva forme TOML scalar/table/inline/dotted/quoted e boundary `[[array-of-tables]]`. `doctor` confronta TOML, CLI effettivo e l'intero artifact agente rigenerato byte-per-byte: segnala v2 attivo, mancante, disabilitato/ignorato per trust-sessione e ogni drift/override di prompt, identità, modello, effort o sandbox.

## [6.36.2] - 2026-07-22

### Fixed

- **Config perso a ogni update quando lo stash pop fallisce (la vera root
  cause del "configure non salva mai")**: l'auto-stash pre-update è blanket;
  se il pop finale conflitta su QUALUNQUE path (tipicamente
  `.baldart/generated/*` sporchi che l'update rigenera) fallisce in blocco e
  si porta via anche le risposte non committate di `baldart.config.yml` — al
  giro dopo il detector le vede mancanti e riscrive i default. Ora entrambi i
  path di pop (finalizer reset-and-reapply + flusso classico) eseguono un
  **recovery selettivo automatico** del config dallo stash: i valori
  dell'utente VINCONO, il file su disco contribuisce solo le chiavi nuove
  (mergePreserving). Best-effort, mai throw. Riprodotto e verificato sul
  consumer reale (capabilities email-architect + has_email_layer persi due
  volte, recuperati dallo stash).

## [6.36.1] - 2026-07-22

### Fixed

- **Version regression su npm `latest`**: la v6.35.1 è stata taggata mentre una
  sessione parallela aveva già rilasciato la v6.36.0 — il dist-tag `latest` è
  regredito a 6.35.1. Questa release ricongiunge la linea: contiene sia la
  wave Codex multi-agent v2 (6.36.0) sia il presence-backfill delle chiavi
  config (6.35.1). Nessun altro cambiamento.

## [6.35.1] - 2026-07-22

### Fixed

- **Nag ricorrente "New config keys in this version" dopo un configure
  annullato**: `update` proponeva `baldart configure` per le chiavi nuove, ma
  se il run interattivo veniva rifiutato o interrotto NULLA veniva scritto —
  le chiavi restavano assenti e il prompt tornava a ogni update. Ora la
  presenza è garantita in ogni uscita: (a) `update`, su decline, backfilla le
  chiavi mancanti coi DEFAULT del template (valori esistenti intoccati);
  (b) `configure`, sul cancel del prompt interno, fa lo stesso presence-
  backfill (solo default template, mai i valori autodetected non confermati).
  Popolarle con valori veri resta un `baldart configure` completo.

## [6.35.0] - 2026-07-22

### Added — registro email "no-grep" (blocco comportamentale) + integrazione wiki

- **Blocco comportamentale nello schema del registro**
  (`email-protocol.md` § registry): `sender`, `send_conditions`, `frequency`,
  `attachments`, `code_refs` (call sites `path:line`), `related` (PRD/ADR),
  `last_verified`. **No-grep principle**: una domanda documentale su un'email
  ("quando parte? cosa contiene? da dove viene inviata?") si risponde dalla
  entry, mai grep-ando la codebase; `code_refs`+`last_verified` sono l'ancora
  che doc-reviewer e la routine `email-align` ri-verificano
  (`EMAIL_CONTEXT_STALE`).
- **Skill `/email` v1.1.0**: `email-registry-check.mjs` parsa le
  block-sequence annidate ed emette advisory **`CONTEXT_GAP`** (non-bloccanti,
  testate) sulle entry `active` prive di `sender`/`code_refs`/`last_verified`;
  `/email init` popola il blocco direttamente dagli angoli dello sweep
  (by-provider → `code_refs`, by-content → `sender`); Step 7 REGISTER lo
  richiede; nuovo **Step 7b Wiki tap** (gated `features.has_wiki_overlay`):
  flussi email cross-document → `synthesis_candidate` in
  `${paths.wiki_dir}/log.md`.
- **wiki-curator v1.1.0**: nuova fonte di candidati sintesi slot-gated
  `{{#has_email_layer}}` — ≥3 entry del registro che condividono un
  lifecycle/flow senza pagina → flow-overview synthesis; entry cambiata dopo
  il `last_updated` di una pagina che la cita → freshness trigger. Consuma il
  registro, mai lo edita.
- Nessuna nuova chiave `baldart.config.yml` (campi opzionali dello schema
  registro + flag esistenti) → schema-change propagation rule N/A. Codex
  parity: portable.

## [6.34.0] - 2026-07-22

### Added — capability `email-architect`: skill madre `/email`, registro email, copy profile-driven, DS email, deliverability v2

- **Nuova capability `email-architect`** (`framework/capabilities/registry.yml`,
  `includes: [design]`; `email-deliverability-architect` condiviso con
  `website`): creazione email end-to-end con disciplina reuse-first.
- **Nuovo modulo SSOT `framework/agents/email-protocol.md`** (`SECTION=`
  dispatch): tassonomia chiusa (transactional/notification/lifecycle/digest/
  marketing), contratto anatomico (subject≤45/preheader/CTA-budget/footer
  compliance), vincoli HTML-email (600px HTML-lite, dark-mode a tre client con
  contrast headroom, bulletproof CTA, plain-text part), contratti
  `PROFILE=` per il copy, disciplina del registro, declinazione DS email,
  invarianti di deliverability field-verified (distillate dall'esperienza di
  produzione: reply-to on-domain mai freemail / `FREEMAIL_FORGED_REPLYTO`
  misurato 7.7→10/10, runbook DMARC a stadi con gate di ricezione — zero
  report ≠ pulito, tracking OFF di default, gate go-live mail-tester ≥9 +
  multi-client, warm-up SNDS/JMRP).
- **Nuova skill `/email` v1.0.0** (gated `features.has_email_layer`): triage →
  **pre-flight sul registro BLOCKING** (REUSE/EXTEND/NEW — mai ricreare una
  email esistente) → cascade DS email (create-if-missing) → copy via
  `email-copywriter` → HTML via `ui-expert` email mode → spam-risk gate ≥85
  via `email-deliverability-architect` → REGISTER nello stesso change. Modi
  `review`/`audit`/`init`. Validator zero-dep
  `scripts/email-registry-check.mjs` (schema + `EMAIL_ENTRY_ORPHANED` +
  `EMAIL_REGISTRY_DRIFT` via `--scan`, lookup keyword per il pre-flight).
- **Registro email centrale** (`paths.email_registry`, default
  `docs/emails/registry.yml`): schema per-entry (key/type/trigger/audience/
  goal/template_path/profile/locales/status), drift codes canonici. Insegnato
  a **codebase-architect** (nuovo § Shared block: Email Registry Lookup in
  `analysis-profiles.md` — registro letto PRIMA di ogni grep), **doc-reviewer
  v2.1.0** (Email Registry Scope slot-gated `{{#has_email_layer}}` — cura +
  `EMAIL_CONTEXT_STALE`), **context-primer v1.2.0** (priming email-shaped).
- **DS email**: declinazione `${paths.design_system}/email/` (subset token
  email-safe + component specs) — regole in `email-protocol.md`
  § design-system; `design-system-protocol.md` aggiunge il pointer (drift
  codes riusati verbatim, modern-CSS escluso). **ui-expert v2.4.0** guadagna
  «Email Mode»; **ui-design v2.5.0** il ramo email vertical (wrapped
  `{{#cap_email_architect}}`).
- **Agenti**: nuovo **email-copywriter v1.0.0** (sonnet/default, Codex
  inherit/medium — scelte utente; PROFILE= chiuso a 6 profili, solo copy);
  **email-deliverability-architect v2.0.0** (sezione BINDING «Field-verified
  invariants», owner del § deliverability). REGISTRY.md aggiornato.
- **Schema-change propagation rule APPLICATA**: `features.has_email_layer` +
  `paths.email_registry` propagate su template + `configure` (prompt flag +
  default path) + `update` (detector generico sul template) + `doctor`
  (advisory `email-registry-missing`) + slot flag `has_email_layer` in
  `agent-slots.js` + CHANGELOG.
- **State-of-the-art wave (stessa release)**: (1) **Gate 0 deterministico**
  `email-craft-check.mjs` (floor zero-dep pre-send: freemail reply-to,
  dark-mode meta, preheader, width>600, alt, link budget/shortener,
  unsubscribe, asset esterni, image-only, subject spam — testato 12/12 su
  fixture negativa, 0 su positiva) PRIMA del review LLM; (2) **Step 6b
  verify-with-eyes**: render triplo light/dark/mobile via Playwright +
  `ui-quality-critic` (generator ≠ evaluator anche sul canale email; degrade
  onesto su runtime non-multimodali); (3) nuova **routine settimanale
  `email-align`** (doc-reviewer: scan deterministico + curation, report
  committato — safety net del gate per-task); (4) **`email-copywriter`
  `memory: project`** (brand voice appresa, solo preferenze confermate);
  (5) **`/email smoke`** opt-in: invio reale via `email-smoke.mjs` (Resend
  HTTP API zero-dep; `RESEND_API_KEY`+`EMAIL_SMOKE_RECIPIENTS` da env o
  `~/.baldart/secrets.yml`, MAI hardcoded nel payload; la verifica del
  ricevuto resta gate umano). Capability `email-architect`: `routines:
  [email-align]` + `secrets: [RESEND_API_KEY]`.
- **Integrazione end-to-end nel flusso PRD → card → /new (stessa release,
  tutta parametrica su `features.has_email_layer` — flag off = no-op)**:
  `/prd` v1.22.0 guadagna l'**Email Pre-flight Gate** in Discovery (dopo ISA:
  REUSE/EXTEND/NEW per ogni email della feature contro il registro, decisione
  + key nel PRD; registry assente → `/email init` prima delle card);
  **prd-card-writer v1.7.0** emette card email SPECIALIZZATE (Rule A «Email
  cards»: template+copy `ui-expert` split dal send-wiring `coder`,
  `links.email_key` + decisione sulla card, AC di REGISTER per le NEW;
  `card-schema.md` aggiunge la riga `links.email_key`); `/new` v3.6.0 guadagna
  lo **step 6a2 email-build delegation gate** (il gemello email del
  mockup-first 6a: la card email delega il build alla pipeline `/email` —
  pre-flight, copy, EMAIL MODE, Gate 0, spam-risk, REGISTER — `owner_agent`
  mantiene solo review/ownership; fallback dichiarato senza skill linkata;
  `new2.js` inline questa release, deferral documentato come per
  /ui-implement). Guard verdi: check-new-parity (27 policy),
  check-new-native-contract (283 obbligazioni), check-codex-parity.
- **Nuova disciplina maintainer VINCOLANTE** (CLAUDE.md § "Cross-surface
  integration discipline"): ogni nuova capability/skill/agente richiede
  l'audit trasversale dei punti di innesto (pipeline PRD→card→/new, retrieval,
  curation, superfici adiacenti, meccanica, guard) NELLO STESSO change —
  lezione di questa release (il layer email era perfetto ma il flusso
  orchestrato non ne sapeva nulla finché l'audit non l'ha innestato).
- Codex parity: **portable** (modulo markdown + script zero-dep + agente
  transpilato in `.codex/agents/email-copywriter.toml`; nessun costrutto
  Claude-only nel flusso — decision gates degradano a prompt piani per
  `runtime-portability-protocol.md`; il render Step 6b degrada a Gate 0-only
  dichiarato su runtime non-multimodali).
## [6.33.6] - 2026-07-21

### Fixed

- **Il gate markdown è diff-scoped per davvero ([#77](https://github.com/antbald/BALDART/issues/77))** — `markdownlint-cli2` UNISCE i glob della config consumer (`**/*.md`) agli argomenti CLI: la run «sui due file» scansionava 70k+ file (kill manuale). `--no-globs` ora obbligatorio in `qa-sentinel` (2.1.4), `/check`, `/new` inline fast-lane e nel collo di bottiglia cross-runtime — il primitivo `AGENTS.md` § Workflow Gates (1.18.5), la prosa che l'orchestratore Codex legge davvero (finding del gate Codex: il body di qa-sentinel non entra nel TOML compatto; nessun markdownlint nel path `/new` Codex-nativo). Gate Codex: CONFERMA al 2° passaggio. Codex parity: portable (primitivo generato cross-runtime).
- **`/bug` impone il routing i18n specializzato ([#77](https://github.com/antbald/BALDART/issues/77))** — skill `bug` 2.8.0: Phase 4 aggiunge il bucket i18n al routing del fix (ownership citata da `i18n-protocol.md`: le traduzioni nei TARGET locale sono di `i18n-translator`, MAI inline dell'orchestratore/fixer) e Phase 5 Tier 0 il check deterministico BLOCKING (locale target toccati senza pass `i18n-translator` registrato). Nella run reale l'orchestratore aveva tradotto EN/FR/ES/DE inline, intercettato solo dall'utente. Codex parity: portable (skill symlinkata raw, prosa orchestratore).
- **`ui-expert`: test scope dal blast radius + report byte-esatto ([#75](https://github.com/antbald/BALDART/issues/75))** — «22/22 PASS» dichiarato con 3 test consumer rossi mai eseguiti: nuova sezione BINDING «Verification & Test Scope» (2.3.0) — su un componente condiviso/contratto di primitive i test si derivano da co-located + consumer diretti (importers), e il completion report elenca i file di test eseguiti byte-esatti («N/N PASS» senza elenco = incompleto). Parity Codex (finding del gate): il duty entra anche nel contratto compatto `CODEX_NATIVE_ROLE_CONTRACTS` del transpiler (1098/1100 byte, golden-parity verdi). Gate Codex: CONFERMA al 2° passaggio.

## [6.33.5] - 2026-07-21

### Fixed

- **`/ui-design` non degrada più a source-review quando il wrapper `playwright-cli` esterno manca ([#73](https://github.com/antbald/BALDART/issues/73))** — 4ª occorrenza del pattern (#51/#58/#61): il fix v6.29.2 viveva in `webapp-testing` e `bug` Phase 1 ma i siti di cattura di `ui-design` (Step B + `evaluation.md` Step 2) dicevano solo «Capture via the Playwright CLI», e su Codex il flow risolveva il wrapper di terze parti. Ora entrambi citano `webapp-testing § Playwright resolution` (repo-first; wrapper esterno rotto = fallback signal, mai blocker; degradazione solo dichiarata). Skill `ui-design` 2.4.0 → 2.4.1. Gate Codex: CONFERMA (gpt-5.6-terra, effort high). Codex parity: portable.
- **`baldart update` non spazza più un registry i18n dirty di un'altra sessione nel commit di reconcile ([#76](https://github.com/antbald/BALDART/issues/76))** — `i18nRegistryPatterns()` classificava managed il registry (e la sua directory) anche tracked+modified, ma la CLI non ne scrive mai il contenuto (coder STEP 9 / `/i18n` lo popolano in sessione): l'inclusione wholesale era un'over-estensione di v4.89.2 contro la sua stessa invariante. Ora semantica research-seed (managed SOLO se untracked) + il commit di reconcile usa il pathspec esplicito `git commit -- <managed>` — un file consumer STAGED da una sessione concorrente resta nell'index (finding del gate Codex, verificato su simple-git). Gate Codex: CONFERMA al 2° passaggio. Codex parity: portable (CLI runtime-neutra).
- **`coder`: divieto BINDING di bypass degli hook ([#78](https://github.com/antbald/BALDART/issues/78))** — nessuna regola `--no-verify` esisteva nel framework (viveva solo nell'harness Claude Code, che Codex non ha e che il messaggio di un hook consumer può contraddire): un coder ha bypassato un gate di produzione «come suggerito dall'hook stesso». Nuova Forbidden Action: `--no-verify` (o qualunque bypass) richiede autorizzazione ESPLICITA dell'utente citata nel briefing; la prosa di un hook non è autorizzazione; hook bloccante → soddisfatto o STUCK-ESCALATE. Coder 2.3.0 → 2.4.0. Codex parity: adapter-generated (transpilata nel .toml).
- **`analysis-profiles` PROFILE=bug: audit di vincoli/schema exhaustive-by-grep ([#78](https://github.com/antbald/BALDART/issues/78))** — l'architect enumerava le FK leggendo la migration «canonica» e ha mancato una quinta FK RESTRICT aggiunta 3 settimane dopo (presa solo dal plan-auditor; una migration di produzione sarebbe fallita a metà). Regola BINDING: vincoli di una tabella = grep dell'INTERA directory migrations per ogni riferimento, mai solo la migration che introduce il modello.
- **`analysis-profiles`: output block `## Repro surface` ([#79](https://github.com/antbald/BALDART/issues/79))** — la mappa file di PROFILE=bug non diceva DOVE riprodurre: con split mobile/desktop dietro shell il binding route↔componente↔viewport non è deducibile dal path (~6 riproduzioni buttate su una URL sbagliata). Obbligatorio per PROFILE=ui, e per PROFILE=bug quando il sintomo è UI: per ogni componente citato la route che lo monta + viewport + URL esatta.
- **`/bug` tratta l'ambiente di riproduzione come volatile ([#79](https://github.com/antbald/BALDART/issues/79))** — Phase 1 registra il fingerprint di ciò CONTRO CUI riproduce (stack/porta + marker dataset; HEAD + dirty dei file toccati); Phase 5 lo ricalcola prima di confrontare (divergenza = verifica non comparabile, mai un verde stantio — run reale: baseline invalidata 2 volte). Phase 4 passa al fixer lo snapshot `git status --porcelain` pre-run come lista chiusa «già dirty → non tuo» (anti-misattribution). `lanes.md`: tecnica patch-revert per baselinare in checkout condivisi senza stash. Phase 6 allineata ai nomi ESATTI dello schema + campi condizionali; nuovo `bug-registry-check.mjs --scaffold [full]` stampa lo scheletro corretto (5 violazioni al primo colpo nella run reale). Skill `bug` 2.6.0 → 2.7.0. Codex parity: portable.
- **Ripristinato il vocabolario chiuso `poll_timeout`/`agent_deadline_exceeded` perso dalla wave 6.33.1–6.33.4** — `check-codex-parity.js` 2b era ROSSO su main: le riscritture anti-polling avevano fatto cadere i token dal primitivo AGENTS e dal modulo runtime-portability. Il primitivo nomina i due stati (wait scaduto PRIMA della deadline con worker sano = `poll_timeout`, MAI un fallimento); il modulo ri-aggiunge la riga nella state machine. Primitivo AGENTS 1.18.3 → 1.18.4. Codex parity: portable.

Deferred to next sweep (budget 5 slot: 4 fix + 1 gate Codex): [#77](https://github.com/antbald/BALDART/issues/77) (markdownlint `--no-globs` + routing i18n in /bug), [#75](https://github.com/antbald/BALDART/issues/75) (elenco byte-esatto dei test nel return contract di ui-expert).

## [6.33.4] - 2026-07-21

### Fixed

- **Righe Codex `Interacted with` ridotte al minimo.** Il primitivo condiviso vieta le
  interazioni lifecycle ridondanti (`list_agents`, messaggi/follow-up di stato, acknowledgement,
  ripetizione del brief): dopo lo spawn restano solo il singolo wait terminale, una correzione
  materiale nuova o il checkpoint alla deadline reale. `primitive_version` AGENTS 1.18.2 →
  1.18.3. Claude invariato; Codex parity: portable.

## [6.33.3] - 2026-07-21

### Fixed

- **`/new` resume dirty-tree (#74).** `setup-worktree.sh` committa sul branch
  dedicato lo snapshot delle sole card selezionate copiate dal main; il planner
  non scambia piu la propria mutazione di bootstrap per dirty esterno. Semantica
  condivisa Claude/Codex: SSOT `worktree-manager`, runtime Codex consumer.

## [6.33.2] - 2026-07-21

### Fixed

- **Nessun annuncio ridondante allo spawn.** Il primitivo `AGENTS.md` vieta anche `Started
  <agent>`, "contesto in caricamento" e status equivalenti: il worker attivo è già esposto dalla
  UI host. Solo checkpoint materiale o risultato terminale. Un unico `wait_agent` terminale,
  dopo il lavoro indipendente, conserva la notifica di completamento senza polling. `primitive_version`
  1.18.0 → 1.18.2. Codex parity: portable.

## [6.33.1] - 2026-07-21

### Fixed

- **Agent lifecycle senza polling visibile.** Il primitivo condiviso `AGENTS.md` e il binding
  Codex in `runtime-portability-protocol.md` vietano `wait_agent`, heartbeat di routine e la
  narrazione `Waiting for agents` / `Finished waiting`. Dopo uno spawn: lavoro indipendente o
  controllo restituito all'host; ispezione e checkpoint solo alla deadline reale. Gli esiti
  terminali espliciti restano fail-closed. `primitive_version` AGENTS 1.17.1 → 1.18.0. Codex
  parity: portable.

## [6.33.0] - 2026-07-21

### Fixed

- **Il transcript non ha MAI viaggiato con un beacon Codex ([#72](https://github.com/antbald/BALDART/issues/72))** — `bat-beacon.mjs` risolve il rollout Codex leggendo il `session_meta` sulla prima riga del file, ma ne leggeva solo i primi **4096 byte**: Codex inlinea `base_instructions` in quel payload, e sui file reali la prima riga è ~18KB, oltre i 4KB su **145/145** rollout campionati. Il `JSON.parse` lanciava sempre → `codexMeta` sempre `null` → cadevano **entrambe** le vie di risoluzione (validazione `CODEX_THREAD_ID` e fallback per `cwd`). Effetto misurato sulle issue upstream: `Transcript | non disponibile` su **27 beacon su 27 con `Runtime host | codex`**, 0 su Claude — cioè il curator triagava senza transcript esattamente sul runtime che produce la maggioranza delle segnalazioni. Le uniche 4 issue Codex col transcript passavano `--session`, che risolve per nome file senza toccare il meta. Ora la prima riga è letta per intero a chunk fino al newline, accumulando **Buffer** e non stringhe (un `toString` per chunk spezzerebbe un carattere UTF-8 multi-byte al confine), con cap di sanità a 4MB; controprova sui rollout reali **836/836 parsati, 0 fallimenti**. Il test non aveva una sola occorrenza di `codexMeta`/`session_meta` — la nuova fixture **asserisce** di superare i 4096 byte, perché una piccola sarebbe passata anche col codice rotto. Un `non disponibile` ora dichiara la causa (nessuna rollout con quel `cwd` / N sessioni ambigue → rilancia con `--session`) invece di tacere. Skill `bat-beacon` 1.7.1 → 1.8.0. Codex parity: portable (script zero-dep dual-runtime; il difetto era esclusivamente sul ramo Codex).

### Added

- **`/batcave`: gate di natività Codex sui fix nati da un beacon Codex** — i bat-beacon nascono in maggioranza su Codex, ma il fix viene progettato e implementato su Claude, e l'unica garanzia di parity era la parola «parity Codex» in prosa dentro la Fase 3: una regola senza codice non è enforcement ([[project_tokens_content_loss_guard_v6260]]). Nuova **Fase 3.6**, con trigger meccanico (`| Runtime host | codex |` nella tabella metadati della issue, non un giudizio): la soluzione implementata viene sottoposta a **Codex stesso** in read-only via `codex-companion.mjs task … --wait`, con modello risolto dal tier `standard` di `framework/runtime/codex-model-map.yml` (mai hardcodato — i nomi dei modelli churnano) ed `--effort high`, entrambi **pinnati** perché un default implicito rende il verbale non riproducibile. Tre uscite tracciate nella issue: conferma → release; refutazione con prova `path:riga` verificata → correggi e ripassa UNA volta, poi valgono le disposizioni della Fase 3.5; plugin assente / exit non-zero / risposta non ancorata → `Gate Codex: NON ESEGUITO (<motivo>)`, che non blocca la release ma **non è un verde**. Non sostituisce il gate adversariale di Fase 3.5 (una issue Codex `refactor-auto` passa da entrambi). Skill `batcave` 2.4.0 → 2.5.0 — maintainer-only, non distribuita. Codex parity: N/A (tooling di manutenzione, gira solo sulla repo BALDART su host Claude).

## [6.32.0] - 2026-07-21

### Fixed

- **Un doc-index consumer opzionale assente degrada, non blocca ([#55](https://github.com/antbald/BALDART/issues/55))** — `framework/agents/project-context.md` § 3 copriva la *chiave mancante* (→ chiedi), ma il caso diverso in cui la **chiave è risolta e il file derivato non esiste** non aveva alcuna regola: `agents/index.md` presentava `${paths.references_dir}/ui/index.md` e `/simplify` presentava `component-registry.md` come sorgenti autoritative, e su un consumer che non li ha creati l'agente restava fra il fermarsi e l'improvvisare. Nuova **§ 3.bis "Resolved-path-missing protocol"** nel collo di bottiglia (il modulo è citato da 54 superfici del payload): preflight di esistenza → degradazione lungo la cascata che il protocollo di dominio già definisce → dichiarazione in una riga → **doc-debt** advisory a `doc-reviewer`; mai chiedere all'utente (la chiave c'è), mai improvvisare in silenzio. Il caso opposto resta distinto: un path MANDATORY/BLOCKING gated su un flag `true` che manca è un difetto di setup vero. Le tre superfici incriminate (`agents/index.md`, `simplify/SKILL.md`, `simplify-protocol.md` § doc-awareness) passano da «authoritative» a «authoritative **when present**» e **citano** la sezione anziché ridefinirla. Skill `simplify` 1.3.0 → 1.3.1. Codex parity: portable.
- **Un nome di agente già occupato non è un fallimento di spawn ([#64](https://github.com/antbald/BALDART/issues/64))** — la state machine di lifecycle in `framework/agents/runtime-portability-protocol.md` copriva il rifiuto su **follow-up** (`agent_thread_limit`) ma non quello sullo **spawn** per nome ancora tenuto da un worker interrotto: l'orchestratore scopriva la collisione a mano e ripiegava su `followup_task`, ereditandone in silenzio lo stato stale. Nuovo stato `agent_name_taken` nel vocabolario chiuso: chiudi l'occupante se terminale, altrimenti spawna con un nome **univoco per invocazione**; `followup_task` solo sulla continuazione del *proprio* task, mai per riciclare un nome. `context-primer` non è toccata di proposito — il bundle è cross-runtime puro e non nomina alcun task name, quindi il difetto stava nel vocabolario di spawn condiviso e il fix vale per ogni skill che spawna su Codex. Codex parity: portable.
- **`/ui-design` esce prima di pagare i Prerequisites quando non è la skill giusta ([#70](https://github.com/antbald/BALDART/issues/70))** — l'esclusione «in-code UI fixes with no design decision» viveva **solo** nella `description`, in coda a sei trigger positivi larghi, e una volta selezionata la skill apriva con le letture BLOCCANTI (INDEX.md + ui-guidelines + tokens-reference + spec) e un workflow di mockup-generation irrilevante: un singolo fix in-code su un controllo esistente pagava tutto prima che l'esclusione potesse applicarsi. Nuovo **Step 0 — Scope gate** in testa al body, prima di ogni lettura: tre domande (direzione visiva? superficie che non esiste ancora? `/prd` Step 3 o live?) — tutte NO → uscita verso `PROFILE=ui` + `ui-expert`; ambiguo → UNA domanda che nomina l'alternativa, mai il default sul path pesante. La `description` non è toccata: `agents/index.md:23` mostra che il routing era già corretto. La selezione resta probabilistica; il costo di una selezione errata no. Skill `ui-design` 2.3.0 → 2.4.0. Codex parity: portable.

## [6.31.0] - 2026-07-21

### Fixed

- **`config_scalar` spogliava solo le double quote — ogni chiave vuota scritta da `configure` scavalcava il proprio default** (#67). `js-yaml` emette le stringhe vuote come `''`, quindi la funzione restituiva il letterale di 2 caratteri `''`: non-vuoto, e ogni fallback `[ -n "$X" ] || X=<default>` a valle veniva saltato. Con `paths.metrics: ''` il classificatore di `merge-worktree.sh` trattava un conflitto additivo su `docs/metrics/*.jsonl` come `code_conflict`, fermando il rebase invece di auto-risolverlo; affette anche `git.trunk_branch`, `git.merge_strategy`, `paths.backlog_dir`. Fix nel collo di bottiglia, in **tutte e tre** le copie parallele (`merge-worktree.sh`, `setup-worktree.sh`, `allocate-id.sh`), con un nuovo `config-scalar.test.mjs` che verifica sia il contratto di parsing sia l'**identità reciproca delle tre copie** — un fix applicato a una sola location ora fallisce il test.
- **`mw-docs --continue` falliva con `final PRD seal rejected: ENOENT`** (#67). La verifica del seal dal filesystem e il clean-tree check sono precondizioni del **primo ingresso**: sotto `--continue` il rebase in pausa non ha ancora riapplicato il commit che porta il seal, e i file appena risolti sono staged. Ora sono condizionate; la riverifica autorevole degli hash post-rebase resta invariata e fail-closed in entrambi i percorsi.
- **La lookup di ricorrenza del bug registry era cieca su parte del registry, in silenzio** (#53). `bug-registry-check.mjs --find` — ciò che `/bug` Phase 0 usa per trovare i pattern già visti — restituiva `no prior records match` su record presenti e leggibili. Tre cause, misurate su un registry reale da 36 record: il reader non gestiva i **blocchi YAML folded/literal** (`symptom` 2/36, `root_cause.detail` + `lesson` 33/36); le sequenze erano accettate **solo a 2 spazi**, quindi `fix.files` era perso in **30/36** record; `bug_id` ed `escaped_gates` non erano nell'haystack, rendendo un record non trovabile **per il proprio nome**. Un indice parziale ora si dichiara (`WARNING — N record(s) not indexed`) invece di leggersi come «nessuna ricorrenza», e un match è sempre **citabile** (`bug_id ?? id ?? filename`) invece di stampare `undefined` in un campo destinato a `recurrence_of`. Non-breaking: i 36 record producono le stesse identiche 10 violazioni prima e dopo.

### Changed

- **Un valore env deployato che rompe una feature in produzione non è più `not-a-bug`** (#65). La clausola di reroute di `/bug` Phase 0 diceva «env issue → report to the user» perdendo il qualificatore *local* che la definizione della riga porta, così un guasto che rompe il prodotto per tutti gli utenti veniva espulso con un report di una riga. Ora è escluso esplicitamente dall'exit e chiude ad **ABSTAIN** («root cause outside the repo») con `root_cause.category: config-env`, portando evidenze ed `escaped_gates` nel registry. Nuovo pattern di triage **«Asymmetric symptom»**: una sola feature muta col codice che passa la review è la firma di un valore *deployato* malformato — uno spazio o `\n` di troppo viene normalizzato via da header e URL (REST/auth sani) ma sopravvive come `%0A` in query string, uccidendo il canale realtime. Osservato: ~161 giorni in produzione prima di una segnalazione utente.
- **`agents/env-reference.md` — igiene di scrittura e validazione di forma delle env** (#65). L'SSOT vendor-neutro verso cui `agents/index.md` già instrada documenta ora le trappole di scrittura (flag esplicito mai via stdin — `echo` appende un newline e `printf '%s'` può salvare un valore vuoto uscendo 0; il sensitive-default rende ingannevole la verifica post-scrittura; confrontare byte, mai l'aspetto) e la validazione al boot passa dalla sola **presenza** alla **forma** (`value !== value.trim()` → throw), l'unico punto di enforcement che impedisce la recidiva.
- **Un rebuild lock di Graphify occupato è BUSY, non un'attesa infinita** (#54). `graphify update` blocca **senza output** mentre un altro processo detiene il lock esclusivo — due agenti sullo stesso checkout bastano, e il chiamante non lo distingue da una build lenta senza ispezionare i processi. Le Fallback rules di `code-graph-protocol.md` coprono ora la contesa (attesa limitata, e allo scadere si degrada come per qualsiasi grafo non disponibile: l'altra sessione sta già ricostruendo l'artefatto condiviso, quindi aspettare non compra nulla), applicata nel punto di esecuzione reale (`doc-graph-aligner` v1.1.0) e in quello interattivo (`/graph-align` v1.2.0). Colmava un buco della policy già scritta in § Gating: «never block a task on the graph layer».
- **Contratto `audit-summary.json` completato** (#67). La tabella «the EXACT contract `prd-final-gate.mjs seal` enforces» documenta ora `assurances.cross_ac` (`status: PASS` + `checked_cards` non vuoto + le 4 `dimensions`) e la **forma bare-key** dei phase receipt (`- <id>: …`): entrambi erano enforced dallo script ma assenti dalla tabella, e il seal falliva con 8 violation costringendo a reverse-engineerare lo script — la failure mode #28 che quella sezione dichiara di aver chiuso. La guidance markdownlint avverte inoltre di non usare `--fix`: il glob `**/*.md` della config repo sovrascrive l'argomento file (osservati 333 file riscritti).

**Codex parity:** portable — bash/Node zero-dep e prosa condivisa; nessuna chiave nuova in `baldart.config.yml`, quindi la schema-change propagation rule non si applica.

## [6.30.0] - 2026-07-21

Giro bat-cave sulle issue deferite — 8 issue chiuse (5 fix, di cui 2 recurrence-cluster), 7 deferite al prossimo giro. Nessuna decisione di classe refactor presa (l'unica `refactor-auto`, #53, richiede il gate adversariale + 2 slot → deferita).

- **bug v2.4.0 — verifica proporzionata + real-oracle + checklist responsive (#50 #59 #60 #62 #69 #71)**:
  hardening di `/bug` che chiude 6 bat-beacon sulla stessa area. Phase 5 Tier 0 ora **scoped-first**:
  i gate del fix (file cambiati + BRT + test dell'area, 2a) restano MUST-green, mentre la suite completa /
  lint repo-wide gira **output-safe** (`<cmd> > artifact 2>&1; exit=$?; tail`) con **attribuzione** di ogni
  failure — intersecante il diff = FAIL reale, estraneo = `unrelated-preexisting` (registrato, non-bloccante)
  → fine del blocco/flood di un fix verificato per un rosso preesistente in un checkout condiviso sporco
  (#50 #62 #69). **Real-oracle fail-closed** sui bug ricorrenti/interaction: su un `recurrence_of` hit o
  sintomo ri-segnalato serve una oracle matrix (route/state autenticata, event target, gesture, before/after)
  e un E2E skippato/irraggiungibile è `verification_blocked`, mai un PASS (#60). Nuova **checklist
  responsive/viewport** in `lanes.md`: fallback SSR per entrambi i viewport, riconciliazione hydration in
  entrambe le direzioni, non-duplicazione dei dialog portal via ispezione del DOM runtime (#59 #71).
  Nessuna nuova chiave config. Codex parity: portable.
- **bat-beacon v1.7.1 — beacon-scan.mjs, fine dei false-positive `review-degraded` (#66)**: la regola tracker
  `degradedReviewers` matchava anche la lista **vuota** `=[]` (stato pulito) e il fallback by-design
  Codex→code-reviewer → una run pulita del watchdog proponeva finding falsi al curator. Ora la regex richiede
  una lista NON vuota; nuovo test dedicato `beacon-scan.test.mjs`.
- **new v3.5.21 — codex-fallback reason + team-mode briefing (#66)**: `new-final-review.js` +
  `new-card-review.js` loggano il **motivo** del fallback in `codexEngine` (`companion-not-found (preflight)`
  vs `no-json (marker=…)`) → distingue un outage Codex da un load-shed sotto wave pesante, prima invisibile.
  `team-mode.md` Step B vincola il coder a **non committare** (il commit è dell'orchestratore al D.5) e a
  dichiarare gli edit fuori MAY-EDIT nel nuovo campo `out_of_scope_edits:` della completion report.
- **i18n-translator v1.1.0 — output formatter-clean (#68)**: l'agente esegue il formatter di progetto sui
  **soli** locale file scritti (mai l'intero repo) prima di ritornare — `toolchain.commands.format` scoped
  quando `has_toolchain`, altrimenti il formatter rilevato → output commit-clean al primo colpo, il
  pre-commit hook di `/new` non respinge più il commit.

Deferite al prossimo giro (budget 5 slot esaurito): #53 (bug-record schema, `refactor-auto` → gate adversariale + 2 slot), #54 (graphify lock — difetto in tool di terze parti, fix framework = routing/guidance), #55 (fallback index consumer assenti), #64 (context-primer restart worker Codex), #65 (doctor env-hygiene — capability nuova, giro dedicato), #67 (mw-docs classifier defect — smoke test), #70 (routing UI over-eager `ui-design`).

Codex parity: portable (prosa skill/agente + zero-dep JS; workflow payload Claude-only per finding 1 di #66).

## [6.29.2] - 2026-07-21

Robin nightly — sweep bat-beacon: portabilità interprete, resolution Playwright dal repo, gate markdownlint provisioned-only. 6 issue chiuse (3 auto-fix + 3 recurrence).

- **webapp-testing v1.1.0 — interprete `python3` + Playwright resolution order (#52 #56 #61 #51 #58)**:
  la skill prescriveva `python scripts/with_server.py` ma lo script è `python3` e gli host moderni
  (macOS/Homebrew) non hanno l'alias `python` → `command not found`; ora tutta la prosa usa `python3`
  con nota interprete + fallback. Nuova sezione "Playwright resolution": il runner del repo vince
  sempre (`npm run test:e2e` → `node_modules/.bin/playwright`/`npx --no-install playwright` → script
  Python); un wrapper esterno `playwright-cli` che esce `command not found` è un **fallback signal,
  non un live-verification blocker**; `npx --package` senza `--no-install` vietato (fetch di rete → hang).
- **bug v2.3.1 — screenshot oracle risolve Playwright dal repo (#61 #58)**: la branch `screenshot`
  di Phase 1 rinvia a webapp-testing § Playwright resolution — mai "live blocked" quando il repo
  espone un runner Playwright.
- **AGENTS primitive v1.17.1 — pre-commit markdown-lint provisioned-only (#63)**: il gate
  prescriveva `markdown-lint on changed .md` senza qualificarlo; su un consumer senza
  `markdownlint-cli2` provisioned, `npx markdownlint-cli2` fa un fetch di rete senza timeout →
  hang indefinito. Ora scatta solo se il linter è provisioned (toolchain/lefthook o `npx --no-install`);
  bare `npx` vietato; assente → skip, mai hang.

Codex parity: portable (prosa skill markdown + primitive; nessuno script toccato). Nessuna nuova chiave `baldart.config.yml`.

## [6.29.1] - 2026-07-20

Batcave #57 — lefthook non clobbera più gli hook versionati del consumer, e il clobber esistente viene diagnosticato.

- **Toolchain / lefthook adapter — incumbent guard su `core.hooksPath` custom (#57)**:
  `postInstall()` ora SALTA `npx lefthook install` (con reason stampata
  dall'installer — mai silenzioso) quando `core.hooksPath` punta a una directory
  esistente con hook standard consumer-authored (contenuto non-shim): registrare
  lefthook lì sovrascriverebbe i gate del consumer con lo shim (backup in
  `<hook>.old`), uccidendoli in silenzio — è la firma osservata su mayo, dove il
  pre-commit schema-doc guard è stato rimpiazzato dallo shim. Stessa policy
  dell'incumbent husky. Fail-safe: ogni errore di probe → comportamento
  precedente.
- **`githooks.js` — detection `clobbered`**: il check di salute degli hook
  versionati (nato dalla PRIMA occorrenza mayo, hooksPath→dir vuota) ora rileva
  anche la SECONDA firma: hook attivo il cui contenuto è uno shim runner
  (lefthook/husky) con sibling `<hook>.old` non-shim → `ok=false` +
  `clobbered[]` + fixCommand di ripristino. Shim SENZA `.old` = adozione
  intenzionale → silenzio (discriminante anti-falso-positivo). WARN-only come
  tutto il modulo (mai muta; il postinstall npm di lefthook ri-clobbererebbe
  comunque a ogni `npm install` — la coesistenza durevole è documentata in
  `TOOLCHAIN-LAYER.md` § 7). `doctor` espone lo stato con label dedicata
  "Git hooks CLOBBERED". Test: 2 casi nuovi in `scripts/test-githooks.sh`
  (7/7 verdi) + smoke verificato sul repo mayo reale.
- Nessuna nuova chiave `baldart.config.yml` → schema-change propagation rule
  N/A. Codex parity: N/A (CLI installer/doctor, runtime-agnostico).

## [6.29.0] - 2026-07-20

Auto-commit al termine di un fix nella skill `/bug`.

- **bug 2.3.0 — Phase 7 AUTO-COMMIT**: quando la skill `/bug` finisce di fixare
  un bug, committa automaticamente il fix insieme al bug record in un commit
  atomico locale. Fires SOLO su `outcome: fixed` con Phase 5 Tier-0 tutto verde
  (repro passa / diff screenshot / rate baseline; toolchain gates; residue-scan
  exit 0; nessun anti-deceptive flag) più il giudice `code-reviewer` sulla lane
  deep. **Ramo else esplicito** (disciplina MUST-contestuale): abstain /
  not-a-bug / reroute / Tier-0 non verde → nessun commit del fix. Guard:
  branch-safety (se il branch corrente == `git.protected_branch`, crea prima
  `fix/<slug>`), stage esplicito dei soli file toccati + il bug record (mai
  `git add -A`), commit `fix(<domain>): <symptom>` con `fix.commit: pending`
  (SHA auto-referenziale), sha reale mostrata all'utente. **Local-only: mai
  push, mai PR** (resta azione esplicita / flusso `/mw`). Phase 6 aggiornata
  di conseguenza. No nuova chiave config (usa `git.protected_branch` /
  `git.trunk_branch` esistenti) → schema-change propagation rule NON si applica.
  Codex parity: portable (git è portabile su entrambi i runtime).

## [6.28.1] - 2026-07-18

Chiusura #27 (decisione maintainer: Opzione A) — il fan-out audit di `/prd` § 6.6
resta named+file-store, con le due guardie che eliminano il silent-failure mode:

- **prd 1.21.0**: Named-teammate channel rule in 6.6c (il return message di un
  teammate named non è un canale — unica fonte il task store) + Findings-presence
  gate in 6.7 (task `completed` con `## FINDINGS` vuota = anomalia HARD → re-spawn
  singolo, poi `audit_status: failed_missing_findings` + STOP; mai silenzio).
  Codex parity: nessun cambio — il ramo Codex era già file-based con validazione
  d'envelope via task-store.mjs.

## [6.28.0] - 2026-07-18

Giro implementativo di Robin sulle 3 raccomandazioni `needs-antonio` approvate
dal maintainer (#3, #2, #39).

- **Release loop meccanizzato — `scripts/robin-release.mjs` + batcave 1.7.0**
  (#3, Raccomandazione A a scope tagliato): sequenza atomica lock
  (`.git/robin-release.lock`, PID+timestamp, stale 15 min) → guard → `fetch
  --tags` → rebase su `origin/main` → collision-check (prima versione libera) →
  bump VERSION → splice CHANGELOG → commit → tag → push, con un solo retry se
  origin avanza nella finestra di push. La skill batcave Fase 5 INVOCA lo
  script invece di descrivere i passi; i changelog satellite restano manuali
  finché la riscrittura cross-file non è provata su ≥2 release. Test su repo
  fixture (`robin-release.test.mjs`). Questa stessa release è il primo
  consumatore reale dello script. Codex parity: N/A (maintainer-only, gira solo
  nella repo BALDART).
- **beacon-scan `--artifact` + doctor→beacon bridge — bat-beacon 1.7.0** (#2,
  opzioni B+C): `beacon-scan.mjs` guadagna il dispatcher ripetibile
  `--artifact <path>` con type-detect su set CHIUSO (transcript e2e-review →
  `review-degraded`/`retry-ceiling`; `design-sync.json` →
  `unresolved-divergence`; record BUG-* del bug-registry → `gate-not-fired` su
  `escaped_gates`; `skill-conflicts.json`; candidati capture-queue) e fallback
  esplicito `not_actionable` sull'ambiguo — mai una classe per congettura.
  `src/utils/beacon.js` guadagna `writeDoctorFindingBeacons`: lo smart doctor
  accoda beacon sui finding framework (drift/misconfig su set chiuso di action
  key, dedup per key, fail-safe) — non più solo sui crash. Codex parity:
  portable (zero-dep, dual-runtime).
- **Merge Heavy sicuro con main condiviso dirty — worktree-manager 1.9.0**
  (#39, Raccomandazione A): `merge-worktree.sh` step 5b — i guard trunk-only
  (es. migration push develop-only, BUG-0028) girano da un checkout EFFIMERO
  pulito e linked al merge commit (helper `run_in_clean_linked_worktree`: copy
  `stack.env_files`, symlink `node_modules` read-use, env `BALDART_*`, timeout
  300s), NON-BLOCKING, esito nel manifest (`migration_push: ran|none|failed`);
  il main condiviso non viene mai mutato dal land. Nuova chiave
  `stack.migration_push_command` con schema-change propagation completa
  (template + probe/prompt in configure + update detector; doctor N/A — chiave
  advisory senza stato diagnosticabile). Coordinato con #37. Codex parity:
  portable (script shell condiviso).

## [6.27.7] - 2026-07-18

- **`/new` prova invarianti su file immutati.** I target tracciati citati dai
  gate diventano evidenza read-only con SHA-256 nella review; il writer non
  ottiene nuovi permessi. Claude parity: non applicabile — scope Codex Model B.

## [6.27.6] - 2026-07-18

Quarto giro di Robin (wave 4 supervisionata): 1 fix dalla coda bat-beacon
(unica issue auto-risolvibile rimasta; le 4 `needs-antonio` restano in attesa).

- **Skill bat-beacon 1.6.0** (#49): minimo contenutistico del beacon — un
  `collect` senza analisi (`--details`) né transcript risolto non parte più
  come stub silenzioso (pattern #19): beacon **degradato** con warning nel
  body, marker machine-readable `degraded:stub` e WARN stderr istruttivo;
  rifiuto hard (exit 1) SOLO su `--details` vuoto/inesistente senza transcript
  (content-loss reale). Degradazione come default così `/wrap` e i watchdog non
  vengono mai soppressi. Nuova hard rule 9 in SKILL.md, test estesi (3 casi).
  Codex parity: portable (script zero-dep dual-runtime).

## [6.27.5] - 2026-07-18

Terzo giro di Robin (wave 3 supervisionata): 5 fix dalle issue bat-beacon
(tutte `optimization` deferite dalle wave 1-2).

- **Primitives AGENTS 1.17.0 / CLAUDE 1.7.0 + skill bug 2.2.0** (#8):
  l'eccezione orchestrated-pipeline (§ Change Tiers) copre ora la lane **light**
  di `/bug` — il grounding inline della lane soddisfa l'obbligo di understanding
  incluso il mandato `PROFILE=ui`; `lanes.md` cita l'eccezione di ritorno.
  Chiuso il conflitto mandato-root ↔ lane-contract. Codex parity: portable.
- **Skill simplify 1.3.0** (#9): `simplify-scan.mjs` scopa i candidati alle
  righe cambiate (`git diff -U0`, default) — un edit di 1 riga non riesuma più
  i cloni storici del file; `--all-windows` ripristina lo scope whole-file,
  campo `window_scope` in output, +1 test fixture. Codex parity: portable.
- **code-reviewer 2.2.0 + plan-auditor 2.2.0** (#11): nuova classe di review
  **invariant layering** — regola di dominio cross-field su campi persistiti
  enforced solo client-side → HIGH; filtro UI che nasconde un valore persistito
  senza normalizzarlo → finding; regola duplicata invece di predicato condiviso
  → finding; checklist §B del plan-auditor chiede i layer di enforcement.
  Codex parity: adapter-generated (transpiler TOML).
- **Skill new 3.5.19** (#20): `gateSpecs()` (runtime Codex) dedupa le
  validation command cross-card per argv finale distinto — lo stesso comando
  dichiarato da N card girava N volte nel gate union (5× ESLint misurati).
  Codex parity: il file È il runtime Codex-native.
- **Skill bug 2.2.0** (#31): `bug-registry-check.mjs --find` passa da conteggio
  piatto a pesi IDF sul registry + requisito multi-token + floor di rilevanza +
  evidenza per-hit (`matched`/`score`/`relevance`) — la specificità del sintomo
  batte l'overlap generico. Codex parity: portable.

## [6.27.4] - 2026-07-18

Secondo giro di Robin (wave 2 supervisionata): 5 fix dalle issue bat-beacon.

- **i18n gate** (#19): gli `ignores` generati (flat + legacy) escludono ora la
  superficie non-prod (`*.config.*`, `scripts/`, `tests/`, `__tests__/`,
  `__mocks__/`, `e2e/`, `*.stories.*`, `.storybook/`) — il gate lintava l'intero
  repo, config e tooling inclusi. Il claim "JSX/TSX non verificati" è stato
  REFUTATO sul codice (flat files-glob + parser TS con jsx già corretti). Codex
  parity: N/A (CLI, config generato identico).
- **`design-sync` 1.2.0 + `design-system-init` 1.3.1** (#22): `--group` di
  `compile-ds-cards.mjs` obbligatorio e fail-loud (derivazione: state → marker di
  una card pubblicata → chiedi; persistito come `group` in
  `.baldart/design-sync.json` via bootstrap/publish); niente più tabella variants
  placeholder con `variants: []`; `finalize_plan` documentato con `deletes: []`
  esplicito; must-measure #4 chiuso in negativo (i due guard indiretti per un
  `--all` autorizzato); diff-scope publish ridefinito su `spec_sha` + nuovi spec.
  Codex parity: portable (script zero-dep; skill Claude-gated com'era).
- **`worktree-manager` 1.8.0** (#37): `setup-worktree.sh` §8 worktree-aware —
  `core.hooksPath` assoluto verso il checkout principale rimappato worktree-local
  (`extensions.worktreeConfig`), lefthook ri-sincronizzato per-worktree;
  caveat SKILL sui gate diff/staged ciechi agli untracked (parte consumer).
  Codex parity: portable (shell zero-dep).
- **`playwright-skill` 1.1.0** (#6): prescrizione anti-riuso-cieco del webServer —
  `reuseExistingServer: true` salta ogni guard di freschezza incapsulato nel
  comando; con un lock/supervisore il riuso va validato contro HEAD, altrimenti
  `reuseExistingServer: false` a working tree cambiata. Codex parity: portable.
- **`runtime-portability-protocol`** (#4): nuovo stato `agent_thread_limit` nella
  state machine + regola dei follow-up bounded (~2-3 per thread; oltre o su
  thread-limit → spawn fresh by-name con re-brief compatto — mai un fallimento
  dell'agente). Codex parity: portable (modulo condiviso, binding per-runtime).

## [6.27.3] - 2026-07-18

Primo giro di Robin (il triage autonomo di `/batcave`, wave 1 supervisionata): 5
fix dalle issue bat-beacon, release completata dal watchdog dopo il deferral per
collisione con release concorrente (#3, 4ª occorrenza).

- **`worktree-manager` 1.7.2** (#25): hook git disabilitati worktree-local nel
  worktree docs (`nw-docs`, create+resume) — il gate Biome di lefthook non blocca
  più i commit doc-only. Codex parity: portable (script zero-dep).
- **`bat-beacon` 1.5.5** (#10): `mark-retro` senza `--session` esplicito registra
  `unknown` + warning — mai più guess del transcript più recente. (#17): hard rule
  di pertinenza per `--bug-record` — allegato solo se il beacon nasce da quella
  run di `/bug`. Codex parity: portable.
- **`update`** (#16): gli mtime dei file dirty sopravvivono all'auto-stash/pop
  (snapshot hash pre-stash, ripristino sui 3 siti di pop) — la forensics
  post-update non è più accecata. Codex parity: N/A (CLI).
- **`bug` 2.1.2** (#21): `fix.commit: pending` è valido end-to-end nel bug
  registry (+ format check sha) — il record si scrive prima che il commit esista.
  Codex parity: portable.
- **`wrap` 1.2.1**: aggiornamento riferimenti mark-retro coerente con #10.

## [6.27.2] - 2026-07-18

- **`/new` non duplica i gate nel writer.** I leaf mutanti usano check mirati;
  `union.gate` esegue una sola volta i comandi dichiarati e le suite complete.
  Claude parity: non applicabile — ownership del runtime Codex Model B.

## [6.27.1] - 2026-07-18

- **`/new` accetta scope card strutturati.** Le entry oggetto
  `files_likely_touched[].path` diventano MAY-EDIT reali; il fingerprint Git
  tollera la sola metadata PR aggiunta da Codex Desktop al config comune.
  Claude parity: il formato card resta condiviso; il fingerprint è Codex-only.

## [6.27.0] - 2026-07-18

- **Le lesson del watchdog `/new` entrano in `/prd` Model B (#48).** L'audit
  adversarial produce una receipt cross-AC esplicita su precedenza,
  nullability, equal-state e specificita; fan-in e final seal falliscono
  chiusi quando manca o fallisce. Il boundary normalizza `path`/azione, il
  companion persiste telemetry esatta e il bridge condiviso gestisce brief
  dash-leading. Matrice di trasferimento nel post-mortem FEAT-0087-01.
  Claude parity: bridge condiviso; assurance nuova nella variante Codex.

## [6.26.18] - 2026-07-18

- **Il companion usa un brief portabile con `claude -p` (#47).** Una riga
  introduttiva non-dash precede i binding canonici, così il valore di `-p` non
  viene reinterpretato come opzione. Claude parity: condivisa.

## [6.26.17] - 2026-07-18

- **Il Claude companion accetta i brief canonici dash-leading (#47).** Il
  prompt viene passato come singolo argomento `--print=<brief>`, evitando che
  `- run_id:` sia interpretato come opzione sconosciuta. Claude parity:
  condivisa — bridge Claude e chiamante Codex usano lo stesso contratto.

## [6.26.16] - 2026-07-18

Giro Robin (batcave nightly) — 5 difetti framework auto-risolti dalle issue bat-beacon.

- **`overlay validate` riconosce gli overlay root-doc (#5).** `deriveTargetFromOverlayPath` trattava `.baldart/overlays/{CLAUDE,AGENTS}.md` come skill (parts.length===1) → rifiuto `no YAML frontmatter` su file che per design non ne hanno. Aggiunto il `kind: root-doc` + un branch in `validate()` che fa il dry-run del vero renderer (`root-primitives.renderPrimitive`, ora con `cfg` di default + `warnings` propagati) e riporta gli orphan-marker. Verificato sul consumer `mayo` (CLAUDE.md/AGENTS.md → exit 0, skill overlay senza regressione). Codex parity: N/A (CLI condivisa).
- **Drift-guard oracoli: confronto nello stesso spazio JSON-escaped (#26).** `validate-card-baseline.js` confrontava un oracolo grezzo contro `JSON.stringify(validation_commands)` — escape su un solo lato, così ogni oracolo con `"` o `\` non matchava MAI. Ora sia il needle sia ogni entry sono confrontati in forma JSON-escaped (copre anche le entry-oggetto). No regressione (`check-card-baseline` verde). Codex parity: portable.
- **`prd-writing-phase.md` allineato all'enum di stato (#29).** Prescriveva `status: cards-writing`, assente dall'enum CHIUSO di `state-template.md` che lo stesso modulo rende BLOCKER → corretto in `backlog`. Codex parity: portable.
- **Contratto di `audit-summary.json` documentato in `validation-phase.md` (#28).** Il seal (`prd-final-gate.mjs`) falliva su summary scritti alla lettera della prescrizione. Aggiunta la tabella-contratto SSOT-derivata: `finding_id` non `id`, `requires_action: false` (non `resolved`), enum CHIUSO `adversarial_reviewer` (Codex→`external`), `schema_version:1` intero, `manual_items` vuoto, i 7 phase-receipt. Codex parity: portable.
- **Path portabile del clone-floor in `simplify` (#18).** La skill prescriveva `framework/.claude/skills/…` (inesistente nel consumer: il payload sta sotto `.framework/framework/…`) → risoluzione a due path `ls … | head -1`, valida sia nel consumer sia nella repo BALDART. Codex parity: portable.

Skill bump: `prd` 1.19.0→1.19.1, `simplify` 1.2.0→1.2.1.

## [6.26.15] - 2026-07-18

- **La union riconcilia le card prima della delivery (#46).** Il runtime Git
  marca ogni card selezionata `DONE`, aggiunge data/provenance e include tali
  file nel commit atomico della union; la verifica remota non trova piu card
  `TODO`. Claude parity: non applicabile — delivery Codex Model B.

## [6.26.14] - 2026-07-18

- **I finding assurance azionabili entrano nel ledger (#45).** Il boundary
  normalizza in memoria `path`→`file` e una stringa non vuota in
  `requires_action`→`true`, preservando l'artifact raw; il prompt richiede ora
  `file` e booleano espliciti. Claude parity: non applicabile — review
  transport Codex.

## [6.26.13] - 2026-07-17

- **Il CompletionTransport accetta teardown riuscito dei discendenti.** Dopo
  l'exit del child, discendenti residui terminati entro grace non generano piu
  `TRANSPORT_TEARDOWN_FAILED`; resta fatal solo un process tree ancora vivo.
  Claude parity: non applicabile — transport Codex.

## [6.26.12] - 2026-07-17

- **L'assurance AC usa digest file, non scope hash (#43).** I prompt review e
  completeness ricevono la mappa immutabile path→SHA-256 del frozen scope e
  distinguono esplicitamente `scope_hash` da `evidence.digest`; il validator
  continua a rifiutare digest non-file. Claude parity: non applicabile —
  review prompt Codex.

## [6.26.11] - 2026-07-17

- **Un retry transport senza telemetry resta budget-safe (#42).** Il budget
  snapshot contabilizza solo un `transport-transient` retryable senza usage al
  `token_budget` immutabile della reservation, marcandolo conservativo; la
  review puo proseguire, l'aggregate cap include l'upper-bound e gli altri
  missing-usage restano fail-closed. Claude parity: non applicabile — broker
  budget accounting Codex.

## [6.26.10] - 2026-07-17

- **I writer Codex recuperano un failure transport dopo mutation scoped (#41).**
  L'implementation riserva un solo retry `transport-transient`; il broker
  conserva il baseline Git originale, verifica tutte le mutation contro
  `may_edit` e passa al worker fresco un brief inspect/preserve/complete sul
  diff cumulativo. Capacity e `UnknownProcessId` non scartano piu lavoro
  valido; mutation fuori scope e failure non-transport restano fail-closed.
  Claude parity: non applicabile — recovery del broker Codex.

## [6.26.9] - 2026-07-17

- **I leaf read-only non eseguono validation (#40).** Il broker impone per
  ogni profilo read-only un contratto evidence-only; la review ribadisce che
  test, build, generatori e comandi con temp/cache write appartengono ai gate
  parent-owned. Nessun `PHASE_PARTIAL` per Vitest EPERM nel sandbox. Claude
  parity: non applicabile — policy sandbox del broker Codex.

## [6.26.8] - 2026-07-17

- **Le review con finding azionabili entrano nel ledger (#38).** Il boundary
  accetta `status: FAIL` solo con finding/assurance fallita, normalizza
  `P0|P1|P2|P3` negli enum canonici e vincola i path al frozen scope relativo.
  Il finding reale non viene piu scartato come artifact invalido. Claude
  parity: non applicabile — boundary del review engine Codex.

## [6.26.7] - 2026-07-17

- **I JobSpec review congelano lo scope dinamico parent-owned (#36).** Il
  broker non sovrascrive piu lo scope di diff con quello del manifest: prompt,
  JobSpec, card contract e validazione risultato usano lo stesso digest. Test
  handler e broker sul caso reale union review. Claude parity: non applicabile
  — JobSpec e broker sono runtime Codex.

## [6.26.6] - 2026-07-17

- **Il launcher Codex `/new` supera il cap host di 300 secondi (#35).** Una
  sola `functions.exec` residente avvia la CompletionTransport CLI via
  host-exec e consuma internamente la stessa sessione: zero turni di polling
  modello, proof dedicata e process tree posseduto fino al terminale.
- **Il leaf worker non emette evidence candidate non vincolate (#34).** Il
  brief materializza gli indici autorizzati; con `evidence_bindings: []`
  impone `artifacts: []`, evitando il `FRAMEWORK_DEFECT` dopo implementazioni
  valide. Claude parity: non applicabile — entrambi i fix sono runtime Codex.

## [6.26.5] - 2026-07-17

- **Il bootstrap worktree non importa card locali fuori scope.** La sync delle
  card non ancora su trunk filtra ora gli ID ricevuti in `--cards`/`--card`;
  una card tracked in un commit locale ma assente da `origin/develop` non rende
  piu dirty il worktree di un'altra run. Script condiviso Claude/Codex; test
  reale su repo remoto + commit locale concorrente.

## [6.26.4] - 2026-07-17

- **Le decisioni dirty-tree Model B conservano il path byte-esatto.** Il parser
  porcelain NUL non passa piu da `.trim()`, che sui file tracked rimuoveva lo
  spazio di stato iniziale e poi la prima lettera del path (`backlog` →
  `acklog`). Regressione su modifica tracked concorrente. Claude parity: non
  applicabile — preflight del motore Codex.

## [6.26.3] - 2026-07-17

- **Il broker Model B non attribuisce piu al worker i ref aggiornati da altri
  worktree.** Lo snapshot Git sorveglia il ref della branch posseduta dal
  worktree, non l'intero `refs/` condiviso; config, hook e ref corrente restano
  fail-closed. Regressioni su ref concorrente e ref corrente a tree invariato.
  Claude parity: non applicabile — isolamento del broker Codex.

## [6.26.2] - 2026-07-17

- **Il grounding Codex Model B non esegue piu validation in sandbox read-only
  e non emette combinazioni status/missing_input incoerenti (#24).** Il prompt
  chiude da evidenza e il leaf brief conserva le invarianti discriminate
  parent-side. Claude parity: non applicabile — fix del broker Codex.

## [6.26.1] - 2026-07-17

- **Il broker Codex Model B non rifiuta piu grounding validi per il replay
  della prompt cache (#23).** Il budget per-job conta solo input non cached,
  cache-write e output; il cap aggregate continua a usare l'intero
  `processed_total`. Test di regressione sul caso mayo 450521/150000. Claude
  parity: non applicabile — il broker e specifico del runtime Codex.
- **`bat-beacon collect` non allega piu il bug record piu recente senza
  richiesta esplicita.** La raccolta standalone resta confinata ai flag del
  chiamante; `--bug-record <path>` conserva l'allegato intenzionale. Copertura
  child-process nel gate reference-integrity.
- `bat-beacon` 1.5.4. Codex parity: **portable** — script Node zero-dep,
  condiviso fra i runtime.

## [6.26.0] - 2026-07-16

- **Un rebuild dei token non cancella più codice scritto a mano (#15, #7).**
  `baldart tokens build` rigenerava l'output dal `.tokens.json` preservando ogni
  VALORE e **restringendo in silenzio la superficie del modulo**: un consumer con
  alias/derivati/helper accanto all'albero dei token si ritrovava il repo non
  compilabile (2 occorrenze identiche, 03-07 e 16-07). Il banner `do not edit`
  rendeva instabile ogni restore a mano → loop rompi/ripara.
  - **Il generatore rifiuta la perdita** (`tokens-generator.js`): `render()`
    calcola `contentLoss` (dichiarazioni top-level presenti su disco che il
    render non riprodurrebbe) e `build()` **non scrive nulla** in quel caso
    (`{ok:false, reason:'content-loss'}`), elencando cosa si perderebbe e
    indicando la barrel. `baldart tokens build --force` è lo scarto deliberato.
    Il guard vive nel generatore, non nel doctor, perché in `--auto` il doctor
    esegue anche le azioni non-`autoOk`: il collo di bottiglia va dove passano
    tutti i writer.
  - **Il doctor smette di essere il coltello** (`doctor.js`): `isStale` distingue
    ora una staleness benigna da una distruttiva (`contentLoss`); la seconda è
    **advisory, mai `autoOk`** — era l'unica azione `autoOk` distruttiva e senza
    prompt interno, in violazione della policy `autoOk` del file stesso. Il
    return code di `tokens build`, prima scartato, è propagato.
  - **Chi ha premuto il grilletto resta indeterminato, e non conta**: `update.js`
    non ha nessun code path ai token (tracciato esaustivamente), ma il `.tokens.json`
    del consumer non era cambiato (mtime invariato) mentre il `.ts` era rigenerato
    per intero → era un **rebuild a freddo innescato dalla staleness**, non una
    build legittima. Candidati: il backfill del doctor, o le skill che istruiscono
    un agente a lanciare `baldart tokens build` (`ds-new`/`ds-edit`/`ui-expert`/
    `design-sync`/routine `ds-drift`). Il file rigenerato porta la firma
    dell'emitter (quoting `"…"` di `JSON.stringify` vs `'…'` in HEAD) → è passato
    da `emit()`/`build()`: **il guard nel generatore li copre tutti**.
    L'attribuzione a `update` nasce dal suo auto-stash/pop, che riscrive ogni file
    dirty stampandoci sopra il proprio mtime — l'update incolpa se stesso.
  - **La barrel diventa la convenzione**: l'output punta a un file
    generated-only (`tokens.generated.ts`), il modulo che l'app importa resta
    hand-owned e lo ri-esporta (`export * from './tokens.generated'` + il layer
    derivato). Gli import site non cambiano. È il default proposto da
    `configure` e in `design-system-init` (skill 1.3.0), documentata in
    `COMPONENT-MANIFEST-LAYER.md` + template.
  - **Meccanizza una regola che esisteva già**: `design-system-init` imponeva in
    prosa un "export-preservation check (ENFORCED)" che nessun codice applicava —
    e infatti non è stato onorato. Ora il check è deterministico nel CLI e la
    skill delega invece di reimplementarlo a occhio. L'unità è la
    **dichiarazione, non l'export** (un `type` non esportato che sparisce rompe
    l'helper esportato che lo usa).
  - Rilevamento via hook opzionale per-emitter `declarations` (`token-emitters/`,
    solo `ts`; nessun hook → nessun guard, scelta valida e documentata).
  - Nessuna chiave nuova in `baldart.config.yml` (`design_tokens.outputs[].path`
    è già libero) → la schema-change propagation rule non si applica.
  - Copertura: `src/utils/__tests__/tokens-generator.test.js` (12 test, incluso
    l'incident replay e la controprova che la staleness benigna resta a un click).
  - Codex parity: **N/A** (CLI Node puro, zero-dep, path identico sui due runtime).

## [6.25.0] - 2026-07-16

- **Codex Model B: bootstrap dirty-tree interattivo e resumable (#14).**
  Il dirty iniziale e `MAIN_CHECKOUT_MUTATED` materializzano ora un
  `NEEDS_DECISION` prima del manifest, con journal append-only, delta path e
  resume sullo stesso run. Le opzioni permettono attesa/riprova, snapshot del
  trunk remoto corrente o abort; nessun dirty esterno entra nel worktree.
- `new` Codex variant 2.1.0. Codex parity: **native**; Claude conserva il gate
  semantico tramite `AskUserQuestion` e non richiede modifiche runtime.

## [6.24.2] - 2026-07-16

- **Codex Model B: grounding ripristinato e diagnostica azionabile (#13).**
  La proiezione OpenAI-strict rimuove ora `uniqueItems`, keyword rifiutata
  dalla Responses API; l'invariante resta nel canonical schema e viene ancora
  validato dal parent. Sugli exit non-zero senza stderr, l'adapter usa l'ultimo
  errore del JSONL Codex invece di persistere `exited 1:` senza causa.
- `new` Codex variant 2.0.1; regression test su schema proiettato e fallback
  diagnostico. Codex parity: **native**; variante Claude invariata.

## [6.24.1] - 2026-07-16

**Due guardie cieche, trovate mentre se ne costruiva una terza.** Nessun
impatto sui consumer: entrambe erano strumenti interni che tacevano.

- **`src/utils/git.js` e `src/commands/push.js` erano invisibili a grep.**
  Contenevano byte NUL GREZZI dentro string literal (`line.split('<NUL>')`)
  dove l'intento era l'escape `'\0'`. JS valido, runtime corretto — ma
  `file(1)` li classificava `data` e **grep li saltava in silenzio**: ogni
  retrieval grep-based (agenti, Explore, `codebase-architect`) era cieco su
  due file core, e durante questa stessa sessione ha quasi prodotto la
  conclusione che `ensureFrameworkNotIgnored` non esistesse (è a
  `git.js:644`). Byte grezzo → `\0`: semantica identica (`'\0' ===
  String.fromCharCode(0)`), file di nuovo testo. Verificato eseguendo il
  parsing `%x00` su output `git log` reale, non solo per ragionamento.
  Ironia: uno dei tre NUL era nel detector di file binari di `push.js:152`.
- **`scripts/check-update-skill-drift.js` era morto dalla v4.88.0** — contava
  le entry di `HOOK_REGISTRY` in `src/utils/hooks.js`, che dalla v4.88.0 è un
  thin shim: 0 contro 4 attese, quindi drift falso a ogni run E cecità totale
  a un drift vero. Ora legge la struttura ESPORTATA da `hooks/registry.js`
  invece di grepparne le righe `id:`: il modulo esporta anche `RETIRED_HOOKS`,
  le cui entry hanno la stessa forma, quindi una regex sul file conterebbe 6
  — scambiando un falso drift per un altro.
- **Snapshot drift riallineato** (`baldart-update` 1.1.1) — `last_verified` era
  fermo a v4.8.0 e il template ha guadagnato `design_tokens`/`graph`/
  `toolchain` da allora. Il decision point #4 tratta lo schema drift come
  silente e non enumera i namespace: la skill è ancora accurata, cambia solo
  lo snapshot. Il check ora è verde e può di nuovo segnalare un drift reale.
- Codex parity: **N/A** — tooling di manutenzione della repo BALDART, non
  payload shippato.

## [6.24.0] - 2026-07-16

**BALDART sporcava il working tree dei consumer e nessuno glielo aveva detto.**
Sintomo osservato in dogfooding: un beacon inviato da un altro terminale resta
untracked in `.baldart/bat-beacon/sent/`, e chi lo trova non sa cosa farne —
"è mio? posso committarlo? se resta qui domani è orfano?". Non lo è: la issue
GitHub è già aperta, il file è solo la ricevuta.

- **Nuovo `src/utils/runtime-ignores.js`** — BALDART scrive sotto `.baldart/`
  due specie di file, e finora non aveva mai spiegato a git la differenza:
  CONTENUTO che il consumer versiona di proposito (`overlays/`,
  `bug-registry/`, `state.json`, i payload generati) e RUNTIME di sessione
  (coda beacon, ricevute, marker retro) che appartiene alla macchina che l'ha
  prodotto. Il modulo possiede la POLICY (quali path) e il MECCANISMO (blocco
  `.gitignore` delimitato da marker, idempotente, che tocca solo il testo tra
  i propri marker — le regole dell'utente restano intatte).
- **Non era solo rumore: era una perdita di integrità.**
  `BALDART_MANAGED_PATTERNS` matcha `^\.baldart(\/|$)` per intero, quindi
  l'auto-commit di `update` classificava come "managed" i residui di un'altra
  sessione e li **committava nella storia del consumer**. Un path ignorato non
  raggiunge `git status`: ignorarli chiude entrambi i buchi con un solo fatto.
  Per questo la scrittura vive dentro `postUpdateAutoCommit` — l'unico choke
  point attraversato da entrambi i path di auto-commit (reset + standard).
- **Nuova azione `doctor`: `ignore-runtime-artifacts`** (`autoOk`) — backfilla
  i consumer installati prima del blocco e lo risana quando la policy cresce.
  Gemella speculare di `unignore-framework`: `.framework/` NON deve essere
  ignorato, il runtime di sessione DEVE esserlo.
- **`.baldart/skill-conflicts.json` deliberatamente NON ignorato** — sembra un
  log, ma è versionato di proposito ("commit ... so the team sees the same
  state", `migrate.js`; "these ARE meant to be" committed, worktree-manager +
  `/new` setup.md). Verificato prima di shippare, non dedotto: una regola
  troppo larga qui cancella lavoro vero dalla storia di un consumer.
- **`bat-beacon` 1.5.3** — hard rule 4 dice ora esplicitamente che `outbox/`,
  `sent/` e `last-retro.json` sono gitignorati e che un file in `sent/` è una
  ricevuta, non lavoro orfano da recuperare.
- Codex parity: **portable** — è uno strato installer/CLI, runtime-agnostico;
  i path di bat-beacon sono identici sui due runtime (script zero-dep).

## [6.23.3] - 2026-07-16

**Il bat-signal si stampava ma restava invisibile** — secondo anello della
stessa catena, emerso su Codex (issue #8) con lo script 1.5.1 già installato.

- **`bat-beacon` 1.5.2** — 6.23.2 ha tolto la soppressione, non ha risolto la
  CONSEGNA. Lo stdout di una skill non lo legge un umano al terminale: lo legge
  l'agente, che poi scrive all'utente un recap in prosa — e il pipistrello
  finisce nel cestino insieme al resto dell'output. La decorazione da terminale
  era progettata per un pubblico che non esiste in un host agentico.
- **Nuova hard rule 8 in `SKILL.md`**: in interattivo l'output di `send` va
  incollato VERBATIM nel messaggio finale (arte + caption + URL in un blocco di
  codice), mai parafrasato o ridisegnato; in autonomo (watchdog `/new`/`/prd`,
  routine, CI) non serve — non c'è nessuno che legge.
- **Limite dichiarato**: è un gate in prosa e resta best-effort. Non c'è modo
  deterministico di imporre cosa un agente scrive all'utente; su Codex i gate
  in prosa reggono meno che su Claude (cfr. v5.14.0).
- Lezione (gemella di 6.23.2): non basta chiedersi *se* un output viene
  prodotto, ma **chi lo legge davvero**.

## [6.23.2] - 2026-07-16

**Fix: il bat-signal di 6.23.1 non si accendeva mai.** Difetto rilevato alla
prima invocazione reale in un consumer (issue #7, repo `mayo`).

- **`bat-beacon` 1.5.1** — l'arte era gatata su `process.stdout.isTTY` per
  distinguere "umano" da "macchina". Assunzione sbagliata e mai misurata: la
  skill gira SEMPRE dentro il tool Bash di un agente, dove `isTTY` è
  `undefined` in ogni caso — invocazione umana e watchdog autonomo sono
  indistinguibili per quel predicato. La guardia sopprimeva il 100% delle
  invocazioni reali. Ora l'arte si stampa sempre (opt-out `BALDART_NO_ART`);
  il colore resta gatato su TTY + `NO_COLOR` per non sporcare i transcript
  pipati con escape ANSI.
- **Gli URL delle issue sono ri-stampati sotto l'arte**: un chiamante che fa
  `send | tail -N` (osservato in produzione) teneva il pipistrello e tagliava
  la riga `SENT` con l'URL — l'unico dato che conta. `SKILL.md` Step 5 ora
  vieta esplicitamente le pipe su `send`.
- Lezione: un gate va misurato nel contesto in cui vive davvero, non in un
  test che forza la condizione che si vuole dimostrare.

## [6.23.1] - 2026-07-16

**Bat-signal ASCII a fine `/bat-beacon send`** — puramente estetico, nessun
cambio di comportamento del beacon.

- **`bat-beacon` 1.5.0** — dopo un `send` riuscito lo script stampa il
  bat-simbolo in giallo (59 col, puro ASCII: nessun glifo a blocchi Unicode →
  resa identica su qualunque font di terminale, nessuna cella a doppia
  larghezza), centrato sulle colonne reali, con caption `N beacon → <repo>`.
  Stampato UNA volta per `send`, non per issue (con `--all` su N beacon resta
  un solo pipistrello).
- **Zero rumore per i consumer macchina**: soppresso quando stdout non è un
  TTY (CI, pipe, watchdog autonomi di `/new` Phase 8.5 e `/prd` Step 7.7) e
  su `NO_COLOR` / `BALDART_NO_ART`. Sotto le 59 colonne l'arte si spezzerebbe
  su più righe → si stampa la sola caption.
- Codex parity: portable (stdout Node zero-dep, identico sui due runtime).

## [6.23.0] - 2026-07-16

**Hook retrospettiva differita RITIRATI** (troppo rigidi, bruciavano token, e
su Codex il SessionStart hook falliva con "invalid session start JSON output"
— stampava testo semplice dove Codex si aspetta JSON). La retrospettiva
terminale resta un MUST, ma torna interamente model-driven via `/wrap`.

- **Rimossi** `session-retro-marker.js` (SessionEnd) e `session-retro-check.js`
  (SessionStart) dal payload e dal registry hook.
- **Nuovo meccanismo `RETIRED_HOOKS`** in `src/utils/hooks/registry.js` +
  prune in entrambi i renderer (`claude.js`, `codex.js`): al prossimo
  `add`/`update`/`doctor` gli hook ritirati vengono RIMOSSI dagli install
  esistenti (`.claude/settings.json` per id o marker, `.codex/hooks.json` per
  marker), preservando hook utente/Graphify — senza prune, un hook uscito dal
  registry restava nei consumer per sempre (registerAll è additivo).
- **`wrap` 1.2.0** — modalità recover solo su richiesta esplicita dell'utente
  (il marker legacy `pending-retro.json` resta leggibile come fonte); mai
  preemptive.
- **Primitives**: `AGENTS.md` 1.16.0 (via il "Deferred recovery", il MUST resta
  con nota esplicita non-preemptive), `CLAUDE.md` 1.6.0 (via la sezione hook).
- Codex parity: adapter-generated (prune renderizzato per entrambi i runtime
  dallo stesso registry SSOT).

## [6.22.2] - 2026-07-16

**Fix trigger retrospettiva differita troppo aggressivo** (il recover di
`/wrap` partiva anteponendosi al task dell'utente, e scattava anche quando la
retrospettiva era in realtà già stata fatta).

- **`session-retro-marker.js`** — due nuovi guard prima di scrivere
  `pending-retro.json`: (1) una retro marcata negli ultimi 60 min copre l'exit
  anche se il session id non coincide (un resume cambia l'id dopo che `/wrap`
  ha marcato quello vecchio — la causa del falso "pendente"); (2) soglia di
  sostanza — una sessione con <3 turni utente reali (tool-result esclusi) non
  merita retro differita, spezzando la catena infinita delle sessioni
  retro-only che generavano a loro volta un pending.
- **`session-retro-check.js`** — marker già coperto da una retro (id match o
  `retro_at` ≥ `ended_at`) o stantio (>7 giorni) → pulito in silenzio, nessuna
  iniezione; quando il marker è genuino il messaggio è ora esplicitamente
  NON-preemptive ("esegui prima il task dell'utente; recover a fine sessione
  o in un momento morto").
- **`wrap` 1.1.0** — la modalità recover codifica la stessa priorità (mai
  anteporre il recover al task; marker coperto → solo `mark-retro`).
- Codex parity: portable (`session-retro-check.js` è già dual-runtime; il
  marker su SessionEnd resta Claude-only come da design v6.21.0).

## [6.22.1] - 2026-07-16

**Fix bat-beacon: transcript risolto anche su host Codex** (issue #4 — un
beacon partito da Codex arrivava con `Transcript: non disponibile`).

- **`bat-beacon` 1.4.1** — `resolveTranscript` acquisisce la risoluzione
  Codex-nativa (approccio validato con parere nativo Codex): via primaria
  `CODEX_THREAD_ID` dall'ambiente → rollout in
  `${CODEX_HOME:-~/.codex}/sessions/` validato contro
  `session_meta.payload.id`; ultima risorsa: rollout più recente con
  `payload.cwd` corrente, esclusi i fork subagent, null se ambiguo (sessioni
  parallele nello stesso cwd). SKILL.md hard rule 4 aggiornata.
- Codex parity: portable (fix interamente nel percorso Codex dello script
  zero-dep condiviso).

## [6.22.0] - 2026-07-16

**Context-economy split dei moduli residenti di `/new`** (il fix più grosso del
deep-audit FEAT-0085, rilasciato a parte per il suo carattere strutturale).

- **`/new` 3.5.0** — `setup.md` e `team-mode.md` non vengono più tenuti
  integralmente residenti per tutto il run (misurati ~10–13M di cache_read
  passivo su 124 turni): il happy path resta residente, la prosa condizionale
  va nel nuovo `setup-conditional.md` (letto solo quando la condizione scatta)
  e i corpi di Step D + sequential fallback nel nuovo `team-mode-review.md`
  (letto al primo ingresso in Step D). Residente combinato −45%
  (108.5KB → 59.7KB). Contenuto spostato byte-verbatim; load-site e routing di
  SKILL.md aggiornati; recovery post-compaction verificata.
- Verifica avversariale indipendente: 1 violazione trovata e fissata
  (grep-verification discipline richiamata allo Step C prima della lettura di
  Step D → ora cita l'SSOT `completeness.md`), pointer 6b gatato sul RUN path,
  citazione stale in `new2` (1.12.1) aggiornata.
- Guard: golden-parity 279 ✓, skill-renders ✓, new-parity 27 ✓.

Codex parity: N/A — il package Codex-nativo di `/new` (thin SKILL.md +
controller con prompt per-fase) ha già l'economia equivalente per costruzione
e non condivide questi moduli (manifest scripts-only); nessun lavoro Codex.

## [6.21.0] - 2026-07-16

**`/wrap` + retrospettiva exit-proof.** La retrospettiva terminale (v6.19.0) è
prosa eseguita dal modello: una sessione chiusa con `exit` secco non aveva mai
il turno per farla. Tre pezzi complementari la rendono inevitabile:

- **Nuova skill `/wrap` (1.0.0, capability `core`)**: il comando di chiusura
  sessione — auto-analisi della run (hard fault + soft signal, include i
  `framework_fault:` degli agenti), beacon aggregato se serve, poi
  `mark-retro` e via libera all'exit. Modalità `recover` per la retrospettiva
  differita della sessione precedente.
- **Due nuovi hook nel registry condiviso**: `session-retro-marker.js`
  (SessionEnd, Claude-only: scrive `pending-retro.json` se la sessione esce
  senza `/wrap`; fail-safe) e `session-retro-check.js` (SessionStart,
  Claude+Codex: inietta in contesto la retrospettiva pendente → `/wrap`
  recover prima del nuovo lavoro).
- **`bat-beacon` 1.4.0**: subcommand `mark-retro [--session <id>]` — scrive
  `last-retro.json` (sessione auto-risolta dal transcript più recente) e
  pulisce il marker pendente.
- **Primitivi**: `AGENTS.md` 1.15.0 (il MUST retrospettiva guadagna veicolo
  `/wrap` + dovere di recupero differito a inizio sessione) e `CLAUDE.md`
  1.5.0 (sezione Claude-native: comando + hook).
- Codex parity: adapter-generated per gli hook (SessionStart bound su
  entrambi; SessionEnd non è nel vocabolario hook Codex verificato →
  deliberatamente non bound, compensato dalla skill `/wrap` portabile e dal
  recovery SessionStart che legge anche i marker scritti da sessioni Claude).

## [6.20.0] - 2026-07-16

**Wave anti-waste dall'audit deep del run FEAT-0085 (mayo)** — 4 fix chirurgici
verificati avversarialmente (~12–18M/run evitabili), meta 50/50 tra prevenzione
in `/prd` e disciplina in `/new`.

- **`/prd` 1.19.0 + `prd-card-writer` 1.6.0 — vietato il grep location-pinned
  in `validation_commands`**: un grep di simbolo pinnato a un file che la card
  crea forza la decomposizione dell'implementazione (misurati ~4.1M di rework
  per "spostare il codice dove lo aspetta il grep") ed è gamabile con un
  commento. Ora: behavior-level o symbol-existence repo-wide; mirror lint WARN
  in `agents/card-schema.md`.
- **`/prd` — check cross-wave test-invalidation** nel wave planner
  (`backlog-phase.md` + variante Codex): test structure-pinning di una card
  wave-1 su superfici che una card wave-2 sostituisce vanno risolti in
  authoring (test semantici o migrazione dei test); la card che sostituisce UI
  enumera le suite colocate pre-esistenti che romperà (~2–2.5M evitabili).
- **`/new` 3.4.0 — grounding repo-wide prima di "evidence invalid"**
  (`review-cycle.md` + `team-mode.md` + variante Codex `review-policy.md`):
  mai scartare un finding con un `ls` di singola directory — ricerca repo-wide
  + esecuzione del test citato (un finding VERO scartato così è costato
  ~3.6–5.6M in un run; near-miss di merge con 20 test rossi).
- **`/new` 3.4.0 + `coder` 2.3.0 — anti-loop formatter vs suppression comment**
  (`implement.md` + coder § Author-Time): auto-fixer prima dei fix manuali,
  rimedio strutturale al primo tentativo sugli a11y custom (~5.5M/run,
  deterministico e ricorrente).

Codex parity: portable (prosa dual-runtime; varianti Mode B aggiornate in
parallelo: `new/runtimes/codex/references/review-policy.md`,
`prd/runtimes/codex/references/phases/backlog.md`; coder via transpiler).
Restano dal report: split SECTION= di setup.md/team-mode.md (refactor
strutturale, release dedicata — tocca gli anchor del golden-parity guard),
baseline epic-wide, trappole 1-riga (1.10), grep -a sui transcript (1.6),
grounding BLOCKER pre-fixer (1.7).

## [6.19.0] - 2026-07-16

**Bat-beacon nativity discipline: BALDART migliora con l'utilizzo — by design.**
Il loop bat-beacon → bat-cave diventa parte del DNA di ogni superficie, guidato
da una mappatura completa di skill/agenti/CLI/hook (16 orchestratori scoperti
senza watchdog).

- **Nuova disciplina VINCOLANTE in `CLAUDE.md`** (gemella della cross-tool
  parity): (1) ogni NUOVA skill/agente nasce con la progettazione bat-beacon
  nativa — SSOT skill in `skill-creator/references/skill-structure.md`
  § "Bat-beacon design (REQUIRED at birth)" (una skill senza blocco
  `## Bat-beacon` non shippa), SSOT agenti in `REGISTRY.md` § Notes;
  (2) ogni sessione termina con l'auto-analisi.
- **Primitivo `AGENTS.md` 1.14.0**: il MUST bat-beacon diventa **retrospettiva
  terminale universale** — a fine sessione l'agente rilegge la propria run
  (hard fault E soft signal: attrito, spreco di contesto, capability mancante)
  e spara UN beacon aggregato anche per la minima opportunità di
  miglioramento; il silenzio è legittimo solo a valle della retrospettiva.
- **`return-contract-protocol.md` § Framework-fault flag**: gli agenti (che non
  spawnano mai) relayano i guasti framework-attribuibili con una riga
  `framework_fault: <class> — <evidence>` (classi = set CHIUSO di
  beacon-scan); l'orchestratore le aggrega nel beacon terminale.
- **Blocco `## Bat-beacon` retrofittato in 16 skill orchestratrici** (bump
  MINOR + changelog ciascuna): e2e-review, codexreview, design-sync, research,
  ui-implement, ui-design, design-system-init, graphify-bootstrap,
  lsp-bootstrap, toolchain-bootstrap, overlay, baldart-update, baldart-push,
  graph-align, i18n, capture — watchdog terminale standalone con classi
  dichiarate, anti-doppio-beacon sotto orchestratore.
- Follow-up proposti via bat-beacon (dogfooding): estensione dei reader
  deterministici di `beacon-scan.mjs` (`--e2e`, `--design-sync`,
  `--bug-record`), beacon non-crash del doctor.
- Codex parity: portable (prosa + protocol module condivisi; il primitivo
  AGENTS.md è cross-tool per costruzione).

## [6.18.0] - 2026-07-16

**bat-beacon: il transcript completo viaggia sempre con la issue.** Massima
visibilità per il triage `/batcave`: ogni beacon porta con sé il `.jsonl`
integrale della sessione, non più solo l'estratto-coda.

- **Skill `bat-beacon` 1.3.0** (`scripts/bat-beacon.mjs`): `collect` risolve
  da solo il transcript (`--transcript` esplicito → lookup `--session` negli
  store Claude `~/.claude/projects/` e Codex `~/.codex/sessions/` → `.jsonl`
  più recente del progetto) e ne snapshotta una copia integrale in outbox
  accanto al beacon (nulla si perde se lo store di sessione ruota prima del
  flush). `send` carica lo snapshot sulla branch dedicata
  `bat-cave-transcripts` della repo upstream via GitHub contents API
  (`gh api`, creazione branch idempotente, base64 via Node) e linka il blob
  nella riga `| Transcript |` della tabella metadati. Fail-open: upload
  fallito → issue inviata comunque con nota, snapshot conservato in `sent/`.
  L'estratto inline (~60KB) resta per il triage veloce. Autorizzato dal
  maintainer (repo consumer tutte sue, repo upstream privata).
- **Skill maintainer `batcave` 1.1.0**: Fase 1 scarica il transcript linkato
  (`gh api …?ref=bat-cave-transcripts`) prima di classificare le issue non
  banali — l'analisi root-cause si fa sull'ordine reale degli eventi, non
  sull'estratto.
- Codex parity: portable (script zero-dep Node + `gh`, lookup store sessioni
  Codex incluso; `batcave` è maintainer-only, fuori dal payload).

## [6.17.2] - 2026-07-15

**Codex agent lifecycle: polling timeout non è invocation timeout.** Corretto il
difetto orchestrativo osservato in `/bug` e altri flussi: un
`wait_agent timed_out` indicava soltanto assenza di eventi nella finestra di
polling, ma la regola condivisa “invocation times out → STOP” permetteva di
interpretarlo come fallimento, interrompere un worker sano e dichiarare BLOCKED.

- **Primitivo `AGENTS.md` 1.13.0**: distinzione normativa
  `poll_timeout` / deadline reale; `list_agents` obbligatorio prima di inferire
  lo stato; worker `running` pre-deadline = `safe_to_wait`; checkpoint + grace
  poll obbligatori prima di `interrupt_agent`.
- **Runtime portability**: nuovo state machine chiuso
  (`poll_timeout`, `agent_deadline_exceeded`, `agent_failed`,
  `agent_unresponsive`, `agent_completed`), telemetria minima e default Codex
  `fork_turns: "none"` con brief autosufficiente. `"all"` solo con motivazione
  esplicita. Heartbeat orchestrator-driven dopo 60s senza evento/checkpoint,
  massimo una richiesta outstanding, con fase/file/next-gate/blocker.
- **`/bug` 2.1.1**: deadline dichiarative per architect/reviewer/fixer
  (300/420/600s), polling non terminale, checkpoint via messaggio e grace 60s.
- **Guard**: `check-codex-parity.js` impedisce regressioni della tassonomia,
  del default di contesto e della sequenza checkpoint/grace.
- Limite esplicito: BALDART non modifica il payload host di `wait_agent` e non
  garantisce il recupero di output solo in memoria dopo interrupt; la policy usa
  `list_agents` e persistenza/checkpoint best-effort.
- Nessuna nuova chiave `baldart.config.yml`. Codex parity: portable; `/new`
  Mode B invariato perché usa già worker process-owned, timeout reali e risultati
  persistiti.

## [6.17.1] - 2026-07-15

**Advisory restart-session su skill nuove in `baldart update`.** Caso reale
(mayo): update a 6.17.0 corretto su disco, ma la sessione Codex aperta non
vedeva `bat-beacon` — i tool AI caricano il catalogo skill all'avvio sessione.
`update` ora confronta `.claude/skills/` prima/dopo il reconcile e, se il
payload ha aggiunto skill nuove, stampa l'avviso di riavviare le sessioni
Claude/Codex aperte. Fail-safe (advisory only). Codex parity: N/A (output CLI).

## [6.17.0] - 2026-07-15

**Copertura bat-beacon estesa a tutta l'infrastruttura (Tier 1+2 del
censimento) + regola nel primitivo AGENTS.md.** Principi mantenuti: beacon
solo a livello orchestratore (gli agenti restano route-back, role boundary),
solo colpe framework-attribuibili, anti-doppio-beacon, silenzio = nessun
beacon.

- **Routine schedulate (il buco più grosso)**: clausola WATCHDOG standard in
  coda al prompt di TUTTE le 10 routine (`code-review` in entrambi i prompt,
  claude+codex) — prima un fallimento in-run non aveva né artefatto né canale;
  su backend senza `gh` auth l'outbox viene inclusa nel commit del report.
- **`new2` v1.12.0**: Step 7 post-hygiene — gemello in-memory della Phase 8.5
  di `/new` (consuma `degraded`/`degradationReasons`/`gateLedger` del return
  del workflow; era il motore di default di `/new -auto` senza copertura).
  Regola nuova nel tripwire `check-new-parity.mjs` (27 policy).
- **CLI crash beacon**: nuovo `src/utils/beacon.js` (write-only, fail-safe,
  formato compatibile con l'outbox della skill) + wiring in `bin/baldart.js`
  (catch del doctor, catch di parse, handler `unhandledRejection` per i
  command async) — un crash dell'installer è per definizione un difetto del
  framework.
- **`framework-edit-gate`**: il fail-open resta fail-open, ma il catch
  terminale ora accoda un beacon `gate-not-fired` (dedup: uno in coda alla
  volta) — prima gli errori interni del gate erano invisibili (`exit 0`
  indistinguibile da "nessuna violazione").
- **Tier 2**: `i18n-adopt` v1.1.0 (beacon sul blocker terminale
  framework-attribuibile, mai su stringhe/test del progetto) e
  `worktree-manager` v1.7.0 (`/mw` standalone, solo sottoclasse
  wrapper-error/registry corrotto — mai `build_fail`/`code_conflict` del
  progetto).
- **Primitivo `AGENTS.md` 1.12.0**: nuovo MUST § Non-negotiables (dopo il
  blocco `/bug`) — framework-fault → bat-beacon, con scope esplicito
  (orchestratore-only / framework-only / anti-doppio-beacon / outbox = esito
  valido). Deliberatamente NON duplicato nel primitivo `CLAUDE.md` (regola
  runtime-agnostic → SSOT nel protocollo cross-tool). `bat-beacon` v1.2.0
  (hard rule 6 speculare).
- Deliberatamente SENZA beacon (degradazioni attese, non anomalie — lezione
  v5.21.0): `e2e-review`, `codexreview`, `ui-implement`, `ds-render`,
  `design-sync`, `design-system-init`, `webapp-testing`; e nessun aggancio
  agent-side (gli agenti non hanno Task tool by design).
- Nessuna nuova chiave `baldart.config.yml`. Codex parity: portable (prosa +
  script zero-dep; le routine coprono entrambi i prompt engine).

## [6.16.0] - 2026-07-15

**Watchdog bat-beacon nelle fasi terminali di `/new` e `/prd` — su ENTRAMBI i
runtime.** Chiude il gap emerso in analisi: i segnali di fallimento
orchestrativo erano già tutti su disco (tracker `## Issues & Flags`,
`degradedReviewers`, recovery KILLED, journal Mode B con eventi `anomaly` e
contatori del circuit-breaker, state file `/prd` con sync markers e re-audit
cap) ma NESSUNO li aggregava a fine run. Ora un aggregatore deterministico li
trasforma in un segnale upstream al curator.

- **`bat-beacon` v1.1.0 — `scripts/beacon-scan.mjs`**: scanner zero-dep con
  tre lettori (`--tracker` /new Claude, `--journal` Mode B JSONL, `--state`
  /prd) e un set CHIUSO di 8 classi framework-attribuibili
  (`agent-death-recovered`, `killed-recovery`, `gate-not-fired`,
  `adapter-fail-closed`, `budget-anomaly`, `review-degraded`,
  `retry-ceiling`, `unresolved-divergence`). Lo scan PROPONE, la prosa
  orchestratrice DISPONE; contratto di silenzio (`count 0` = nessun beacon);
  dedup per evidence + cap 5/classe; UN beacon aggregato per run; i
  fallimenti del progetto consumer non sono mai findings.
- **`/new` v3.3.0**: Phase 8.5 (Claude, in `metrics.md`, prima
  dell'emissione terminale — INV-12 e report-ultimo-messaggio intatti) +
  fase `8b-watchdog` (Codex Mode B: contract 25 fasi, runner, evidence
  `watchdog` scritta anche a count 0, prosa nativa) ancorata in
  `PHASE_ANCHORS`.
- **`/prd` v1.18.0**: Step 7.7 (Claude, in `validation-phase.md`, prima del
  Final output; la HARD RULE no-manual-actions resta intatta — la riga
  `Beacon:` documenta un'azione già presa) + fase `7.7-watchdog` (Codex
  Mode B: contract 23 fasi, playbook deletion-list compliant) + 3 anchor
  nuovi in `check-prd-native-contract.mjs` § J.
- Guard verdi: new-native 279 obligations (25 fasi ancorate), prd-native 651,
  new↔new2 tripwire (nessuna policy duplicata aggiunta — `new2.js` resta
  deliberatamente senza watchdog: sperimentale Claude-only, follow-up alla
  promozione), codex-parity, skill-renders, capabilities; test runner Mode B
  e matrice prd-contracts passano.
- Nessuna nuova chiave `baldart.config.yml` → schema-change propagation rule
  non applicabile. Codex parity: portable + adapter-generated (script
  condiviso zero-dep; fasi native dedicate nei due Mode B).

## [6.15.0] - 2026-07-15

**Primo wiring del watchdog bat-beacon: `/bug` Phase 6.** La skill `bug`
(v2.1.0) guadagna la **framework-fault escalation** nel REGISTER: quando
l'investigazione mostra che la root cause (o un gate mancato in
`escaped_gates`) appartiene a BALDART stesso — skill/agente/hook/script del
framework, o un workaround a un difetto del framework — il coordinatore DEVE
invocare `bat-beacon` con `--bug-record <record appena scritto>` prima di
chiudere. Il record locale resta la SSOT del progetto; il beacon è il segnale
upstream al curator. Root cause solo-progetto → nessun beacon (nessun rumore).
Il wiring nella fase terminale di `/new`/`/prd` resta il prossimo incremento
(tocca i guard di golden-parity). Codex parity: portable (prosa + script già
portabili).

## [6.14.1] - 2026-07-15

**Rename `batbeacon` → `bat-beacon`** (leggibilità). Directory skill, script,
label GitHub `bat-beacon`, outbox `.baldart/bat-beacon/`, capability registry,
README e riferimenti in `/batcave` (skill bump 1.0.1). Rilasciato a minuti
dalla v6.14.0, prima di qualsiasi install consumer — nessuna migrazione.
Codex parity: portable (solo rename).

## [6.14.0] - 2026-07-15

**Loop di miglioramento autonomo — lato segnalazione (`/bat-beacon`) + lato
curator (`/batcave`).** Fino a oggi il feedback consumer→curator era un
"messaggio per il curator" incollato a mano: incompleto, senza versioni, senza
transcript, non standardizzato. Questa release apre il canale strutturato,
usando le **issue GitHub della repo BALDART** (privata) come bus asincrono.

- **Nuova skill `bat-beacon` (v1.0.0, capability `core`)** — il bat-segnale.
  Il contesto meccanico è raccolto da uno script deterministico zero-dep
  (`scripts/bat-beacon.mjs collect`): versione BALDART da `.baldart/state.json`,
  runtime host (Claude/Codex autodetect), superficie+componente in fault,
  `tools.enabled` + feature flags, stato git, ultimo record del bug-registry,
  estratto transcript (cap 60KB, collassato in `<details>`). Il modello
  narra solo il sintomo — mai i fatti (lezione v4.53.0: il modello
  fabbrica/dimentica, lo script no). Invio con `send` via `gh issue create`
  (repo risolto da `state.framework_repo`, labels idempotenti); **outbox
  persistente** `.baldart/bat-beacon/outbox/` come coda offline — un invio
  fallito non perde mai il payload (`send --all` flusha). Classi:
  `defect` / `optimization` / `refactor-proposal`. Attivazione: manuale
  (`/bat-beacon`) + clausola watchdog by-intent per gli orchestratori (modello
  `/bug` Phase 6): a fine sessione un malfunzionamento riconducibile al
  framework DEVE produrre un beacon. Distinzione netta da `/bug` (bug del
  progetto → registry locale; causa nel framework → beacon con
  `--bug-record`).
- **Nuova skill maintainer-only `batcave` (repo BALDART, non distribuita)** —
  il triage settimanale del curator: raccoglie le issue `bat-beacon` aperte,
  classifica (`triage:defect|optimization|refactor-proposal|recurrence|
  not-actionable`; la recurrence cross-repo è il segnale prioritario),
  auto-risolve defect+optimization (max 5/giro) tracciando diagnosi, piano e
  diff nella issue, scrive il piano operativo + label `needs-antonio` per i
  refactor da discutere insieme, e chiude le issue risolte alla release
  citando la versione. Tag+push della release cumulativa restano **gated su
  conferma umana**.
- Wiring profondo del watchdog dentro `/new`/`/prd` (fase terminale) e la
  schedulazione domenicale del giro curator: follow-up deliberato — toccano i
  guard di parity e vanno misurati a parte.
- Nessuna nuova chiave `baldart.config.yml` (repo target da
  `state.framework_repo`) → la schema-change propagation rule NON si applica.
  Codex parity: portable (script zero-dep Node + `gh` CLI, autodetect runtime).

## [6.13.0] - 2026-07-15

**Fix — falso verde i18n su object-literal props.** Post-mortem `mayo`: durante
l'estensione del primitivo `SearchBar` con `accessoryAction`, la migrazione ha
introdotto `accessoryAction={{ label: 'Filtri' }}` — stringhe user-facing
hardcoded dentro object-literal TypeScript. Il pre-commit i18n è partito ma è
`jsx-only`: `eslint-plugin-i18next` intercetta testo/attributi JSX, **non** le
proprietà di un object literal passato come valore di prop. Risultato: gate
verde, label non tradotta. Due buchi concentrici chiusi, entrambi in modo
**generico** (per *nome di prop*, mai per componente) e senza passare a
`mode: 'all'` (troppo rumore):

- **Gate deterministico (primario).** La flat config generata
  (`src/utils/i18n-gate.js`, `FLAT_BODY`) ora definisce una seconda regola
  BALDART-owned, `local-i18n/no-hardcoded-ui-object-prop`, che flagga un valore
  stringa su chiave user-facing (`label`/`ariaLabel`/`aria-label`/`title`/
  `placeholder`/`alt`) quando l'object literal vive sotto un JSX expression
  container. La regola è **self-contained e serializzata via `.toString()`** nel
  config (SSOT unica: la stessa funzione `objectPropRuleCreate` è importata e
  unit-testata). Skippa glyph/punteggiatura (`\p{L}`), chiavi computed, valori
  non-stringa (es. `t()`), e object literal fuori da JSX (config). *Limite*: la
  regola è JS → vive nella **flat** config (ESLint 9); una **legacy** rc/JSON
  (ESLint 8) non può ospitarla e resta sul backstop semantico.
- **Backstop semantico (secondario).** `code-reviewer` assumeva "prop JSX → già
  coperta dal linter" ed escludeva la classe: l'object-literal-prop non era né
  lì né tra i `toast()/throw` del suo scope → cadeva nel vuoto anche in review.
  `i18n-protocol.md` § backstop ora la include esplicitamente (enforcement in
  legacy, conferma in flat).
- **Rollout consumer esistenti.** `writeConfig` non sovrascrive mai un config
  user-owned, quindi la regola non raggiunge chi ha già il gate. `baldart
  doctor` guadagna l'advisory non-distruttivo `i18n-gate-objectprop-stale`
  (`I18nGate.isObjectPropRuleMissing`) che rileva un flat config pre-v6.13.0 e
  suggerisce la rigenerazione — senza toccare il file.
- **Test.** Nuovo `src/utils/__tests__/i18n-gate.test.js` (zero-dep, `node:test`):
  replay dell'incidente (`{ label: 'Filtri' }` → 1 finding), copertura di ogni
  chiave allowlist, no-FP su `t()`/glyph/prop-tecnica/computed/config, guardia
  anti-drift lista inline ↔ export, contratto del `FLAT_BODY` e round-trip
  `writeConfig` → `isObjectPropRuleMissing`.

Nessuna nuova chiave `baldart.config.yml` (la regola cavalca il config generato
esistente) → la schema-change propagation rule NON si applica. Codex parity:
**portable** (ESLint gira identico; la strategia i18n-gate è già runtime-agnostica).

## [6.12.0] - 2026-07-14

**Feature — `/new` Codex Modello B residente.** Implementate W1–W3 del deep
post-mortem delle sessioni mayo `019f6020-2f07-7751-adcc-e868ee15d962` e
`019f6027-505f-7170-834d-314088e12b9f`. Il root Codex non orchestra più il
batch con il loop conversazionale `next/record`: una singola
`mcp__node_repl__js` avvia `CompletionTransport`, che attesta il binding,
attende event-driven il processo residente e restituisce solo un confine
`NEEDS_DECISION` o un terminale tipizzato. Nessun polling, heartbeat,
`wait`/`write_stdin` o agente host durante il run.

- **W1 — containment e forensics.** Modello A è default-off e rollback-only
  (`BALDART_CODEX_NEW_MODEL_A_ROLLBACK=1`). L'adapter consuma lo stream JSONL
  di `codex exec`, persiste raw/stderr/terminale per attempt e normalizza usage
  senza ricontare reasoning. Auditor v2 conta tool custom/nested, child rollout
  host, adapter telemetry e backtrack PRD; fixture hashate riproducono entrambe
  le sessioni incidentali.
- **W2 — control plane residente.** Planner/preflight atomico, manifest
  congelato, journal append-only hash-chained, projection ricostruibile, lease
  e crash reconciliation. Un broker unico applica JobSpec/WorkerResult/
  ResultEnvelope strict, ownership del worktree, scope git prima/dopo, usage e
  receipt immutabili; EvidenceRegistry valida le prove lato parent.
- **W3 — pipeline eseguibile.** Handler concreti per grounding,
  implementazione, review/finding lifecycle, gate, commit, delivery e terminal
  transaction. Scope freeze, invalidazione post-review, gate environment-bound,
  outward policy e report terminale digestato sono verificati dal runtime,
  non dichiarati dal worker.
- **Cross-model + hard stops.** Final union `code-reviewer` con Claude Tier 1
  bounded e probe cachata once-per-run, un solo fallback Codex; abort atomico
  finalizzato dal lease owner. `COMPLETE` fallisce su usage ignota/non misurata,
  processi orfani o cap aggregate/per-worker/wall-time oltre soglia.
- Guard di parità `/new`, test regressivi Model B e diagnostica `doctor/status`
  aggiornati. La variante Claude resta invariata. Il canary consumer same-card
  è deliberatamente il prossimo `/new` reale su card nuove.

## [6.11.0] - 2026-07-14

**Hardening `/prd` Mode B — post-mortem Codex session
`019f6096-869f-7a10-80bb-085ed46e3709`.** Aggiunti: equivalence gate sui
primitive riusati; ownership parent-only degli output read-only; parser state
section-local; validator whole-pack; audit summary fresco e final seal con
hash; commit finale delle evidenze; merge docs senza build con delivery
configurata obbligatoria. `mw-docs` verifica il seal, il generic merge rifiuta
i worktree docs, landa prima localmente e prova l'esatto SHA remoto prima del
cleanup. Test regressivi riproducono cap
overflow e level drift FEAT-0084. Codex parity: native + shared gates; Claude
compatibility preservata.

## [6.10.0] - 2026-07-14

**Feature — card-audit cross-model parity su host Codex per `/prd` + `/prd-add`.**
Chiude un'asimmetria di parità reale: su host **Claude** l'audit adversariale
delle card di `/prd` (Step 6.6d) gira già **cross-model** (il companion Codex
rivede il piano di Claude, automaticamente); su host **Codex** restava
same-model. Ora lo slot adversariale whole-set del 6.6 è a **due tier**:

- **Tier-1 `claude-companion` (cross-model)** — Claude Code headless rivede il
  piano scritto da Codex (famiglia di modelli genuinamente diversa) — SOLO
  quando `tools.enabled ∋ claude` AND la probe once-per-run passa (sandbox
  Codex network-OFF → probe fallisce → fallback ATTESO, mai escalation della
  sandbox). Riusa il bridge esistente `framework/scripts/claude-companion.mjs`
  (nessun engine nuovo), eseguito come `exec` read-only (`--permission-mode
  plan`), con costo `CLAUDE_COMPANION_META cost_usd` sempre loggato (spesa
  Anthropic da sessione Codex).
- **Tier-2 fresh-context same-model** — l'unico path quando `tools.enabled` non
  ha `claude` o la probe fallisce; label `fresh-context adversarial
  (same-model)`, mai cross-model.

`/prd-add` eredita la parità attraverso gli stessi native phase playbooks (un
CR con REDO che ri-raggiunge la card-audit); non aggiunge uno slot proprio.

Questo **ribalta consapevolmente** la decisione measure-first di v5.21.0 (che
teneva i check a livello di piano same-model finché un delta misurato non
giustificasse la spesa) — scelta esplicita del maintainer per la parità piena;
l'header di `claude-companion.mjs` è aggiornato di conseguenza.

`prd-audit-fanin.mjs` guadagna `--adversarial-reviewer
<same-model|claude-companion|external>` per etichettare onestamente lo slot
(`--external-reviewer` resta alias). Aggiornati playbook nativi
(`phases/audit`, `audit-policy`, `invariants`), guard
`check-prd-native-contract.mjs` (+3 anchor, 637 obbligazioni, `--enforce`
verde) e freeze matrix `audit-validation.yml`. **Codex parity: portable** (un
bridge zero-dep condiviso, già usato da `/new` + `codexreview`). Nessuna nuova
chiave `baldart.config.yml` (cavalca `tools.enabled`) → la schema-change
propagation rule NON si applica.

## [6.9.0] - 2026-07-14

**Feature — `/prd` + `/prd-add` Codex-native (Mode B variant).** Le due skill di
orchestrazione più pesanti guadagnano un package nativo `runtimes/codex/`
([`docs/references/PRD-CODEX-NATIVE-IMPLEMENTATION-REPORT.md`](docs/references/PRD-CODEX-NATIVE-IMPLEMENTATION-REPORT.md),
9 wave, verificato su `main` § 4.9). Il percorso Claude è **invariato per
costruzione**: i rami host-Codex del bundle base sono avvolti in `{{#rt_codex}}`
(droppati dal render Claude), il resto è byte-identico. A differenza di `/new`,
`/prd` resta una **conversazione multi-turn** — il package nativo è
playbook + contratto di verifica + helper atomici, MAI un controller che avanza
le fasi: lo state file Markdown resta il recovery SSOT.

- **Progress + HITL nativi.** Plan projection a 5 macro-item derive-only
  (`prd-plan-view.mjs`, capability-gated su `update_plan`); decision files
  atomici e replay-safe (`prd-decision.mjs` + `decision.schema.json`) per i
  ~21 gate chiusi via `request_user_input`, con fallback plain; le 11 discovery
  dimension restano una domanda aperta per turno. State template esteso in modo
  ADDITIVO (`Runtime`, `## Runtime Capabilities` con chiavi neutre,
  `## Runtime Decisions`, `## Phase Evidence`) — stati legacy coerenti.
- **Esecuzione agenti condivisa.** Core process/lifecycle estratto da `/new` in
  `framework/scripts/codex-agent-adapter.mjs` (capability proofs fail-closed,
  exact-role TOML, `codex exec --ephemeral` con output schema, barrier full-set,
  teardown, lifecycle inventory); `/new` ne diventa wrapper (regression 51/51);
  `/prd` lo binda alla sua fleet chiusa di 9 ruoli § 12.3.
  `hyper-gamification-designer` pinnato `standard`/read-only (era `inherit`).
- **Dipendenze de-nestate.** `context-primer` → brief diretto a
  `codebase-architect`; `worktree-manager` → nuovi `docs-worktree.mjs` (create)
  + `mw-docs.mjs` (merge docs-mode con return object verificato, worktree
  intatto su fallimento, main HEAD mai mosso). `ui-design` interno
  capability-gated con degradazione DICHIARATA (full parity = follow-up).
- **Audit nativo.** `prd-audit-fanin.mjs` deterministico (coverage, dedup che
  preserva le evidence quote, `[MANUAL]`, classificazione complete|partial|
  blocked, nulla droppato in silenzio); prompt teammate Codex-nativo (zero
  task-spine Claude, un worker per ruolo, barrier process-level); apply-findings
  con drift guard + holistic stamp `agents_run` non fabbricato.
- **Provenance dual-runtime.** Template PRD/card con `planning_runtime` +
  `planning_session_source` (backward-compatible; `unavailable` esplicito, mai
  id inventati) — nessuna env var host nel template condiviso.
- **Golden guard BLOCKING in CI** (`scripts/check-prd-native-contract.mjs
  --enforce`, 634 obbligazioni, zero worklist) + 16 unit-suite; matrice di
  parità congelata in `docs/references/prd-codex-parity-matrix/`; fix drift
  soglia ICIAS `/prd-add` (SKIP≤2). Le fixture end-to-end § 22 e la transcript
  regression § 23 restano consumer-run su host Codex live (vedi DoD § 27).
- Codex parity: **adapter-generated / native package** (un `.md` SSOT per
  runtime, rami Claude invariati, nessuna nuova chiave `baldart.config.yml`).

## [6.8.0] - 2026-07-14

**Feature — wave anti-turn-tax + prevention dal deep post-mortem FEAT-0078.**
Post-mortem verificato (analisti sonnet + verify adversariale opus) di due run
`/new` reali (epic FEAT-0078, mayo: 101M + 182M token fatturati, 0 detector FAIL):
la coda cerimoniale dell'orchestratore valeva il 28-35% del suo input (lone
tracker-Edit 35 turn/10.8M, turn di prosa no-op 47/15.8M, Read ridondante del
Final Report), un FAIL del gate test full-tier non aveva rotta verso il fix pass
off-context, e le classi di rework più care avevano cause di authoring a monte in
`/prd`. Ogni fix cita il numero misurato; ogni claim dei finding è passato dalla
verify adversariale (che ha refutato/ridimensionato ~metà delle stime headline).

- **Skill `new` v3.1.0** (variante Claude): Final Report emesso dalla copia
  in-context (Read dal tracker solo su compaction); "End the turn" = zero testo
  ai barrier; tracker flush come tool call parallele in UN messaggio + nuovo caso
  allowlist per i write di dispatch/close-out; disciplina di attribuzione
  test-failure (isolamento + 3× rerun + repro grezza, mai diagnosi-come-fatto);
  trust discipline dei gate overlay in merge-cleanup (segnale per-item autorevole
  + recovery deterministica per env-gap noti) + step 5f push del commit di
  riconciliazione in entrambe le modalità.
- **Workflow `new-card-review.js`**: i test FAIL full-tier di Discovery
  sintetizzano finding VERIFIED routati al coder fix pass off-context, con
  re-verify diff-scoped post-fix che supersede le righe FAIL (fail-closed);
  scope boundary nel task Codex (mai BLOCKER su stato infra pre-esistente).
  `new-merge-release.js`: il reconcile-6b agent pusha il commit di
  riconciliazione (FF-only).
- **Variante Codex `new` v1.1.0**: disciplina di attribuzione in
  `review-policy.md`; push 6b nella transition table SSOT (`new-contracts.mjs`)
  con `phase-contract.json` rigenerato.
- **Agente `coder` v2.2.0**: red-flag sui pass ottenuti con flag non-default
  (timeout alzato / parallelismo disabilitato ⇒ presunzione di colpa dei propri
  file nuovi + caveat obbligatorio nel completion report); mai backgroundare
  lint veloci single-file; self-check di batching sugli Edit pendenti stesso file.
- **Skill `prd` v1.12.0**: disgiunzione file dei gruppi paralleli verificata
  deterministicamente prima di persistere il launch plan; Complexity Gate con
  audit trail per card boundary; pre-check 8b della divergenza del main checkout
  prima di `mw-docs`; +3 invarianti di authoring in `card-schema.md` (read-path
  AC sulle card edit-di-valore-esistente, seedability degli oracle E2E,
  meccanismo atomico risolto in authoring).
- Guard: `check-new-native-contract` / `check-skill-renders` / `check-new-parity`
  / `check-codex-parity` / `check-agent-runtimes` / `check-capabilities` PASS;
  test del runner Codex verdi. Nessuna nuova chiave `baldart.config.yml` (la
  schema-change propagation rule non si applica). Codex parity: portable +
  adapter-generated (dettaglio per-fix nei CHANGELOG delle skill).

## [6.7.0] - 2026-07-13

**Change — `graphify-mcp` non più materializzato di default (CLI-first).**
Follow-up diretto di v6.6.0: il fix rendeva l'MCP di Graphify *funzionante*, ma
l'analisi ha mostrato che non serve. BALDART usa Graphify **solo via CLI**
(`graphify query/path/explain/affected` + il `GRAPH_REPORT.md` nativo) — nessun
agente/skill consuma i tool del server MCP. Eppure il registro capabilities
(v6.0.0) dichiarava `graphify-mcp` nel `mcp_catalog` + nella capability
`development`, così il capability-gate lo scriveva in `.mcp.json` di ogni
consumer: un processo `graphify-mcp` avviato a ogni sessione **senza che nulla lo
usi** (pattern "materializzato ma non consumato").

- **`framework/capabilities/registry.yml`**: rimosso `graphify-mcp` da
  `capabilities.development.mcp_servers` (ora `[]`) e dall'`mcp_catalog` (ora
  `{}`), con commenti che documentano il perché. Il tool CLI `graphify` resta in
  `external_tools`. Su un consumer con `owned: [graphify-mcp]` nel ledger
  `.baldart/mcp-servers.json`, il prossimo `baldart update` **prune**
  automaticamente l'entry da `.mcp.json` + ledger (`mcp-config.js`
  `syncFromCapabilities` → resolved set vuoto → prune degli owned non più
  desiderati).
- **L'MCP resta opt-in** per chi vuole i tool del server (usi futuri:
  `god_nodes`/`get_community`/`triage_prs`) via il meccanismo del graph layer
  `graph.register_mcp: true` (path indipendente, `claude mcp add`); il fix `[mcp]`
  di v6.6.0 resta valido per quel path.
- **Docs**: `CODE-GRAPH-LAYER.md` (riga MCP: "opt-in, not materialized by
  default") allineata; `MCP-INTEGRATION.md` / `PROJECT-CONFIGURATION.md` già
  descrivevano l'MCP come opt-in.
- Codex parity: **adapter-generated** (il prune copre anche l'advisory TOML Codex
  via lo stesso resolved set). **Nessuna nuova chiave `baldart.config.yml`** → la
  schema-change propagation rule non si applica.

## [6.6.0] - 2026-07-13

**Fix — Graphify MCP: install con l'extra `[mcp]` + backfill automatico.**
L'installer di Graphify (`src/utils/graphify-installer.js`) installava
`graphifyy` **senza** l'extra `[mcp]`, ma il flusso registrava comunque il
server MCP (`claude mcp add … graphify-mcp`). Risultato su ogni consumer con
`features.has_code_graph: true` (es. `mayo`): un MCP registrato che crashava
all'avvio con `ModuleNotFoundError: No module named 'mcp'`, mentre la CLI (il
percorso che gli agenti usano davvero, via `graphify query/path/explain/affected`)
funzionava. Il gap peggiora sulla serie upstream **0.9.x**, dove `mcp` è un extra
gated (verificato: `graphifyy` PyPI latest 0.9.14, `requires_dist` `mcp; extra == "mcp"`).

- **`graphify-installer.js`**: `install()`/`upgrade()` usano lo spec
  `'graphifyy[mcp]'` (costante `PKG_SPEC`, single-quoted contro il glob della
  shell). Due nuovi metodi best-effort/never-throw: `mcpExtraPresent()` (probe via
  `pipx runpip graphifyy show mcp` → fallback `python3 -c "import mcp"`, ritorna
  `true|false|null`) e `ensureMcpExtra()` (backfill chirurgico: `pipx inject
  graphifyy mcp` per pipx — non tocca la versione della CLI — o
  `pip install --user 'graphifyy[mcp]'`).
- **`doctor.js`**: nuova detection `graphMcpBroken` (helper `graphifyMcpRegistered()`
  legge `.mcp.json`) — flagga SOLO quando il server è registrato E il probe è
  conclusivamente negativo (mai nag su probe indeterminabile). Nuova azione
  `graph-mcp-fix` (`autoOk:false`, delega a `ensureMcpExtra()`) + riga di stato
  "MCP server broken — missing [mcp] extra".
- **`configure.js`**: quando `graph.register_mcp: true`, se l'extra manca su
  un'install preesistente lo ripara prima di registrare l'MCP (auto-heal di
  install pre-fix come mayo). Messaggi UI aggiornati allo spec `[mcp]`.
- **Docs/skill/template**: `graphify-bootstrap` (skill v1.1.0 + CHANGELOG),
  `CODE-GRAPH-LAYER.md`, `MCP-INTEGRATION.md`, `PROJECT-CONFIGURATION.md`,
  `baldart.config.template.yml` aggiornati allo spec `'graphifyy[mcp]'`.
- Codex parity: **portable** — lo spec d'install è identico su ogni runtime
  (l'extra `mcp` è innocuo su Codex); il backfill MCP è Claude-specifico by design
  (la registrazione MCP di Graphify passa da `claude mcp add`, su Codex l'MCP è
  advisory), quindi `graph-mcp-fix` non appare su un consumer Codex-only.
- **Nessuna nuova chiave `baldart.config.yml`** (cavalca `features.has_code_graph` +
  il blocco `graph.*` esistenti) → la schema-change propagation rule non si applica.

## [6.5.0] - 2026-07-13

**Feature — read-economy discipline nella flotta, misurata sui transcript reali.**
Follow-up data-driven della valutazione sul `coder`: invece di rifattorizzarne la
definizione (bocciato — peso non-outlier, ~2% di guadagno), un audit dei transcript
`mayo` (431 run coder / 314 card, nuovo tool `scripts/coder-read-audit.mjs`) ha
scomposto il read del coder (`read_bytes_share=0.74`, il vero driver di costo) in:
re-read recuperabile **10–17%** (intra-run + intra-card cross-run, rischio ~zero) ·
whole-file **80%** (di cui 58% file >15KB, bacino range-read ma con rischio-correttezza
FEAT-0042). I top offender hanno rivelato i frutti bassi: `README.md` 2.7MB (#1),
screenshot PNG letti in contesti di codice, doc-hub riletti interi.

- **`framework/agents/code-search-protocol.md`** (§ Large-file read discipline, SSOT
  condiviso): 3 regole nuove — (4) doc-hub per sezione via router, mai whole/ripetuto;
  (5) mai Read di asset binari (PNG/screenshot → dominio `visual-fidelity-verifier`/
  `ui-quality-critic`); (6) anti-re-read memo entro la run. + blocco Grounding coi
  numeri misurati. **Ereditato per citazione** da 13 agenti (coder, ui-expert,
  codebase-architect, …) senza toccare i loro file.
- **`framework/.claude/agents/coder.md`** (v2.1.4): 1 riga di salienza nel Context
  Loading Protocol (il target misurato) che nomina i 3 comportamenti + cita l'SSOT.
- **`scripts/coder-read-audit.mjs`** (nuovo, maintainer-only, Codex parity N/A —
  transcript Claude-specifici): twin di `agent-telemetry.mjs`, scompone il read in
  necessario/recuperabile per il before/after ripetibile.

Nessuna nuova config key → la schema-change propagation rule NON si applica. Codex
parity: portable (modulo markdown ereditato + agent transpilato; il tool è
maintainer-only). Misura di controllo: ri-lanciare l'audit dopo alcune settimane di
run per verificare la discesa del read.

## [6.4.0] - 2026-07-13

**Feature — asse leggibilità/self-documentation nel simplify layer (classe A10 +
prevenzione author-time nel `coder`).** La leggibilità semantica (naming,
magic-literal, argument-opacity) era coperta solo *implicitamente* e sparsa
(Beck rule 2, comments-as-deodorant), senza una classe di tassonomia propria.
Grounded in ricerca 2025-26 (report:
[`docs/research/simplify-readability-self-explanatory-lens.md`](docs/research/simplify-readability-self-explanatory-lens.md))
che ha **refutato la versione ovvia** (un readability *score* / un *fixer* di
rename autonomo): il giudizio di leggibilità di un LLM è manipolabile dal naming
(fino a ~78% di score solo rinominando — arXiv:2507.05289), non esiste un floor
deterministico affidabile (i model Buse-Weimer/Scalabrino correlano poco sul
codice AI), e il refactoring iterativo di readability oscilla (arXiv:2602.21833).
La forma adottata è **advisory, evidence-grounded, no nuovo spawn**:

- **`framework/agents/simplify-protocol.md`**: nuova classe **A10 —
  Intent-opacity** in `SECTION=taxonomy` (ortogonale ad A9 strutturale; solo
  segnali verificabili, mai uno score; advisory LOW/MEDIUM), folded nel Quality
  lens di `SECTION=rubric`, e nuovo **naming-caution guard** in
  `SECTION=bias-guards` (il buon naming non prova la correttezza / il naming
  criptico su codice corretto non è un blocker; no comment-addition reward; fix
  one-shot, mai loop di rename).
- **`framework/.claude/agents/coder.md`** (v2.1.3): 1 riga author-time che
  operazionalizza "readable" (naming intent-revealing · no magic-literal · no
  boolean/positional-arg opachi · nome/struttura prima del commento) — il fronte
  a ROI più alto (prevenzione, zero churn).
- **`framework/.claude/agents/code-simplifier.md`** (v1.0.6): enum `taxonomy`
  esteso ad A10; Quality lens dello Step 2 cita A10 (mappa sulla category
  `maintainability` già esistente).
- **`framework/.claude/skills/simplify/SKILL.md`** (v1.2.0): Agent 2 punto 10
  (Intent-opacity) + naming-caution guard.

Deliberatamente **NON** incluso: un micro-floor grep-based magic-literal (vero
script deterministico con reali rischi di falso-positivo su enum/i18n/config →
card separata con analisi di ridondanza, coerente con lo YAGNI). Nessuna nuova
config key → la schema-change propagation rule NON si applica. Codex parity:
portable (prosa di protocollo + agent transpilati, nessun sub-spawn).

## [6.3.1] - 2026-07-13

**Fix — l'hook `agent-discovery-info` e il `doctor` ora sono capability-aware.**
Un consumer che disabilita una capability (es. `machine-learning`,
`gamification`, `psychology`) non genera i relativi agent (nessun `.md` symlink
su Claude, nessun `.toml` su Codex): è corretto. Ma le tre enumerazioni che
confrontano gli agent *sorgente* con quelli su disco non applicavano il
capability-gate, così segnalavano gli agent legittimamente esclusi come
"mancanti/rotti" — all'avvio di Codex `agent-discovery-info` avvisava su
`deep-human-insight`, `hybrid-ml-architect`, `hyper-gamification-designer`, e il
`doctor` proponeva un `baldart update` inutile (che li riprunerebbe). Fix:

- **`src/commands/doctor.js`**: un helper `capFilterAgents()` risolve la
  selezione una volta (memoizzata) e filtra sia `agentSymlinksBroken` (Claude)
  sia `codexAgentsMissing`/`codexAgentsStale` (Codex) via
  `capability-gate.filterItems`. Fail-open: qualunque errore → filtro identità
  (comportamento pre-v6.0.0). Gli agent non dichiarati nel registry restano
  sempre attesi (trattati come core).
- **`framework/.claude/hooks/agent-discovery-info.sh`**: prima di segnalare i
  mancanti, interroga la nuova CLI `node .framework/src/utils/capability-gate.js
  gated-off agents <cwd>` (SSOT della risoluzione, `includes:` transitivi
  inclusi) e salta gli agent gated-off. Fail-open: niente node / niente
  `js-yaml` risolvibile / qualunque errore → set vuoto → riporta esattamente
  come prima (nessuna regressione; la protezione reale sugli agent *abilitati*
  ma mancanti resta intatta).
- **`src/utils/capability-gate.js`**: aggiunta CLI `gated-off <kind> [cwd]` che
  stampa i nomi dichiarati ma esclusi dalla selezione (fail-open: stampa nulla).

Codex parity: portable — l'hook è dual-runtime (`--runtime codex`), il doctor e
la CLI sono runtime-agnostici. Nessuna nuova config key → la schema-change
propagation rule NON si applica.

## [6.3.0] - 2026-07-13

**Capability slot flags (`{{#cap_*}}`) — le superfici condivise si parametrizzano
sulla selezione delle capability.** Il capability-gate installava/prunava gli
item ma le skill/agenti multidisciplinari continuavano a portare gate verso item
di capability non abilitate (es. il MUST-invoke `hyper-gamification-designer`
di `/prd` su un consumer senza `gamification` → contesto morto + spawn di un
agente fantasma, che su Codex non onora nemmeno il gate in prosa). Ora i gate
cross-capability sono avvolti in `{{#cap_<name>}}…{{/cap_<name>}}` (+
`{{^cap_*}}` per il ramo else), risolti a install-time:

- **Terza dimensione del renderer skill** (`src/utils/skill-renderer.js`): i
  token `cap_*` mettono il bundle in render mode (detection-based, come `rt_*`)
  e si risolvono nello stesso fill; i flag derivano da
  `capability-gate.capabilityFlags()` (ON = raggiungibile da `enabled`+core via
  `includes:`; trattini → underscore; identici per Claude e Codex → parity per
  costruzione); il `config_sha` del marker copre i cap flag → il doctor guarisce
  il drift al cambio di selezione. **Fail-open**: senza blocco `capabilities:`
  ogni ramo positivo resta (payload pieno, prosa invariata). Il lint
  forbidden-token resta legato alla SOLA dimensione runtime: un bundle
  capability-only non dichiara nativeness e mantiene lo stile dual-runtime a
  rami in prosa (`/prd` è il primo caso: entra in render mode senza lint).
- **Stessi flag negli agent slots** (`SymlinkUtils._mergeBulkDir` + probe
  doctor `generatedAgentsStale`): un agente con blocchi `cap_*` diventa
  generated per-consumer come gli slot v5.0.0.
- **Gate parametrizzati in questa release** (mappatura completa: 4 gate
  operativi reali, il resto già risolto da item condivisi/DAG `includes:`/
  fallback esistenti): `/prd` skill v1.11.0 (Specialist Audit gamification) +
  agente `prd` v1.1.0 (regola 7, Phase 0 check, Phase 1.5 con no-op line
  quando off, template §12, riga card-owner `legal-counsel-gdpr` →
  `cap_website`); `plan-auditor` v2.2.0 (riga auto-spawn `hybrid-ml-architect`
  + enum `source:` → `cap_machine_learning`); `onboarding-architect-lead`
  v1.1.0 (delega MUST verso coder/ui-expert/visual-designer → `cap_design`,
  ramo else degradato `CAPABILITY_UNAVAILABLE`).
- **Guard maintainer**: `check-skill-renders.mjs` renderizza ogni skill in
  render mode sotto ENTRAMBE le selezioni estreme (all-on / core-only) — un
  ramo else rotto fallisce la release; `check-capabilities.mjs` #7 esteso a
  tutti i `.md` dei bundle (non solo SKILL.md) e reso cap-block-aware, nuovo
  #8 sui corpi degli agenti (riferimento fuori set non coperto da blocco
  `{{#cap_*}}` → WARN con suggerimento del wrapper; 13 warning pre-esistenti =
  worklist advisory-then-fail).
- **Residual agent-slots ristretto ai token a blocchi** (`{{#…}}`/`{{^…}}`/
  `{{/…}}`): uno scalare model-facing legittimo (es. `{{git.trunk_branch}}` in
  prd.md) non genera più warning quando il file entra in slot-mode.
- **Nessuna nuova chiave** in `baldart.config.yml` (riusa il blocco
  `capabilities:` v6.0.0) → la schema-change propagation rule NON si applica.
- Codex parity: **adapter-generated** — un solo SSOT, render/transpile
  per-runtime dagli stessi flag.

## [6.2.0] - 2026-07-13

**`baldart capabilities` — editor da terminale per abilitare/disabilitare i
cluster di capability.** Nuovo comando (alias `caps`) con un menu a **checkbox
interattivo** (SPACE per toggle, pre-compilato con la selezione corrente): alla
conferma persiste la whitelist in `baldart.config.yml`, **ri-mergia l'INTERO
payload** (skill / agenti / comandi / routine / workflow / output-style) + sincronizza
`.mcp.json`, chiede gli eventuali **secret condivisi** mancanti (salvati in
`~/.baldart/secrets.yml`) e stampa il **diff** degli item aggiunti/rimossi.
Scriptabile in non-interattivo: `--list`, `--enable a,b`, `--disable c`, `--all`,
`--json`, `-y`.

- **Componente centralizzato** `src/utils/capabilities-ui.js` (read/choices/persist/
  apply/diff/secrets) riusato da TRE superfici — il comando, `configure` e `doctor` —
  così menu, regola di persistenza (all-on → whitelist rimossa) e re-merge/prune
  sono identici ovunque. Persistenza attraverso il serializzatore canonico di
  `configure` (`loadExisting`/`serialize` esportati) → un solo SSOT per il file.
- **Gap chiuso in `configure`**: la ri-mergia finale ora passa da `applySelection`
  e riconcilia anche commands/workflows/output-style/MCP (prima solo skill+agenti,
  quindi i cambi di capability su quei kind valevano solo al successivo `update`).
- `doctor` `reconcile-capabilities` delega allo stesso `applySelection`.
- Nuovo helper `UI.checkbox()` (multiselect inquirer) in `src/utils/ui.js`.
- **Nessuna nuova chiave** in `baldart.config.yml` (il blocco `capabilities:`
  esiste da v6.0.0) → la schema-change propagation rule NON si applica.
- Codex parity: **N/A / portable** — plumbing CLI dell'installer, tool-agnostico
  per costruzione: modifica solo la config + ri-mergia il payload, che il
  capability-gate risolve già a set identici per Claude e Codex (il merge itera
  `tools.enabled`).

## [6.1.0] - 2026-07-13

**BALDART Console — la UI di controllo interna diventa una piattaforma unica.**
Nasce `scripts/console/` (maintainer-only, NON distribuito): un solo comando
(`node scripts/console/server.mjs`, porta 4678, skill `/baldart`) accende tutti
gli editor interni come moduli di una sola SPA zero-dep — **Overview** (versione,
inventario payload, stato git, validator on-demand su allowlist), **Model Map** e
**Capability Map** (logica portata 1:1 dai server standalone, semantica di
salvataggio invariata; in più la matrice ora rende visibile la membership
EREDITATA — `✓ core`/`✓ <cap>` attenuato nelle celle: core è sempre-on e gli
`includes:` propagano, quindi l'item non va ri-dichiarato). Architettura modulare pronta a scalare: un modulo nuovo =
1 file API in `modules/` + 1 vista in `ui/views/` + 1 entry di registry. Il design
è governato dal **design system interno** `scripts/console/DESIGN-SYSTEM.md`
(direzione "Tactical HUD"; token SSOT in `ui/tokens.css`, componenti in
`ui/console.css` — vincolante per ogni vista futura). Ritirati
`scripts/model-map-server.mjs` (porta 4680) e `scripts/capability-map-server.mjs`
(porta 4681); le skill `/model-map` e `/capability-map` (v2.0.0) restano come
redirect deprecati verso la console. Codex parity: N/A — tooling maintainer-only
del repo BALDART, non fa parte del payload installato nei consumer.

## [6.0.1] - 2026-07-11

- Codex `/new`: prompt e brief outcome-first allineati alla guida ufficiale GPT-5.6,
  con completion bar e stop condition espliciti e meno prosa duplicata.
- Adapter agenti Codex: output schema OpenAI-strict proiettato per fase, identita
  vincolata, path evidenza assoluti e failure closed; aggiunti guard di budget prompt.
- Documentata la guida ufficiale come riferimento obbligatorio per futura prosa Codex.

## [6.0.0] - 2026-07-11

**Capabilities — il payload diventa selezionabile per cluster.** Skill, agenti,
comandi, routine, workflow, output style, tool esterni e server MCP sono ora
raggruppati in **capabilities** dichiarative (SSOT:
`framework/capabilities/registry.yml` — v1: `core` sempre-on, `development`
(includes `design`), `design`, `website` (includes `design`), `creative-media`,
`machine-learning`, `gamification`, `psychology`; item condivisibili,
composizione via `includes:`). Il consumer seleziona via il nuovo blocco
`capabilities:` in `baldart.config.yml` (`enabled` + `extra_*` + `exclude`;
**blocco assente = TUTTE** — retrocompatibilità piena, nessuna rimozione
silenziosa; `core` intoccabile). Il filtro è il PRIMO gate, ortogonale e
composto col feature-gate `requires_feature` esistente; risolto UNA volta nello
strato di merge condiviso (`src/utils/capability-gate.js`, fail-open ovunque) e
applicato a skill (incl. render per-runtime), agenti (incl. prune dei TOML
Codex), comandi, workflow, output style, routine (filtro nel catalogo) e MCP.
**Tool esterni & credenziali condivise**: `tools_catalog` (probe + install_via)
e `mcp_catalog` (placeholder `${secret:NAME}` only — mai valori); store
user-level `~/.baldart/secrets.yml` (chmod 600, condiviso tra progetti, env
vince); writer `.mcp.json` marker-owned (ledger `.baldart/mcp-servers.json`,
mai clobber di entry utente); su Codex advisory paste-block (invariante: mai
scrivere `~/.codex/config.toml`). **Propagazione schema completa**: template +
prompt multi-select in `configure` (interactive-only; tutte-on ⇒ chiave omessa;
prompt secrets mancanti) + notifica capabilities disponibili in `update` +
doctor (`reconcile-capabilities` self-heal, advisory unknown-names /
credential-missing / Codex MCP). **Maintainer tooling**: validator
`scripts/check-capabilities.mjs` (FAIL su item fantasma/orfani/cicli
includes/catalog refs/valori secret letterali; WARN su agenti delegati fuori
set) + skill interna `/capability-map` (UI zero-dep porta 4681, matrice
item×capability, orfani/ghost, validator al save — non distribuita). MAJOR:
nuova superficie di install selettiva (nessun breaking per config esistenti —
il default resta identico). Codex parity: adapter-generated (registro SSOT
unico, gate nello strato condiviso; MCP = documented-advisory). Design doc:
`docs/references/CAPABILITIES-DESIGN.md`.

## [5.22.5] - 2026-07-11

**Tier‑1 Claude strutturato per `/new` Codex.** Il bridge usa il vincolo nativo
`claude --json-schema`, adatta deterministicamente il meta-schema canonico al
dialetto CLI, normalizza JSON esatto o fenced e rifiuta output malformato.
Il controller persiste direttamente l'envelope e `record` lo valida fail-closed.
I brief includono il `run_id`; finding e relativi enum sono validati
ricorsivamente. Un fallimento Tier‑1 a task avviato viene materializzato e
reindirizzato una volta al reviewer Tier‑2, senza perdere il gate.
Smoke reale: review read-only, finding enum conformi, nessuna modifica al codice.
Variante Claude `/new` invariata.

## [5.22.4] - 2026-07-11

**Shared-script resolution Codex.** `new-runner` risolve ora
`final-sync-gate.mjs` e `revert-files.sh` dal bundle base condiviso anche quando
`import.meta.url` attraversa il symlink e punta a `runtimes/codex/scripts/`.
Regression test sull'esistenza reale del comando F.6. Claude invariato.

## [5.22.3] - 2026-07-11

**Entry guard Codex import-safe.** Il confronto symlink-safe degli entrypoint
ora tratta `node -`/eval come import, senza chiamare `realpathSync("-")`.
Controller, contracts e adapter restano eseguibili dal mount e importabili nei
tool di simulazione/test. Claude invariato.

## [5.22.2] - 2026-07-11

**CI runtime guard render-aware.** `check-runtime-asset-resolution.mjs` non
tratta più il bundle base Claude di una skill Mode-B come percorso Codex:
verifica invece la variante `runtimes/codex/` installata realmente. Elimina i
tre falsi positivi `~/.claude/plugins` introdotti dal passaggio per-runtime,
senza cambiare un byte della variante Claude.

Codex parity: **maintainer guard**. Runtime payload invariato rispetto a 5.22.1.

## [5.22.1] - 2026-07-11

**`/new` Codex realmente eseguibile sull'host reale.** Patch post-audit della
prima variante renderizzata: il bundle non usa più la collaboration API host,
che non espone selezione ruolo/modello, barrier completa o close.

### Fixed

- **Adapter worker Codex-native** — `codex-agent-adapter.mjs` esegue ruoli con
  `codex exec --ephemeral`, caricando il TOML esatto e pinning esplicito di
  modello, reasoning effort e sandbox; output JSON Schema, fan-out concorrente
  con barrier `Promise.all`, timeout/abort PID-owned e inventory closure.
- **Controller installato** — entrypoint symlink-safe, path `.agents/skills/new/`
  corretti, input invalido fail-closed, flag `full|stats|auto|auto-ship|effort`
  persistiti. Le fasi senza comando reale sono checkpoint root-owned; ogni
  `exec` porta un `command[]` eseguibile.
- **Agenti generati nativi** — i 15 ruoli `/new` ricevono contratti Codex-native
  compatti (circa 1-2 KB, non il body Claude fino a 67 KB), mantenendo modello,
  effort e sandbox dal frontmatter. Nuovo guard anti-meccaniche Claude; suite
  controller **50 pass, 0 fail**, golden parity **228/228**.

Codex parity: **native CLI adapter**. Claude variant: byte-invariata.

## [5.22.0] - 2026-07-11

**Skill native per runtime + `/new` Codex-nativo.** Il major switch "un SSOT logico → pacchetti nativi per runtime": una skill può ora essere installata come bundle GENERATO per-runtime invece che come symlink unico cross-runtime, e `/new` è la prima skill renderizzata — la variante Claude (bundle base v3.0.0, zero rami host-Codex, gate Codex-plugin intatto) e un **pacchetto Codex-nativo completo** costruito dalla spec verificata da Codex (`docs/references/NEW-CODEX-NATIVE-IMPLEMENTATION-REPORT.md`; piano: `docs/references/skill-runtime-render-plan.md`).

### Added

- **`src/utils/skill-renderer.js` + test** — il renderer per-runtime delle skill, due modalità detection-based: **Modalità A** (condizionali `{{#rt_*}}`/`{{^rt_*}}` + frontmatter `runtimes:` declinato per runtime, sidecar `agents/openai.yaml`, campi install-time strippati dai render) e **Modalità B** (variant set: `runtimes/<rt>/` = pacchetto nativo + `manifest.json` dei file condivisi montati come symlink — mai fork). Output byte-stable in `.baldart/generated/skills/<rt>/<name>/` con marker; **lint per-runtime hard-fail** integrato (`adapter.skillRenderProfile()`: deletion-list §19 per Codex, meccaniche host-Codex per Claude — mai la parola "Codex" in sé). Design doc: `framework/docs/SKILL-RENDERING.md`; convenzione in `skill-creator/references/skill-structure.md` (v1.1.0).
- **`/new` pacchetto Codex-nativo** (`framework/.claude/skills/new/runtimes/codex/`) — SKILL.md sottile (~195 righe: 12 invarianti, root loop a 7 passi, contratti decisione/recovery/terminal-report) + 6 reference + **controller deterministico**: `new-runner.mjs` (state machine `init|next|record|decide|abort|status`, envelope chiuso `exec|spawn|join|decision|checkpoint|complete|blocked`, fail-closed, riconciliazione da git), `new-contracts.mjs` (transition table 24 fasi → `phase-contract.json` generato, F.0 fast-lane a 6 predicati, dispositions con rigetto di DEFER/SKIP nudi), `new-state.mjs` (3 autorità di stato, AC ledger `implemented|deferred`, scritture atomiche), `new-budget.mjs` (circuit breaker SOLO su counter live-enforceable; unsupported = `unavailable`, mai stimato), `codex-agent-adapter.mjs` (capability matrix fail-closed, `fork_turns:"none"` sempre, `fork_context` mai, lifecycle `active_batch_threads==0` a ogni exit), brief generator cache-friendly, review fan-in con writer serializzati; 4 JSON schema; team-mode a wave DAG (max 3 owner, una barrier full-set). Gate adversarial a 2 tier: `claude -p` cross-model via `claude-companion.mjs` quando configurato+probato, fallback fresh-context `code-reviewer` etichettato onestamente same-model. 43+ unit test.
- **Agenti `security-finder` + `doc-finder`** (v1.0.0) — i finder READ-ONLY richiesti dal report §10 (Domain Partition esplicita vs i writer `security-reviewer`/`doc-reviewer`); Claude opus/high e sonnet/medium, Codex deep/high e fast/medium con `sandbox: read-only`. Fleet → **35 agenti**.
- **Launch profile Codex opzionale** — template `framework/templates/codex/baldart-new.config.toml` (root economico `gpt-5.6-terra/low` + auto-compact 150k) offerto da `configure` in opt-in interattivo default-NO (mai riscrittura silente di config globali); advisory non-bloccante in `doctor`.
- **Guard maintainer**: `scripts/check-skill-renders.mjs` (render in-memory claude+codex + lint + zero prosa legacy nel bundle base) e `scripts/check-new-native-contract.mjs` (**golden parity §20.2**, ~200 obbligazioni: 24 fasi ancorate nella variante Claude, 12 invarianti in ENTRAMBE le varianti, dispositions/F.0 eseguiti davvero, deletion-list, `phase-contract.json` in sync, anti-inherit sulla flotta `/new`) — la prova primaria che le due varianti restano semanticamente allineate ora che la SSOT non è più un testo unico.

### Changed

- **`/new` (skill v3.0.0)** — il bundle base è ora la variante Claude NATIVA: rimossi tutti i rami host-Codex (~100 righe su 10 file, zero cambi di policy, verifica avversariale no-normative-loss); frontmatter `runtimes:` con `codex: {variant: runtimes/codex}`. Il gate di review cross-model via plugin Codex su run Claude-hosted RESTA integrale (il discriminante è la semantica host-runtime, mai la parola "Codex").
- **Declinazioni `runtimes.codex` allineate alla tabella §10 del report** — 8 divergenze corrette (tra cui `plan-auditor` erroneamente workspace-write → read-only; `merge-conflict-resolver` da inherit → deep/high).
- **`symlinks.js`/`doctor`/`update`** — `_mergeSkillsForTool` a tre vie (prune / render / symlink) con conversioni legacy in entrambe le direzioni; probe `renderedSkillsStale` + azione `regenerate-skills` + advisory; allowlist reconcilable estesa esplicitamente ai bundle generati. `agent-slots.js`: negazione `{{^flag}}` + fill key-scoped (backwards-compatible).
- **`check-codex-parity.js`** — invarianti aggiornati alla nuova architettura: `/new` base NON cita più `runtime-portability-protocol.md` (single-runtime by design; il modulo resta per `/prd` e le altre skill), il pacchetto nativo deve esistere con manifest interamente risolvibile.

Codex parity: **adapter-generated / native package** (un bundle SSOT, render per-runtime; la parity semantica è garantita dal golden-parity guard, non più dalla prosa condivisa). Nessuna nuova chiave `baldart.config.yml` (tutto su `tools.enabled`) → la schema-change propagation rule non si applica. Fuori scope dichiarato prima release: Responses API adapter, fixture funzionali §20.4 su host reale, A/B §21.

## [5.21.1] - 2026-07-10

**claude-companion: model/effort espliciti.** `claude-companion.mjs` ora pinna modello ed effort delle chiamate headless invece di ereditare il default dell'account utente (lo smoke test v5.21.0 girava su Opus per caso): task = `--model sonnet --effort high` (il finder non necessita opus — l'FP-gate valida i suoi finding; mai haiku, ban fleet-wide), probe = `sonnet --effort low`; entrambi overridabili per-call via passthrough `--model`/`--effort` (escalation a `opus` sui trigger high-risk, documentata nei due gate). Verificato end-to-end con chiamate reali. Nota onesta: il risparmio è contenuto (il costo per chiamata è dominato dalla cache-creation del system prompt, non dal tier). Skill `codexreview` v1.1.1, `new` v2.9.1. Codex parity: inline-fallback (invariata).

## [5.21.0] - 2026-07-10

**claude-companion: cross-model review anche su host Codex.** Lo speculare del companion Claude→Codex: su un run Codex-hosted lo slot cross-model dei gate di review (finora degradato a same-model fresh-context, v5.20.0) può essere riempito da **Claude Code headless** (`claude -p`, read-only). Design passato per refutazione avversariale leggera e ridimensionato: 2 gate su 5, gating su intent esplicito, probe reale, zero escalation sandbox. Validato con smoke test reale end-to-end (review corretta di un bug seminato, read-only rispettato, costo tracciato).

### Added

- **`framework/scripts/claude-companion.mjs` + `claude-companion.test.mjs`** — bridge zero-dep BALDART-owned (il verso opposto non ha un plugin da riusare, a differenza di `codex-companion.mjs` che è del plugin openai-codex): `probe` (check REALE binario+auth+network con un round-trip economico — un sandbox Codex network-OFF fallisce QUI, by design; refutato il probe solo-binario: binario presente ≠ intent né connettività) e `task` (`--permission-mode plan` + `--disallowedTools Write,Edit,NotebookEdit,…`, env scrubbing `CLAUDE_CODE_*`/`CLAUDECODE` con `CLAUDE_CONFIG_DIR` preservato per l'auth, timeout interno — macOS non ha `timeout(1)` —, parse dell'envelope JSON, `CLAUDE_UNAVAILABLE:<classe>` con classi `no-binary|no-auth|network|timeout|bad-output`). Contratto output: testo review su `--output`, meta-riga `CLAUDE_COMPANION_META cost_usd=…` su stderr che il chiamante DEVE loggare (spende il piano Anthropic dell'utente da una sessione non-Anthropic). THIN by contract: spawn/timeout/parse/verdict, zero logica di review. Test CI-safe (dry-run + classificazione, nessuna chiamata reale).

### Changed

- **`codexreview` (skill v1.1.0)** — il binding Codex-hosted dell'agente #4 diventa a DUE TIER: Tier 1 `claude-companion` **cross-model** SOLO quando `tools.enabled ∋ claude` (intent dichiarato — un consumer Codex-only non è MAI addebitato; niente nuova chiave config) AND il probe once-per-run è passato (verdetto cachato nel tracker; network-OFF → fallback ATTESO, mai chiedere escalation della sandbox); Tier 2 = il fallback v5.20.0 (`code-reviewer` fresh-context, mai etichettato cross-model). `source: "claude"` per i finding Tier 1; `Method: codex=claude-companion(cross-model)`.
- **`/new` (skill v2.9.0)** — final-review F.3: stesso two-tier sul ramo Codex-hosted, stesso `$REVIEW_FILE` (F.4 invariato), stesso probe cachato del gate 3.7; `codex-gate.md` puntatore aggiornato.
- **`runtime-portability-protocol.md`** § "Adversarial vs cross-model": la riga Codex documenta il companion come cross-model condizionale; consumer odierni SOLO agent #4 + F.3 — i check a livello piano (cross-card, 6.6d, discovery) restano same-model finché un delta misurato non giustifica la spesa (data-driven sopra threshold arbitrarie).
- **`doctor`** — advisory `claude-companion` quando `tools.enabled` contiene sia codex sia claude: check GRATUITO del binario (mai il probe pagato, di cui stampa il comando manuale); binario assente = warn informativo, mai failure (il fallback è sempre valido).

Codex parity: **inline-fallback** (il companion è l'arricchimento condizionale del ramo Codex; il fallback same-model resta il percorso garantito; Claude-hosted byte-invariato). Nessuna nuova chiave `baldart.config.yml` (gating su `tools.enabled` esistente) → la schema-change propagation rule non si applica.

## [5.20.0] - 2026-07-10

**Codex-native `/new`+`/prd` — fascia P0: chiusura dei buchi invocabili.** Prima tranche del piano di allineamento Codex-native (report `baldart-codex-native-new-prd-alignment-plan.md`, validato in-repo): su un run Codex-hosted il gate di review per-card era materialmente morto (`Skill: codexreview` invocava un command Claude-only) e 5 gate risolvevano `codex-companion.mjs` da path plugin Claude inesistenti su Codex. Nessun kernel/engine in questa release (le fasce P1+ passano prima dalla refutazione avversariale).

### Added

- **Skill portabile `codexreview` (v1.0.0)** — la semantica canonica della deep multi-agent card review (Step -1 → 4, numerazione INVARIATA: il lean contract di `/new` la referenzia) spostata VERBATIM da `framework/.claude/commands/codexreview.md` nella nuova `framework/.claude/skills/codexreview/`. Aggiunte SOLO le superfici runtime: tabella "Runtime bindings" (spawn / decision gate / agent #4) che CITA `runtime-portability-protocol.md`, e il **binding Codex-hosted per l'agente #4**: nessun probe del companion, `code-reviewer` by name come finder ADVERSARIAL fresh-context (`source: "adversarial"`, `Method: codex=host-model→code-reviewer`), mai etichettato cross-model. Il command Claude diventa un **thin wrapper** (il modello `/baldart-push`) — un solo SSOT su disco.
- **`framework/scripts/task-store.mjs` + `agent-result.schema.json` + `task-store.test.mjs` (P1 ridimensionata)** — la fascia "contratti e task store" del report è passata per la refutazione avversariale ed è sopravvissuta a ~1/4: NIENTE `runtime-plan.mjs` (le capabilities per-runtime sono statiche nella binding table), NIENTE capabilities schema, NIENTE pipeline-contract YML (terzo SSOT della gate policy + i workflow Claude non possono leggere file), NIENTE migrazione del tracker `/new` (resta il recovery SSOT markdown con write-allowlist chirurgica). Sopravvive il **minimo meccanico**: un task store file-backed zero-dep (`init` / `claim` O_EXCL / `complete` con validazione dell'envelope risultato contro lo schema / `status --json` con deadline→`timed_out`) che mechanizza SOLO ciò che il modello-nel-loop fabbrica o stalla (claim atomico, validazione, aggregazione) — barrier/timeout/gate policy restano all'orchestratore + prosa. **Primo consumatore cablato nello stesso change**: la queue audit Codex di `/prd` (Steps 6.6–6.7 — la nota "no new script" è ritirata; la consolidazione semantica resta inline) + la riga state-spine di `runtime-portability-protocol.md`. Unit test in CI.
- **`scripts/check-runtime-asset-resolution.mjs`** — guard deterministico zero-dep in CI (`check-reference-integrity.yml`, insieme a `check-codex-parity.js` ora anch'esso in CI): (1) ogni `~/.claude/plugins` nel payload deve stare sotto un guard `Claude-hosted` entro 20 righe; (2) ogni `Skill: <name>` deve risolversi a una skill installabile; (3) ogni `.codex/agents/<name>.toml` citato deve avere il sorgente `framework/.claude/agents/<name>.md`; (4) `codexreview` resta skill portabile + command thin (<60 righe).
- **`framework/scripts/prd-state-check.mjs` (P3 ridimensionata — unico residuo sopravvissuto alla refutazione, col report durevole sotto)** — diagnostica READ-ONLY sullo state file `/prd` prima di un resume: enum chiusi, sezioni obbligatorie non duplicate, guard once-per-session, placeholder residui, worktree path coerente oltre discovery. Non scrive mai — lo state markdown resta il recovery SSOT; il refuted "PRD engine" (sidecar JSON, phase state machine, research track, verify/record-merge) NON è stato costruito. Cablata in `/prd` HARD RULE 5.

### Fixed

- **Report audit `/prd` durevoli** — gli output di plan audit 6.6d e discovery-completeness check si spostano da `/tmp` a `sessions/<slug>-audit/` accanto allo state file: l'evidence di audit sopravvive a riavvii e cleanup di `/tmp`.

### Changed

- **`/new` (skill v2.8.0)** — `codex-gate.md` Step C con binding per-runtime dell'invocazione (Claude: Skill tool, invariato; Codex: esegue la skill INLINE come root, agenti foglia by name, max_depth 1); i 3 siti companion (`setup.md` cross-card, `final-review.md` F.3, `team-mode.md`) marcati **Claude-hosted only** con ramo Codex-hosted operativo (spawn by name, stesso prompt/output file, `fresh-context adversarial (same-model)`).
- **`/prd` (skill v1.10.0)** — Discovery Completeness e Adversarial Plan Audit 6.6d: il runtime guard dichiarativo (S4–S6) diventa OPERATIVO (ramo Codex-hosted che spawna `plan-auditor` by name sullo stesso output file; 6.6d applica il dedup 6.6c). Binding Codex espliciti per `senior-researcher` (research, background/foreground) e `api-perf-cost-auditor` (Step 4.5).
- README + CLAUDE.md + `skills-mapping.md`: inventario a 43 skill, `codexreview` registrata.

Codex parity: **portable** (skill unica linkata su entrambi i runtime; rami Codex additivi, Claude byte-compatibile; command = wrapper Claude-only dichiarato). Nessuna nuova chiave `baldart.config.yml` → la schema-change propagation rule non si applica.

## [5.19.0] - 2026-07-10

**`/bug` v2.0.0 coordinatore a profili + bug registry + routine `bug-mine`.** La skill di debugging passa da protocollo monolitico (ogni bug pagava l'intera macchina a 5 fasi) a **coordinatore a lane** che scala il flusso sul profilo del bug, e il framework guadagna il **registro dei bug** — il knowledge loop che chiude risolvi→registra→ritrova→mina. Research-grounded (Agentless structure-beats-autonomy; SWT-Bench/AssertFlip BRT-first; cascate cost-aware con escalation solo su segnale verificabile; Abstain-and-Validate; Huang ICLR'24 no-self-judging; Agent KB +12pp da knowledge reuse).

### Changed

- **Skill `bug` v2.0.0** (`framework/.claude/skills/bug/`): Phase 0 = triage dominio (tassonomia chiusa: `visual|logic|structural|data|flaky|perf|build|security|not-a-bug`) × oracolo (`failing-test|screenshot|statistical|none`) → lane provvisoria `light|balanced|deep` (vocabolario Change Tiers/Rule C, citato non duplicato). **Tier Confirmation Checkpoint** post-evidenze = unico punto di downgrade (criteri meccanici) + uscite ABSTAIN e reroute `not-a-bug`; **escalation trigger** sempre armati, upgrade-only, su segnali verificabili (Tier-0 rosso, envelope breach, Heavy trigger, loop detection). Riproduzione a 3 rami (BRT-first AssertFlip / screenshot pre-fix / statistica flaky). **Fix per delega** (partizione `/e2e-review`: visual→`ui-expert`, funzionale→`coder`; mai inline per codice sostanziale) con **model routing per lane** (light → `model: sonnet`, pattern fix-pass v5.0.0, kill-switch overlay). **Verifica stratificata**: Tier 0 deterministico su ogni lane + check anti-deceptive-fix; giudice separato `code-reviewer` solo su deep (bounded, candidati 2-3 patch sul non-banale). Nuova reference `references/lanes.md`; description con trigger visivi + anti-trigger espliciti.
- **`skill-improver` v2.1.0**: nuovo `MODE: bug-miner` (aggregazione `root_cause × escaped_gates × recurrence_of` + divergenze di lane + cluster abstain + A/B `fix.model`; classi esistenti + `discard-as-one-off`; write surface invariata). Entry gemella in `framework/.claude/agents/CHANGELOG.md`.
- **Primitivo `AGENTS.md` v1.11.0**: nuovo MUST by-intent — ogni segnalazione di comportamento rotto/inatteso passa da `/bug` (mai fix diretto improvvisato); eccezioni esplicite (finding in-run `/new`/`new2` → review lanes; change request estetica → `/ui-design`/`ui-expert`). Trigger keyword-literale deliberatamente rifiutato.
- `framework/agents/skills-mapping.md` (sezione `/bug`, decision tree, chain) e README (tabella routine) aggiornati.

### Added

- **Bug registry per-progetto** `.baldart/bug-registry/BUG-<YYYYMMDD>-<slug>.yml` — append-only, consumer-owned (mai toccato da `update`), path fisso per convenzione (**nessuna nuova chiave `baldart.config.yml`** → schema-change propagation rule non applicabile). Schema con vocabolari chiusi in `framework/templates/bug-record.template.yml`; i campi generativi per il mining sono `escaped_gates`, `detection_channel`, `recurrence_of`, `lane_initial/final`. Design doc autorevole: [`framework/docs/BUG-REGISTRY.md`](framework/docs/BUG-REGISTRY.md) — include la decisione cross-progetto v1 (per-progetto + aggregazione maintainer; `push --registry` esplicitamente NON costruito in questa release).
- **Script zero-dep** (Node ≥18, propose-not-dispose, exit 0/1/2, `--json`, 12 test co-locati `bug-scripts.test.mjs`): `debug-residue-scan.mjs` (gate deterministico dei residui `// DEBUG:`/`# DEBUG:`/`[DEBUG:*]` sul diff — sostituisce il grep manuale hardcodato su `src/*.ts`; ignora i `.md` che citano i marker) e `bug-registry-check.mjs` (validate schema + `--find` lookup di recurrence per Phase 0).
- **Routine settimanale `bug-mine`** (`framework/routines/bug-mine.routine.yml`, lun 04:00 UTC, `agent: skill-improver`, `optional: true`, graceful su registry vuoto) — registrata in `routines/index.yml`; output `docs/reports/{{YYYYMMDD}}-bug-mine.md`, commit `[BUG-MINE]`. Il gemello bug-side di `finding-mine`.

Codex parity: **portable** (skill/moduli markdown + delega by-name + YAML + script zero-dep identici sui due runtime; la routine si deploya via i 3 backend adapter esistenti; il model override per-lane è best-effort dove il runtime lo supporta, come il fix-pass v5.0.0).

## [5.18.3] - 2026-07-10

**`AGENTS.md` primitive: la domanda-di-verifica è un trigger di grounding di prima classe.** Fix di salienza sullo skeleton spedito — nessuna policy nuova, nessun cambio di layout/comando.

### Changed

- **`framework/templates/primitives/AGENTS.md` (`primitive_version` 1.9.0 → 1.10.0).** Il mandato `codebase-architect` + `PROFILE` già copriva una "verification/question whose correct answer needs grounding", ma la clausola era sepolta come coda dell'esempio UI e leggeva come UI-only → una richiesta conversazionale di *confermare comportamento del codice* (es. "verifica che l'email ordine metta in To: tutte le mail fornitore + referente e in CC: gli utenti selezionati") veniva risolta con un excerpt di `Explore` invece di un grounding profilato, e un singolo excerpt può confermare una verità **parziale**. La regola ora dichiara esplicitamente che una verifica/domanda sul comportamento del codice è un trigger di prima classe — **non solo un cambio di codice** — quando la risposta corretta richiede tracing cross-file, con un esempio **non-UI** (i destinatari To/CC assemblati tra collector + selezione UI + mailer) accanto all'esempio UI esistente. Motivato da una sessione Sonnet-5 (consumer mayo) che ha instradato esattamente questa verifica a `Explore`, che `CLAUDE.md` vieta per il grounding.

Codex parity: **portable** (skeleton markdown condiviso, reso identico su entrambi i runtime al write-time).

## [5.18.2] - 2026-07-09

**Refresh modelli Codex GPT-5.6 (Sol/Terra/Luna) + backfill `version:` su tutti gli agenti + dropdown modelli in `/model-map`.** Manutenzione salvata via l'editor `/model-map` (v5.18.0) più due migliorie di tooling — nessun nuovo agente, nessun breaking change.

### Changed

- **Tier map `framework/runtime/codex-model-map.yml`** aggiornata ai modelli GPT-5.6 (lancio pubblico 2026-07-09, [OpenAI](https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna)): `deep` → `gpt-5.6-sol` (flagship), `standard` → `gpt-5.6-terra` (balanced), `fast` → `gpt-5.6-terra`, `spark` → `gpt-5.6-luna` (fast/cost-efficient). Gli agenti ereditano il nuovo modello Codex al prossimo transpile senza toccare i 33 sorgenti.
- **Declinazione per-runtime** ritoccata su diversi agenti via `/model-map` (Codex effort/tier/sandbox; `prd` → `opus` + Codex `deep`). Entry gemelle in `framework/.claude/agents/CHANGELOG.md`.

### Added

- **`version: 1.0.0`** sui 17 agenti che ne erano privi (contratto di authoring: ogni agente è versionato) — metadata-only, backwards-compatible. Ora tutti i 33 agenti dichiarano `version:`.
- **Dropdown modelli in `/model-map`**: il campo modello della tier-map è ora un `<select>` popolato da un catalogo `CODEX_MODELS` (5.6 inclusi) invece che testo libero — union col valore corrente per non perdere id custom. Catalogo in un unico punto (`scripts/model-map-server.mjs`), da aggiornare a ogni generazione di modelli.

Codex parity: **adapter-generated** (stesso `.md` SSOT; il tier risolve a un modello Codex diverso al prossimo transpile). Il tool `/model-map` e il suo server restano maintainer-only (non distribuiti).

## [5.18.1] - 2026-07-09

**`/model-map` tuning: `codebase-architect` Codex tier `standard` → `fast`.** Aggiustamento della declinazione per-runtime salvato tramite l'editor `/model-map` (v5.18.0) — nessun nuovo agente, nessuna funzionalità aggiunta, nessun breaking change.

### Changed

- **`codebase-architect`** (v2.0.3): `runtimes.codex.model` `standard` → `fast` (Codex risolve a `gpt-5.4-mini` invece di `gpt-5.4`); Claude declination invariata (sonnet/medium). Entry gemella in `framework/.claude/agents/CHANGELOG.md`.

Codex parity: **adapter-generated** (stesso `.md` SSOT, il tier risolve a un modello Codex diverso al prossimo transpile).

## [5.18.0] - 2026-07-09

**Declinazione per-runtime degli agenti (model / effort / sandbox per Claude E Codex) + tool interno `/model-map`.** Fino a v5.17 il transpiler Codex ometteva sempre `model` (inherit) e clampava l'effort — nessuna possibilità di dare a un agente un modello Codex diverso. Ora il singolo `.md` (SSOT invariato) dichiara la declinazione per OGNI runtime e l'installer la rende nel formato nativo di ciascuno.

### Added

- **Blocco frontmatter `runtimes.codex`** su tutti i 33 agenti: `{model: <tier|inherit>, effort: <minimal|low|medium|high|xhigh|inherit>, sandbox?: read-only|workspace-write}`. I top-level `model:`/`effort:` restano la declinazione CLAUDE (letti nativamente via symlink — byte-invariati per Claude). `inherit` è una scelta esplicita e legittima ("non pinno nulla"); un blocco MANCANTE è un agente dimenticato.
- **Tier astratti anti-churn** (`framework/runtime/codex-model-map.yml`): gli agenti dichiarano `fast|standard|deep|spark`, mai un nome OpenAI concreto (churnano ogni settimana); la mappa risolve tier→modello (snapshot: gpt-5.4-mini / gpt-5.4 / gpt-5.5 / gpt-5.3-codex-spark) e si aggiorna in UN punto per generazione di modelli. Override consumer per-tier: **nuova chiave `tools.codex.model_map`** in `baldart.config.yml` — **schema-change propagation applicata**: template + backfill in `configure` + detector nested in `update` + advisory doctor + questo CHANGELOG. Tier irrisolvibile → `model` omesso (inherit) con warning, mai un abort.
- **Transpiler esteso** (`src/utils/codex-agent-transpiler.js`): emette `model` (tier risolto), `model_reasoning_effort` (vocabolario Codex ufficiale `minimal|low|medium|high|xhigh` — la pagina subagents che elenca solo low/med/high è INCOMPLETA, fa fede la config-reference; il legacy `mapEffort` ora passa i livelli condivisi e mappa il `max` Claude → `xhigh`), e `sandbox_mode` (enforcement nativo del role-boundary read-only, prose-only su Claude). Slot-fill e overlay-merge restano PRIMA del transpile. Warning di validazione fluiscono nei merge warnings via `onWarning`.
- **Tool interno `/model-map`** (repo-root `.claude/skills/`, NON distribuito — modello `/new-audit`): web UI locale zero-dep (`scripts/model-map-server.mjs`, porta 4680) con griglia agente×runtime (model/effort/sandbox, bulk-set per colonna) + editor della mappa tier. Il salvataggio fa edit CHIRURGICI dei frontmatter (mai re-dump YAML), bumpa `version:` patch, appende l'agents-CHANGELOG e ri-esegue il validator in pagina.
- **Guard maintainer `scripts/check-agent-runtimes.mjs`**: FAIL se un agente non dichiara il blocco completo, usa valori fuori vocabolario, un nome Codex concreto, un tier irrisolvibile, o `model: haiku` (ban fleet-wide). Gemello di `check-new-parity.mjs`.
- **Regola di authoring BINDING** (CLAUDE.md + REGISTRY.md § Notes + CODEX-AGENTS.md): alla creazione/revisione sostanziale di un agente si CHIEDE all'utente model+effort per ogni runtime — mai dedotti in silenzio.
- **Doc di riferimento runtime** [`docs/references/agent-runtime-frontmatter.md`](docs/references/agent-runtime-frontmatter.md): snapshot verificato dei frontmatter agent Claude vs TOML Codex (campi, vocabolari, delta per il transpiler) + metodo (re-fetch della doc ufficiale prima di estendere).

### Changed

- **`baldart doctor`**: l'advisory Codex-agents ora segnala anche i `.toml` **stale** (generati pre-v5.18: sorgente con `runtimes.codex` ma TOML senza `model`) → fix `baldart update`.
- **33 agenti**: frontmatter aggiornati con la declinazione Codex scelta via `/model-map` (distribuzione: deep/gpt-5.5 per implementer+reviewer pesanti, fast/gpt-5.4-mini per doc/wiki/i18n, standard/gpt-5.4, inherit esplicito per la coda specialist) — version bump + entry per-agente in `framework/.claude/agents/CHANGELOG.md`.

Codex parity: **adapter-generated** (un solo `.md` SSOT; rendering per-runtime — Claude legge i top-level nativamente, Codex riceve il TOML transpilato con la declinazione dedicata). Consumer Codex esistenti si allineano col prossimo `npx baldart update`.

## [5.17.0] - 2026-07-09

**Wave di prevenzione dal post-mortem `/new FEAT-0071 -full -auto` (mayo, sessione `ec0e6b98`, 83.73M).** Un run `-auto` su un epic di 6 card CSS-token banali, delegato a `new2`, ha prodotto tre sintomi con un'unica causa di fondo: **il ledger di `new2` si fidava del RITORNO dell'agente invece che di `git HEAD`**, e le sue macchine di grounding/compressione giravano nel punto o col predicato sbagliato. Analisi via `/new-audit --deep` (5 analisti + verifica avversariale — che ha refutato la diagnosi "return droppato": il crash era un self-report `epic:true` intatto). Report: [`docs/post-mortem/FEAT-0071-new2-ec0e6b98.md`](docs/post-mortem/FEAT-0071-new2-ec0e6b98.md).

### Fixed

- **P0 — false-skip di una card committata (`new2.js`).** Card FEAT-0071-02 era stata implementata e committata (`5f0d78385`, 8 file) da un ui-expert di fallback (architect unavailable), ma il ledger l'ha registrata `epic-skipped`/`reviewers=[]` perché l'owner ha auto-riportato `epic:true` e `new2.js:709` si fidava ciecamente del flag, scavalcando il pre-flight deterministico (`node.isEpic=false`) e il commit reale in HEAD → merge NON revisionato + report "non implementata" + turni forensici. Due fix: **(P0-B)** `impl.epic` onorato solo se l'owner non ha prodotto `scopeFiles`; con lavoro presente → fall-through al path normale (review girano, commit adottato) + riga ledger `EPIC-MISREPORT`. **(P0-A)** nuovo `reconcileStatusAgainstHead()` dopo lo scheduler: un batch-level read-only spawn riconcilia lo status delle sole card sospette (`epic-skipped` non-pre-flight/`followup`/`pending`/`failed`, esclusi gli epic veri) contro `TRUNK..HEAD` — un commit che referenzia la card ⇒ override a `committed` + flag `reviewersUnknown`. Rende strutturalmente impossibile la classe "codice in HEAD registrato skipped".
- **P1-A — merge senza copertura reviewer (`new2.js`).** La review di `new2` è per-card-ledger-based ma il merge è branch-based: una card committata-ma-non-instradata raggiungeva il trunk non revisionata senza che alcun gate lo asserisse. Nuova asserzione al merge-integrity gate: una card committata con `reviewersRun=[]` (o git-reconciled) e `review_profile != skip` blocca il merge quando la final-review non è girata (residuo HIGH `unreviewed-committed`), altrimenti è coperta dalla final-review sull'intero diff.
- **P1-B — ordering bug del reconcile (`new2.js`).** `reconcileLedgerAgainstHead()` + dedup girano ora PRIMA di `phase('Merge')` (calcolo `closureBlockedEpics`), non a fine flusso: un residuo ritrattabile spariva dal blocklist solo DOPO che il merge-agent aveva già lasciato l'epic aperto, garantendo un secondo `/new` per chiuderlo.
- **P1-C — reconcile ora cattura anche i residui INFONDATI (`new2.js`).** Il contratto passa da "già-soddisfatto in HEAD" a "già-soddisfatto OR mai-esistito" (ispeziona il `file:line` citato; se il codice era già corretto, il residuo era un falso positivo), e l'agent passa da `general-purpose/sonnet` a `codebase-architect` per il ragionamento semantico (quale barra monta una pagina, quale token è canonico) — proprio il grounding che nel run reale il secondo `/new` fece a mano. Chiude il flood di follow-up phantom (2 su 3 erano infondati).

### Changed

- **Tier-guard: la larghezza NON è un trigger Heavy (`AGENTS.md` primitive 1.9.0 + `prd-card-writer.md` + `card-schema.md`).** File/module/card-count non rendono Heavy una modifica: un refactor **omogeneo** (un dominio, es. tutto CSS/style/token) con zero trigger Heavy reali è **Light a prescindere da quante card/file tocca** — "wide or cross-module refactor" significa ora una modifica ETEROGENEA/logic-bearing. `prd-card-writer` deriva il tier dal card-set su invocazione diretta (floor self-contained) + merge-check "type + suo unico consumatore = una card" + regola "mechanical sweep → epic-closer, mai peer in conflitto". `card-schema` § completeness: un file-set eterogeneo deve enumerare ogni sotto-caso in un AC con `WHERE`-branch (il sotto-caso barless non enumerato aveva forzato un token inventato a runtime + una card di review). Questa è la leva macro sul costo anomalo (il batch tutto-CSS classificato Heavy aveva instradato lavoro banale nella pipeline pesante due volte).

### Deferred (documentati, non silenziosi — vedi il post-mortem §8)

- `/new`-classico de-escalation diff-based + collasso Final su batch multi-card light: rischio di buco di copertura doc/review (il light multi-card deferisce la doc coverage alla Final) — richiede design coverage-gated proprio; non-bloccante perché il Tier-guard è già la leva macro. `new2.js` grounding gate su `resolve(blocker)`: il false-blocker misurato (5.80M) era nel `/new` classico, non in new2, e P1-C aggiunge già il grounding founded-ness.

**Codex parity: adapter-generated / N·A.** I fix `new2.js` sono Claude-only (il workflow è Claude-only) e correggono un meccanismo del ledger che il path `/new` classico Codex non possiede → nessuna degradazione Codex. I fix prose (`AGENTS.md`, `prd-card-writer.md`, `card-schema.md`) sono SSOT condivisa runtime-agnostica, portabile su entrambi (l'agente si transpila in `.codex/agents/*.toml`). Nessuna nuova chiave `baldart.config.yml` → la schema-change propagation rule NON si applica.

## [5.16.0] - 2026-07-08

### Changed

- **`/prd`, `/new`, `/new2` non pinnano più l'effort — ereditano l'orchestratore.** Rimossa la chiave `effort:` dal frontmatter delle tre skill (`prd` 1.8.0→1.9.0 era `high`, `new` 2.6.0→2.7.0 era `medium`, `new2` 1.10.0→1.11.0 era `high`). Senza baseline in frontmatter, il livello di reasoning **e il modello** sono ereditati dal livello di sessione/orchestratore secondo la precedenza documentata in `framework/agents/effort-protocol.md` (`effort: frontmatter > /effort session > model default`). Il blocco `## Effort` di ciascuna skill è riscritto per documentare l'ereditarietà; l'inline override `effort=<low|medium|high|xhigh|max>` resta onorato come escalation per-run (in `new2` continua a essere forwardato al workflow via `args.flags.effort`). `skill-creator`/`quick_validate.py` trattano `effort` come opzionale (warn non-fatale), quindi la rimozione è conforme. Codex parity: portable (skill = stesso `SKILL.md`; effort è nativo del runtime, nessuna chiave `baldart.config.yml` → la schema-change propagation rule NON si applica).

## [5.15.0] - 2026-07-08

### Added

- **`/ui-design live` — live in-browser iteration mode (ui-design v2.1.0)**: `/ui-design` gains **Step F-Live**, an optional high-bandwidth alternative to chat-driven iteration and a standalone entry point for restyling a running app. The user selects an element directly in the browser, picks a design action (Bolder/Quieter/Colorize/Layout/Typeset/…/Freeform) or steers by text/voice; the skill generates 3 variants hot-swapped in place via the dev server's HMR, and Accept persists the chosen variant into source (with carbonize cleanup, generated-file fallback to true source, and a durable journal for recovery). Implementation **vendors the `live` subsystem of [impeccable](https://github.com/pbakaus/impeccable) (Apache-2.0)** — helper server, browser overlay, poll loop, session store, SvelteKit adapter, CSP detector — verbatim-as-possible under `ui-design/scripts/live-mode/` (NOTICE.md records provenance + re-sync policy) plus its 11 action references under `references/live-actions/`. The BALDART adaptation is concentrated in the new `references/live-iteration.md`: registry-first cascade BLOCKING before the first `generate`, token-first variants, accepted variants flow through Step H Post-Intervention Coherence; upstream `DESIGN.md`/`PRODUCT.md` sidecars map to `tokens-reference.md`+`${paths.ui_guidelines}` / `identity.design_philosophy`; no dedicated manual-edit-applier agent (inline apply, same JSON contract). Second supported target: the skill's own `option-*.html` mockups — and (ui-design v2.2.0) this is the **DEFAULT Step E presentation channel**: the 3 options open already live-instrumented (mockup dir = isolated live project root, `DESIGN.md`/`PRODUCT.md` sidecars generated from the Design Read, poll loop started immediately), so iteration is in-browser from the first second; plain static `open` is the announced fallback only. Step G gains a pre-persist guard (exit live + no leftover `impeccable-*` markers). Runtime state stays in the upstream-compatible `.impeccable/live/` at the consumer root. No new `baldart.config.yml` key → the schema-change propagation rule does NOT apply. Codex parity: portable (zero-dep Node scripts identical on both runtimes; poll loop foreground on Codex per the harness policy, background task on Claude; documented in the skill's runtime-portability section).

## [5.14.0] - 2026-07-08

**Codex `/new` audit wave — five fixes distilled from a real Codex-hosted `/new` run (mayo FEAT-0070, two sessions, 35.6M + 19.0M tokens), each grounded in the framework's actual state and two of them re-designed by an adversarial-refutation pass before any code was written.**
The audit (Codex reading its own rollout transcripts) surfaced eight findings + seven Codex-specific asks. Validation against the framework showed most were **already covered** (doc-invariants, the SYNC prose gate, worktree branch-lock, the effort→`model_reasoning_effort` transpiler map, AGENTS.md under the 32KB Codex cap) — the through-line of the *real* gaps was that **the Codex orchestrator does not reliably execute prose gates that already exist**, so the leverage is deterministic enforcement, not new prose. Five ship:

- **C6 — Codex-aware `new-audit` resolver (maintainer tooling).** `scripts/new-session-audit.mjs` gained a dedicated Codex branch: a `/new` run driven by Codex leaves no `~/.claude` transcript (it writes `~/.codex/sessions/**/rollout-*-<uuid>.jsonl`), so a Codex UUID / `codex-last` now falls through to a Codex auditor that computes deterministic telemetry from the event stream (tool histogram incl. `spawn_agent`, agent-type breakdown, running `token_count` total, compactions, `turn_aborted`) + Codex-native detectors (`agent-teardown` = `spawn_agent` > `close_agent`; `sync-gate` = hard `SYNC-*` present; `resume` = aborts). This resolver **independently reproduced the audit's own numbers** (396 tool calls, 29 spawns, 35.63M / 19.0M tokens, 2 aborts) — and its `agent-teardown`/`sync-gate` detectors fired on the real run, confirming F1 (4 leaked agent threads) and F7 (unresolved SYNC markers) with hard data.

- **C5 / F7 — deterministic final SYNC / terminality gate.** New `framework/.claude/skills/new/scripts/final-sync-gate.mjs` (+ `.test.mjs`, zero-dep twin of `doc-invariants.mjs`): reads the batch tracker and reports unresolved `[SYNC-NEEDS-DECISION]`/`[SYNC-BLOCKER]` markers + any card still parked under `## Current Card`. `final-review.md` step 4 runs it before the verdict — exit `3` forces the floor to ⚠️ PARTIAL and names the residuals; it never hangs (PARTIAL-not-block, coherent with merge-cleanup's AUTONOMOUS rule) and SKIPs on a missing tracker. **The proposed native Codex `Stop` hook was adversarially REFUTED** (inert until user-trusted → zero guardrail in CI; fail-closed vs the PARTIAL rule the prose deliberately chose; a hook can't know which per-batch tracker `$(git-common-dir)/baldart/run/batch-tracker-<FIRST-CARD>.md` to read) — the surviving design is this prose-invoked script. `new2.js` needs no change (its G19–G23 merge gate map is already deterministic code). `new` skill → 2.6.0.

- **F5 — nav-reachability gate (opt-in, deterministic).** e2e-review Phase 3 emits a `nav reachability` scenario for each card `entrypoints.ui[].nav: true`: a plain DOM assertion that a primary-nav `<a href>` points at the route (desktop + mobile), closing the "implemented but unreachable" gap. Missing desktop link → `nav-unreachable` (major, gating). **The original design was cut ~⅔ by an adversarial pass**: route-HTTP-status was dropped (already in `render_health`), and **entitlement / licence locked-state was refused from the base** (a domain concept BALDART must not hard-code → an `.baldart/overlays/e2e-review.md` reach-state recipe or an `ac_verification` oracle). Scoped strictly to `nav: true` (absent = not checked → no false-fail on deep-links / admin / sub-routes). New opt-in field `entrypoints.ui[].nav`; e2e-review → 1.4.0, prd-card-writer → 1.4.0.

- **F4 — behavioural-AC oracle rule.** When an AC asserts a *behaviour* (perf property, N+1 absence, per-record/per-config/per-tenant resolution, call-count bound), a static presence-`grep` is no longer an acceptable primary oracle — it proves code exists, not that it behaves (a per-config email lookup can grep-green while resolving globally). The oracle must EXERCISE the behaviour (mock call-count, multi-record/multi-config/multi-tenant fixture, measured bound). SSOT in `card-schema.md` § ac_verification; authoring cite in `prd-card-writer`.

- **C2 — effort differentiation on adversarial review agents.** `security-reviewer` (→ v1.1.0) and `api-perf-cost-auditor` (→ v1.0.0, first tracked version) now declare `effort: high` — both are intrinsically deep-reasoning (threat-model/attack-path; N+1/fanout/index modelling), mapped to `model_reasoning_effort = high` on Codex by the transpiler (verified). **Restraint:** `doc-reviewer` / `codebase-architect` were deliberately NOT lowered — doc-reviewer also authors RAG docs and architect is PROFILE-driven; differentiating them needs telemetry first.

**Codex parity: portable** — every artifact is a zero-dep Node script or runtime-agnostic prose/DOM assertion; the audit resolver, the sync gate, the nav check and the behavioural-oracle rule all run identically on Claude and Codex, and the effort headers are transpiled. No new `baldart.config.yml` key (F5's field is a card sub-field; nothing gates on config) → the schema-change propagation rule does not apply.

## [5.13.0] - 2026-07-08

**Per-worktree service teardown — a merged worktree no longer orphans its ephemeral services (docker / supabase-local / testcontainers).**
A user reported that after every `/new` batch (and Codex-hosted runs) the OrbStack instances a worktree spun up — networks named after the branch, e.g. `supabase_network_mayo-feat-feat-0069-00-…`, `…-chore-0052-…`, `…-feat-0070-…` (the last created by Codex) — stayed **running for the rest of the session**, one leaked stack per batch. Root cause: BALDART has a rich worktree *setup* side (`setup-worktree.sh` + `stack.env_files` bring gitignored state INTO a fresh worktree) but **no symmetric teardown** for the services a worktree starts. `merge-worktree.sh` removed the worktree dir, and any container scoped to that branch orphaned. The existing Phase 6c teardown only ended the **Codex broker** + idle subagents — never project services (correctly: core is stack-agnostic and never names supabase/docker/orbstack).

The fix is the **teardown twin of `env_files`**, single-sourced in the one deterministic merge script all three runtimes share, so `/new`, `new2` AND `/mw` are covered from ONE place with zero per-caller duplication:

- **New config key `stack.worktree_teardown_command`** (string, default `""` → silent no-op) — a project-owned command run to stop a worktree's ephemeral services. **Stack-agnostic**: core never names docker/supabase/orbstack; the command (or a committed script path) lives in the project.
- **The hook lives in `merge-worktree.sh` (step 7a)** — the SSOT deterministic merge script consumed identically by `/new` (Phase 6), `new2` (Merge phase, via its merge subagent) and `/mw`. It runs the command **in the worktree root, immediately BEFORE `git worktree remove`** (cwd still valid, so `supabase stop` / `docker compose down` find their config), under a **120s timeout**, **NON-BLOCKING** (a failure/timeout WARNs; the merge already landed and is never undone), and **only on the successful merge path** (a worktree preserved after a build fail / code conflict keeps its services for debugging). Env vars `BALDART_WORKTREE` / `BALDART_BRANCH` / `BALDART_TRUNK` / `BALDART_MAIN` are exported (use `$BALDART_BRANCH` to scope container/project names). New manifest field `services_teardown: <ran|none|failed>`; surfaced in the `/new` Phase 6c report.
- **`/mw` inline fallback mirror** (`worktree-manager/SKILL.md` step 6) — the human-readable SSOT the script mirrors gains the same pre-removal teardown, so the script-absent path (old subtree) stays in sync.
- **Activation on FIRST install**: `configure` (which `baldart add` runs) prompts for the key when it detects a **per-worktree local service stack** — `supabase/config.toml` → default `supabase stop`, a `docker-compose.yml`/`compose.yaml` → default `docker compose down` — OR when parallel worktrees already exist. Keyed off files present at configure time, NOT off worktrees existing right then (a first install has none — that would have skipped the prompt every time). No service stack detected → not asked; the key stays `""` (silent no-op), so CI / non-service projects are never nagged.
- **Full schema-change propagation**: template (`baldart.config.template.yml`) + `configure` (the detection-gated prompt above) + `update` detector (auto-covered by the existing `stack.*` scalar diff; comment updated — a pre-5.13 consumer is told "new config key" on `baldart update` and re-runs `configure`) + `PROJECT-CONFIGURATION.md`. No `doctor` diagnostic (opt-in, empty default is valid).

**Codex parity: portable** — the hook is a shared deterministic bash script with no runtime-specific code; all three callers (`/new`, `new2`, `/mw`, including the Codex fallback) route through it, so a `tools.enabled: [codex]` consumer gets the identical teardown. Scoped strictly to the successful-merge path — the early-exit path deliberately preserves a worktree and its services intact for debugging. **No `doctor` diagnostic** (opt-in, empty default valid). `worktree-manager` skill → 1.3.0.

## [5.12.0] - 2026-07-08

**FEAT-0069 post-mortem wave — ten adversarially-verified fixes from the deep audit of two real `/new` runs + their `/prd` origin (427M tokens, ~21–27% measured waste) and a live DS-drift incident.**
A `/new-audit --deep` pass over the FEAT-0069 "gestione licenze" feature on the `mayo` consumer (runs `ba1e674e` 148.45M + `d84c6e77` 238.16M + `/prd` `d192df73` 40.65M) produced 42 findings → 14 confirmed after adversarial verification, plus a user-reported incident: `baldart ds-gate` returned green while `AppShell.md`/`Sidebar.md` were stale vs their sources. The resulting improvement plan was then itself adversarially reviewed (11 skeptics, one per item): 1 item REFUTED integrally (SECTION= dispatch of the `/new` modules — the modules are already phase-lazy-loaded and module text is not the bloat driver), several sub-items dropped as redundant-with-existing or mis-targeted, and every surviving item reshaped. What shipped:

- **Destructive-migration content gate (A1/B1 — the prod-incident class).** `production-readiness.md` gains a "Destructive-migration content gate" (qualified candidate extraction from pending migrations; `git grep -I` on BOTH `origin/${git.protected_branch}` and `origin/${git.trunk_branch}` with exclusion pathspecs; PROPOSES/DISPOSES — evidence + deferral/user-choice, never a hard block on grep counts; fetch-first fail-open; shared-DB-project probe from `stack.env_files` → destructive deploy human-only when the project-ref is shared) + a "Migration sync verification" subsection (parse the per-migration local-vs-remote list, NEVER a wrapper's exit code — A14). Card side: `db_migration {required, expected_file, remote_push, class: additive|destructive|contract}` promoted from consumer-overlay vocabulary into `card-schema.md` (— / C / C), with destructive/contract requiring a machine-checkable EARS precondition AC (`ac_verification` oracle, non-manual); `prd-card-writer` authors it, `plan-auditor` checks it, `validate-card-baseline.js` WARNs (fixtures updated). Cited (not restated) by `/new` setup Phase 0 1b and `new2` Step 3.5/Phase 7 (blocked deploys → existing `schemaDeploysDeferred`). Codex parity: portable.
- **`ds-gate` deterministic staleness check (A19 — the reported incident).** `src/utils/ds-reuse-gate.js` gains a change-scoped, sha-first `DS_COMPONENT_STALE` layer: `git hash-object(source)` vs the spec HEAD's `source_sha`; on mismatch the extractor classifies (identical fields → warn `sha-only-resync`; divergent `props`/`variant_prop`/`variants`/`composes` → **block**, naming component/spec/fields; extractor unavailable → warn; source missing → block; pre-stale at baseline → warn). `--all` opt-in sweep (warn-only) wired into the `ds-drift` routine as its first deterministic pass. Extractor resolved consumer-first (`.framework/…/extract-one.mjs`, FEAT-0042 skew guard) then bundle-relative. Prose callers (qa-sentinel 2.1.0, code-reviewer 2.1.0, ui-expert 2.2.0, both review workflows, implement.md, design-system-init 1.1.0, ds-new 1.2.0) now say "closed-set violation OR stale component spec". 12 new unit tests incl. an end-to-end incident replay (the exact AppShell case now blocks). Codex parity: portable (Node CLI, zero-dep extractor fallback).
- **Coder turn-economy (A7).** `agent-operating-protocol.md` SECTION=tool-budget item 5: batch all independent edits of ONE sub-task in one turn (independence test included; Claude = multiple Edit/Write per message, Codex = one `apply_patch`), new file = one complete Write, no narrate-only turns (text-only legitimate only for final report + blocking escalation). 1-line BINDING citation in `coder` (2.1.0) + reminder in the `/new` coder briefing. Codex parity: portable (prompt-level).
- **Verify-before-ask + recorded decisions (A2).** `/new` SKILL.md BINDING: a decision-grade `AskUserQuestion` resting on a factual on-disk claim must verify it first (context evidence or ONE cheap Read, never a spawn) and cite `path:line`; non-verifiable → `unverified:` label. Both review workflows + `new2` now receive `decisions[]` (derived at spawn time from tracker `## Issues & Flags` prefixes, 1-line each, wave-scoped) injected into finder/fix briefs: a fix reversing a recorded decision is never applied — it returns as `NEEDS_MANUAL_CONFIRMATION` citing the decision (never silently dropped). Codex parity: inline-fallback (prose briefs = Codex SSOT).
- **Fix-lane hardening (A9/A16/A17).** FIX_SCHEMA in all THREE siblings (`new-card-review.js`, `new-final-review.js`, `new2-resolve.js`) gains optional `escalated[]` (fixer-declared `needs-card-context` — wiring beyond a finding's evidence is never half-applied; escalated ids stay in `unresolved` and route to a follow-up card, never a blind respawn) and `testEdits[]` (every test-file edit declared with why the ORIGINAL expectation was wrong). A qa-sentinel adjudication judge reconciles declared vs actual test-file diffs after every fix pass — adverse or undeclared edits raise a synthetic `test-expectation-integrity` FAIL gate row (the GATE_DISABLED_BY_FIX/DoD-erosion class, 62 signals in the audited run). The bundle-handoff sub-item was REFUTED (fixBrief already carries a stronger anti-re-read contract). Codex parity: inline-fallback (prose SSOT updated in the same change).
- **Value-absent lane + shared-surface non-regression ACs (B3/A5 — the nav-lock regression class).** `card-schema.md` authoring invariants + `prd-card-writer` (1.3.0) + `plan-auditor` (2.1.0): a card gating behavior on a NEW data field/flag must carry an EARS AC for the value-absent lane (pre-auth/demo/legacy rows; else-branch for atomic backfill → `assumptions`); MODIFY on a shared shell/nav/layout surface → `test_plan` lists the existing suites exercising it + a "SHALL CONTINUE TO" AC with a non-manual oracle (`manual:` fallback when not executable in-worktree). Integration/E2E oracles must declare runtime prerequisites — missing prerequisite = visible ENV-SKIPPED (≠PASS, ≠FAIL), never a mocked pass (B7). The wave-gate Playwright sub-item was REFUTED (wrong SSOT — belongs to `/e2e-review` if ever). Codex parity: portable.
- **Launch-plan decision record (B2).** `backlog-phase.md`: BOTH partition branches persist a durable `## Launch Plan / decision:` record (single-run vs waves: N) in the session state; `validation-phase.md` numbered pre-commit CHECK (decision exists + epic carries ≥2 distinct `wave_name` when waves) + render-side persistence-failure guard; `/new` setup pre-flight 1-line non-blocking WARN when the invoked card-set straddles waves (silent no-op when keys absent); `new2` gate-map log-and-proceed. Codex parity: portable.
- **Merge/release deterministic subset off-context (A3/A4/A13 — 16.6M→40.8M merge-tail scaling).** New Claude workflow `new-merge-release.js` hosting ONLY the deterministic subset (merge-script relay, optional 1-line CI verdict, Phase 6b reconciliation, Phase 7 DETECTION) returning compact `{mergeStatus, checklist, needsDecision[]}` — 6c gates, deploys and user text stay in the skill; delegation gate at the head of `merge-cleanup.md` Phase 6 with byte-conserved inline prose (= Codex fallback/SSOT); checklist persisted to the tracker `## Final Report` (v5.11.2 persist-then-emit invariant). New `skills/new/scripts/deploy-status.sh` (stack-dispatched 1-line probe; Vercel identity-gated on `.vercel/project.json` before any listing). `/new` SKILL.md context-economy PRINCIPLE (with the worktree-manager inline-fallback carve-out). The "retire auto-deploy" behavior change is deliberately NOT shipped — pending explicit user sign-off. Codex parity: inline-fallback.
- **Final doc-reviewer delta-scope (A8 — 23.45M triple doc coverage).** Optional `postDocPassFiles`/`waveDocResiduals` args on `new-final-review.js` (absent → byte-identical full scope; `new2` untouched): when provided, the final doc audit covers cross-card invariant surfaces + wave residuals + docs joined to CODE changed after each wave's doc pass (tracker-derived, never git SHAs). Codex parity: portable (rule lives in the shared prose).
- **Hygiene.** Follow-up cards authored standalone under a PRD epic now set `group.parent` to the epic's `-00` id so closure greps see them (B6-redirect, prd-card-writer); background-task idle notifications are never by themselves a reason for an orchestrator turn (A12-half); coder briefings inline the RESOLVED `toolchain.commands.*` strings (A18-g4); 5 new `check-new-parity.mjs` tripwires (`needs-card-context`, `testEdits`, `Recorded user decisions`, `new-merge-release`, `postDocPassFiles`) — 26 policies green.

Dropped with recorded rationale: SECTION= dispatch of `/new` modules + e2e off-context twin + tracker digests (P7 — refuted on telemetry + existing rules); rg-first duplication, worktree pre-commit hook guard, epic children/total_cards mutation, token-pipeline duplication, serializer rule duplication, batch-ceiling threshold (already existing / refuted); fix-in-cluster + pooled security finder deferred pending telemetry. **No new `baldart.config.yml` key anywhere** → the schema-change propagation rule does NOT apply. Skill bumps: `new` 2.5.0, `new2` 1.10.0, `prd` 1.8.0 (+ per-skill CHANGELOGs); agents-CHANGELOG entries for all 6 bumped agents. Parity guards (`check-new-parity.mjs`, `check-codex-parity.js`) and the full test suites green.

## [5.11.2] - 2026-07-07

**`/new` final report — including the v5.8.0 wave-aware launch plan — is now the SINGLE terminal message on BOTH runtimes; fixes the "buried wave summary" on Codex.**
A user reported that despite the v5.8.0 wave-launch-plan update, a **Codex-hosted** `/new` run ended not on the mandated `## Prossimi Passi` / `### Prossime wave da sviluppare` section but on an ad-hoc operational status (`db:check-sync OK`, `registry:sync OK`, `Repo allineate`, `card 01-05 DONE`). Root cause, confirmed after adversarially refuting the obvious diagnoses: the `/prd` wave-tag persistence is engine-agnostic (backlog-phase.md — fine) and `new2`'s wave line is Claude-only (not the Codex path), so the defect is entirely in `/new`'s **terminal reporting structure**. The wave-aware Next Steps lives only in Final Review **F.6 step 4b**, which the phase order places **before** Phase 6 (merge) / Phase 7 (production readiness) / Phase 8 (metrics). On **Claude** the near-silent rule suppresses that trailing Phase 6/7/8 narration, so the mid-run report survives *by accident* as the last block. On **Codex** the runtime does **not** honor that suppression (it narrates `db:check-sync` / `registry:sync` / `git status`), so the report gets buried and the run closes on operational noise + a self-composed summary. F.6 4b also had **no `Codex:` branch**, against the runtime-portability discipline.

The fix makes the Final Report **engine-independent** by turning it into a genuine terminal emission (no reliance on suppression):
- **Compose + persist, don't emit mid-run** — Final Review F.6 (steps 4/4b, `final-review.md`) now composes the verdict + per-card result block + actionable residue + **wave-aware Next Steps** and writes them to the tracker under a `## Final Report` section instead of emitting them in place; Phase 7 (`production-readiness.md`) appends its checklist to the same block instead of rendering standalone.
- **New terminal step** (`metrics.md` § "Terminal — Final Report emission") — runs LAST, after every Phase 8 step, and **ALWAYS** (even if the non-blocking metrics write SKIPPED/failed): it reads `## Final Report` from the tracker and emits it verbatim as the ONE terminal user-facing message, with a fallback that re-derives the wave summary if the tracker block is missing (legacy / compaction) rather than closing on a bare status line.
- **Explicit `Codex:` reinforcement** on F.6 4b + the terminal step: on Codex the run MUST still END on the `## Final Report` (wave-aware Next Steps included) regardless of any preceding operational narration.
- **New binding-table row** in `framework/agents/runtime-portability-protocol.md` § "Terminal report emission" (the SSOT the three skill modules now cite) — Claude relies on near-silent suppression; Codex must compose+persist+emit-terminally. SKILL.md § near-silent case 3 updated to name the terminal anchor.

**Codex parity: portable** (prose-only; the persist-then-emit-terminally design is what makes it engine-independent — it is the fix's whole point). `new2` is unaffected (Claude-only; its report is the workflow return, terminal by construction). **No new `baldart.config.yml` key** → the schema-change propagation rule does NOT apply. `/new` skill version → 2.4.2. Parity guard (`scripts/check-new-parity.mjs`) green.

## [5.11.1] - 2026-07-07

**`/new` now tears down its idle background subagents at end of run — closes the "Background work is running" leak on terminal exit.**
A user reported that after every interactive `/new` run, closing the terminal warns `Background work is running` and lists the batch's teammates (`## AUTONOMOUS CARD IMPLEMENTATION — FEAT-00NN`, the `i18n-translator` fill, the doc closer) as if they were still executing. They are not — they **finished**. Root cause: `/new` spawns its coder + closer subagents with `run_in_background: true`, and a background subagent that completes its work does **NOT** terminate — it goes idle ("rests") so the orchestrator could `SendMessage`/resume it (deliberate: the empty-result gate reads their transcripts from disk, and the framework explicitly forbids messaging a rested teammate). Phase 6c's teardown ended the **Codex broker** (a Bash process) + reaped orphan MCP servers, but there was **no `TaskStop` anywhere in the framework** — so the idle teammates stayed alive until Claude Code's `SessionEnd` (whenever the user closes the terminal). Harmless to the committed work (already merged), but noisy and wasteful of session resources.

- **New Phase 6c step 5c** (`framework/.claude/skills/new/references/merge-cleanup.md`) — after the merge and after every teammate's transcript has been consumed, `TaskList` → `TaskStop` the batch's idle teammates, scoped to THIS run only: matched by the `coder-<CARD-ID>` name prefix OR any card ID in the tracker `## Card Queue` (so the match survives a compaction). NEVER stops a task outside the batch signature (the user's unrelated background work in the same session). Best-effort + non-blocking + one log line (`Background subagents teardown: <stopped N | none alive | N/A (Codex) | skipped>`).
- **Mirror on early exit** (`framework/.claude/skills/new/SKILL.md` § "Terminal hygiene") — the same `TaskList`→`TaskStop` runs when the batch ends before Phase 6c (unrecoverable HALT / killed final-review workflow), exactly as the Codex-broker early teardown already did.
- **Pointer** in `references/team-mode.md` — the parallel `coder-<CARD-ID>` fan-out is the widest source of the leak; noted that Phase 6c 5c tears them down after merge.
- **`/new2` is unaffected** — its agents run inside the dynamic-workflow runtime and die with the workflow; only the classic `/new` (Task-tool background subagents) leaked.
- **Codex parity: N/A** — no observable Agent/Task tool on Codex, team-mode is forced sequential there, and Codex named-agent spawns do not persist as idle background tasks (`framework/agents/runtime-portability-protocol.md`). **No new `baldart.config.yml` key** → the schema-change propagation rule does NOT apply. `/new` skill version → 2.4.1.

## [5.11.0] - 2026-07-06

**Seamless framework-drift reconciliation in `baldart update` — a trivially-drifted shared framework file no longer dead-ends the whole update.**
A `mayo` consumer stuck at v5.7.0→v5.9.0 surfaced the failure: five ancient `.framework/` commits (an add-then-revert of the v4.77.0 extractor, a superseded responsive-prose delete, a FEAT-0054-03 gating feature) were all flagged `custom-other` → class `mixed`/`real-custom` → **hard block**. Root cause, proven by diffing HEAD vs upstream file-by-file: of the **14** framework files those commits touched, **13 were already byte-identical to upstream** — only ONE, `design-system-protocol.md`, still diverged, and its *entire* divergence was **2 cosmetic characters** (a `+ ` vs `- ` markdown continuation bullet, a merge artifact). Because `isAbsorbedAgainstUpstream` is **all-or-nothing per commit** (a commit demotes to `absorbed` only when EVERY touched payload file matches upstream), that single 2-char drift in a file all five commits share poisoned the whole set. Five commits blocked by two cosmetic characters, on every release.

The fix encodes BALDART's own contract — **direct `.framework/` edits are blocked by the `framework-edit-gate` and sanctioned customizations live in overlays, so an uncovered residual is drift upstream OWNS** — into three additive classifier/updater pieces (no new agent, no new file):

- **New `liveDriftAgainstUpstream(touched)` + `reconcilable` category** (`src/utils/git.js` `classifyDivergence`). A would-be `custom-other` commit whose SURVIVING drift (`absorbablePayloadSubset` with HEAD≠upstream, computed by cheap blob-SHA compare) is confined to framework-owned files that (a) still exist upstream and (b) are NOT captured by any existing overlay is demoted from `custom-other` to `reconcilable`, carrying the exact file list. Overlay-COVERED drift routes to the capture path (intentional, tracked); net-new/deleted framework files keep flagging as real drift (an overwrite can't cleanly reconcile them).
- **Revert-pair cancellation** (`GitUtils.cancelRevertPairs`, pure/static). A `Revert "<subject>"` commit and the commit it reverts are a net no-op on the tree — both demote to `absorbed`. Kills the add-then-revert class that re-triggered the gate every release.
- **Seamless auto-align in `update`** (`src/commands/update.js`). New non-blocking `reconcilable` divergence branch: it lists the framework-owned files that drift, aligns each to its upstream (FETCH_HEAD) blob via the new `reconcileFrameworkDrift` helper and commits ONCE (so the subtree pull has no local hunk to conflict on), then proceeds — with the pre-pull backup tag making it fully recoverable. `git.on_divergence: abort` still lets a cautious consumer opt out; a genuine `custom-other` blocker (net-new/deleted/overlay-covered) still dominates to `mixed`/`real-custom` and stops for `/baldart-push` or `--on-divergence pull`.
- **`aggregateDivergenceClass`** gains the `reconcilable` class (blockers dominate → `mixed`/`real-custom`; else overlay-able → `overlay-able`; else reconcilable → `reconcilable`; else `all-noise`). `doctor`/`version` need no change (they delegate to `update`, which now self-heals the drift).
- Tests: `src/utils/__tests__/classify-divergence.test.js` gains 5 `reconcilable` aggregate fixtures + 4 `cancelRevertPairs` fixtures (61 pass, 0 fail). End-to-end validated on a synthetic subtree consumer reproducing the exact cosmetic-drift cascade → class flips from blocking to `reconcilable`, and `reconcileFrameworkDrift` aligns the file byte-for-byte to upstream.
- **Codex parity: portable** — pure CLI (Node), runtime-agnostic; benefits `[claude]`, `[codex]`, and `[claude, codex]` consumers identically. **No new `baldart.config.yml` key** (rides on the existing `git.on_divergence` policy) → the schema-change propagation rule does NOT apply.

## [5.10.1] - 2026-07-06

**Per-wave Codex relay backported to the v5.6.0 blocking-foreground pattern — no more bogus TIMEOUT at 83s that kills a live Codex.**
The v5.6.0 FEAT-0068 fix (final-review relay: FOREGROUND blocking Bash call with `timeout: 600000`, no background+poll) was never backported to the **per-wave** relay in `new-card-review.js`, so `/new`'s Discovery phase still ran the old `run_in_background:true` + poll relay. A live FEAT-0069 run reproduced the exact failure: the structured-output enforcement force-answered the polling relay at ~83s of its instructed 10-minute window (an idle poll turn is a forceable turn), it returned `marker:"TIMEOUT"` while Codex was still executing, and the Codex process died with the relay → cold `code-reviewer` fallback for that wave with no cross-model diversity. Diagnosed by a parallel budget-verification + adversarial-refutation pass (the refutation killed the tempting "launch-in-background-and-don't-poll" redesign: a dynamic-workflow script has no Bash, so Codex is always a child of the launcher agent — returning early *guarantees* the kill; the poll is what keeps it alive, so the fix is a **blocking** call, not a non-polling one).

- **`new-card-review.js`** — the per-wave `codexPrompt` now mirrors `new-final-review.js:343-351` verbatim (adjusted only for the wave-scoped `/tmp` file names): a single FOREGROUND `node … task … --wait` call with the Bash tool `timeout` at 600000ms, `marker` derived by grepping the file *after* the call returns, `RELAY_EXIT` 124/137/143 → real `TIMEOUT` + a worktree-scoped orphan reap (`pkill -f -- "--cwd <worktree>"`, never a bare pkill), clean exit with no sentinel → `TURN_COMPLETED`. The header comment is corrected (no longer "polls"). The fan-in (awk extract, prose salvage, deterministic re-extract, `code-reviewer` fallback) is unchanged.
- **Effect**: Codex now gets its full ~10-minute Bash-timeout budget per wave instead of dying at 83s; a genuinely slow/hung Codex still degrades cleanly (real timeout → reap → fallback), so the graceful-degradation contract is preserved.
- **Codex parity: N/A** — dynamic workflows are Claude-only (`codex.js` returns `null` for `workflowsDir()`); Codex consumers use the skill's inline review fallback, which was never affected. No behaviour change for Codex, no new `baldart.config.yml` key → the schema-change propagation rule does NOT apply.

## [5.10.0] - 2026-07-06

**Dedicated `code-simplifier` agent + `simplify-protocol.md` module + a deterministic clone floor — the hyper-specialized reuse & simplicity layer that catches the agentic defects generic linters miss (reuse-misses, duplication, dead code, wrong-altitude abstractions, missed optimizations).**
`/simplify` had been stuck since v4.57/v4.58: an LLM-eyeballing pass with no deterministic grounding, and — the real anomaly — it was the ONLY review lane with no dedicated agent (it ran on bare `general-purpose` in `new-card-review.js`, while security/qa/code all have a named specialist). Motivated by web research on the agentic-era state of the art (Abbassi et al. arXiv:2503.06327: ~90% of LLM-generated code has ≥1 inefficiency; self-review is unreliable — Huang ICLR'24 arXiv:2310.01798, Xu ACL'24 arXiv:2402.11436; the industry pattern is deterministic-detectors-first → LLM-triage — Semgrep Assistant, CodeRabbit, Sourcery) and hardened by **four adversarial refutation passes** that cut the original design roughly in half (killed a false-coupled external-tool tier, a redundant 4th parallel lens, and premature module relocation; forced the domain-partition contract and the clone-floor spec).

- **New agent [`code-simplifier`](framework/.claude/agents/code-simplifier.md) (v1.0.0)** — `model: sonnet` (Haiku BANNED), `memory: project`, slotless, analysis-only (Can Edit Code: No — the orchestrator/`simplify` skill applies fixes; `bypassPermissions` is execution access, not an implementer role). Closes the agentless-lane anomaly. Carries a binding **Domain Partition** vs `code-reviewer`: code-reviewer owns correctness/security/DS-as-HIGH/DRY-as-a-blocker (findings BLOCK), code-simplifier owns reuse/duplication/dead-code/wrong-altitude/missed-opt as **non-blocking quality cleanups** — no double-reporting in the shared `/new` cluster.
- **New protocol module [`framework/agents/simplify-protocol.md`](framework/agents/simplify-protocol.md)** — the reuse & simplicity verification engine, the simplify twin of `review-protocol.md`. `SECTION=` dispatch: `taxonomy` (agentic defects A1–A9 with a detectable signal each), `canonical-vocab` (Fowler smells / Four Rules of Simple Design / AHA·YAGNI·Rule-of-Three / Beck tidyings + calibrated complexity thresholds — cite the name, never invent a number), `clone-vocab` (Type-1/2/3/4), `rubric` (the lens checklists), `deterministic-tier` (the clone-floor contract), `doc-awareness` (retrieve-then-verify reuse retrieval over the registry/design-system/code-graph/LSP/API/wiki/ADR — docs are a high-recall index, code is the SSOT; doc-drift is an advisory handoff to `doc-reviewer`), `bias-guards` (a simplify judge is **inverted** — longer code is a NEGATIVE signal; evidence over opinion; independence, not self-grading). It **EXTENDS** `review-protocol.md`, never duplicates it.
- **New deterministic clone floor [`skills/simplify/scripts/simplify-scan.mjs`](framework/.claude/skills/simplify/scripts/simplify-scan.mjs)** — a zero-dependency, language-agnostic Node rolling-hash near-duplicate detector (whitespace/comment-normalized, min-6-line windows, ignore-list + import/trivial suppression, NO identifier normalization so boilerplate does not explode into false positives), run BEFORE eyeballing to ground the Reuse lens with an external clone signal the LLM lacks. It PROPOSES anchored `path:line` candidates; the reviewer DISPOSES. Twin of `ui-design/craft-check.mjs`, identical on Codex. Fixture test [`simplify-scan.test.mjs`](framework/.claude/skills/simplify/scripts/simplify-scan.test.mjs) proves it is silent on boilerplate and catches real cross-file clones.
- **No external-tool tier (deliberate).** An adversarial pass proved gating a jscpd/lizard/ast-grep booster on `has_toolchain` is a **false coupling** (that flag installs Biome/Vitest/tsc/Lefthook — none of them), reintroducing the silent capability-drift the design had removed; and per-stack npm/pip/native dispatch breaks Codex portability + is refused-scope (cf. the ds-handoff "no new parser" precedent). The always-on zero-dep floor ships; any future external booster is its own card with its own flag + full propagation + a documented Codex fallback + a measured redundancy analysis vs the floor.
- **`simplify` skill (v1.1.0)** — new Step 1.5 (run the clone floor), Agent 1 consumes the anchored candidates, Agent 2 gains two advisory design-defect bullets (wrong-altitude abstraction A4, sibling-pattern inconsistency A6 — folded into Quality, NOT a 4th parallel lens), an inverted-verbosity bias guard, and it now CITES the module instead of owning the taxonomy inline.
- **Wiring**: `new-card-review.js` routes the simplify finder to `agentType: 'code-simplifier'` (+ clone-floor + design-altitude in `simplifyPrompt`); `/new` inline `review-cycle.md` Phase 2.55 spawns one `code-simplifier` covering all lenses (was three anonymous `general-purpose` agents); `coder` § Author-Time gains a design-altitude reflex and cites the module. (`new2.js` has no simplify lane — nothing to rewire.)
- Routing wired in `agents/index.md` (routing rule + Modules list); `REGISTRY.md` (Agent Map + Model Matrix + decision tree + note); `agents/CHANGELOG.md`. README agent inventory updated.
- **Codex parity: portable** — the module is markdown read at runtime by both; the scan is zero-dep Node; the agent transpiles to `.codex/agents/code-simplifier.toml` automatically (one persona, no sub-spawn — respects `agents.max_depth: 1`). **No new `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [5.9.0] - 2026-07-06

**`security-reviewer` upgraded to the 2026 AppSec state of the art — benchmarked against Codex `codex-security` + Claude Code `security-guidance`, imported as a single-sourced protocol module.**
The `security-reviewer` was a high-quality *monolithic* agent (senior persona, dual review/apply, pooled YAML, generic Challenge/CoVe via `review-protocol.md`) but stuck at a single-pass "read → list findings by severity" model. A benchmark against two real corpora — Codex's phased `codex-security` workbench (threat-model → discovery → validation → attack-path → severity, coverage ledgers, instance-preserving proof) and Claude Code's `security-guidance` agentic reviewer (adversarial self-refute, attacker≠victim, modern high-miss patterns) — found BALDART at ~60% of current best practice. This wave closes the gap **without a monolithic rewrite**, following the agent-spine 3-layer model.

- **New protocol module [`framework/agents/security-review-protocol.md`](framework/agents/security-review-protocol.md)** — the AppSec verification engine, the security twin of `review-protocol.md`. `SECTION=` dispatch: `boundary-gate` (read `SECURITY.md` + classify `product_surface`/`source_trust`/`boundary_crossed` before confirming), `threat-model` (repository-scoped, never diff-biased), `discovery-lens` (the modern **high-miss** taxonomy: fail-open state drift, allowlist semantic escape, control regression, gate/action field mismatch, parser/validator differential, sensitive-to-observability, CI/CD trust, IaC omitted arg, over-broad grant, stale identity mapping, security-registry fanout, resource-bound placement), `proof-tuples` (class-specific validation tuples + confidence anchored to validation **method** not bug class + instance-preserving), `attack-path` (structured facts → **mechanical** severity), `adversarial-refute` (SURVIVES-unless-refuted, the **attacker≠victim** privilege-boundary test first, precondition-vs-counterevidence discipline, diff newly-introduced-only), `coverage-ledger` (surface dispositions so a clean verdict is honest). It **EXTENDS** `review-protocol.md`'s generic passes, never duplicates them — no dual-SSOT.
- **`security-reviewer` agent (v1.0.0)** — gains a versioned frontmatter + a "Security Review Passes (BINDING)" block that cites the module with 1-line binding statements (agent-spine: core inline + module on demand). Mode A pooled YAML schema gains `vuln_class`, `cwe: [CWE-###]`, `validation_method`, and a structured `attack_path` block; the summary header carries a compact `coverage:` block. New **secret-redaction** BINDING (report a leaked credential as first 2–4 chars + `****`, never the live value — a review artifact reproducing a secret is itself a leak).
- **`framework/agents/security.md`** — the consumer-facing security template modernized from **OWASP Top 10:2017 → 2021** (XXE folded into Security Misconfiguration, insecure deserialization into Software & Data Integrity Failures, XSS into Injection; new A04 Insecure Design + A10 SSRF), with a 2025-draft note.
- Routing wired in `agents/index.md` (security routing rule + Modules list). README security-reviewer inventory line updated.
- **Codex parity: portable** — the module is markdown read at runtime by both runtimes; the agent-body edits transpile to `.codex/agents/security-reviewer.toml` automatically. **No new `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [5.8.0] - 2026-07-06

**Suggested launch plan at the end of `/prd` — waves with a NAME, persisted on the epic card.**
`/prd` used to close with a progress bar and a bare summary (PRD path, card IDs, branch, PR/SHA)
but never told you **how to launch `/new`** — one run or several waves, where the boundaries are,
which engine. You had to reason it out by hand. Now every `/prd` run ends with a mandatory
**"🚀 Piano di implementazione suggerito"** section, and the plan is **persisted on the epic
card** so `/new`/`/new2` can, at the end of a subset run, remind you which waves remain to close
the epic — each wave carries a short, memorable **name** ("Wave 1 · Fondamenta server").

The design was hardened by a two-reviewer adversarial pass that refuted the obvious v1
(a separate `implementation_plan` block, an `engine` field on the card, a new reference module,
the mechanical card-writer computing the heuristic). The shipped shape:

- **Drift-free persistence — no new block, no new module.** The wave partition is three OPTIONAL
  sub-keys (`wave` / `wave_name` / `wave_rationale`) on the EXISTING `execution_strategy.groups[]`
  entries — the single card-ID partition already consumed by `/new` for parallelism. No parallel
  `waves[]` array re-listing card IDs (that would drift). Because they are sub-keys of an
  already-`R` field, the baseline validator never sees them: **no new field-state-matrix row, no
  migration of legacy epics**. Schema SSOT: `framework/agents/card-schema.md`.
- **Compute-by-orchestrator / persist-by-writer** (the `provenance.planning_session` split):
  the `/prd` orchestrator (opus/high) computes the single-run-vs-waves judgment + names + rationale
  — it needs the discovery context (migration-early, `review_profile: deep` pressure, server↔UI
  cut) the mechanical `prd-card-writer` never holds — and persists it with a targeted Edit on the
  epic card. The card-writer emits the keys ONLY when handed them; else it omits them (explicit
  else branch → no fabrication). Single-run is the default (no wave tags).
- **Engine + hard rules are render-time, never persisted.** The `/new` vs `/new2` recommendation
  is computed at render and is **runtime-aware** — `/new2` needs the Workflow tool and is
  Claude-only, so on Codex the plan recommends `/new` only. A durable, cross-tool artifact must
  not carry a runtime-capability recommendation a Codex consumer cannot execute.
- **`/new` + `/new2` remaining-waves.** `/new`'s Step F.6 "Prossimi Passi" and `/new2`'s
  `buildReport()` "Epic goal" render the remaining waves BY NAME (fallback to the flat card list
  when the epic has no wave tags — no regression). `/new2` runs fs-less, so the `wave_name` is
  threaded through `cardGraph` (new `waveName` node field derived by the preflight agent).
- Project-specific between-wave cautions (migration/deploy gate, compile check, stash) stay in
  `.baldart/overlays/prd.md` and are surfaced at render-time — no stack names in the base.

No new `baldart.config.yml` key (rides on the existing `execution_strategy` + `features.*`) → the
schema-change propagation rule does NOT apply. **Codex parity: portable** (runtime-neutral field;
engine computed at render-time; `/new2` remaining-waves via `cardGraph`).

Skills bumped: `prd` 1.6.1→1.7.0, `new` 2.3.1→2.4.0, `new2` 1.8.1→1.9.0.

## [5.7.0] - 2026-07-06

**Feature-gated skill linking — stop wasting Codex's 2% skill-description budget.** Codex
reserves ~2% of its context window for skill descriptions; BALDART's ~40 skills alone
already exceed it (~6.5k tokens), so Codex warns *"Skill descriptions were shortened to fit
the 2% skills context budget"* and truncates them (degrading skill routing). Until now the
installer linked EVERY framework skill into a consumer — including ones whose gating
`features.*` flag is OFF (e.g. the `ds-*` cluster, `i18n`, `graph-align`, `e2e-review`),
which then merely refuse at runtime: dead weight on the budget. Now a skill is linked only
when its declared feature is enabled, trimming ~2.5k tokens (~40%) on a typical project
without design-system/i18n/code-graph/e2e.

- **`requires_feature` frontmatter (SSOT) + `src/utils/skill-gate.js`**: an OPTIONAL
  `requires_feature: has_x` (string or list = AND) on a skill's `SKILL.md`. Makes
  machine-readable a dependency the skills already declared in prose (their "Project
  Context" header) — no second SSOT, same pattern as `version:`/`effort:`. `resolveSkillGate`
  reads `features.*` **directly** from `baldart.config.yml` (NOT `agent-slots.resolveAgentFlags`,
  which omits `has_prd_workflow`/`has_e2e_review`) and **fails open** globally (config
  absent/unreadable → link everything) and per-skill (malformed frontmatter → link that skill).
- **Applied to BOTH runtimes from one code path** (`mergeSkills` in `src/utils/symlinks.js`):
  the gate is shared SEMANTICS (which skills are relevant), not runtime mechanics, so Claude
  (`.claude/skills/`) and Codex (`.agents/skills/`) stay identical by construction. Adds the
  **prune** direction — a feature flipped off (or a fail-open first install) now removes stale
  framework symlinks, never touching a user skill dir or a user-override symlink pointing
  elsewhere.
- **17 skills gated**: `has_design_system` → ui-design, ui-implement, ds-new, ds-edit,
  ds-render, design-sync, ds-handoff, motion-design, gamification-design · `has_i18n` → i18n ·
  `has_code_graph` → graph-align · `has_e2e_review` → e2e-review · `has_prd_workflow` → prd,
  prd-add · `has_backlog` → new, new2 · `has_wiki_overlay` → capture. The five bootstrap
  skills (`design-system-init`, `i18n-adopt`, `graphify-bootstrap`, `lsp-bootstrap`,
  `toolchain-bootstrap`) stay ALWAYS-linked so a feature can always be turned on — then
  `configure`/`update` re-merge and the cluster appears.
- **Lifecycle wiring**: `configure` now re-runs `mergeSkills` after writing config (first
  install linked fail-open before the config existed — this is what actually prunes);
  `update` already re-merges (prune committed via `BALDART_MANAGED_PATTERNS`); `doctor` gains
  a non-blocking `reconcile-skills` advisory when a gated-off skill is still linked, and its
  skill sample-probe no longer samples gatable skills (`prd`/`capture`). `skill-creator`:
  `quick_validate.py` allow-lists `requires_feature` (it previously HARD-FAILED on unknown
  keys) + soft-warns unknown flags; `references/skill-structure.md` documents the field.
  `new2.js` binding-gate now guards on `has_design_system`.
- **Behaviour change for Claude consumers**: a skill whose feature is off now *disappears*
  from `.claude/skills/` instead of appearing and refusing at runtime. Mitigated: bootstraps
  stay visible, and re-enabling the feature (then any merge) brings the cluster back.
- **Rider — two skill descriptions made `quick_validate`-clean** (pre-existing failures, both
  among the heaviest so this also shrinks the Codex budget): `ds-edit` description trimmed
  1080→1014 chars (was over the 1024 max), `ds-handoff` angle-bracket placeholder
  (`<screens-or-feature>`) rephrased bracket-free — no trigger phrases lost.
- Verified: unit-level gate partitioning (design_system-off/backlog-on → 15 gated off, `new`
  linked, bootstrap present), YAML validity of all 17 frontmatters, `quick_validate` PASS.
- **No new `baldart.config.yml` key** (consumes existing `features.*`) → the schema-change
  propagation rule does NOT apply; it is skill-frontmatter metadata like `version`/`effort`.
- **Codex parity: portable** — single semantic filter, both runtimes, install-time only;
  `requires_feature` is ignored by both skill loaders (they read only name+description, as
  they already do for `version`/`effort`).

## [5.6.0] - 2026-07-04

**Outcome-integrity wave (FEAT-0068 post-mortem — mayo, 72.78M token, DONE falso).** Il
batch chiuse 6/6 card + epic come "✅ DONE, merged" mentre il goal dell'epic era invisibile
sul ~90% delle schermate: il root cause era in `/prd` (inventario ISA sui consumer del
meccanismo invece che sui proprietari della superficie; recipe ancorata su un componente
`@deprecated`; due gate phantom-stampati), attraversava card che nessun gate poteva salvare
(coverage passata vacuamente su un inventario sbagliato), e veniva sigillato da `/new`/`new2`
(chiusura epic al commit dell'ultima card — PRIMA della final review — con predicato
conta-figli, nessun verdetto PARTIAL, safety-commit che spazzava file non gated su develop,
commit agent in detour sul main checkout). 28 finding crit/high confermati adversarialmente;
ogni fix è ancorato a un'evidenza del run.

- **Chiusura epic v2 (`/new` 2.3.0 + `new2` 1.8.0)**: SOLO in merge phase (il gate
  commit-time v4.99.0 è ritirato — aveva anche un anchor bug `^parent:` che lo rendeva
  no-op); predicato esteso (figli **e follow-up** DONE + nessun residual bloccante +
  **AC epic verificati con oracoli**); blocklist JS + guardia post-merge (F-029 twin) +
  ramo **REOPEN** (il ratchet è bidirezionale); classe finding **closure-invalidating**
  (un "goal not achieved" è un difetto di STATO — il fix di prosa non chiude).
- **Verdetto PARTIAL calcolato** (`final-review.md` + `buildReport`): residual HIGH+ /
  epic aperta / AC epic `[ ]` ⇒ headline `⚠️ PARTIAL` col gap nominato; sezione
  `## Epic goal`; severity sui residui; "nulla perso" mai come surrogato di "non finito".
- **Igiene git orchestrata**: `merge-worktree.sh --safety-sweep allowlist:` (worktree-manager
  1.2.0 — il codice non gated va in quarantena etichettata, mai nel trunk sotto un
  "[safety]" opaco); isolamento worktree del commit agent esteso ai TOOL con evidenza
  porcelain + guardia JS (era 7.1M di detour + stash dangling); divieto stash + audit
  post-run dei commit non-card e degli stash del batch (mai "non miei da droppare").
- **Review**: relay Codex della final review in foreground `--wait` (il polling veniva
  ucciso dall'enforcement structured-output a 83s/600s → fallback freddo da 3.68M);
  esito **refuted** ratificato in `new2-resolve` (premessa falsa = zero edit) + divieto
  di fix che disattivano verifiche mandate da AC (`verification-posture-change`);
  suite aggiunte dalla card ESEGUITE targeted pre-commit (`completeness.md` 1b).
- **Routing**: `hasMockup` semantico (la card COSTRUISCE la UI mockata — un bare
  `links.design` su card docs va all'owner router; era 6.55M di ui-expert su una card
  docs+ADR).
- **`/prd` 1.6.0 (il root cause)**: categoria ISA **UI ROLE FAMILY (registry-first)**
  (owner della superficie dal registry + flag BLOCK su `@deprecated`); §0 LOAD CONTRACT
  su `discovery-phase.md` (vista PARTIAL ⇒ offset-Read fino in fondo — la coda persa
  conteneva il Codex completeness cross-check mai eseguito); `module_read` stamps
  ESEGUITI + item 0-ii Phase-evidence audit (2 gate phantom-stampati sul run reale);
  `step_3_mode` enum chiuso + contratto `skipped` (mai design hand-authored fuori
  `/ui-design`) + Reconciliation-lite senza mockup + trigger di ri-enumerazione ISA;
  gate quantificatore-universale sugli AC epic (`invariant_owners:`); stamp holistic
  **evidence-gated** (`stamp-holistic-audit.js --audit-report` → `agents_run`; senza
  artefatto → `status: skipped` = ASSENTE per `/new`); card-writer 1.2.0
  (self-validation eseguita, epic `ac_verification` obbligatoria, mai count hardcodato,
  `links.design` solo su card UI, invariante baseline snapshot); HARD RULE mw-docs
  (+ guardia `/mw` su `kind: docs`); EDIT CONTRACT dello state file (5 retry ciechi
  da ~250k); `research-protocol.md` pre-flight libreria in difesa.
- **Tooling maintainer**: `new-session-audit.mjs` guadagna **EPIC_AC_UNMET** (FAIL —
  retro-validato: il run FEAT-0068 ora fallisce l'audit invece di passarlo) e
  **GATE_DISABLED_BY_FIX** (WARN) + fix del digest-writer su id workflow con `/`;
  `check-new-parity.mjs` esteso a 21 policy (chiusura epic, PARTIAL, mockup semantico —
  con token FORBIDDEN sui meccanismi ritirati).
- Deliberatamente rinviati (MED/LOW non verificati dal pass avversariale): digest
  per-batch dei moduli di riferimento (tassa cache ~3-5M), card shape
  VERIFY-ONLY/MECHANICAL-MIGRATION, sampling visivo del goal epic in final review,
  script `update-state.mjs` per `/prd` (mitigato dall'EDIT CONTRACT).
- Codex parity: **portable** — le policy vivono nei moduli prose SSOT (`/new`, `/prd`)
  e in script zero-dep/shell; i workflow Claude-only (`new2*`, `new-final-review`)
  hanno ricevuto le STESSE policy dei twin prose (guardia di parity estesa a 21 regole).
  Nessuna nuova chiave `baldart.config.yml` → la schema-change propagation rule non si
  applica.

## [5.5.0] - 2026-07-03

**Migration deploy chain hardened (FEAT-0067 post-mortem — mayo).** The batch's DB
migration was fully declared in the cards (`db_migration` + `db_migrations_summary`,
per the consumer's prd-card-writer overlay) but the remote `db:push` was never
executed nor surfaced: new2's Migration Gate probed ONLY `migration_plan` (structural
vocabulary mismatch → silent no-op) and the end-of-batch `db-migration-deploy`
residual was lost by the mid-run DONE-flip, closing the epic against an unmet DoD.

- **Vocabulary aliases** (`/new` 2.2.1, `new2` 1.7.0): the `-auto` delegation gate,
  `/new` Phase 0 1b and new2 Step 3.5 now detect a declared migration under
  `migration_plan.required` (epic) OR `db_migrations_summary` (epic) OR
  `db_migration.required: true` (per-card overlay convention).
- **new2 Step 5.1b Migration Sync Tail (NEW)**: after the autonomous workflow returns
  (post-merge, on trunk), the SKILL runs the project's post-merge sync check
  (overlay "branch B"); on red it asks Push now / Defer / Override interactively —
  never pushes autonomously. Triggered by the Step-3.5 DECLARATION, not only by the
  residual ledger (belt-and-braces vs a dropped residual). Telemetry
  `migration_sync_tail`.
- Codex parity: portable (prose-only skill/reference edits; `new2` is Claude-only by
  design — `/new` classic remains the Codex path and gains the same alias probe).

## [5.4.1] - 2026-07-03

**Minimal light lane — the roster for a simple card is now literally: grounding (inline),
coder, code review, doc review, commit, merge. Nothing else.** Completes v5.4.0 per the
maintainer's directive ("per cose così semplici voglio due passaggi in croce"): the compact
lane cut the CARD COUNT; this cuts the per-run AGENT COUNT. FEAT-0067-class run: ~33
agents → ~7.

- **Per-card doc review at light (single-card batch)**: the light review matrix adds the
  scoped `doc-reviewer` per-card (review-cycle.md deferral-gate exception) because the
  Final pass it used to defer to is now skippable.
- **Final Review skip gate (`final-review.md` § F.0 + `new2.js` `skipFinalLight`)**: ONE
  light card + zero actionable per-card findings (→ zero fixes applied) → skip F.2–F.5
  entirely. Honors the v4.56.1 refutation (per-card PRE-fix + final POST-fix are
  non-redundant) by scoping the skip to runs where there IS no post-fix state; any
  finding, fix, second card, or degraded reviewer → Final runs as always.
- **i18n fill spawn cut at light**: the owner briefing instructs the coder to fill every
  target locale itself; the anti-hardcoded lint + locale-parity gates still enforce.
- **Production-readiness spawn cut for all-light batches**: no deployable surface exists
  by Rule C (no migrations/indexes/cron/schema) — the agent only ever reported zero items.
- `check-new-parity.mjs` rule 18 locks the prose↔workflow lockstep for every cut.
- Skills: `new` 2.2.0 · `new2` 1.6.0.

Codex parity: portable (prose modules are the cross-tool SSOT; `new2.js` Claude-only by design).

## [5.4.0] - 2026-07-03

**Elastic pipeline — `/prd`+`/new` compress to the minimum for trivial features and
scale up only on Heavy triggers.** Follow-up of the FEAT-0067 post-mortem: a label
feature was decomposed into an epic + 2 balanced cards and burned ~54M billed tokens in
`/new` (+8M in `/prd`) — every child card pays a fixed pipeline cost the feature size
cannot amortize, and nothing bounded decomposition from below. Both skeptic-refuted and
adjusted before shipping (epic wrapper kept per HARD RULE 15; PRD compresses PROSE, never
its section inventory; the grounding obligation is untouched — only its delegation scales).

- **`/prd` Step 2.1 Tier Gate** (skill 1.5.0, `discovery-phase.md`): at Discovery exit
  the feature is classified TRIVIAL/LIGHT/HEAVY per AGENTS.md § Change Tiers (SSOT,
  cited not restated) with one user confirm → `lane: compact|full` in the state file.
  Compact lane: research skipped by default, Step 3 routed toward mockup-only/skipped,
  compressed-prose PRD (full section structure, 1–2 paragraphs each, ≤3 pages), ONE
  child card + the mandatory epic tracker, no added audit ceremony (agent depth already
  scales with `review_profile`). Full lane byte-unchanged.
- **`prd-card-writer` v1.1.0 — Proportionality lower-bound + Rule C fixes**: atomicity
  now has a FLOOR (TRIVIAL/LIGHT feature → default ONE child card; split only on an
  isolable Heavy trigger or different-owner halves each >3 files; sizing rules are upper
  bounds, never split targets — the tier travels in the Step 5 briefing). Rule C branch
  (b) gains DS-spec doc-sync transparency (the touched component's spec `.md` counts
  like a locale file; never for `components_primitives` paths or HEAD-API changes) and a
  misapplication guard ("tocca i18n / la spec" is never by itself a disqualifier — the
  exact FEAT-0067-02 misclassification).
- **AGENTS.md primitive 1.8.0 (+ CLAUDE.md 1.4.3)**: § Change Tiers Heavy trigger
  refined — an ADDITIVE index-only migration is no longer Heavy by itself (keeps Rule C
  `db_indexes` depth + the api-perf gate); and a new orchestrated-pipeline exception:
  inside `/new`/`new2` a `review_profile: light|skip` card satisfies the (unchanged)
  understanding obligation via the OWNER grounding inline instead of a dedicated
  `codebase-architect` spawn.
- **`/new` 2.1.0 light lane (`implement.md` step 2d) + `new2.js` mirror**: light/skip
  cards skip the Phase-1 retriever spawn; the owner briefing carries the inline-grounding
  instruction. Guarded by `check-new-parity.mjs` rule 17 (prose + workflow + primitive
  must stay in lockstep).

Codex parity: portable — the Tier Gate binds `AskUserQuestion` vs plain-prompt via the
existing `runtime-portability-protocol.md` citation in `/prd`; primitives/AGENTS.md is
the cross-tool SSOT; `new2.js` stays Claude-only by design (documented gap).

## [5.3.2] - 2026-07-03

**AGENTS.md primitive 1.7.4 — commit-format MUST scoped to card-bound work.**
The commit rule was unconditional ("MUST use commit format `[FEAT-XXXX]
description`"), so a Codex session doing card-less chore work on a consumer
(mayo) fabricated the missing datum: it grabbed the last card ID committed in
`git log` (assuming correlation) and typed the commit `feat`. Same failure
family as 5.3.1/1.7.3 — an unscoped constraint pushes the model to invent data
to satisfy it — and it contradicted § Change Tiers ("Light work → no backlog
card").

- **Primitive `AGENTS.md` 1.7.4**: the card reference is scoped to commits that
  implement THAT card (picked, `IN_PROGRESS`, branch carries its ID); card-less
  work uses plain `<type>: description`; copying a card ID from `git log` / the
  previous commit "for consistency" is explicitly forbidden; the type must
  describe THIS change's nature (`chore:` for a chore — never `feat` unless the
  commit adds a feature), never inherited from neighboring commits.
- Codex parity: portable (the primitive is the cross-tool SSOT).

## [5.3.1] - 2026-07-03

**AGENTS.md primitive 1.7.3 — anti-fabricated-policy guard on the grounding spawn.**
A live Codex session on an up-to-date consumer (mayo) reconstructed the
`codebase-architect PROFILE=ui` mandate correctly and then self-exempted by
inventing a runtime constraint ("subagent tool disponibile ma policy runtime lo
vieta senza delega esplicita") that exists nowhere — the only real Codex limit
is `agents.max_depth: 1`, which forbids *nested* spawns from inside a subagent,
never the top-level session. Recurrence of the 1.7.0/v4.98.0 misreading class:
the AGENTS.md delegation rules inverted into "spawn only on explicit user
delegation".

- **Primitive `AGENTS.md` 1.7.3**: the grounding MUST now states in place that
  **the mandate IS the delegation** — claiming a runtime policy requires prior
  user delegation is named as non-compliant, and the degrade precondition is
  tightened to "Task tool absent AND no
  `.codex/agents/codebase-architect.toml`". Declaring the tool available and
  degrading anyway is a protocol violation, not a fallback.
- Codex parity: portable (the primitive is the cross-tool SSOT; the guard's
  Codex branch cites the real `max_depth` semantics from `CODEX-AGENTS.md`).

## [5.3.0] - 2026-07-03

**Haiku banned fleet-wide + FEAT-0067 post-mortem fixes.** The FEAT-0067 audit
(mayo, session `a8b6bc75`) found a haiku `general-purpose` commit agent running
its first git command with `cd <worktree> &&` and its SECOND without — so the
card-DONE bookkeeping commit executed in the MAIN checkout and landed on the
trunk mid-run (`13aa63b62`), forking `ssot-registry.md`, raising a merge
BLOCKER (a 5.3M-token repair), and blocking epic closure. Second haiku-class
failure after the v4.42.0 worktree-setup fabrication → the plumbing carve-out
is revoked.

- **Model policy — haiku eliminated, sonnet is the floor** (`REGISTRY.md`
  § Rules rewritten): no named specialist AND no `general-purpose` plumbing may
  use haiku. Replaced in every spawn site: `new2.js` (`commit:`,
  `integrate:commit`, `rollback:`, `review:*:codex` driver),
  `new-final-review.js` + `new-card-review.js` (`codex-preflight`), and the
  `seo-analytics-strategist` agent (→ sonnet, v1.1.0). Where an op is truly
  deterministic the fix stays a script (`setup-worktree.sh`,
  `revert-files.sh`), not a cheaper model.
- **`new2.js` commit agents — cwd HARD RULE + placement verify**: every shell
  command must be prefixed `cd <worktree> &&` in the same Bash invocation (the
  cwd resets between calls), and the per-card commit agent now verifies the
  returned sha equals `git -C <worktree> log -1 --format=%H`, reporting
  `committed:false` with the stray sha instead of silently corrupting the
  trunk.
- **`new2.js` Phase-1 retriever — schema + drift false-positive fix**: the
  prompt now states the exact required StructuredOutput shape (`ok` REQUIRED —
  the FEAT-0067 run burned 5+3 schema-validation retries), and `missingPaths`
  excludes paths the card itself CREATES (new migrations/tests were flagged as
  drift and disproved downstream).
- **`scripts/new-session-audit.mjs` — workflow-hosted agents counted**: the
  analyzer now traverses `subagents/workflows/<wf>/agent-*.jsonl`; a new2 run
  previously reported the orchestrator alone (FEAT-0067: 3.88M reported vs
  53.89M real — 93% of the run invisible to TOTAL, baselines, and telemetry).

Codex parity: portable / N/A — the three workflows are Claude-only by design
(documented gap with inline fallback); the agent `model:` field is omitted in
the Codex TOML transpile, so the model-policy change is Claude-side only; the
audit script is maintainer tooling, not payload.

## [5.2.1] - 2026-07-03

**`/mw` local-trunk sync robustness.** The worktree merge's local-trunk sync
(step 4c "Common", both `pr` and `local-push` strategies) failed to
fast-forward on the majority of real checkouts, needlessly leaving
`SYNC-DEFERRED` / `SYNC-NEEDS-DECISION`. Two targeted fixes, kept single-sourced
across the executed SSOT (`framework/.claude/skills/worktree-manager/scripts/merge-worktree.sh`)
and the prose (`worktree-manager/SKILL.md` step 4c) — **without `git stash`**,
which stays forbidden here because `refs/stash` is shared across all worktrees
via `git.commondir` (the FEAT-0522 incident): a stash from the main checkout can
pop into the wrong worktree, and two parallel `/mw` runs race on the shared
stash list.

- **Fix 1 — safe ref-advance when main is on another branch.** Previously, if
  the main repo's HEAD was not on `$TRUNK`, sync did a fetch-only and always
  emitted `SYNC-DEFERRED`, never advancing the local `$TRUNK` ref. Now it
  fast-forwards `refs/heads/$TRUNK` via a ref-only, ff-only fetch
  (`git fetch origin $TRUNK:$TRUNK`) that **never touches any working tree** and
  never checks out — turning a large class of deferrals into a clean sync
  (`sync_marker: none`). If `$TRUNK` is checked out in another worktree, git
  refuses and it falls back to fetch-only + defer.
- **Fix 2 — accurate ff-fail disambiguation.** When the ff failed with a
  tracked-clean tree and no owned/foreign dirt, the code fell to a misleading
  "clean (diverged?)" marker (the dirty scan uses `--untracked-files=no`, so an
  untracked-file collision was invisible to it). Now it disambiguates via
  `merge-base --is-ancestor` between a **genuine branch divergence** (needs a
  rebase/merge decision) and an **untracked-file collision** (the blocking files
  are enumerated), emitting the correct actionable marker.

Behaviour is strictly safer: cases #2 (foreign uncommitted work) and #3 (real
divergence) still defer to ONE explicit user gate — never auto-committed, never
stashed — which is the correct behaviour on a shared checkout.

**Codex parity: portable.** The change is pure git plumbing inside the
dual-runtime `merge-worktree.sh` (executed identically under Claude and Codex)
plus its prose twin; no runtime-specific mechanic, no config key. The
schema-change propagation rule does not apply (no `baldart.config.yml` key).

## [5.2.0] - 2026-07-03

**The local design studio.** `/ui-design` is refounded from the root as the
top-of-the-line skill for designing, prototyping and restyling UI locally —
the internal twin of the Claude Design handoff — absorbing design
intelligence distilled from three state-of-the-art open repos
([impeccable](https://github.com/pbakaus/impeccable) Apache-2.0,
[taste-skill](https://github.com/Leonxlnx/taste-skill) MIT,
[open-design](https://github.com/nexu-io/open-design) Apache-2.0, attributed
in each derived reference). The superseded `frontend-design` /
`ui-ux-pro-max` / `huashu-design` dependencies are cut, and the skill-routing
failure ("agents pick frontend-design for UI work") is fixed at every layer.

- **`ui-design` v2.0.0** (public step contract A–H and the `/prd` 3a–3e
  mapping UNCHANGED):
  - NEW **Step C.0 — Design Read & Direction Lock** (`references/design-brief.md`):
    register (brand/product), regime (in-context/exploration), dials
    VARIANCE/MOTION/DENSITY, two-altitude slop test, color-commitment axis,
    dark/light scene sentence, anti-references, ONE-question ambiguity rule,
    cross-run rotation check.
  - NEW **`references/design-directions.md`** (12-direction aesthetic menu +
    font/palette reflex-reject lists: LILA rule, premium-consumer beige ban,
    serif discipline), **`references/craft-standards.md`** (always-on craft;
    numeric ranges cite `design-system-protocol.md` § Reference Tables; the
    **canonical new home of Performance Gates/CWV 2026 + Modern CSS + Form
    quality**, migrated from frontend-design), **`references/anti-slop.md`**
    (the Tells catalog — absolute bans, count-bounded rations, copy tells;
    every ban with a declared-override path).
  - NEW **`scripts/craft-check.mjs`** — deterministic anti-slop/craft gate
    (~22 rules, zero-dep, `--json`, `--strict-palette`); generator self-check
    + Step D Gate 0; identical on Claude and Codex.
  - **Step D rewritten as dual-lens**: Gate 0 craft-check (fail-fast, hidden
    from lenses) → Lens A **`ui-quality-critic`** (true evaluator diversity —
    different agent definition than the generator; rubric SSOT unchanged) →
    Lens B fresh `ui-expert` conformance (B1–B5, evidence-cited scoring,
    worst-sustained-band, no grade inflation). Honest **degraded mode** on
    non-multimodal runtimes (mandatory `⚠️ DEGRADED` banner).
  - Generation is **token-first** (project tokens embedded in `:root` so the
    mockup speaks the project token language, mirroring the Claude Design
    channel) and **`data-region`-annotated** (surgical iteration +
    deterministic inventory join into `component_bindings`).
- **`frontend-design` v2.0.0 — retired to a ROUTER**: the body reroutes
  deterministically (design → `/ui-design`, approved mockup →
  `/ui-implement`, direct in-code UI → `ui-expert`, motion →
  `/motion-design`); the broad legacy triggers are KEPT on purpose as a
  honeypot so reflexive mis-routing gets corrected instead of landing on
  stale guidance. Consumer overlays with aesthetic mandates should port to
  `.baldart/overlays/ui-design.md`.
- **Routing fix at every layer**: `ui-design` description rewritten to claim
  design/restyle/prototype triggers (IT+EN) with an explicit NOT-list;
  `agents/index.md` gains a DESIGN-DECISION routing rule;
  `agents/skills-mapping.md` UI section + decision tree rewritten
  (`ui-ux-pro-max` marked superseded, never route to it).
- **Agents**: `ui-quality-critic` v1.1.0 (gating extended: invoked by
  `/e2e-review` Phase 4b OR `/ui-design` Step D — never ad-hoc; mockup
  payload flavor documented); `ui-expert` v2.1.0 (CWV pointer re-targeted to
  `craft-standards.md`; Linked Skills drop `frontend-design`/`ui-ux-pro-max`,
  `ui-design` references become the standalone design-intelligence library);
  `coder` v2.0.2 + `remotion-animator-orchestrator` v1.0.1 (pointer
  re-targets). REGISTRY rows updated.
- **Docs**: `design-system-protocol.md` consumer lists updated
  (`frontend-design` → `ui-implement` where apt), `PROJECT-CONFIGURATION.md`,
  `configure.js` design-system hint, README + CLAUDE.md conventions;
  `e2e-review` 1.3.1 / `design-system-init` 1.0.1 / `gamification-design`
  1.0.1 pointer fixes.
- **Fix (`update`) riding in this release**: `isAbsorbedAgainstUpstream` now
  scopes the absorbed-commit divergence check to the `.framework/framework/`
  payload only (`GitUtils.absorbablePayloadSubset`) — mixed commits and
  subtree bookkeeping no longer mis-flag net-zero divergence as blocking
  `custom-other` (repro: mayo v5.0.1→v5.1.0 update).
- **No new `baldart.config.yml` key** (rides on `features.has_design_system`
  + existing keys) — the schema-change propagation rule does NOT apply.
- **Codex parity: portable** — same SKILL.md bundles (two symlinks), zero-dep
  script, spawn/decision-gate bindings cited from
  `runtime-portability-protocol.md`, and the visual lenses degrade honestly
  (banner) instead of silently on non-multimodal runtimes.

## [5.1.0] - 2026-07-03

**The research layer.** Research becomes a first-class, reusable asset: the
`senior-researcher` agent is restructured to the v5.0 authoring contract with
**research profiles**, every report lands in a consumer-owned, indexed
**research library**, and the **source matrix** that routes each research
nature to the right sources is versioned and grows project after project.
Design refined by an adversarial refutation pass (13 findings integrated —
ID-race BLOCKER, /prd poll determinism, lifecycle circularity, opt-out
resurrection, auto-commit scoping).

- **NEW protocol module `framework/agents/research-protocol.md`** (`SECTION=`
  dispatch): `profiles` (`PROFILE=<decision|deep|compare|regulatory>` output
  contracts; absent → `deep` = legacy semantics; deep's dzhng-style iterative
  loop + STORM perspectives + CoVe citation spot-check are opt-in via
  `DEPTH=<1..3>`, default 1 = today's cost), `library` (layout, tag-rich report
  frontmatter as source of truth, INDEX.md as regenerable view, `RES-<date>-
  <slug>` ids — no max+1 race, archive-not-delete), `reuse` (pre-flight
  `FULL_REUSE`/`DELTA`/`NEW` + per-category TTLs + anti-false-reuse guard +
  the always-write-a-file polling contract), `sources` (matrix schema,
  SOURCES.md → framework-template resolution — the default matrix lives in ONE
  file, and the `SOURCE_MATRIX_CANDIDATE` growth loop emitting a ready-to-paste
  maintainer prompt).
- **`senior-researcher` v2.0.0** (agents CHANGELOG + heading map): description
  4-examples → 2 lines; `version:`/`effort:` frontmatter; `model: sonnet` KEPT
  and the REGISTRY model-tier row corrected (it claimed "always opus",
  contradicting the frontmatter since day one); NEW Profile Dispatch / Research
  Library (total delivery-path chain: prompt → library → `paths.docs_dir`) /
  Source Matrix sections; Operating Protocol cites incl. `SECTION=
  injection-guard` — the agent reads more untrusted web content than any other
  and previously had NO injection guard; FIRST MESSAGE TEMPLATE scoped to
  interactive direct invocation.
- **NEW skill `/research` (42nd)**: THIN interactive orchestrator — scoping
  (max 3 questions) → profile+nature routing → library reuse check BEFORE
  paying for research → `senior-researcher` launch (parallel fan-out on
  Claude; sequential by-name on Codex per `runtime-portability-protocol.md`)
  → filing + INDEX ownership → matrix growth loop (apply locally to
  `SOURCES.md` `## Local additions` and/or print the upstream BALDART prompt).
  Ships per-profile report templates (`assets/`) + deterministic
  `scripts/rebuild-research-index.mjs` (INDEX regenerated from frontmatter).
- **NEW config key `paths.research_dir`** (default `docs/research`, proposed on
  FIRST encounter only — an explicitly emptied key is a respected opt-out,
  never re-proposed, never resurrected). Schema-change propagation:
  template + `configure` (detect probe + prompt + post-write library creation)
  + `update` (new-key detector already covers `paths.*`; post-pull seed
  backfill into an EXISTING dir only; auto-commit scoped to UNTRACKED seeds —
  never sweeps a consumer-edited SOURCES.md/INDEX.md) + `doctor`
  (`create-research-dir` / `backfill-research-seeds` autoOk actions — seeds are
  generic, unlike the i18n registry; `research-sources-advisory` when the
  framework matrix version moves ahead, never overwrites) +
  `PROJECT-CONFIGURATION.md` (§ 4.2.2 lifecycle + opt-out) + this CHANGELOG.
  New shared helper `src/utils/research-library.js` (idempotent, silent,
  returns `created[]`).
- **Versioned source matrix**: `framework/templates/research-sources.template.md`
  (`matrix_version: 1.0.0` + sibling `research-sources.CHANGELOG.md`) seeds the
  consumer-owned `SOURCES.md`; 5-nature default matrix (scientific / dev /
  regulatory / ux / market × primary / secondary / avoid / notes). Growth loop:
  agent emits run-evidence-only candidates + the copy-paste BALDART prompt.
- **`/prd` 1.4.0**: Step 2.5 spawns with `PROFILE=decision`
  (`PROFILE=regulatory` on signal S2) — resolving the two-conflicting-SSOTs
  bug where the prompt asked for a light per-topic format while the agent's
  contract mandated the full §0–§11 report; /prd now PRE-ALLOCATES the output
  path (library when configured, legacy `prd_dir` fallback otherwise) and the
  agent always writes a file there (`FULL_REUSE` → pointer stub) so the
  10-minute exit-check poll stays deterministic. Discovery-phase background
  launch updated in the same change (parallel location). `hybrid-ml-architect`
  deliberately NOT touched (vertical, on-request) — protected by the defaults.
- Inventories: 42 skills / 33 domain modules (README count was already stale
  at 28); REGISTRY Agent Map + decision tree + model-tier rows updated.
- OSS recon (inspiration only, zero dependencies adopted): GPT Researcher
  (planner/executor, local-first hybrid research), dzhng/deep-research
  (breadth/depth loop), STORM (multi-perspective questioning),
  open_deep_research (scoping phase), Karpathy markdown-first KB (INDEX as the
  single map, archive-not-delete).
- Codex parity: **portable** — skill = same SKILL.md bundle (two symlinks),
  agent = transpiled `.codex/agents/*.toml` from the same markdown, module +
  templates = plain markdown; the only Claude-only mechanic (parallel
  subagent fan-out in `/research`) has an explicit sequential-by-name Codex
  branch citing `runtime-portability-protocol.md`.

## [5.0.1] - 2026-07-03

**Adversarial-verification fixes on the 5.0.0 wave.** A full release verification
(distribution machine dogfooded end-to-end on 4 configs — all green; normative-loss audit
over 4 core agents; pipeline-coherence audit) found the headline mechanisms inert on the
happy path plus a handful of slot-boundary and consistency defects. All fixed; no gate
removed, no H2 moved (consumer overlays unaffected), no `baldart.config.yml` change.

- **FIXED (HIGH) — `/new` step 11b clean branch now proceeds to 11c**: previously only the
  violation branch reached the review-bundle/doc-invariants build, leaving the v5.0.0
  anti-re-read + PASS-fast machinery unused on every clean card (`implement.md`).
- **FIXED (HIGH) — `/codexreview` consumes the full lean contract**: Step -0.5 now parses +
  forwards `review_bundle_path` (anti-re-read phrase in every Step-2 prompt) and
  `specialist_dispatch` (orchestrated-mode suppression, `dispatch_deferred:<agent>`) — both
  were produced by `codex-gate.md` but dead-wired on the `/new`→`/codexreview` hop. Optional
  keys: pre-5.0 contracts keep working unchanged.
- **FIXED (HIGH) — coder slot boundaries**: `vercel-react-best-practices` un-slotted (no
  config flag can express React-non-Next → structural normative loss); the NEVER-ad-hoc
  browser-scripts rule hoisted out of `{{#e2e_playwright}}` (universal pre-5.0). Agents
  CHANGELOG: coder/doc-reviewer/qa-sentinel v2.0.1.
- **FIXED (MEDIUM)** — doc-reviewer per-card cited exceptions (`topological-generation`,
  `schema-drift`) restored + `doc-audit-protocol.md` Contract no longer contradicts the agent
  core; qa-sentinel bundle `gate_logs` demoted to context-only (own-tier gates always re-run);
  `build-review-bundle.mjs` diff scoped to `--may-edit` (multi-card batches no longer leak
  prior cards into card N's bundle/docinv; `new2.js` implement brief mirrors the scoping);
  `new-card-review.js` + the 2.5x delegation thread `reviewBundlePath` into every brief;
  Phase 2.55 pre-5.0 fallback uses `git diff "$TRUNK...HEAD"` (bare `git diff` post-commit is
  vacuous); cap-and-handoff checkpoint-missing branch defined; `/ui-implement`-delegated cards
  explicitly run 7c–11c; PASS-fast Own-Doc Closure claim corrected; "rule 8" pointers
  re-anchored to the renumbered Post-Intervention Coherence rule 7 (prd-card-writer, upgrade
  doc, new2-resolve comment).
- **FIXED (LOW)** — `doc-invariants.mjs` evaluates `documentation_impact` and `canonical_docs`
  independently; deferral KEEP list re-includes `balanced`; stale `coding-standards.md`
  cross-refs to the removed coder H2 updated; `${paths.references_dir}` literals in
  `doc-audit-protocol.md`.
- Skills: `new` 2.0.1, `new2` 1.4.1 (per-skill CHANGELOGs). Codex parity: portable — prose +
  pure-Node scripts, no Claude-only primitive added; the `/codexreview` contract keys are
  optional and ignored by pre-5.0 readers.

## [5.0.0] - 2026-07-03

**The agent-spine major.** Measured driver (1,934 subagent runs / 191 sessions on the
`mayo` consumer, one month, deduplicated by API-call id): 5.0B context tokens, of which
coder 53% and the whole review chain ~10%; the fixed per-spawn preamble (~45k, agent
definition 6–10k of it) is re-paid on EVERY API call; Read results are 66–75% of
tool_result bytes with the same hot files re-read at every chain link; the top-10
monster runs (200–440 calls) alone are 14% of total spend. 5.0 restructures the core
agents and the `/new` pipeline around those numbers — no gate is removed anywhere.

- **MAJOR — agents install as generated files when slot-bearing.** Core agent
  definitions are restructured to 3 layers: core inline (I/O contracts + 1-line binding
  rules) + shared protocol modules read on demand + `{{#flag}}` install-time slots
  resolved per-consumer from EXISTING `baldart.config.yml` keys (a consumer ships ONE
  `stack.database` variant, not four). A slot-bearing agent is GENERATED per-consumer
  (`.baldart/generated/agents/` + symlink, marker gains trailing `config_sha`,
  byte-stable rewrites); slotless agents stay plain symlinks; H2 headings stay stable
  (no-op body when a flag is off) so existing overlays keep matching. First
  `baldart update` after 5.0 flips the affected agents symlink→generated in one
  auto-committed reconcile (allowlist already covers it). Machinery:
  `src/utils/agent-slots.js` (new) + `symlinks.js` slot detection (fill → overlay-merge
  → marker → transpile) + doctor `generatedAgentsStale` probe / `regenerate-agents`
  action + `configure` re-merge after config writes. NO new config key (schema-change
  propagation rule: N/A). Codex parity: adapter-generated — the TOML transpile consumes
  the config-effective markdown, so Codex agents get the same resolved variant.
- **7 core agents restructured** (`coder`, `code-reviewer`, `doc-reviewer`,
  `plan-auditor`, `qa-sentinel`, `codebase-architect`, `ui-expert`) — source cuts
  30–49% with ZERO normative loss (two independent adversarial verification passes:
  no LOST, no WEAKENED rules). All carry `version:` frontmatter (own SemVer line,
  history in the new `framework/.claude/agents/CHANGELOG.md` — NOT installed, excluded
  from every agent enumerator incl. the discovery hooks) and ≤2-line descriptions (the
  old multi-example descriptions were paid by the orchestrator every session). Codex
  parity: adapter-generated (transpiler unchanged, consumes the effective markdown).
- **3 new protocol modules** (`agents/agent-operating-protocol.md`,
  `agents/review-protocol.md`, `agents/doc-audit-protocol.md`) — the deduplicated SSOT
  of the shared operating blocks (injection guard / retrieval / memory / tool budget),
  the reviewer verification engine (Challenge+Actionability / Simulation diff+plan /
  CoVe / absolute risk scoring / specialist-spawn discipline incl. the new
  `specialist_dispatch` orchestrated-mode suppression), and doc-reviewer's audit-mode
  procedures (read ONLY without card context). `SECTION=` dispatch — agents read only
  the matching section. Codex parity: portable (markdown modules).
- **`/new` pipeline wave (skill `new` 2.0.0, `new2` 1.4.0)** — six additive mechanisms,
  each with pre-5.0 graceful degradation and `check-new-parity.mjs` tripwire rules
  (16 total): review bundle (`build-review-bundle.mjs` — deterministic per-card
  handoff + binding anti-re-read phrase in every reviewer briefing, Hot-File Map
  honored downstream); doc-invariants PASS-fast (`doc-invariants.mjs` — the invariant
  table as executable rules; zero blocking rows + no doc signals + non-deep → per-card
  doc-reviewer spawn skipped, Final F.3 backstop unchanged; measured headroom: 84/127
  per-card runs returned PASS); cap-and-handoff (soft 80 / hard 100 tool calls,
  checkpoint + fresh continuation ×2 — attacks the monster-run tail); explicit i18n
  fill pass (step 7c, `i18n-translator` BY NAME — closes the measured general-purpose
  leakage); fix-pass model routing A/B (fix-pass coder → `model: sonnet` best-effort,
  overlay kill-switch `## [OVERRIDE] Fix-pass model`, `model=` in Fix Application Log
  for recurrence telemetry; initial builds and specialists untouched). Codex parity:
  scripts portable (.mjs), prose additive; the new2 workflow mirrors are Claude-only
  with the /new prose as the inline fallback.
- **`finding-mine` routine (monthly) + `skill-improver` 2.0.0 `MODE: miner`** — the
  anticipatory loop: 30 days of pooled findings / QA reports / batch trackers /
  reviewer memories distilled into CLASSIFIED proposals (author-time coder rules,
  never-demote rows, deterministic-gate candidates, card-schema proposals —
  proposal-only, upstream candidates consolidated). Routine count 8 → 9. Codex parity:
  adapter-generated (3 routine backends; agent transpiled).
- **`scripts/agent-telemetry.mjs` (maintainer tool)** — fleet-level per-family
  telemetry (calls p50/p90/max, ctx, read-share, verdicts, rework ratio, cap-hit rate,
  fix-pass models) with `--baseline save|compare`; the pre-5.0 baseline is saved.
  Dedup by `message.id` keeping the LAST record (the JSONL double-count trap). Codex
  parity: N/A (maintainer-only, Claude transcripts).
- **Post-release measurement plan** (2 weeks on a live consumer, vs the saved
  baseline): coder ctx/run −30%+, Read-share <66%, p100 calls down, doc-reviewer
  PASS-fast rate ≈ measured headroom, fix-pass recurrence sonnet ≤ opus, reviewer
  verdict distributions unchanged.

MAJOR — changes the on-disk layout of installed agents (symlink → generated file for
slot-bearing agents). No new `baldart.config.yml` keys anywhere in the release.

## [4.99.0] - 2026-07-03

**Consumer-lessons wave: four recurring failure patterns surfaced by mayo's nightly/weekly routine reports (20260628-skill-improve + 20260701-doc-graph-align) upstreamed into the base framework.** The consumer had fixed each locally via overlays, but three of the four root causes live in the base payload — every future consumer would have re-discovered them. The overarching meta-defect (#2) is that the lesson pipeline itself had no upstream channel: overlay fixes died in the consumer.

1. **Epic-closure gate in `/new` (skill `new` 1.4.0 → 1.5.0).** `commit.md` step 28c said only "add/update the entry for this card's feature area" — nothing about the card being the LAST open child of a parent epic, so epics stayed `IN_PROGRESS` after closure (3 confirmed occurrences across two audit windows, caught only by the nightly). New deterministic sub-step **28.c-bis**: when the card YAML declares `parent:`, count non-DONE siblings; zero → flip the parent's ssot-registry entry (and its own YAML if present) to DONE in the SAME commit. Idempotent, non-triggers enumerated. Mirrored in `new2.js`'s mechanical commit prompt (step 4b + ROLE BOUNDARY widened to the parent-epic YAML status).
2. **`skill-improve` routine + `skill-improver` agent modernized (pre-overlay-era → current model).** The routine prompt (since_version 2.1.0) instructed editing `.claude/skills/**` / `.claude/agents/**` directly — symlinked framework files that the `framework-edit-gate` (v3.3.0+) DENIES in consumers — and read only `*-code-review/doc-review/qa` reports (absent in real consumers; the agent had to improvise onto the nightly reports). Now: write surface = **`.baldart/overlays/`** (skills / agents / root primitives, with section markers); sources = **ALL routine report types** (`doc-graph-align`, `wiki-review`, `i18n-align`, `design-sync`, `ds-drift`, `full-sweep`, + prior skill-improve for context); graceful degradation on missing metrics/report types; and a new mandatory **`## Upstream Candidates`** report section — when a pattern's root cause is in the BASE skill/agent, the overlay fix is applied locally AND the base defect is flagged for the maintainer to carry via `baldart push` (closing the loop this very release exercises).
3. **`API_BODY_CONSTRAINT_DRIFT` — contract-value body check in `doc-reviewer`.** A review/security fix that changes a validation constraint (`.max()`/`.min()`/length/regex/enum/default on a route handler, server action, or API schema) landed with the API reference *body* still stating the old value (real case: DoS cap 500 → 100); frontmatter-freshness checks don't catch body-level factual drift. New invariant-table row: verify the documented value EQUALS the code value; the doc-reviewer is the canonical writer and fixes the body directly. Covered per-card (Phase 3) and at the batch-Final doc pass (fixes applied after the per-card doc pass are caught there).
4. **`update_triggers:` — declarative synthesis freshness (opt-in).** Coverage-map syntheses (query catalogs, schema coverage maps, endpoint inventories) silently rot when new code lands (3 missing entries across 3 cards in one week, all caught only by the nightly; the consumer hand-wrote a per-project reviewer rule). New OPTIONAL frontmatter field on wiki synthesis pages — path globs + `content:` added-line patterns that OBLIGATE updating the page when a diff matches, else `WIKI_SYNTH_STALE`. SSOT in `agents/llm-wiki-methodology.md § Update triggers`; consumed by three layers: `doc-reviewer` invariant table (per-task, fixes directly), `wiki-curator` (nightly, updates or flags), `doc-graph-align` routine (nightly commit-range cross-check, propose-don't-author). No retroactive obligation — pages without the field are exempt.

**MINOR** — adds capability to existing skills/agents/routines (no new agent/skill/command/routine, so README counts unchanged; no install-layout change). **Codex parity: portable** — all changes are markdown skill/agent/protocol/routine bodies shared by both runtimes (`skill-improver`, `doc-reviewer`, `wiki-curator` auto-transpile to `.codex/agents/*.toml` on install/update; `new2.js` is the Claude-only accelerator whose inline `/new` path carries the same gate). **Not a `baldart.config.yml` key** (`update_triggers` is wiki-page frontmatter, not config) → the schema-change propagation rule does NOT apply.

## [4.98.3] - 2026-07-02

**Bugfix: `baldart update` re-flags `multi_tenant_theming` as a "new config key" on every run.** A user reported the schema-drift detector nagging `⚠ New config keys in this version: multi_tenant_theming` at every `baldart update`, even after configuring it. Root cause: `multi_tenant_theming` was the **only** `features.*` key whose autodetected default was hardcoded `null` (`configure.js`), because there is no on-disk signal for per-tenant theming. `mergePreserving` fills it to `false` from the template, but the **non-interactive** configure branch strips every feature whose *detected* value is `null` (the "always-ask, never-assume" contract) — so a `--non-interactive` / `update --auto` / CI / auto-flow configure wrote a config **without** the key, and `update`'s drift detector (`update.js`) then re-flagged it as new on every subsequent run. The interactive path persisted it, so it appeared to "stick sometimes then forget" — a single auto flow in between re-dropped it.

- `src/commands/configure.js`: `detected.features.multi_tenant_theming` default `null → false`. Safe because **every** consumer (`ui-design`, `prd`, `bug`, `code-reviewer`, `/design-review`, …) gates on `features.multi_tenant_theming === true` — so `false` and absent are semantically identical (no skill prompts at use-time when it is absent), and the always-ask contract does not practically apply to this key. Now non-interactive configure WRITES `false` (matching the template), keeping it out of the null-strip, so the detector never re-flags it.

**PATCH** — one-line detection-default fix, no behavior change for skills (they check `=== true`), no new surface, no install-layout change. **Codex parity: portable** — `configure.js` is the shared CLI installer that writes `baldart.config.yml` for both runtimes; nothing runtime-specific. **Not a NEW `baldart.config.yml` key** (the key already exists end-to-end) → the schema-change propagation rule does NOT apply.

## [4.98.2] - 2026-07-02

**Negative guard: the codebase-understanding step is `codebase-architect`, not the generic built-in search agent.** A third live session — a data-wiring task (replace a hardcoded `"unità"` label with the real unit-of-measure in an order-wizard `ProductRow`) — delegated the code-path trace to the built-in `Explore` agent instead of `codebase-architect`, bypassing the profile engine, Reuse Analysis, Canonical Evidence, and the Return Contract. Diagnosed carefully (and deliberately *not* over-fitted): this was **not** a `PROFILE=ui` miss — the task's substance is data/logic wiring (product → UoM → loader → label), whose fitting profile is `feature`, not `ui`; the profile axis was left untouched. The real deviation was on the **agent-selection** axis, and it matches the recurring v4.97.1 shape: a positive mandate ("invoke `codebase-architect`") with no explicit "and NOT the generic explorer" guard (the existing `AGENTS.md` "never auto-route to a generic agent" rule is scoped to the *failure-fallback* case, not the initial choice).

- `AGENTS.md` primitive `1.7.1 → 1.7.2` (§ Non-negotiables): the understanding step goes through `codebase-architect` with the matching profile — **not the generic built-in search / `general-purpose` agent** (runtime-generic phrasing for cross-tool parity).
- `CLAUDE.md` primitive `1.4.1 → 1.4.2` (§ Subagents): names the Claude-native built-in explicitly — use `subagent_type: codebase-architect` + `PROFILE`, **not `Explore` / `general-purpose`**; `Explore` stays fine for a throwaway file-location grep, never for grounding.

**PATCH** — one-clause guard, no new surface, no config key, no install-layout change. The `ui`-vs-`feature` profile refinement (a UI-*framed* but data-*substance* task should route `feature`/`impact`, not `ui`) was **considered and deferred** to avoid churn on the just-stabilized v4.98.0/1 rule. **Codex parity: portable** — AGENTS.md carries the runtime-generic guard, CLAUDE.md names the built-in. **Not a `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [4.98.1] - 2026-07-02

**UI *verification* also routes `PROFILE=ui`, not just UI *change* — closing the one gap v4.98.0 missed.** A live post-update session confirmed v4.98.0 works where it matters: on the actual change ("migriamo") Claude correctly spawned `codebase-architect (PROFILE=ui)`. The only borderline case was the *preceding* verification ("mi verifichi che /products/new su desktop e tablet sia già sul nuovo page layout") — answered with an inline Grep. Root cause was in the v4.98.0 wording itself: the UI grounding mandate was scoped entirely to the verb "change/build", so a literal reader treated a read-only UI verification as outside it. Fixed by broadening "any UI **change**" to "any UI **work** — a change OR a verification/question whose correct answer needs design-system-spec or responsive/state grounding" across the same three surfaces:

- `AGENTS.md` primitive `1.7.0 → 1.7.1` (§ Non-negotiables + § Change Tiers carve-out + Light bullet), `CLAUDE.md` primitive `1.4.0 → 1.4.1` (Plan Mode paragraph).
- `agents/analysis-profiles.md`: no-token inference rule #1 now includes "verifying/inspecting the design-system adoption or responsive state of" a component/page/layout; § 3 routing row broadened to "ad-hoc UI change OR adoption/responsive verification outside a skill → `ui`".
- **A bare symbol-presence check** ("does file X import Y" — a pure Grep, no correctness/responsive judgment) is now explicitly exempt alongside the Trivial one-liner, so narrow lookups keep their fast path and the proportionality concern (no full architect spawn for a yes/no import check) is respected.

**PATCH** — same-day refinement of the v4.98.0 rule, no new surface, no config key, no install-layout change. **Codex parity: portable** — pure prose in the shared skeleton + the `agents/` module both runtimes read. **Not a `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [4.98.0] - 2026-07-02

**Codebase understanding is tier-orthogonal — and every UI change grounds via `PROFILE=ui`, even a "quick" one.** A live Claude session took a small ad-hoc UI request ("modifichiamo al volo questa pagina su desktop e tablet, serve il back button sull'header") and went straight to `find`/`grep`/`Read` instead of spawning `codebase-architect` with `PROFILE=ui` — whereas Codex, on the same framework, routes the profile correctly. An adversarial pass (3 skeptics) **refuted the obvious diagnosis** ("Codex reads a stricter doc than Claude"): `CLAUDE.md` states AGENTS.md is loaded automatically alongside it, so Claude reads the *same* unconditional `codebase-architect` MUST that Codex reads — this was a Claude-side *misreading* of an already-universal rule, not a weaker rule. The real drivers, confirmed with `path:line`:

- **`CLAUDE.md` conflated "Light skips plan mode / delegation" with "Light skips understanding."** The architect + `PROFILE` bullet lived *only* inside the "For a Heavy task:" block, so a Light-task reader treated grounding as Heavy-only machinery. Fixed: the Plan Mode section now states the invocation is **tier-orthogonal** and applies to Light too; UI → `PROFILE=ui` always; only a genuine Trivial one-liner skips it. (CLAUDE primitive `1.3.0 → 1.4.0`.)
- **`AGENTS.md` § Change Tiers never carved understanding out of the tier-gated "heavy machinery" list**, so it got swept into the Light exemption by association. Fixed: § Non-negotiables now says the architect MUST applies **regardless of tier** (UI → `PROFILE=ui`, Trivial one-liner exempt, inline Grep is not a substitute); § Change Tiers explicitly excludes understanding from the tier-gated machinery; the Light bullet reiterates it grounds first when the change needs real understanding. (AGENTS primitive `1.6.0 → 1.7.0`.)
- **The "which profile, when" mapping for ad-hoc UI requests lived only in `/new` `implement.md`**, so a request *outside any skill* had no routing rule. Fixed in `agents/analysis-profiles.md`: the no-token inference rule #1 now catches generic UI work (component / page / header / layout / styling / responsive), not just a named mockup; and the § 3 routing table gains explicit rows for **ad-hoc UI change outside a skill → `ui`** and ad-hoc non-UI grounding → `feature`.

The hook-based enforcement one skeptic proposed (a `PreToolUse` gate that blocks edits until architect ran) was **rejected**: it can't distinguish a typo from a UI change, would block Trivial one-liners and follow-up edits, and reproduces the already-refuted "do-X-every-time gate without a ground truth" antipattern (v4.78.2). The Trivial exemption is preserved throughout, so genuine one-liners keep their fast path — no latency regression.

**MINOR** — shipped-primitive behavior clarification (regenerated into consumers on `baldart update`); no new surface, no config key, no install-layout change. **Codex parity: portable** — pure prose in the shared skeleton + the bulk-symlinked `agents/` module both runtimes read; AGENTS.md is the cross-tool SSOT (already correct, now unambiguous), CLAUDE.md is the Claude-native stub where the regression lived. **Not a `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [4.97.1] - 2026-07-02

**`/ds-new` + `/ds-edit`: locate/reuse-first is a deterministic registry lookup, never a search agent.** A live `/ds-edit` run went off-script — instead of the skill's deterministic spine (resolve the component spec → `ui-expert` source edit → `extract-one` → `serialize-spec` → `doc-reviewer` → `ds-gate`), the model improvised ad-hoc `general-purpose` agents to "find" the component and verify a formatter. Root cause: the skills said *what* to do ("resolve the existing component") but never forbade *how not to*. Reaffirmed (adversarially) that the v4.94.0 `codebase-architect` analysis profiles correctly do NOT belong in the `ds*` skills — those operate on a single, already-named element (no scope→registry fan-out), and `ds-new`'s reuse-first is a **deterministic signature detector** that an LLM `ui`-profile pass would only make fuzzier and duplicate; the `ui` profile lives in the *planning* skills (`/prd`, `/ui-design`, `/new`, `/ds-handoff`, `context-primer`) and, when it matters, already ran upstream before `/prd` delegates to `ds-new`/`ds-edit`.

- **`ds-edit` P1** now states explicitly: read `components/<Name>.md` directly and take its `source` — a plain file resolution, **never** a `general-purpose` / `Explore` / `codebase-architect` spawn (grep the exported symbol if only the source path is uncertain). Skill `1.0.0 → 1.1.0`.
- **`ds-new` step 1** now states explicitly: reuse-first stays the deterministic `ds-reuse-gate.js → signatureDuplicates` / `baldart ds-gate` detector, **never** routed through an LLM search/architect agent. Skill `1.0.0 → 1.1.0`.

**PATCH** — skill-body clarification, no new surface, no install-layout change, no behavior change to the deterministic spine. **Codex parity: portable** — pure prose in the shared skill bodies (both skills already degrade `ui-expert`/`doc-reviewer` delegation to TODO on Codex; this hardening is runtime-agnostic). **Not a `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

## [4.97.0] - 2026-07-02

**Change Tiers — a real Light lane, so ordinary edits stop triggering the heavy machinery.** The root primitives forced plan mode on *any* non-trivial task, a `plan-auditor` + `doc-reviewer` review gate on *every* non-trivial plan, mandatory delegation to a domain agent, and (amplified by project overlays) a worktree + backlog card whenever more than a few files were touched — so even a small direct edit spun up the full orchestration. Root cause: there were only two tiers (Trivial vs everything-else-is-Heavy); no genuine middle lane.

- **New `## Change Tiers` section in the `AGENTS.md` primitive** (cross-tool SSOT, read by Claude + Codex) defines **Trivial / Light / Heavy**, and the *tier* — not raw file count — decides whether the heavy machinery fires. Default is **Light; when in doubt → Light**.
- **Light** (the new default for most work): any change that hits **no Heavy trigger**, even multi-file within one coherent area → implemented **directly on the current branch**, no plan mode, no plan-review, no worktree, no card.
- **Heavy** (aggressive calibration): escalate ONLY for a new DB schema/migration, a new external dependency, a new API contract/route or domain entity, a wide/cross-module refactor, or work parallel to other sessions on shared files → full machinery.
- The plan-review MUST, the delegation MUST, and the Git Workflow worktree rule in `AGENTS.md` are now scoped to **Heavy**; the `CLAUDE.md` primitive Plan Mode section likewise fires only for Heavy (Light/Trivial skip plan mode + review + mandatory delegation). Primitives bumped: `AGENTS.md` 1.5.0 → **1.6.0**, `CLAUDE.md` 1.2.0 → **1.3.0**.

**MINOR** — behavior/posture change to shipped primitives, no install-layout change. **Codex parity: portable** — pure prose in the shared `AGENTS.md` skeleton; Codex reads the same tiers to decide agent-spawn / worktree / card. **Not a `baldart.config.yml` key** → the schema-change propagation rule does NOT apply. Consumers pick it up on `baldart update` (regenerated root primitives); project overlays may add project-specific Heavy triggers.

## [4.96.0] - 2026-07-02

**Protected branch decoupled from the integration trunk — `git.protected_branch`.** The root `AGENTS.md` primitive expressed the "NEVER push directly to X" MUST via `{{ git.trunk_branch }}`, conflating two distinct roles that in a two-branch gitflow are different branches: the *integration trunk* (worktree base + feature PR target in `/prd` + `/new`) and the *protected branch* (production/release line that must not receive direct pushes). A consumer with `trunk_branch: develop` therefore had its agents (Claude + Codex) refuse a direct `develop` push — even though `develop` was the integration branch they push to and `main` was the truly-protected line. `trunk_branch` could not simply be flipped to `main` because it also drives worktree creation (`update.js`), which must branch from `develop`.

- **New optional config key `git.protected_branch`** — the branch the "never push directly" rule protects. **Defaults to `git.trunk_branch`** when omitted, so single-trunk repos are byte-unchanged; a two-branch gitflow sets it explicitly (e.g. `protected_branch: main` while `trunk_branch: develop` stays the worktree base + PR target).
- **Primitive `AGENTS.md` skeleton → 1.5.0**: the push MUST now resolves `{{ git.protected_branch }}` and explicitly permits direct pushes to other branches (including the integration trunk when it differs) when the owner authorizes it.
- **Slot resolution** (`src/utils/root-primitives.js`): `git.protected_branch` resolves to `git.protected_branch || git.trunk_branch || 'main'`.
- **Schema-change propagation** (per the rule): template (`baldart.config.template.yml` — new commented key under `git:`), `configure` (new `detectProtectedBranch` probe + serialized `git.protected_branch` + `_probes`), `update` detector (already diffs all `git.*` keys → surfaces `git.protected_branch`; added a dedicated advisory), CHANGELOG. No doctor diagnostic needed (the update advisory + configure autodetect cover it).

**MINOR** — new optional config key, back-compat default, no install-layout change. **Codex parity: portable** — pure slot resolution in the shared `AGENTS.md` skeleton (transpiled to `.codex/agents/*` unchanged; the generated `AGENTS.md` is the cross-tool SSOT read by both runtimes).

## [4.95.0] - 2026-07-02

**`/prd` — applier economico + EARS+verify (i due interventi strutturali dal forensics della deep analysis).**

- **Apply dei finding delegato** (`prd-card-writer` `MODE: apply-findings` — un agente + token, pattern analysis-profiles v4.94.0): l'orchestratore `/prd` non edita più YAML a piena finestra (misurato: le fasi meccaniche erano il 48% del costo di un run reale, ~48 Edit a ~663k cache-read/chiamata). L'applier gira in contesto fresco (Claude: override `model: sonnet`; Codex: by-name), riceve tutto BY PATH, applica solo campi `[Target:]` con grounding obbligatorio, ESEGUE `validate-card-baseline.js --prd` + `stamp-holistic-audit.js`, ritorna compact; `audit-phase.md` 6.9 = delega thin + drift guard + fallback inline visibile. Verificato con spawn reale su fixture.
- **AC in grammatica EARS** (child/standalone; SSOT `card-schema.md` § "Acceptance-criteria grammar"): binari per forma (5 pattern + `SHALL CONTINUE TO`), ban-list verbi vaghi, BLOCKER in `--prd`; card legacy mai retro-bloccate.
- **`ac_verification`** — oracolo eseguibile per-AC definito a planning time (o `manual:` esplicito), authoring groundato (mai comandi inventati), eseguito da `/new` Phase 2.5 (step 4b — exit status = evidenza primaria); `qa-sentinel` intatto per contratto. Enforcement WARN-first; fixtures CI (Check D).

**MINOR** — capability su skill/agent/moduli condivisi; nessun cambio di layout install. **No new `baldart.config.yml` key**; la propagazione dello schema CARD viaggia completa nella release (matrice + template + writer + validator + consumer + attack-surface — pattern v4.35.0). **Codex parity: portable.**

## [4.94.0] - 2026-07-02

**Analysis profiles — `codebase-architect` specializzato per-tipo-di-analisi, un solo agente, un solo SSOT.** L'assessment dei ~50 callsite reali ha mostrato che "capire il codebase" copre 6 lavori diversi (grounding pre-implementazione, tracing di un bug, blast-radius, contesto di planning, inventario UI/design-system, baseline di review) serviti da UN prompt monolitico (~455 righe caricate a ogni spawn) con la conoscenza di specializzazione dispersa nei prompt dei chiamanti (dual-SSOT: il template di `context-primer` ri-dichiarava la search strategy dell'agente). Fix architetturale, stesso meccanismo di `OUTPUT=terse`:

- **Nuovo modulo SSOT [`framework/agents/analysis-profiles.md`](framework/agents/analysis-profiles.md)** — il gemello retrieval-side di `effort-protocol` + `return-contract-protocol`: contratto del token `PROFILE=<feature|bug|impact|discovery|ui|baseline>`, piano di retrieval + skip-list + budget + output-block per profilo, inference deterministica quando il token manca (fallback `full` = comportamento legacy), tabella di routing normativa per i chiamanti, shared block (Reuse Analysis · UI Resolution Path · Documentation Reliability Scan) estratti dal corpo agente.
- **`codebase-architect` slim + dispatcher**: nuova sezione "Analysis Profile Dispatch" (parse token → Grep `### PROFILE: <x>` → Read SOLO quella sezione); Reuse Analysis / cascade UI / Reliability Scan spostati nel modulo (il corpo tiene l'engine: router canonico, tier strutturale, read-size, Hot-File Map v4.91.0, memoria); dedup della sezione "Your Approach". Il boundary multimodale è ESPLICITO: pixel/mockup/screenshot → `ui-expert` read-only (precedente `/prd` 1.6.5), mai l'architect.
- **Chiamanti migrati al token deterministico** (mai lasciato all'inferenza dell'agente): `/new` implement.md step 3 + `new2.js` B7 → `ui` se la card ha mockup (stesso bit `hasMockup` del routing build v4.89.0) else `feature`; `new-final-review.js`/`new-card-review.js`/`/codexreview` → `baseline`; `plan-auditor` FULL → `impact`; `/prd` discovery-loop → `discovery`, ISA → `impact`; `/prd-add` PATCH → `impact`; `/bug` → `bug`; `/ui-design`/`/ds-handoff` → `ui`; `/e2e-review` reverse-lookup → `impact`; `context-primer` 1.1.0 mappa i task-type sui profili e il suo template smette di ri-dichiarare la strategia (dual-SSOT chiuso).
- **Primitivi root aggiornati**: skeleton `AGENTS.md` 1.4.0 (il MUST understand-before-implement nomina i 6 profili e cita il modulo) + `CLAUDE.md` 1.2.0 (Plan Mode passa il token, cita senza ripetere).
- **Tripwire CI**: nuova regola `analysis-profile contract` in `scripts/check-new-parity.mjs` (modulo + dispatcher + implement.md ↔ new2.js ↔ review workflow ↔ context-primer come obbligazioni a locazione parallela).
- Routing aggiornato: `REGISTRY.md` (riga architect + decision tree), `agents/index.md` (regola + manifest), `skills-mapping.md` (context-primer = canale, non metodologia). Skill bumped: new 1.4.0, new2 1.3.0, prd 1.3.0, prd-add 1.1.0, bug 1.1.0, ui-design 1.1.0, ds-handoff 1.2.0, e2e-review 1.3.0, context-primer 1.1.0.

**MINOR** — nuova capability su agent/skill/protocol condivisi; nessun cambio di layout install. **No new `baldart.config.yml` key** (i profili gate-ano su flag esistenti) → la schema-change propagation rule NON si applica. **Codex parity: portable** (token prompt-level; il modulo viaggia nel payload `agents/` bulk-symlinkato letto da entrambi i runtime; il corpo agente transpila in TOML invariato).

## [4.93.0] - 2026-07-02

**`/prd` gate-hygiene wave — dai driver misurati (forensics di una sessione reale: 61.6M token, 48% del costo in fasi meccaniche, una famiglia di card con zero audit-trail, 96% del wall-clock in attesa umana).** Sette interventi chirurgici, nessuna nuova chiave di config:

- **Tool di misura onesto**: `scripts/new-session-audit.mjs` deduplica l'usage per `message.id` (sovracontava ~2.85× — 175M riportati vs 61.6M reali su FEAT-0061) e riconosce le sessioni `/prd` (`mode=prd`, anche quando la skill è invocata via Skill tool dopo un preambolo; backtrack `self`; confronto col reference FEAT-0041 saltato).
- **Severità assoluta evidence-anchored** (`prd` skill 1.2.0 + `plan-auditor`): via la quota posizionale "top 20% → HIGH must-apply" — fabbricava cicli di apply su piani puliti e demotava finding reali sui batch grandi. Zero-finding è un esito legittimo.
- **Validation eseguita, non recitata**: `validate-card-baseline.js --prd` (requirements condizionale PO1, epic `skip`, marker `[NEEDS CLARIFICATION]` bloccanti) + nuovo `framework/scripts/stamp-holistic-audit.js` (stamp `metadata.holistic_audit` deterministico, idempotente; chiamato da 6.9.4 e dal backstop 5b — i due gate-prosa che erano falliti insieme su FEAT-0064).
- **Context economy dell'orchestratore /prd**: contratto teammate audit estratto in `references/audit-teammate-prompt.md` e passato BY PATH (mai caricato dall'orchestratore; ~−9KB da audit-phase + regola generale per gli spec agent-facing incl. `api-perf-gate.md`); analisi mockup su file delegata a `ui-expert` read-only (misurati ~208KB di mockup in cache per ~300 eventi); Progress Bar solo ai phase-gate con riga compatta intra-fase (~20 righe × 30+ messaggi risparmiate; unifica la policy Codex già esistente).

**MINOR** — capability change su skill/agent/script condivisi; nessun cambio di layout install. **No new `baldart.config.yml` key** → la schema-change propagation rule NON si applica. **Codex parity: portable** (script node zero-dep; prosa skill condivisa ai due runtime; il branch file-backed Codex del teammate contract è documentato in-file).

## [4.92.1] - 2026-07-02

**Review avversariale post-release del motore new2 (coerenza interna).** Tre reviewer indipendenti a contesto fresco (dataflow / edge-case / contratti) sull'onda v4.90–4.92; ogni finding ri-verificato contro il codice prima di agire: **11 confermati e fixati, 9 refutati con evidenza empirica** (tra i refutati: la "race" del worker-pool adattivo — simulata: 0 buchi; `verify-item.sh` su bash 3.2 reale — funziona; `parseInt` con CRLF; un CAP attribuito a `new2.js` che non esiste).

**PATCH** — hardening senza cambi di superficie. **Nessuna nuova chiave config**. **Codex parity: portable/N/A** (prosa + workflow JS Claude-only).

### Fixed

- **`new2.js`**: verifier strutturale che si auto-dichiara `status:"skipped"` = UNDERRUN (mai PASS — era l'ultimo canale di under-run mascherato); **re-verify deterministica** una-tantum dopo un fix strutturale risolto (il self-report del fixer non è la conferma del verifier); **`bindingCheck` consumato** (violazioni sopravvissute al pass dell'agent → resolve ui; report assente → ledger fail-open, mai silente); `${MAIN}` quotato dentro `$(git -C …)` (path con spazi).
- **`/e2e-review` 1.2.1**: stringa lane canonica `skipped:render_unavailable_structural_covered` scritta esplicitamente da Phase 3.8 (l'assertion di Phase 5 valida l'enum chiuso); chiave di convergenza del self-heal estesa con `viewport_role`.
- **`new2` 1.2.1**: Step 3.5 ramo AUTONOMOUS (CI/`BALDART_AUTONOMOUS`/origine delegata → mai una domanda, mai un apply command autonomo); **origine delegata = autonoma ovunque** (escape hatch 3b SKIPPED anche con TTY — chiuso il rischio di ciclo `/new`→`new2`→`/new` a costo pieno); contatore resume **durabile nel batch tracker** (una compaction non resetta il bound) + teardown broker Codex su ogni uscita per stallo.
- **`structural-compare.mjs`**: arbitrary value Tailwind con funzioni (`grid-cols-[repeat(2,_minmax(0,1fr))]`) contati con la `countTracks` parens-aware (era `split('_')` → 3 invece di 2, falso segnale di divergenza).
- **`ui-implement` 1.1.1**: Return Contract documenta i campi `bindings`/`lane_coverage`/`lanes` (schema SSOT in integration.md).

## [4.92.0] - 2026-07-02

**Tier 3 del programma "deep analysis /new+/ui-implement": parity tripwire, bound dei loop, verifica payload new2.** Chiude il programma a 3 tier (4.90.0 → 4.92.0).

**MINOR** — nuovo guard CI + bound espliciti. **Nessuna nuova chiave `baldart.config.yml`**.

**Codex parity: N/A / portable** (script maintainer + prose skill; nessuna superficie runtime nuova).

### Added — T3.A Parity tripwire /new ↔ new2

- **`scripts/check-new-parity.mjs`** (maintainer, CI-hard in `check-reference-integrity.yml`): 9 policy a location parallela codificate come obbligazioni (file, token) che devono valere su TUTTI i lati — mockup-first, router clamp, binding gate, lane strutturale, Verification-Lane Ledger, path tracker durabile (con FORBIDDEN sul path /tmp), verify-item runner, cap adattivo, delegation gate. La classe di difetto "fix propagato a metà" (v4.89.0, B4, singleCard…) ora fallisce la build invece di emergere in un post-mortem.

### Added — T3.C Loop bounds

- **`new2` 1.2.0**: resume loop HARD-bounded (`resumeMaxAttempts = 3` + stall detection sullo stesso set di card incomplete).
- **`/e2e-review` 1.2.0**: convergence detection nel self-heal (fix che genera finding gating NUOVI → stop immediato + override path, `convergence: "diverging"`); spawn del solo fixer-bucket non vuoto. Deliberatamente NIENTE re-verify selettiva delle lane (sarebbe un under-run auto-inflitto contro la Lane-Ledger assertion).

### Verified — T3.B new2 prompt payload (measure-first, esito documentato)

- La claim "i prompt di new2 embeddano 40–80KB di moduli per agente" è **REFUTATA alla verifica**: i prompt passano i moduli **per path** (`per ${REF}/implement.md Phase 2` — l'agent li legge da disco) e l'audit dei run reali mostra i moduli letti 1× ciascuno. Nessun cambio necessario; il "briefing-extract" resta non implementato di proposito (sarebbe un secondo SSOT distillato).

## [4.91.0] - 2026-07-02

**Tier 2 del programma "deep analysis /new+/ui-implement": hot-file economy, right-sizing del review fan-out, robustezza (B4/NUL/verify-item).** Interventi mirati dai numeri delle stesse 3 run reali del Tier 1.

**MINOR** — nuove capacità (script, sezioni protocollo), zero breaking. **Nessuna nuova chiave `baldart.config.yml`**.

**Codex parity: portable** (prosa protocollo/agent transpilata + script bash/zero-dep; i due workflow JS toccati sono Claude-only by design, path inline invariato).

### Added — T2.A Hot-file economy

- **`code-search-protocol.md` § "Large-file read discipline"** (SSOT): file > 800 righe = hot file; misurato un client da 2110 righe (~29k tok) riletto INTERO **28×** in un run (i moduli framework si leggono 1× — il repeated-read vero è questo). Regole: Hot-File Map → range-read; brief per path, mai body.
- **`codebase-architect`**: il contratto terse emette `## Hot-File Map` MANDATORY (`path (N lines): symbol — Lstart–Lend`) per ogni file in scope > 800 righe; `coder`/`ui-expert` citano la disciplina; `implement.md` Step 7 passa il baseline per PATH (guardia `[BASELINE-READ-FAIL]`).

### Added — T2.B Review fan-out right-sizing

- **Empty-scope pre-invocation guard** (Phase 2.5x + team-mode D.1.6): card/wave a diff vuoto non schedulano un cluster `new-card-review` no-op (misurati 2 cluster `agents=0, dur=0` in un batch reale).
- **Cap adattivo** in `new-card-review.js` + `new-final-review.js`: fisso 3 → parte 5, **scende a 2 al primo transient rate-limit** per il resto del run (rispetta il guardrail v4.53.6/v4.66.0 eliminando la queue-wait tax sui run senza pressione). Deliberatamente NON toccata la policy di profondità review (refutazioni slimCodex / review-gate additions).

### Fixed — T2.C Robustezza

- **B4 tracker LOST — propagazione completata**: il fix del path durabile (`$(git rev-parse --git-common-dir)/baldart/run/…`) era atterrato solo in `new/SKILL.md`+`metrics.md`; **`new2.js` creava ancora il tracker in `/tmp`** (la causa dei FAIL B4 nei run di audit) e i prompt fixer di `/e2e-review` citavano il path volatile. Entrambe le location parallele allineate.
- **Sweep anti-NUL**: sezione "Search discipline (MUST)" (prefer `rg`; 0-match ≠ assente) aggiunta ai 5 reviewer che ne erano privi — `qa-sentinel`, `code-reviewer`, `security-reviewer`, `markup-fidelity-verifier`, `api-perf-cost-auditor` (84–149 sintomi binary-grep per run nonostante il fix B3 su coder/architect).
- **`scripts/verify-item.sh` (skill /new)**: la Grep-verification discipline del completeness è ora CODE — pattern come singolo argv `-F` (metacaratteri innocui), exit 2 = `UNVERIFIED` mai `Missing`, rg NUL-safe con fallback `grep -a`. Testato incluso il file-con-NUL (la trappola FEAT-0041 B3). Hand-rolled shell per quei check = protocol violation.
- Versioni skill: `new` 1.3.0 · `new2` 1.1.1 · `e2e-review` 1.1.1.

## [4.90.0] - 2026-07-02

**Tier 1 del programma "deep analysis /new+/ui-implement" (lug-2026): fidelity chain a prova di under-run, build-with-eyes, e la delega off-context del batch `-auto`.** Fondato sui dati di 3 run reali (mayo `a73a52d0`/`a946415b`/`6882a2b7`): (a) il 53–59% dei token fatturati è la turn-tax dell'orchestrator inline; (b) nel run FEAT-0064 la catena di fedeltà è stata SOTTO-eseguita in silenzio (zero `visual-fidelity-verifier`/`ui-quality-critic` spawnati) e l'escape è diventata BUG-0017 lavorata a mano; (c) i moduli reference vengono letti 1× (economy OK) — il costo ricorrente è altrove.

**MINOR** — nuove capacità (gate, script, delega), zero breaking. **Nessuna nuova chiave `baldart.config.yml`** → la schema-change propagation rule NON si applica.

**Codex parity: portable / inline-fallback** — script zero-dep + prosa skill/agent (trasp. TOML) portabili; il delegation gate e la lane new2 sono Claude-only by design con il path classico inline come fallback documentato (ogni run Codex è invariata).

### Added — T1.A Fidelity chain (verification-lane ledger)

- **`design-system-protocol.md` § 6 "Verification-Lane Ledger"** (nuovo check SSOT): riconciliazione planned-vs-executed delle lane di verifica; una lane pianificata mai eseguita = **`verification-coverage-gap`** (Critical, unfixable, bypassa il self-heal, → `blocked`). Nessun verdetto self-served "accepted (harness limit)".
- **`/e2e-review` 1.1.0**: Lane Matrix (chiusura Phase 2) + assertion Phase 5 step 0; bundle HTML Claude Design SEMPRE sondato con `extract-mockup-design.mjs` e threaded come `structural_source` (decode fallito ≠ lane persa); viewport **tablet @768px** scope-driven (`plan.tablet_in_scope` → cattura + pass 2c `viewport_role: tablet`); report JSON additivo `lanes`/`lane_coverage`.
- **`/ui-implement` 1.1.0 — binding gate**: nuovo script deterministico `skills/ui-implement/scripts/check-bindings.mjs` (Step 5b) — binding `reuse`/`reuse-variant` non usato nel codice = `binding-not-honored` (Critical, 1 retry, poi blocked). Fail-open.
- **`new2.js` (location parallela)**: lane strutturale route-free (`markup-fidelity-verifier`) per card con mockup + obbligo Lane-Ledger (verifier assente → residual `verification-coverage-gap`, card mai auto-DONE); binding gate nel brief; cardGraph `designSrcDir`/`designHtml`/`hasBindings`.
- `visual-fidelity-verifier`: `viewport_role: "tablet"` (responsive-integrity + wrong-band collapse con cite dell'AC).

### Added — T1.B Build-with-eyes (costruisci-bene a monte)

- **`framework/scripts/structural-compare.mjs`**: pre-gate strutturale deterministico (grid cols per breakpoint / responsive machinery / landmark, confidence graduata). Semantica onesta: SEGNALI, mai verdetti (le primitive composte sono invisibili al regex) — hint `deterministic_signals` che il `markup-fidelity-verifier` conferma o dismette esplicitamente; un `match` non autorizza mai lo skip del verifier.
- **`ui-expert.md` § "Build-with-eyes"**: self-check in-build bounded (max 2 iter) — strutturale via script + visivo con `npx playwright screenshot` e Read multimodale del mockup; degrado onesto e SEMPRE riportato. Il trio verifier resta il giudice indipendente.
- **`/e2e-review`**: validazione dimensioni viewport prima del pixel-diff (chiude il falso-negativo cross-width); **quality pre-filter skip** (design system + structural pulita + pixel-diff sotto threshold → `ui-quality-critic` saltato come skip autorizzato `quality_prefilter_skip`; mai su `harness-render`/no-mockup).

### Added — T1.C Turn-tax (delega off-context)

- **`/new` 1.2.0 — AUTONOMOUS OFF-CONTEXT DELEGATION GATE** (kickoff, pre-Phase-0): un batch `-auto` è delegato di default al motore `new2` (Workflow tool + `new2.js` + skill presenti, batch migration-free, non `-auto-ship`). Delega MANDATORY a condizioni verdi; escape hatch **`-inline`**; il classico resta SSOT (interattivo / Codex / `-auto-ship` / migrazioni).
- **`new2` 1.1.0**: promossa a motore di default di `/new -auto` (Step 1.5 = contratto della delega + handback su migrazione dichiarata); `docs/WORKFLOWS.md` riframato ("new2 is the engine, /new is the SSOT").

## [4.89.2] - 2026-07-01

**`baldart update` auto-commit now covers the FULL payload — Codex artifacts included.** A consumer reported that `baldart update` "never commits anything" even when authorized, and that other clones keep seeing the same files untracked. Root cause: `postUpdateAutoCommit` (`src/commands/update.js`) auto-commits only paths matching `BALDART_MANAGED_PATTERNS`; everything else is classified `userOwned` and left in the tree forever. The allowlist had drifted behind the payloads BALDART actually writes — it was missing every Codex-native artifact (`.codex/agents/*.toml` transpiled agents v4.80.0, `.codex/hooks.json` v4.88.0, `.agents/skills/*` skill symlinks v3.7.0), `.claude/output-styles/` (v4.50.0), and the i18n gate files (`eslint.i18n.config.mjs` / `.eslintrc.i18n.json` / the `paths.i18n_registry` directory, v4.52.0). On a Codex-enabled consumer the 32 `.toml` files are regenerated as real files on every update, classified `userOwned`, never committed → churn in every other clone.

**PATCH** — bugfix, no install-layout or command-surface change; existing consumers get correct auto-commit on their next `update`. **No new `baldart.config.yml` key** (the i18n directory is *derived from* the existing `paths.i18n_registry`) → the schema-change propagation rule does NOT apply.

**Codex parity: adapter-generated / N/A** — the fix is precisely the parity closure (the Codex-written payloads are now committed exactly like the Claude ones); the single classifier serves both runtimes.

### Fixed

- **`BALDART_MANAGED_PATTERNS` extended** to the full written payload: `.codex/` (agents + hooks), `.agents/skills/`, `.claude/output-styles/`, `eslint.i18n.config.mjs`, `.eslintrc.i18n.json`. A prominent INVARIANT comment now warns that any new written payload MUST add its pattern here in the same change (parity twin of the schema-change propagation rule).
- **New `i18nRegistryPatterns(cwd)`** derives the i18n context-registry file + its containing directory from `paths.i18n_registry` (fallback `docs/i18n/registry.yml`), so the configurable path is auto-committed without hardcoding. `isBaldartManagedPath(p, extra)` now accepts these config-derived patterns.
- **`.claude/settings.local.json` stays excluded** by design (personal, gitignored — the output-style activation target), so widening the `.claude/*` group to `output-styles` does not sweep in personal settings.

## [4.89.1] - 2026-07-01

**`/new` tracker-write economy — remove the phase-entry write prompts.** An audit (mayo FEAT-0064) measured **20 orchestrator turns whose only tool call was a tracker `Edit`** (~5-11M tokens of pure `cache_read` replay, ~3-6% of the run). The § "Context economy" rule already forbids lone tracker-`Edit` turns (and the metric is trending down: 38 → 28 → 20), so this is a compliance gap, not a missing rule. The non-redundant fix: the per-phase reference modules carried ~9 numbered `**Update tracker**: phase = "X"` **phase-entry** steps (no data) that actively prompt a standalone write, contradicting the global co-emit rule. Removing that prompt at the source is more effective than more prose.

**PATCH** — context-economy refinement, no functional behavior change. `/new`-only (`new2` writes its tracker via the JS `ledger()` helper, which costs no orchestrator turn — no parallel change). **No new `baldart.config.yml` key**.

### Changed

- **Phase-entry markers removed** across `implement.md` / `review-cycle.md` / `completeness.md` / `commit.md`: the standalone `phase = "1-claim|2-implement|2.55-simplify|2.6-e2e-review|3-doc-review|3.5-qa|2.5-completeness|2.5b-ac-closure|4-commit"` steps are reworded to "**Enter phase X — no standalone tracker write**" pointing at the allowlist. The card-claim (moves the card to `## Current Card`) and the Phase-4 entry-assertion are preserved but explicitly co-emitted. Step numbers are unchanged (no renumbering / no broken cross-refs).
- **Data-carrying DONE-markers kept** (findings / retry counts / `qa_first_attempt` / AC-closure ledger — Phase-8 telemetry producers), now governed by the allowlist to co-emit with the phase's last tool call.
- **`SKILL.md` § Context Tracking** gains an explicit **tracker-write allowlist**: the ONLY writes are `phase_module_loaded`, a data-carrying DONE-marker, a commit record, a gate verdict, and a blocker/flag — all co-emitted; there is NO "entering phase X" write. Recovery is unchanged (§ "Context recovery protocol" reconstructs the phase from commit + `phase_module_loaded` + backlog).

## [4.89.0] - 2026-07-01

**Mockup-first build routing — a card with a mockup ALWAYS builds via `/ui-implement` → `ui-expert`.** A `/new` post-mortem (mayo FEAT-0064) found two UI cards with a `links.design_src` mockup built by `coder` instead of `ui-expert`, causing layout drift (caught by the v4.78.0 fidelity gate) + three avoidable rework passes + a full `balanced` review (~174M tokens). Root cause: the v4.82.0 `/ui-implement` delegation gate was positioned **inside** `implement.md` step 7's briefing, structurally **subordinate** to the `owner_agent` router (step 6b) — so a card whose `owner_agent` is `coder` (which `prd-card-writer` Rule B correctly assigns to a **mixed UI+logic** card, since `ui-expert` has UI-only scope) resolved `spawn=coder` and the delegation note never fired. `new2.js` was worse: it dispatched `owner_agent` with no delegation gate at all.

**MINOR** — routing behavior change, backwards-compatible (no-mockup cards unchanged; the coarse `ui-expert` inline fallback is preserved for Codex / older installs). **No new `baldart.config.yml` key** (rides on the existing `/ui-implement` + `features.has_design_system`) → the schema-change propagation rule does NOT apply.

### Changed

- **`implement.md` step 6a (NEW — PRIMARY, precedes the `owner_agent` router)**: a card with `links.design` / `links.design_src` has its BUILD routed **mockup-first** to `/ui-implement` (coarse fallback: force `ui-expert` inline + Phase 2.6), *regardless of `owner_agent`*. The epic guard runs first inside 6a. Step 6b (`owner_agent` router) is now explicitly scoped to cards **without** a mockup; the delegation-gate prose is **retired** from step 7's "### Design Reference" (it lives once, at 6a). `owner_agent` still governs review/ownership semantics.
- **`new2.js`**: `cardGraph` gains `hasMockup`; the deterministic `OWNER_SPAWN` clamp forces `spawn=ui-expert` (ledger `CLAMPED-TO-UI (mockup-first)`) for a mockup card. new2 keeps its **inline** `ui-expert` build this release (full `/ui-implement` off-context delegation remains the documented follow-up).
- **`/new` SKILL.md § Agent Routing** + **`prd-card-writer.md` Rule A** + **`design-system-protocol.md`** § "Route-Independent Fidelity" — one aligned note each: the build-agent selection is mockup-first; `owner_agent: coder` on a mixed card means coder owns review/logic semantics, never the UI build. Cited from the single SSOT (implement.md step 6a), not duplicated.

## [4.88.3] - 2026-07-01

**Context-economy trim of the Codex runtime-portability citations.** The top-level RUNTIME PORTABILITY note in `/new` and the HARD-RULES Runtime-portability note in `/prd` (added in v4.88.2's S4/S5 wave) restated content that already lives in `framework/agents/runtime-portability-protocol.md`. Trimmed both to a one-liner (cite, don't restate) — ~406 chars off `/new` SKILL.md and ~363 off `/prd` SKILL.md, i.e. the core each skill drags in context every turn after activation. **No behavior change**: the module citation, the "don't probe `Workflow`/`AskUserQuestion`/the task spine on Codex" rule, and the ADDITIVE guarantee are preserved; RULE 4/11 Codex branches are untouched; the parity guard (`scripts/check-codex-parity.js`) passes. **PATCH.** Skills `new`/`prd` → 1.1.2. No new config key.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** + **`framework/.claude/skills/prd/SKILL.md`** — runtime-portability citations trimmed to one-liners (skill version 1.1.1 → 1.1.2).

## [4.88.2] - 2026-07-01

**`/prd` reliably stamps `holistic_audit` provenance → `/new`/`new2` stop spawning the plan-auditor on every freshly-authored card.** The v4.7.0 skip already existed: `/new` (`implement.md` Phase 1 step 4) skips the per-card plan-auditor grounding when **P1** (the card carries `metadata.holistic_audit`) **∧ P2** (no git drift on its claimed paths since the audit). But P1 was frequently false: the stamp was written by a **soft, model-driven per-card step** (`audit-phase.md` Step 6.9.4) framed as an "optional optimization hint" — the single most-omitted step under context pressure. When omitted, every card ships without the stamp and `/new`/`new2` fall back to spawning the plan-auditor on **every** card even when the batch is implemented immediately after `/prd` (nothing changed since the audit) — the reported waste. Diagnosis was confirmed on a real run: the FIRST card of a batch (HEAD at trunk, zero intra-batch commits ⇒ P2 trivially holds) still ran the plan-auditor, which is only possible if P1 failed. Fix strengthens the **existing** provenance (no new "freshness" signal, per the chosen direction):

- **`audit-phase.md` Step 6.9.4 + Fail-safe contract** — reworded from "optional hint" to **mandatory whenever the audit ran**. If git is genuinely unavailable, write `audited_commit: ""` **explicitly** (never omit the block). The stamp remains an OPTIMIZATION hint for `/new`, never a *correctness* gate.
- **`validation-phase.md` Step 7 item 5b (new) — deterministic backstop.** Right before commit staging (Step 6 complete ⇒ the audit genuinely ran, so the stamp legitimately certifies it), recompute the three run-level values ONCE and backfill `metadata.holistic_audit` on every **non-EPIC** Step-5 card that lacks a non-empty `audited_commit`. Idempotent (a card already stamped at 6.9.4 is untouched), additive (never overwrites other `metadata` keys), EPIC-excluded (`card-schema.md`). The backfilled YAMLs are staged by item 6's existing `FEAT-XXXX-*.yml` glob. Blocks only *silently shipping cards without the stamp*, **never** the PRD.
- **`implement.md` Phase 1 step 4** — a clarifying note (no logic change): the skip is now the NORMAL path for a card implemented right after `/prd`, and P1's fail-safe (missing provenance / drift / git error → RUN the plan-auditor) is unchanged. In autonomous mode the skip already auto-applies (it is not gated on interactivity), matching the requested behavior.

**Correctness guard:** the stamp also lets `/new` Step 3d skip the Codex cross-card check, so it must vouch only for an audit that actually ran — the backstop lives *after* the Step 6 gate, so it never over-claims. A `/prd` run that skips the audit writes no stamp → `/new` correctly runs the plan-auditor. **PATCH** — robustness fix to an existing mechanism, no new capability surface, no new `baldart.config.yml` key (→ schema-change propagation rule N/A). **Codex parity: portable** — the stamp is YAML + git, and `audit-phase.md` / `validation-phase.md` are runtime-neutral prose consumed by both the Claude and Codex `/prd` paths; the shared Step 7 commit/merge covers both. Per-skill: `prd` → 1.1.1, `new` → 1.1.1.

### Changed

- **`framework/.claude/skills/prd/references/audit-phase.md`** — Step 6.9.4 + Fail-safe contract: stamp write is mandatory-when-the-audit-ran, not optional.
- **`framework/.claude/skills/prd/references/validation-phase.md`** — new Step 7 item 5b: deterministic `holistic_audit` provenance backstop before commit staging.
- **`framework/.claude/skills/new/references/implement.md`** — Phase 1 step 4: clarifying note that the provenance skip is the normal fresh-batch path (no logic change).

## [4.88.1] - 2026-07-01

**`baldart update` backfills hooks on the aligned/self-heal path too.** The "already up to date" branch of `update` reconciled newly-shipped payload (agents / skills / workflows / output-styles + root primitives) but did **not** backfill hooks. Same gap class the payload reconcile already fixed: a consumer who framework-updated with an **older CLI** (before a given hook shipped — e.g. the v4.88.0 Codex hooks) and then aligned the CLI would hit the "already current" early-exit and **never get the hooks registered** (recoverable only via `doctor`). Now the aligned path registers both Claude (`.claude/settings.json`) and — when `tools.enabled ∋ codex` — Codex (`.codex/hooks.json`) hooks, **idempotently and additively**: SILENT when nothing is missing, drift left to `doctor` (never prompts on the self-heal path), non-fatal on any hiccup. The `--json` payload gains `backfilled_hooks`. **PATCH** — robustness fix, no behavior change on a healthy install. Codex parity: adapter-generated (the same registrar). No new config key.

### Changed

- **`src/commands/update.js`** — the aligned "already up to date" branch now backfills missing Claude + Codex hooks (silent when none missing) and reports `backfilled_hooks` in `--json`.

## [4.88.0] - 2026-07-01

**Native Codex hooks — BALDART now manages `.codex/hooks.json` in parallel with `.claude/settings.json` (Codex-parity foundation, S1–S3).** Kicks off a phased program to make BALDART **universally tool-agnostic**: everything it ships must work natively on both Claude Code and Codex. Two governance rules are now binding on all future development (CLAUDE.md § "Cross-tool parity discipline"): (1) every NEW tool/surface is first checked for native Codex compatibility (portable / adapter-generated / inline-fallback / N/A, outcome recorded in the CHANGELOG); (2) every change travels in parallel across Codex and Claude in the same change. This release delivers the **hook foundation**:

- **S1 — hook registry split (no behavior change).** `src/utils/hooks.js` is split into a runtime-neutral registry (`hooks/registry.js`, one entry per logical hook with per-runtime bindings) + a Claude renderer (`hooks/claude.js`, `.claude/settings.json`); `hooks.js` becomes a thin compatibility shim. Verified **byte-identical** to the pre-split module across 7 fixtures (getStatus / verifyAll / registerAll keep+replace / unregisterAll + the written settings.json).
- **S2 — Codex hook registrar.** New `hooks/codex.js` renderer manages `.codex/hooks.json` (same schema as Claude, verified against the OpenAI Codex hooks spec, CLI 0.142.5). `CodexAdapter.supportsHooks()` is now `true`. Two hooks ship on Codex: `framework-edit-gate` (PreToolUse; the SAME script made `apply_patch`-aware, **fail-open on parse doubt**) and `agent-discovery-info` (SessionStart; the SAME script with `--runtime codex` → inspects `.codex/agents/*.toml`). Two are deliberately **not bound on Codex** with documented rationale + compensating mechanism: `agent-discovery-gate` (no observable Agent tool → the `/new`+`/prd` Codex preflight) and `overlay-telemetry` (no reliable Read event). Ownership is matched by a command **marker** (Codex entries carry no `id`); registration is **additive** (preserves Graphify + user hooks). **Hook trust is never faked** — BALDART writes `.codex/hooks.json` but never `~/.codex/config.toml`; `add`/`update` surface that Codex asks for one-time approval, and `doctor` reports presence and trust **separately** (untrusted-only is not a failure).
- **S3 — cross-tool delegation docs.** `skills-mapping.md` reframed for portable (Claude + Codex) skills, citing `AGENTS.md` as the delegation SSOT (the domain-ownership core — code→coder / UI→ui-expert, no-silent-fallback — already landed cross-tool in the AGENTS.md primitive), forward-referencing the Codex preflight, and noting that non-BALDART feature-entry skills (e.g. `nf`) must not drive BALDART backlog execution.

**MINOR** — additive capability (native Codex hook management). **Codex parity: adapter-generated** (one registry SSOT + a Codex renderer; the two edit/session hook scripts are single-source dual-runtime). **No new `baldart.config.yml` key** (rides on `tools.enabled` containing `codex`) → the schema-change propagation rule does NOT apply. **Must-measure-first (ships safe-by-default):** one live Codex hook fire + the trust-approval flow remain to confirm on a real Codex consumer — the gate fails open, registration is additive, and trust is never faked, so S2 cannot break a consumer even before that confirmation. Authoritative design: [`framework/docs/CODEX-HOOKS.md`](framework/docs/CODEX-HOOKS.md).

### Added

- **`src/utils/hooks/registry.js`** — runtime-neutral hook set with per-runtime bindings (`runtimes.claude` / `runtimes.codex`) + `defsFor(runtime)`.
- **`src/utils/hooks/claude.js`** — the Claude Code renderer (`.claude/settings.json`).
- **`src/utils/hooks/codex.js`** — the Codex renderer (`.codex/hooks.json`) with additive register/verify/unregister/status/drift + a `~/.codex/config.toml` trust probe.
- **`framework/docs/CODEX-HOOKS.md`** — authoritative design + parity table + must-measure-first.

### Changed

- **`src/utils/hooks.js`** — now a thin compatibility shim re-exporting the Claude renderer (historical `require('../utils/hooks')` surface preserved, byte-identical).
- **`src/utils/tool-adapters/{codex,claude}.js`** — `supportsHooks()` (Codex `false`→`true`) + `hookTargetPath()`.
- **`framework/.claude/hooks/framework-edit-gate.js`** — runtime-aware: parses the Codex `apply_patch` envelope (fail-open); Claude Edit/Write/MultiEdit/NotebookEdit path byte-unchanged.
- **`framework/.claude/hooks/agent-discovery-info.sh`** — accepts `--runtime codex` → inspects `.codex/agents/*.toml`.
- **`src/commands/{add,update,doctor}.js`** — register/backfill/report Codex hooks gated on `tools.enabled ∋ codex`; doctor reports presence and trust separately.
- **`framework/agents/skills-mapping.md`** — cross-tool framing + delegation SSOT note (cites `AGENTS.md`, `CODEX-HOOKS.md`).
- **`framework/docs/CODEX-AGENTS.md`** — cross-ref to `CODEX-HOOKS.md`.
- **`CLAUDE.md`** — new "Cross-tool parity discipline" + "Codex-native hooks" conventions; fixed the stale "hooks remain Claude-only" line.

## [4.87.0] - 2026-07-01

**Plan-review gate promoted to the cross-tool SSOT (AGENTS.md primitive `1.3.0`).** The "have the plan reviewed by `plan-auditor` + `doc-reviewer` before presenting it" rule lived only in the `CLAUDE.md` stub's Plan Mode — so Codex never ran it. Now that both agents exist on Codex (transpiled to `.codex/agents/`), the tool-agnostic principle is promoted into `AGENTS.md`: **MUST have any non-trivial plan reviewed by `plan-auditor` + `doc-reviewer` (in parallel), fold in their feedback, note the review gate in the plan, and present the ALREADY-reviewed plan — never a raw one.** The Claude-only `EnterPlanMode`/`ExitPlanMode` mechanics stay in `CLAUDE.md`; `AGENTS.md` carries the principle so both tools honor it.

**MINOR** — cross-tool enrichment of the shipped skeleton; no config key, no CLI/layout change, contamination-clean. Consumers pick it up on the next `npx baldart update`. `primitive_version` (AGENTS) 1.2.0 → 1.3.0; CLAUDE stub unchanged (1.1.0).

### Changed

- **`framework/templates/primitives/AGENTS.md`** — added the plan-review MUST after the `codebase-architect` MUST. `primitive_version` → 1.3.0.
- **`framework/templates/primitives/AGENTS.CHANGELOG.md`** — `1.3.0` entry.

## [4.86.0] - 2026-07-01

**AGENTS.md primitive regains the universal working-discipline rules (primitive `1.2.0`).** The lean v4.83.0 rewrite of the skeleton dropped a handful of everyday-habit MUSTs to stay short — but those are exactly the "conventions an agent can't infer and gets wrong" the file exists to carry. Restored as compact, stack-agnostic rules in the `AGENTS.md` skeleton (the cross-tool SSOT):

- **File navigation**: MUST find real paths via Glob/Grep before reading — never guess a path from naming conventions.
- **Investigation before claiming**: check git history (`git log --all --grep`) + the actual source before asserting a feature does or does not exist.
- **Implementation completeness**: MUST verify every plan item / acceptance criterion is wired and functional before marking work DONE — cross-check against the code, not the intent.
- **Testing convention** (Workflow Gates): when a test fails, fix the *code* not the test (unless the test is wrong); for a new feature, write the failing test first.

**MINOR** — enriches shipped skeleton content; no config key, no CLI/layout change; contamination-clean (no project tokens). Consumers pick it up on the next `npx baldart update` (byte-stable on everything else; adopted overlays untouched). `primitive_version` (AGENTS) 1.1.0 → 1.2.0; CLAUDE stub unchanged (1.1.0).

### Changed

- **`framework/templates/primitives/AGENTS.md`** — added the file-navigation + investigation + implementation-completeness MUSTs and the testing convention in Workflow Gates. `primitive_version` → 1.2.0.
- **`framework/templates/primitives/AGENTS.CHANGELOG.md`** — `1.2.0` entry.

## [4.85.0] - 2026-07-01

**Root-file primitives gain universal delegation + quality rules (primitive `1.1.0`).** The v4.83.0 skeletons pointed to `.claude/agents/REGISTRY.md` for agent routing but never stated the single highest-frequency rule in the always-loaded file: *which agent writes what*. The `AGENTS.md` primitive (the cross-tool SSOT Codex also reads) now carries a **MUST**: plan-and-delegate, never write substantial code inline — **code / logic / tests → `coder`**, **UI → `ui-expert`** (dual-tool: Claude Task `subagent_type` / Codex custom agent by name; every other agent routes via `REGISTRY.md`). The `CLAUDE.md` stub's Plan Mode names both agents to match (was `coder` only).

Alongside it, the **universal, stack-agnostic** subset of rules a mature project accumulates — previously discoverable only after a project authored them by hand — is promoted into the skeleton: a **security-hygiene MUST** (no hardcoded secrets, no exposed stack traces, no PII in logs, validate/parse external inputs), a **Quality bar (SHOULD)** (bound/paginate list queries — no unbounded/N+1 — + Core Web Vitals for web UIs + baseline accessibility), and explicit **worktree isolation** in the Git Workflow (new branch → `/nw`; never `git switch`/`checkout -b`/`branch` on the shared main checkout). **Deliberately NOT promoted** (they stay project overlay, and would trip the contamination scanner): exact KPI numbers, database/ORM specifics, CSS-framework choices, and project-specific git rituals.

**MINOR** — enriches the shipped skeleton's protocol content; no new `baldart.config.yml` key, no layout/CLI change. Consumers pick it up on the next `npx baldart update` (the writer regenerates `AGENTS.md`/`CLAUDE.md`, byte-stable on everything else). `primitive_version` 1.0.0 → 1.1.0.

### Changed

- **`framework/templates/primitives/AGENTS.md`** — added the implementation-delegation MUST (`coder`/`ui-expert`), the security-hygiene MUST, a `## Quality bar (SHOULD)` section (performance + accessibility), and worktree isolation in Git Workflow. `primitive_version` → 1.1.0.
- **`framework/templates/primitives/CLAUDE.md`** — Plan Mode delegation now names `coder` (code) + `ui-expert` (UI) + `doc-reviewer` (docs). `primitive_version` → 1.1.0.
- **`framework/templates/primitives/AGENTS.CHANGELOG.md` + `CLAUDE.CHANGELOG.md`** — `1.1.0` entries.

## [4.84.0] - 2026-07-01

**`/ds-handoff` decision & preference pass — anticipa le domande di Claude Design invece di rimandargliele.** Il prompt field-level di `ds-handoff` copriva un solo asse (correttezza dei *campi* via field-extraction + coverage gate) e lasciava scoperto il **secondo asse — le decisioni di design aperte + le preferenze di consegna** — che è esattamente ciò su cui Claude Design rimanda le domande all'utente (nel caso reale "Hub fornitore", 6 domande su 6 erano su questo asse, zero sui campi). Una era **auto-inflitta** (il PRD diceva *"da proporre"* e la skill l'ha inoltrata testualmente), un'altra su un fatto **già in mano alla skill** (enum categoria con `color`) mai tradotto in istruzione.

Il fix è **combinato**: (1) un nuovo **Step 5.5 — Decision & preference pass** (`references/decision-pass.md`) che risolve PRIMA di emettere le scelte aperte / formato consegna / scope stati / canvas / varianti — **triggered, non always-on** (stessa filosofia del BLOCK del coverage gate: parte solo se c'è un irrisolto, UNA `AskUserQuestion` consolidata, no-op quando nulla è aperto; per le decisioni aperte offre "decido io" vs "fai proporre N opzioni a Claude Design"); (2) **arricchimento** che pre-compila i fatti rilevabili — nuova **§ 0 "Direttive di consegna & decisioni"** nel template + regola **enum presentazionale** in `field-extraction.md` (metadata `color`/`icon`/`badge`/`label` → render hint automatico, non una domanda). Attivo standalone **e** delegato da `/prd` (Branch B ora inoltra le preferenze già risolte in Discovery → la skill riconcilia e chiede solo il residuo). **DEGRADE** su Codex/no-TTY: bake dei default DICHIARATI in § 0, mai abort.

Resta un *thin orchestrator*, **nessun dual-SSOT** (delivery/placement/varianti non sono in `design-system-protocol.md`). **MINOR** — capability additiva su una skill esistente. **Nessuna nuova chiave `baldart.config.yml`** (rides on `features.has_design_system`) → la schema-change propagation rule NON si applica. Skill versionata `1.0.0 → 1.1.0` (bump `version` + entry nel `CHANGELOG.md` per-skill, stesso edit).

### Added

- **`framework/.claude/skills/ds-handoff/references/decision-pass.md`** — il pre-emit Step 5.5: trigger inventory, `AskUserQuestion` consolidata, DEGRADE bake-defaults, riconciliazione in modalità delegata, false-positive guards.

### Changed

- **`framework/.claude/skills/ds-handoff/SKILL.md`** — nuovo Step 5.5 nel workflow, Step 6 popola anche § 0, Invocation + Inline fallbacks aggiornati; `version: 1.0.0 → 1.1.0`.
- **`framework/.claude/skills/ds-handoff/assets/claude-design-handoff-prompt.template.md`** — nuova § 0 "Direttive di consegna & decisioni"; § 7 rimanda a § 0 per il formato export.
- **`framework/.claude/skills/ds-handoff/references/field-extraction.md`** — regola enum presentazionale (render hint automatico) nel mapping + sezione dedicata.
- **`framework/.claude/skills/prd/references/ui-design-phase.md`** — Branch B inoltra a `/ds-handoff` le preferenze di consegna/decisione già emerse in Discovery.
- **`framework/.claude/skills/ds-handoff/CHANGELOG.md`** — entry `1.1.0`.

## [4.83.0] - 2026-07-01

**Versioned root-file primitives — BALDART now generates `AGENTS.md` + `CLAUDE.md` from SOTA skeletons.** Previously `AGENTS.md` was a bulk symlink into the framework payload (not customizable without losing auto-update) and `CLAUDE.md` was not shipped at all. Now both are **generated real files**, produced on every install/update from **versioned skeletons** under `framework/templates/primitives/`.

Each skeleton mixes three tiers: **STATIC** (universal SOTA prose, shipped verbatim) · **`{{ slot }}`** (project facts filled *deterministically* at write-time from `baldart.config.yml` + `package.json`/README probes — concrete values, so Codex/Cursor/Aider read real commands, never raw `${…}`) · **`[OVERLAY]`** (project opinions authored in `.baldart/overlays/{AGENTS,CLAUDE}.md`, merged via the existing OVERRIDE/APPEND/PREPEND engine). This split follows the 2026 research consensus (short hand-authored skeleton + project depth in the overlay; auto-generating a whole file with an LLM measurably reduces agent task success).

**`AGENTS.md` is the cross-tool SSOT** (read natively by Claude Code + Codex + others). **`CLAUDE.md` is a thin Claude-native stub** that deliberately does **not** import `AGENTS.md` (Claude Code loads both automatically → an `@AGENTS.md` would double-load the protocol) — it holds only Claude-only mechanics (Plan Mode / `AskUserQuestion` / Task tool / output-style / hooks). The writer (`src/utils/root-primitives.js`) reuses `overlay-merger.js` and the same `baldart-generated` marker, so the `framework-edit-gate` already blocks direct edits to the generated files (edit the overlay).

**Existing hand-written files are adopted zero-loss:** a consumer's hand-authored `AGENTS.md`/`CLAUDE.md` (no marker) is preserved **verbatim** into `.baldart/overlays/<name>.md`, then the root is regenerated as skeleton + overlay. Adoption is **interactive-gated** — in autonomous/CI mode a hand-written file is never mutated (the writer emits a blocker and skips it). Wired into `add` (post-configure), `update` (aligned + post-pull paths, idempotent + byte-stable), `doctor` (`regenerate-root-primitives` / `adopt-root-primitives` actions), and `migrate` (legacy symlink → generated file). Each skeleton carries its own `primitive_version` + a sibling `*.CHANGELOG.md` kept out of the generated file.

**MINOR** — additive capability. **No new `baldart.config.yml` key** (slots resolve from existing config + repo probes) → the schema-change propagation rule does NOT apply. Degrades safely: missing commands omit rows, non-multimodal/Codex consumers ignore `CLAUDE.md`, autonomous runs never mutate hand-written files. Authoritative design: [`framework/docs/ROOT-PRIMITIVES.md`](framework/docs/ROOT-PRIMITIVES.md).

### Added

- **`framework/templates/primitives/AGENTS.md` + `CLAUDE.md`** — the versioned SOTA skeletons (STATIC + `{{slot}}` + `[OVERLAY]` hook), each with a `primitive_version` banner.
- **`framework/templates/primitives/AGENTS.CHANGELOG.md` + `CLAUDE.CHANGELOG.md`** — internal per-primitive changelogs (kept out of the generated file).
- **`src/utils/root-primitives.js`** — the writer: slot resolver + minimal zero-dep template engine (`{{ }}` / `{{#flag}}` / `{{> partial}}`) + overlay merge + `baldart-generated` marker + state machine (generated / symlink / handwritten / absent) + `rootPrimitiveStatus()` for doctor.
- **`framework/docs/ROOT-PRIMITIVES.md`** — consumer-facing design/lifecycle/adoption doc.

### Changed

- **`src/utils/symlinks.js`** — `createBulkSymlinks()` no longer symlinks `AGENTS.md`; `createAllSymlinks()` calls the new `writeRootPrimitives()`; `verifySymlinks()` no longer expects `AGENTS.md` to be a symlink.
- **`src/commands/add.js`** — writes root primitives after `configure` populates the config (final concrete values).
- **`src/commands/update.js`** — regenerates root primitives on both the aligned-install and post-pull paths (`reconcileRootPrimitives`); adds `CLAUDE.md` to `BALDART_MANAGED_PATTERNS`.
- **`src/commands/doctor.js`** — detects stale/adoptable root primitives and offers `regenerate-root-primitives` / `adopt-root-primitives`.
- **`src/commands/migrate.js`** — converts a legacy `AGENTS.md` symlink into the generated file (Step 2b), idempotent.
- **`README.md`** — install description + repo-tree updated (AGENTS.md/CLAUDE.md generated, not symlinked).

### Removed

- **`framework/AGENTS.md`** — the old bulk-symlink source; its STATIC content now lives in `framework/templates/primitives/AGENTS.md` (single SSOT).

## [4.82.0] - 2026-07-01

**Mockup→code extracted into a dedicated self-verifying skill `/ui-implement` + skills are now versioned.** The mockup→code fidelity knowledge (v4.73→v4.78) was **inlined inside `/new`**: the "Design Reference" briefing (image/HTML/code-src branching), the design-source declaration + Step 7b grounding, and the Phase 2.6 `/e2e-review` invocation. This release **externalizes** it into one specialized, self-verifying skill and, separately, introduces **per-skill versioning + a canonical structure** across all skills.

`/ui-implement` takes an ALREADY-APPROVED mockup and (1) **dispatches on the input channel** via `references/playbook.md` — a Claude Design handoff link (→ pull the screen's source read-only via the DesignSync MCP into `mockups/_src/`, then build-from-code — the primary standalone channel), finished code already on disk (`links.design_src`/`mockups/_src/` → build-from-code, reuse-first), image (→ load-image, pixels are the target), HTML (→ read-html, decode a Claude Design offline bundle via `extract-mockup-design.mjs`); **no Figma channel** (this project never uses Figma); (2) runs the registry-first cascade + component reconciliation; (3) builds via the **`ui-expert`** agent + grounds the design-source declaration; (4) reconciles the Post-Intervention Coherence; (5) **self-verifies fidelity** by folding in `/e2e-review` (structural + visual + design-quality trio + its own self-heal). A **THIN orchestrator** that cites `design-system-protocol.md`, `ui-expert.md`, `/e2e-review` and the verifier trio, never duplicating them (no dual-SSOT). Its scope is **fidelity ONLY** — `/new` still owns code-quality/doc/QA/ownership for a mixed card (no review hole); a pure-graphics card already skips those. `/new` **delegates** a UI card's implement+fidelity to it (programmatic, COMPACT JSON, same `status` contract as `/e2e-review`), retiring `implement.md`'s mockup detail to a thin pointer + **coarse inline fallback** (delegate-or-else-inline, the `/prd`→`/ds-handoff` model). Usable standalone (`/ui-implement <mockup-path-or-CARD-ID>`). `new2.js` keeps its inline path this release (parity is a follow-up after the context-economy measurement).

Every shipped skill now carries a **`version:` SemVer** (starting `1.0.0`, the skill's own line of evolution) + a **separate per-skill `CHANGELOG.md`** (deliberately NOT a `SKILL.md` section, so version history never burns context on load). Canonical structure: `SKILL.md` + `CHANGELOG.md` required; `references/`/`assets/`/`scripts/` on demand only. `skill-creator` is the SSOT (`init_skill.py` scaffolds `version` + `CHANGELOG.md` and drops placeholder-dir noise; `quick_validate.py` allows `effort`/`version` and warns non-fatally on missing `version`/`effort`/`CHANGELOG.md`; full rule in `references/skill-structure.md`).

**MINOR** — adds a skill (40 → 41) + a skill-metadata convention. No new `baldart.config.yml` key (`version` is skill-frontmatter metadata, `/ui-implement` rides on `features.has_design_system` + `features.has_e2e_review`) → the schema-change propagation rule does NOT apply. Additive + degrades to no-op / coarse fallback everywhere it can't run.

### Added

- **`framework/.claude/skills/ui-implement/`** — new skill: `SKILL.md` (thin orchestrator), `references/playbook.md` (input-type → build-workflow dispatch), `references/integration.md` (delegation contract from `/new`/`new2` + standalone), `CHANGELOG.md`.
- **`framework/.claude/skills/skill-creator/references/skill-structure.md`** — the canonical-structure + versioning SSOT.
- **`framework/.claude/skills/*/CHANGELOG.md`** — a per-skill changelog for all 41 skills (baseline `1.0.0`).

### Changed

- **`framework/.claude/skills/*/SKILL.md`** — added/normalized `version: 1.0.0` in every skill's frontmatter.
- **`framework/.claude/skills/new/references/implement.md`** — the Design-Reference block now leads with a **delegation gate** to `/ui-implement` and compresses the mockup-specific detail into a coarse fallback that cites the playbook + `ui-expert.md`; registry-first cascade + Post-Intervention Coherence left intact (they apply to every UI card).
- **`framework/.claude/skills/new/references/review-cycle.md`** — Phase 2.6 skips when the card was delegated to `/ui-implement` (the skill already ran the fidelity verify) + a maintainers' drift note (new2 keeps the inline path).
- **`framework/.claude/skills/skill-creator/scripts/init_skill.py`** — template frontmatter gains `effort` + `version`; generates `CHANGELOG.md`; no longer scaffolds placeholder `scripts/`/`references/`/`assets/` dirs; updated next-steps.
- **`framework/.claude/skills/skill-creator/scripts/quick_validate.py`** — `effort`/`version` are allowed frontmatter keys; non-fatal warnings when `version`/`effort`/`CHANGELOG.md` are missing.
- **`framework/.claude/skills/skill-creator/SKILL.md`** — `CHANGELOG.md` is now the documented exception to "what not to include" (required); Step 3/4 describe the new scaffold behaviour; frontmatter section adds `version` + the no-angle-brackets rule.
- **`framework/agents/skills-mapping.md`** — new `### ui-implement` entry (UI/UX Skills) + a "Skill Versioning & Canonical Structure" section.
- **`CLAUDE.md`**, **`README.md`** — skill count 40 → 41, `ui-implement` in the inventory, two new convention entries (the dedicated skill + skill versioning).

## [4.81.0] - 2026-06-29

**Playwright capture is CLI-only — no Playwright MCP anywhere in the framework.** BALDART already used the **Playwright Test CLI** for its main capture pipeline (`/e2e-review` Phase 3→4, `playwright-skill`, `webapp-testing`'s headless Python scripts), but three surfaces still drove the **Playwright MCP** interactively: the `bug` skill (Phase 1 UI reproduction via `browser_navigate`/`browser_take_screenshot`/`browser_console_messages`/`browser_network_requests`/`browser_evaluate`), the `/design-review` command (`allowed-tools` declared the four `mcp__playwright__*` tools + relied on `localhost:9221`), and the `/qa` DEEP tier ("Use MCP Playwright if available"). For **deterministic capture** the MCP is pure token overhead — every action round-trips a result (often a full a11y-tree snapshot) into context, accumulating per-step; a headless CLI script enters context once and writes screenshots/console/network to disk, read back only when needed. This release converts those three surfaces to the CLI so **no Playwright MCP is ever loaded into the agent context**.

The `bug` skill's interactive `browser_evaluate` instrumentation (fetch interceptor, SWR mutation tracker) is preserved with **zero loss of capability** by moving it into `page.addInitScript(() => { … })` inside the headless capture script — same injection, headless and reproducible, with the forwarded `console.*` collected via `page.on('console', …)` and read from disk.

**MINOR** — behaviour change across `bug`, `/design-review`, `/qa` (the Playwright MCP dependency is removed; capture is now CLI). No new config key, no install-layout change. The Codex broker's leaked Playwright-MCP orphans (spawned from the user's `~/.codex/config.toml`, independent of BALDART) remain handled by the `baldart doctor` orphan reaper — untouched.

### Changed

- **`framework/.claude/skills/bug/SKILL.md`** — frontmatter, "MCP dependencies" (Playwright MCP required → "None required for browser capture"; CLI via webapp-testing), Phase 0 triage tool, Phase 1 UI reproduction (browser_* steps → headless capture script writing screenshot/console/network to disk), Phase 2 client-side logging (`browser_evaluate` → `page.addInitScript`), Phase 3 evidence collection, Phase 5 verify, Tools Quick Reference.
- **`framework/.claude/skills/bug/references/logging-patterns.md`** — "Playwright MCP Debug Patterns" → "Playwright CLI Debug Patterns"; `// browser_evaluate:` → `page.addInitScript()` in the capture script.
- **`framework/.claude/commands/design-review.md`** — removed the four `mcp__playwright__*` tools from `allowed-tools`; the route capture is now a headless Playwright CLI screenshot (Bash) read back as a PNG, instead of opening the MCP browser on `localhost:9221`.
- **`framework/.claude/commands/qa.md`** — DEEP-tier UI quality checks use a headless Playwright CLI screenshot script (when `@playwright/test` is installed) instead of MCP Playwright.
- **`framework/.claude/skills/playwright-skill/SKILL.md`** — the cross-reference to `bug` no longer says it "requires `mcp__playwright__browser_*`"; both use the CLI.
- **`framework/docs/MCP-INTEGRATION.md`** — "Browser automation" catalog entry: Playwright MCP → **Playwright CLI**, with the token-economy rationale and the note that the Codex-broker orphan reaper is independent; the section-authoring example tool prefix swapped off `mcp__playwright__browser_navigate` (now a Firebase example) so the doc no longer implies BALDART uses the Playwright MCP.

## [4.80.1] - 2026-06-29

**Self-heal payload drift on an already-current install + a Codex-correct "agent not found" message.** Shipping v4.80.0 surfaced two gaps in the field, both from the same incident: a consumer's Codex session emitted the AGENTS.md STOP message for `codebase-architect`, yet the framework was already at v4.80.0. Two root causes:

1. **CLI↔framework skew left the consumer stuck even after the framework reached 4.80.0.** The consumer's subtree was pulled to v4.80.0, but the `update` that pulled it ran with an **older global CLI** (4.76.x) whose `codex.agentsDir()` was still `null` — so the new Codex-agent payload was silently skipped, and `.codex/agents/` was never created. Worse, *re-running* `update` after aligning the CLI hit the **`Already up to date!` early-exit** (line ~893), which returned BEFORE the symlink/merge reconciliation block — so the version-aligned install never self-healed the missing payload (recoverable only via `baldart doctor`). Fix: the `isAligned` branch now runs the idempotent, additive per-item merges (skills/agents/commands/workflows/output-styles) **before** declaring "up to date". They stay silent when nothing is missing and emit a one-line "Reconciled N newly-shipped item(s)" when they heal real drift — so a CLI-skew victim self-heals on the next `update`, not only via `doctor`. This generalises: ANY future payload-kind addition (a new `MERGE_KINDS` entry) is now reconciled on version-aligned installs too.
2. **The "agent not found" STOP message told a Codex user to "riavvia Claude Code".** The v4.80.0 AGENTS.md edit gave the Codex branch the Claude verbatim message. Fix: a **Codex-specific** message — *"Agent `<name>` non trovato tra i custom agent di Codex (`.codex/agents/<name>.toml`) — allinea la CLI (`npm i -g baldart@latest`), esegui `npx baldart update` (o `npx baldart doctor`), assicurati che il progetto sia trusted in Codex, poi riavvia la sessione Codex"* — encoding the exact remediation (CLI align → update → trust → restart; Codex loads `.codex/agents/` only in **trusted** projects, and discovers agents at session start).

**PATCH** — script bugfix (update self-heals payload drift) + AGENTS.md message correction; no behaviour change on the happy path, no new config key. Existing v4.80.0 consumers should `npm i -g baldart@latest` (kill the skew at the source) then `npx baldart update`; the Codex session must be restarted to discover the newly-installed agents.

### Changed

- **`src/commands/update.js`** — the `Already up to date!` (`status.isAligned`) branch now runs the per-item merges before the early exit (self-heal of payload drift from a prior CLI-skew update), reports the healed count, and adds `reconciled_items` to the `already-current` JSON event.
- **`framework/AGENTS.md`** — the no-silent-substitution rule's Codex branch now carries its own verbatim "agent not found" remediation message (was reusing the Claude "riavvia Claude Code" text).

## [4.80.0] - 2026-06-29

**Codex-native subagents — the 32 BALDART agents are now installed for Codex, not just Claude.** A portability bug surfaced in the field: a skill running on Codex delegated to `codebase-architect` (*"invoke the codebase-architect agent"*), but the agent did not exist in the Codex tree and there was no obvious mechanism to spawn it, so the model **improvised** a generic explorer sub-agent. Root cause (confirmed against the OpenAI Codex docs, June 2026): the framework treated subagents as Claude-only — the Codex tool adapter declared `supportsSubagents() === false` / `agentsDir() === null`, so the agent definitions under `framework/.claude/agents/` were **never installed into a Codex tree**. But **Codex now has custom agents** (subagents) — they were simply in a different shape than Claude's, and BALDART had not yet mapped them. This release closes that gap.

Because a Codex custom agent is a **different on-disk artifact** than a Claude subagent — a TOML file at `.codex/agents/<name>.toml` with `name` / `description` / `developer_instructions` (+ optional `model_reasoning_effort`), spawned **by name** — the skill-style "one source, two symlinks" model does not apply. Instead the same `framework/.claude/agents/<name>.md` stays the **single SSOT** and BALDART **transpiles** it to `.toml` on install/update (new `src/utils/codex-agent-transpiler.js`): frontmatter `name`/`description` → same, body → `developer_instructions`, `effort` (low|medium|high|xhigh|max) → `model_reasoning_effort` (xhigh/max clamp to high), `model` **omitted on purpose** (opus/sonnet/haiku have no 1:1 Codex model → inherit the parent session). The generated `.toml` is a real file (you cannot symlink `.md` → `.toml`) carrying a `# baldart-generated:` marker, so an `update` regenerates it idempotently and a user fork (no marker) is never clobbered. Overlays apply **first** (the markdown `## [OVERRIDE]/[APPEND]/[PREPEND]` overlay is merged, then transpiled) — Codex agents honour `.baldart/overlays/agents/<name>.md` exactly like Claude. The transpile rides on the **existing** `MERGE_KINDS` machinery via one new adapter hook, `transpilerFor(kind)` (`codex` returns a transpiler for `agent`; `claude` returns null), so nothing else in the merge path changed and Claude is byte-for-byte unaffected. Existing Codex consumers pick it up on the next `npx baldart update`. AGENTS.md's "MUST invoke codebase-architect" + no-silent-fallback rules are now tool-aware (the agent exists on both tools; spawn by name on Codex). Verified end-to-end: all 32 real agents transpile to valid TOML (parsed with `@iarna/toml`), idempotency / source-change-regen / user-fork-protection / overlay-aware-transpile all pass.

**Standing convention** (the user's explicit ask: *every agent from now on is natively Codex-compatible*): any new agent is auto-portable — drop a `.md` under `framework/.claude/agents/`, keep `name`+`description` frontmatter, put behaviour in the body (→ `developer_instructions`), and don't rely on an agent spawning a further agent (Codex `agents.max_depth: 1`, consistent with BALDART's `## Role boundary`). Skills were already Codex-native (identical format). Authoritative design: [`framework/docs/CODEX-AGENTS.md`](framework/docs/CODEX-AGENTS.md).

**MINOR** — additive capability (Codex consumers gain 32 agents; existing Claude installs unchanged). **No new `baldart.config.yml` key** — rides on `tools.enabled` containing `codex`, so the schema-change propagation rule does NOT apply. Slash commands / dynamic workflows / output styles stay Claude-only (`commandsDir`/`workflowsDir`/`outputStylesDir` remain null for Codex — no Codex equivalent).

### Added

- **`src/utils/codex-agent-transpiler.js`** — pure `.md` → `.toml` transpiler (`mdAgentToToml`, `isGeneratedToml`, `mapEffort`). Emits `developer_instructions` as a TOML multi-line **literal** (`'''…'''`) so arbitrary markdown survives byte-for-byte, with an escaped `"""…"""` fallback only when the body itself contains `'''`.
- **`framework/docs/CODEX-AGENTS.md`** — design + field-mapping table + the per-agent authoring contract.

### Changed

- **`src/utils/tool-adapters/codex.js`** — `agentsDir()` → `.codex/agents` (was null), `supportsSubagents()` → true (was false), new `transpilerFor(kind)` hook returning the TOML transpiler for `agent`. Header comment corrected (Codex DOES have custom agents now).
- **`src/utils/tool-adapters/claude.js`** — added `transpilerFor()` → null (explicit no-op: Claude consumes every payload in its native format).
- **`src/utils/symlinks.js`** — `_mergeBulkDir` routes a kind through the new `_installTranspiledFile` (generate tool-native file, overlay-merged first when present) when the adapter declares a transpiler; `verifySymlinks` recognises transpiled `.toml` (via `isGeneratedToml`) as framework-managed alongside symlinks and overlay-generated `.md`.
- **`framework/AGENTS.md`** — the "MUST invoke `codebase-architect`" and no-silent-substitution rules are now tool-aware: the agent exists on both tools (Claude → Task/`subagent_type`; Codex → spawn by name at `.codex/agents/codebase-architect.toml`); the #56869 specifics stay Claude-scoped, the principle stays universal.
- **`framework/.claude/agents/REGISTRY.md`** + **`CLAUDE.md`** — document the Codex-native subagent convention + the authoring contract.

## [4.79.0] - 2026-06-29

**New skill `/ds-handoff` — the field-level, 1:1 Claude Design handoff brief.** When `/prd` reached the "hand off to Claude Design" branch, the prompt it produced described each screen only coarsely — *purpose / states / actions / data shown* — and never the exact fields. With no field-level contract, Claude Design **imagines** control types, required-ness, validation, defaults and enum values to cover the gap, so the result is never 1:1. The root cause: BALDART had **no mandatory phase** that enumerated a screen's real fields, and **no coverage gate** verifying them before the prompt was emitted (field discovery was manual discovery-questions). `/ds-handoff` fixes this end-to-end: it (1) identifies each screen's backing data entities, (2) **grounds the field inventory in the project's real schemas** (Zod / TS types / form schemas) by composing existing retrieval (code-graph → LSP → `codebase-architect` → Grep — **no new AST/Zod parser**, an over-scoped fragility the design refused), (3) reconciles fields↔screens (which fields show / are editable / group into which card), (4) runs a **coverage gate** that **BLOCKS** emission in interactive mode when a backed screen's `{control, required}` are unverified — degrading gracefully to user-confirmation when no formal schema exists (portable / Codex), (5) renders the **mandatory field-level template** (a 9-column field-inventory table per screen + actions / states / responsive). It becomes the **single SSOT** for the Claude Design prompt: `/prd` Step 3.0 Branch B now **delegates** to it (graceful inline-coarse fallback when the skill is unavailable), and the old inline `/prd` template is **retired**. Usable standalone via `/ds-handoff <screens-or-feature>` (e.g. an ad-hoc orchestrator + `ui-expert` run). A THIN orchestrator — it **cites** `design-system-protocol.md` (Functional Traceability Gate `backed/decorative/orphan` vocabulary + Reference Tables numbers) and **reads** `discovery-phase.md`'s `mockup_analysis` as input, never duplicating their rules (no dual-SSOT).

**MINOR** — new skill (40 portable skills). **No new `baldart.config.yml` key** — rides on the already-present `features.has_design_system` (BLOCKING gate; REFUSE + suggest `/design-system-init` when `false`) + `features.has_prd_workflow` when delegated, so the schema-change propagation rule does NOT apply. Claude + Codex portable (the gate degrades, subagent steps fall back to Grep). No agent / template / routine outside the skill changed; `discovery-phase.md` (`mockup_analysis`) and `card-schema.md` (`component_bindings`) are deliberately **untouched** (the field layer is a skill-local artifact, not a card/state schema field).

### Added

- **`/ds-handoff`** (`framework/.claude/skills/ds-handoff/`) — `SKILL.md` (thin 8-step orchestrator), `assets/claude-design-handoff-prompt.template.md` (the mandatory field-level template — replaces the retired `/prd` coarse template), `references/field-extraction.md` (retrieval order + control-type mapping table + readonly detection), `references/coverage-gate.md` (BLOCK-vs-degrade logic + the worked "Fornitori" example). Gated on `features.has_design_system`.

### Changed

- **`/prd`** (`framework/.claude/skills/prd/references/ui-design-phase.md`) — Step 3.0 Branch B now **delegates** the Claude Design handoff prompt to `/ds-handoff` (passing the resolved context so it skips its own screen-resolution), presents what it returns, and persists `mockups.field_inventory_path`. Graceful **inline-coarse fallback** (stated explicitly) when `/ds-handoff` is unavailable.
- **`framework/agents/skills-mapping.md`** — expanded from a curated 11-skill subset to the **full 40-skill routing map** (Design System, Internationalization, Retrieval/Toolchain Bootstrap, Framework Maintenance categories added; per-entry `features.*` gating noted; decision tree + chains updated with the design-system + Claude-Design-handoff flows), so the routing doc the agent reads covers every shipped skill including `ds-handoff`.

### Removed

- **`framework/.claude/skills/prd/assets/claude-design-handoff-prompt.template.md`** — the coarse `/prd` handoff template, superseded by the field-level template in `ds-handoff/assets/` (single SSOT; coarse content survives inline in the Branch B fallback).

## [4.78.2] - 2026-06-29

**The "visual-verify every `ui-expert` intervention" hard rule was investigated and deliberately NOT built.** The ask sounded obvious — wire an invariant `UI_VISUAL_UNVERIFIED` forcing a visual check after every `ui-expert` implementation. Three adversarial reviewers (distinct lenses) refuted the diagnosis: (1) standalone skills already verify — `/ui-design` has its generator/evaluator Playwright loop, `/ds-new` + `/ds-edit` delegate to `ui-expert`'s own BLOCKING pre-work + `ds-gate` + the coherence report, `/ds-render` feeds `/e2e-review` Phase 4b; (2) the invariant would be redundant with the existing `DS_*` coherence codes + the `coverage-gap` finding, and would fire **constantly** on every Storybook-less project — the gate-with-no-exit structural false-positive the framework explicitly avoids — plus risk the assertion-fitting the verifier trio forbids; (3) the real driver the obvious diagnosis missed: **visual verification is deliberately deferred to where a ground-truth exists** (a composed route / a mockup) — an inert scaffolded primitive has nothing to be judged against, so verifying it *now* judges the void. The single useful residue (a deterministic `PostToolUse` hook would beat prose if one ever wanted to harden `/new`'s own prose-driven Phase 2.6) was logged, not built. So this release ships only the **zero-cost, non-regressive nudge**: a documentation pointer making explicit *where* route-level fidelity lives.

**PATCH** — doc-only pointers; no new capability, no agent, no invariant, **no `baldart.config.yml` key**. The nudge is a pointer to the existing `/e2e-review` Phase 4, never a blocking obligation; skipped/irrelevant where no render surface exists.

### Changed

- **`/ds-render`** (`framework/.claude/skills/ds-render/SKILL.md`) — added a "Where route-level fidelity lives" note: the harness verifies a primitive in isolation (quality only, no ground-truth); the full fidelity check happens downstream via `/e2e-review` Phase 4 when the primitive is composed into a real route. A pointer, never a blocking obligation.
- **`/ui-design`** (`framework/.claude/skills/ui-design/SKILL.md`) — added a "Downstream visual verification" note after Step H: the design is verified visually at the mockup stage (Step D evaluator); the fidelity of the *implemented* result is verified downstream via `/e2e-review` Phase 4 (auto-run by `/new`, or manual `/e2e-review CARD-ID` outside the pipeline).

## [4.78.1] - 2026-06-29

**`/ds-new` + `/ds-edit` now close the closed-loop (the `DS_MIRROR_STALE` nudge).** The mirror advisory — check #6 of the Post-Intervention Coherence Check, which nudges `/design-sync publish` when the in-repo registry has moved ahead of its Claude Design satellite — was cited by **every** UI-touching surface (`code-reviewer`, `ui-design`, `/design-review`, `/new` implement, `/prd` ui-design-phase, the `ds-drift` backstop) **except the two skills whose whole job is to create/modify a canonical element**: `/ds-new` and `/ds-edit`. So creating a primitive with `/ds-new` on a satellite-bound project never reminded the user to publish it — the one event that most directly makes the mirror stale was the one that skipped the nudge. This release adds the citation to both skills' step-8 coherence report.

**PATCH** — doc-only completion of an existing advisory check; no new capability, no agent, **no `baldart.config.yml` key** (the check rides on the already-present `features.has_design_system` + the satellite-bound gate). The definition stays the single SSOT in `design-system-protocol.md` § Post-Intervention Coherence Check check #6 — the skills **cite, never redefine** it. Skipped when no satellite is bound (the vast majority of consumers → no-op); never blocks creation/edit.

### Changed

- **`/ds-new`** (`framework/.claude/skills/ds-new/SKILL.md`) — step 8 coherence report now surfaces `DS_MIRROR_STALE` (advisory, non-blocking) nudging `/design-sync publish` when a Claude Design satellite is bound (`.baldart/design-sync.json` has a `project_id`); added a `See Also` pointer to `/design-sync`.
- **`/ds-edit`** (`framework/.claude/skills/ds-edit/SKILL.md`) — same step-8 mirror advisory, with the breaking-edit note that satellite drift is the same class as the migration; added the `/design-sync` `See Also` pointer.

## [4.78.0] - 2026-06-29

**Route-independent per-card mockup fidelity + graphics fast-lane.** On a `/new` run a desktop "Edit product" page shipped as **1 column** against a **2-column** mockup, marked DONE: the fidelity gate never ran on a rendered page because the auth-gated route returned 500 under the empty-Supabase demo lane, the views were "soft-skipped", and the verdict came back `ACCEPTED (harness limit)` — **a skip masked as a pass**. Two adversarial passes found this is not one bug but two distinct holes hitting two cards, plus the fact that the fidelity check depended on rendering a live, data-bearing route at all. This release inverts the priority: per-card UI fidelity is verified **route-independently, before proceeding** — a code-structural diff first, an isolated render as confirmation — and an un-runnable required check now **blocks instead of silently passing**. Bets validated against 2026 best practice (structural TED + perceptual diff as complementary dimensions; MLLM-as-judge reliable at comparison, weak at absolute scoring → fidelity gates, quality stays advisory).

**MINOR** — adds an agent (`markup-fidelity-verifier`, 30→31) + capability. **No new `baldart.config.yml` key** (rides on `features.has_e2e_review` + `features.has_design_system`), so the schema-change propagation rule does NOT apply. Claude + Codex portable (the structural pass is Read-only; the image-load/isolated-render degrade to no-op where unavailable).

### Added

- **`markup-fidelity-verifier` agent** (`framework/.claude/agents/markup-fidelity-verifier.md`, `model: sonnet`) — the **route-independent structural twin** of `visual-fidelity-verifier`. Compares the implementation's code structure against the mockup expressed as code (`links.design` HTML or `links.design_src`/`mockups/_src/`) with no browser/auth/data — the DOM tree-edit-distance analogue, complementary to the visual verifier's pixel/perceptual diff. Deterministically catches the **2-column-mockup → 1-column-build** class of divergence. READS code (mockup-vs-impl, no assertion-fitting risk); same canonical taxonomy + JSON output; fixes route to `ui-expert`.
- **`/e2e-review` Phase 2.7 — Structural Fidelity (route-free)** — runs the `markup-fidelity-verifier` BEFORE the browser passes, on every card with an HTML/`design_src` mockup. The workhorse of per-card fidelity.
- **`/e2e-review` Phase 3.8 — Comparable-render gate (state-aware)** — the reviewer never accepts "opened it, it's empty, skip". Detects a degenerate render (non-2xx / empty-state / wrong tenant) and runs a bounded **reach-state loop** (overlay recipes `seed_data` / `select_populated_context` / `goto_populated_entity`, with generic fallbacks: real auth, MSW/API-mock, deterministic factory/snapshot, heuristic context auto-seek) before diffing. New `data-state-mismatch` (Critical) finding.
- **Coverage obligation** — new `coverage-gap` (Critical) finding synthesized when an image-only mockup route could not be rendered in any lane and its state could not be reached; maps to the existing `blocked` verdict, bypasses the self-heal loop. Guards exclude `skip`/`harness-render`/`compliance-only`/mobile-pass. This turns the FEAT-0061 "skip masked as pass" into a real block.
- **Starter overlay** `framework/templates/overlays/e2e-review.fidelity-example.md` — schema for project-specific reach-state recipes.

### Changed

- **`visual-fidelity-verifier`** — `coverage-gap` + `data-state-mismatch` added to the canonical severity taxonomy (SSOT); new "Data-State Awareness" rule: never compare a degenerate render, emit `data-state-mismatch` instead of a silent skip/pass.
- **`/e2e-review` Phase 2** ingests **HTML mockups** (`links.design` `.html` → headless `file://` render to PNG for the visual pass + retained as code for the structural pass) and recognizes `links.design_src`. **Phase 5b self-heal** routes ALL fidelity fixes (structural/presentational/quality) to `ui-expert`; `coder` only for genuine wiring.
- **`/new` graphics fast-lane** — a **pure-graphics card** (`owner_agent: ui-expert` + `review_profile: light` + `areas == [ui]` + zero Step-A triggers) skips Simplify (Phase 2.55) and per-card Codex (Phase 3.7); qa is the build/lint/tsc floor; the `/e2e-review` fidelity trio is the review. Safety net preserved by the batch-wide Final FULL Codex gate. Wired in `review-cycle.md`, `codex-gate.md`, `prd-card-writer.md` (emit `owner_agent: ui-expert` on Rule C branch-(b) cards).
- **`/new` no cross-card fidelity deferral** — `completeness.md` Phase 2.5b Step 4 fidelity-AC carve-out: a visual-fidelity AC cannot be silently deferred ("Approva il deferral" withheld) — only implement-now (→ `ui-expert`), explicit follow-up card, or halt. `review-cycle.md` Phase 2.6: a fidelity AC closes only on a real `passed`, never a `skipped`.
- **`design-system-protocol.md`** — new SSOT section "Route-Independent Fidelity & Coverage Obligation" (structural fidelity, visual confirmation, state-aware fidelity, coverage obligation, graphics-card rule, gating calibration) cited by `/e2e-review`, `/new`, `/design-review`.

## [4.77.0] - 2026-06-29

**TS-aware design-system extractor + the `variant_prop` HEAD field.** The component-manifest extractor was pure-regex (zero-dep) and could not resolve TypeScript: a `type Variant = 'a'|'b'; variant: Variant` captured the alias *name* `"Variant"`, not its members, so on a TS-strict consumer (mayo) every HEAD came out with `variants: []` and ~25 with `props: {}`. Since the HEAD is the SSOT every UI agent reads for discovery, `variants: []` on a 6-variant Button induces wrong code. Separately, the variant axis was **hardcoded** to the prop name `variant` in three places — wrong for a `Chip`/`Pill` whose axis is `tone`. This release makes the extractor **optionally TS-aware** (it borrows the consumer's own `typescript`, resolved from the repo root) and adds a deterministic **`variant_prop`** field naming the real axis.

**MINOR** — adds capability + a new deterministic HEAD field. **No new `baldart.config.yml` key** (rides on `features.has_design_system`), so the schema-change propagation rule does NOT apply. **Zero regression**: when typescript is absent (Codex / greenfield / non-TS stack) or a file fails to parse, the extractor falls back to the existing zero-dep regex path per file; legacy HEADs (pre-v4.77.0, no `variant_prop`) default the render axis to `'variant'` so existing registries render identically until a one-time re-extract backfills the field.

### Added

- **`extract-ts.mjs`** (`framework/.claude/skills/design-system-init/scripts/extract-ts.mjs`) — optional TS-aware resolver. Probes `typescript` via `require.resolve('typescript', { paths: [process.cwd()] })` (the **consumer's** dependency, not BALDART's), parses syntactically with `ts.createSourceFile` (no tsconfig / Program / TypeChecker — fast, in-file, deterministic), and resolves type **aliases**, **unions**, **intersections**, **`Omit`/`Pick`**, interface `extends`, and `forwardRef`/`memo`/`React.FC` wrappers. Opaque / imported types (`ComponentPropsWithoutRef<'button'>`, `HTMLAttributes`) stay **empty** — never a DOM-attribute dump. Per-file try/catch → regex fallback. The typescript-free `computeVariantProp` / `pureStringLiteralUnion` / `literalMembers` are **shared** with the regex path so the axis ranking is byte-identical on both.
- **`variant_prop` deterministic HEAD field** — which prop IS the variant axis, ranked `variant > tone > intent > kind > appearance > state > status > color > severity` (else the sole literal-union prop, else `''`). `variants` is the literal members of `variant_prop`'s type. Emitted by `serialize-spec.mjs`, consumed as the axis key by `render-manifest.mjs` (`entriesFor`) and `compile-ds-cards.mjs` (row label) instead of the hardcoded `'variant'`. Documented in `component-manifest-schema.md` (SSOT), the canonical `component-spec.template.md`, the `/design-system-init` `/ds-new` `/ds-edit` skills, `COMPONENT-MANIFEST-LAYER.md`, and `design-system-protocol.md`.

### Fixed

- **`render-manifest.mjs parseHead` now reads block-style `variants`** — the serializer emits non-empty `variants` as a YAML **block list** (`variants:\n  - a`), but `parseHead` only matched flow `variants: [a, b]`, so on any real serialized spec the variant matrix read `[]` (every component collapsed to a single `--default` entry). It now accepts both block and flow forms; with `variant_prop` this makes the render/publish axis work end-to-end (a `Chip` mounts `{ tone }` per variant).

### Changed

- **`extract-manifest.mjs extractFile`** — TS-path-first with per-file regex fallback; emits `variant_prop`; `extractVariants(props)` generalized to `extractVariantsFor(props, variantProp)` (the axis is no longer hardcoded to `variant`).
- **`/design-system-init` UPGRADE mode** (`SKILL.md`) — a one-line **pre-v4.77.0 backfill** note: re-extract all components ignoring `source_sha` once to populate `variant_prop` + resolve aliases, agentic fields preserved by the existing non-clobber rule.

## [4.76.1] - 2026-06-26

**`/design-system-init` offers a render surface (opt-in Storybook scaffold).** The render harness (Stage C) needs a Storybook, and a registry-only project has none — so the harness was a silent no-op with nowhere to go. Now `/design-system-init`, after scaffolding the registry, OFFERS (default-NO, JS/TS stacks only) to run the **official `npx storybook init`** so the harness has a surface. It **never hand-writes a per-stack render surface and never auto-wires the project's providers** — that "auto-provider-shim" is the refuted false-fidelity trap; instead it leaves a `.storybook/preview` decorator TODO so **the user owns the context-shimming** (theme/i18n/router), which is how Storybook does it correctly. `/ds-render`'s no-op message now points here. Declined / Codex / non-JS → skipped (publish still works via static-spec cards).

**PATCH** — opt-in, no behaviour change for projects that decline or already have Storybook; no new config key.

### Changed
- **`/design-system-init` step 7b** (`framework/.claude/skills/design-system-init/SKILL.md`) — opt-in `npx storybook init` offer + the user-owned-decorator TODO; cross-ref `/ds-render`.
- **`/ds-render` no-Storybook message** (`framework/.claude/skills/ds-render/SKILL.md`) — points to the init offer instead of a dead no-op.

## [4.76.0] - 2026-06-26

**Closed-loop design system: BALDART ⇄ Claude Design.** BALDART owns the design system; Claude Design is where designers draw mockups. The two **diverge** over time — v4.75 made every new mockup born on the project's tokens, but that alignment decays as the in-repo registry evolves and the satellite stays frozen. This release adds the machinery to keep them in sync with **ONE authority** (the in-repo registry — `.tokens.json` + `components/<Name>.md` HEAD + INDEX/Selection-Policy + `ds-gate`); Claude Design is a **satellite mirror**, never a 2nd SSOT. Planned via a multi-agent workflow (map → per-axis design → adversarial refutation → completeness/SSOT critic → synthesis) which **refuted** the obvious-but-wrong pieces before any code.

**MINOR** — adds capability; **no new `baldart.config.yml` key** (rides on `features.has_design_system`), so the schema-change propagation rule does NOT apply. Claude-only (needs the DesignSync MCP); everything degrades to a clean no-op on Codex / no satellite / no auth / no Storybook. Shipped as a **staged program** (each stage an independent, safe-by-default increment); the write-direction stages carry their safety guards (refuse-on-divergence, `--all` hard-blocked, report-only routine) so the live calibrations (one `report_validate` write, the satellite token-format, `create_project` shape, the per-file clobber signal, Playwright-isolation, i18n-in-render) become the **consumer's first-run verification checklist**, not blockers to shipping the framework.

### Added

- **Stage A — foundations** (`framework/.claude/skills/design-system-init/scripts/render-manifest.mjs` + `src/utils/design-sync-state.js`): a zero-dep generator that reads each component HEAD → `render-manifest.json` (variant-axis default; `--full` bounded cartesian, capped); and `.baldart/design-sync.json` — a **content-free** synchronized baseline (only `source_sha`/`spec_sha`/`tokens_sha`, reusing the existing `git hash-object` anchors), `source_of_truth:code` invariant enforced on read+write. (A content-bearing baseline would be the 2nd SSOT every other piece avoids.) Tested on synthetic HEAD fixtures + state round-trip.
- **Stage C — render harness** (`src/utils/render-adapters/{index,storybook,claude-design-seed}.js` + `src/commands/render.js` + `baldart render build|shot` + `/ds-render`): mounts registry primitives in ISOLATION (today BALDART documents components but only ever renders full routes). **Storybook-only v1** — reuses the project's real `.storybook` decorators for context-shimming; the framework-native preview-route generator + the auto-provider-shim were **refuted** (false-fidelity + unbounded per-stack surface). PNGs feed `e2e-review` Phase 4b/4c **quality** ONLY.
- **Stage D — publish** (`compile-ds-cards.mjs` static-spec `@dsCard` compiler + the `/design-sync` skill): the `/design-sync` skill (`publish`/`bootstrap`/`drift` modes) **delegates every write to the DesignSync MCP** — it never reimplements `finalize_plan`/`write_files`/`report_validate` (the `/baldart-push` model); incremental, human-gated at `finalize_plan`, **refuse-on-divergence** (never clobbers a designer edit), interactive-auth.
- **`DS_MIRROR_STALE` — check #6** (advisory, non-blocking) in the Post-Intervention Coherence Check (`framework/agents/design-system-protocol.md`), defined ONCE and cited everywhere; gated on a bound satellite (skipped for the majority → no MCP/auth per UI task); the interactive twin of `ds-drift`.

### Changed

- **Stage B — seed** (`framework/.claude/skills/design-system-init/SKILL.md` new `--mode seed`): when the registry is born in Claude Design (the mayo case), an upstream read-only pull materializes the sources, then the existing pipeline runs unchanged; the pulled CSS custom-prop token sheet is fed to the existing Token-SSOT-bootstrap (CSS→DTCG); cut-the-cord (no `synced_from_satellite` in the HEAD → no authoritative re-pull).
- **Stage E — reconcile** (`/design-sync drift`): a designer's satellite edit is a governed proposal → `/ds-edit`/`/ds-new` candidate (detector-not-applicator; tombstone guard for deliberate deletions); the registry stays authority.
- **Flagship wiring** — `DS_MIRROR_STALE` cited at `/new` (`implement.md` + `code-reviewer.md` rule 8), `/prd` (`ui-design-phase.md` Step 3e) + a **seed-trigger** at `/prd` `discovery-phase.md` 1.6.4 §1b (DS-project handoff + `has_design_system:false` → `/design-system-init --mode seed`), `/ui-design` (Step H + a render-harness preference at Step B), `ds-drift` (report-only mirror backstop + `DS_RENDER_BROKEN`), `/design-review` (`ds_drift_code` enum). **`e2e-review` `harness-render` guard**: a new `mockup_source.level: "harness-render"` that Phase 4 **fidelity SKIPS** (anti-assertion-fitting by image) and Phase 4b/4c **quality** consumes (with an `i18n-incomplete` disclaimer for raw-key renders).

## [4.75.0] - 2026-06-26

**Build UI from the finished design CODE, not from a re-interpreted screenshot.** A Claude Design mockup is not a picture — it is finished code (JSX + CSS + the project's own design tokens) authored against `src/ui/tokens.ts`. The pipeline was rendering that code to a PNG and reinterpreting the pixels, throwing away exact spacing, radii, motion, and breakpoints that were already written. v4.73/v4.74 made the implementer *load the image*; this release makes it **build from the source**. Verified live against a real `claude.ai/design` project via the `DesignSync` MCP: `get_file` returns readable `.jsx` components + `mayo-tokens.css` (the project's design system as code), not an encoded bundle.

The channel matters: **DesignSync MCP handoff** (the user pastes the Claude Design handoff URL → `/prd` pulls the screen's `.jsx`/`.css`/tokens **read-only** via `list_files`+`get_file` and archives them under `mockups/_src/`) is the richest input; a **Download-as-.zip** export is the no-auth equivalent; the **Standalone HTML (offline)** export is an encoded self-rendering bundle and only a fallback (decoded by a new deterministic script). `ui-expert` then builds **reuse-first**: it maps the source's tokens/classes to the registry primitives (near-1:1, since the mockup speaks the project's token language) and copies structure/spacing/`@keyframes`/`@media` faithfully, introducing a genuinely-new component only via `/ds-new`. The fidelity is **proven** (a declaration gate + deterministic grounding) and **verified** (code-reviewer), not hoped for.

**MINOR** — adds capability; **no new `baldart.config.yml` key** (rides on `features.has_design_system`), so the schema-change propagation rule does NOT apply. `links.design_src` is an OPTIONAL `links` sub-field (validator/matrix untouched). The `visual-fidelity-verifier` is deliberately UNCHANGED (giving it the mockup's declared styles would reintroduce the assertion-fitting bias it forbids — fidelity is enforced on the *input* side). Degrades to no-op for image-only mockups (PNG/Figma), for Codex (the archived `.jsx`/`.css` is plain text), and if Claude Design changes the bundle format. The DesignSync MCP auth (`/design-login`) is interactive → the pull happens at `/prd` time and the source is archived to disk so the autonomous `/new` run reads it without the MCP. `new2.js` unchanged (reads `implement.md` for the briefing semantics).

### Added

- **`extract-mockup-design.mjs` deterministic fallback extractor** (`framework/scripts/`) — decodes a Claude Design "Standalone HTML (offline)" bundle's `__bundler/template` blob into readable HTML+CSS under `mockups/_src/`. Form-dispatcher: offline-bundle → decode; already-readable `.jsx`/`.css` → copy; image/no-markup → `{status:"no-code"}`; any error → `{status:"skipped"}`, **always exit 0** (never aborts the PRD, per `canvas-intake-recipe.md:111`). Verified on a real offline export (decoded 121 KB HTML + 119 KB CSS with exact tokens).
- **`links.design_src` card sub-field** (`framework/agents/card-schema.md`) — OPTIONAL on-disk path to the screen's archived design source; the fidelity source `ui-expert` builds from.

### Changed

- **`ui-expert` gains a "Design Reference — When the mockup is finished CODE" protocol** (`framework/.claude/agents/ui-expert.md`) — finished source ⇒ build from it (reuse-first + token-map, copy-faithful only for the genuinely new → `/ds-new`, honor exact `@keyframes`/`@media`); image-only ⇒ the established image-load behaviour.
- **`/prd` intake pulls + archives the design source** (`framework/.claude/skills/prd/references/discovery-phase.md`) — new `claude-design-handoff` format detection + a read-only `DesignSync` `list_files`/`get_file` pull into `mockups/_src/` (1.6.4 step 1b); the bundle/ZIP path (1.6.4b step e) mirrors its `.jsx`/`.css` into `_src/` and sets `mockups.source_kind: claude-design-code`. The format prompt now recommends handoff/ZIP over offline-HTML. Component Reconciliation (`ui-design-phase.md`) builds `/ds-new`/`/ds-edit` from the source and sets `links.design_src`.
- **`/new` briefing prefers the source code + a declaration-grounding gate** (`framework/.claude/skills/new/references/implement.md`) — the "Design Reference" section makes `links.design_src`/`mockups/_src/` the primary fidelity source (image = secondary cross-check); `ui-expert` must declare which source files it read, the symbols it reused, and the registry primitives it mapped to; a new orchestrator **Step 7b** grounds that declaration with `grep` against `mockups/_src/` + the registry and rejects a hollow one.
- **`prd-card-writer` writes `links.design_src`** (`framework/.claude/agents/prd-card-writer.md`) when `mockups.source_kind: claude-design-code`.
- **`code-reviewer` verifies design-source reuse** (`framework/.claude/agents/code-reviewer.md` rule 4b) — may read `mockups/_src/` (the assertion-fitting prohibition is the verifier's, not the reviewer's) and flags a diff that ignored the available source and redrew the screen from scratch (HIGH, routed to `ui-expert`).

## [4.74.0] - 2026-06-25

**Update in autonomy: a standing `git.on_divergence` policy so `npx baldart update` stops asking every release.** A consumer that keeps local commits on `.framework/` which the updater cannot auto-classify as overlays — the canonical case being a project-specific edit to a `framework/agents/*.md` **protocol module** (e.g. a design-system review rule naming the project's own files), which unlike `.claude/{agents,skills,commands}/` has **no overlay channel** — was stopped on EVERY update with "Custom commits touch non-overlayable paths … cannot be auto-resolved", forcing a manual `--on-divergence pull` each time. The non-interactive escape hatch existed only as a per-run CLI flag; there was no way to persist the choice. Now `git.on_divergence` in `baldart.config.yml` records a standing policy the interactive `update` honours, so a consumer opts in once and never passes the flag again. The `--on-divergence` CLI flag still wins for a single run, and the default (unset/null) preserves the current "stop once and ask" behaviour exactly.

**MINOR** — adds capability via a new config key. The **schema-change propagation rule applies** and is satisfied end-to-end: template (`git.on_divergence`) + configure prompt + update detector (already diffs all `git.*` keys, so the new key auto-surfaces on existing configs) + the gating code in `update` + CHANGELOG. Backward compatible: an existing config without the key resolves to null → no policy → unchanged behaviour; a config typo fails loud up-front (validated against the same `DIVERGENCE_STRATEGIES` set as the flag).

### Added

- **`git.on_divergence` standing divergence policy** (`framework/templates/baldart.config.template.yml` § `git:`). Values: `null`/unset (DEFAULT — stop once and ask), `pull` (keep the local commits and merge upstream over them; non-destructive, the autonomy choice), `scaffold-overlays` (auto-scaffold the overlay-able subset, then stop). Read by `update` via a new `readDivergencePolicy()` helper; surfaced interactively (`Applying persisted divergence policy …`) so an auto-merge is never a surprise. The `configure` git section now prompts for it (`src/commands/configure.js`), writing the key explicitly (as null when unset) so the schema-drift detector does not re-offer it every release.

### Changed

- **`update` resolves the divergence strategy as flag-or-config** (`src/commands/update.js`). `--on-divergence` (explicit per-run override) wins; absent a flag, the resolver falls back to `git.on_divergence`. Up-front validation now reports the correct source (`--on-divergence` vs ``git.on_divergence`` in baldart.config.yml) on a typo. The stale-CLI self-relaunch needs no change — the newer child re-reads the same config and resolves the policy itself.

### Fixed

- **`links.design` now resolves to a path that EXISTS — making the v4.73.0 image-load fire in the mockup-only case.** v4.73.0 made the `/new` implementer `Read` the mockup image when `links.design` is a `.png`, but the card field was still hardcoded to `design.html` (`framework/.claude/agents/prd-card-writer.md`, `framework/.claude/skills/prd/assets/card-template.yml`). In **mockup-only** runs — exactly the "I hand `/prd` Claude Design / Figma mockups" flow — no `design.html` is generated (`ui-design-phase.md` sets `design.html_path: n/a (mockup-only)`), so the card pointed at a **non-existent** file: the implementer got a dead path AND the image branch never fired (extension was `.html`). `prd-card-writer` now resolves `links.design` to an on-disk path — design.html when generated, ELSE the screen's canonical mockup PNG from `mockup_analysis.screens[].mockup_ref` / `prerender_png` under `mockups/` (chat-only `chat://image-N` images are never persisted, so such a card is flagged visually-unreferenced rather than pointed at a path that won't resolve). The `/new` briefing (`framework/.claude/skills/new/references/implement.md` § "Design Reference") also falls back to the sibling `mockups/` PNG when a legacy card's `.html` path is absent. No behaviour change for generated-design runs.

## [4.73.0] - 2026-06-25

**Mockup→code fidelity: let the implementer actually SEE the mockup, and carry the intent a flat list loses.** A recurring complaint: a detailed mockup (Claude Design / Figma) is handed to `/prd`, then `/new` ships something that is not 1:1 with the design. The obvious diagnosis — "the PRD's visual mapping is too weak, add a mockup-understanding agent" — was **refuted** by an adversarial pass: the visual data largely already exists (`mockup_analysis` captures affordances + traceability + states + DS violations; `component_bindings` propagates region→component to the implementer), and a parallel machine-readable `design-spec` artifact would be a second SSOT that drifts from the mockup and would feed the verifier a textual proxy instead of the pixels (a bias `visual-fidelity-verifier` explicitly forbids). The **actual driver**: the `/new` UI briefing passes the mockup as a *path* and says "Read the design.html file" — but for a binary PNG/Figma mockup nothing ever instructs the (multimodal) `ui-expert` to **load the image into its context**. The implementer was building from text, never looking at the design.

**MINOR** — adds capability; **no new `baldart.config.yml` key** (rides on the existing `features.has_design_system`), so the schema-change propagation rule does NOT apply. The new `component_bindings[].props` is an **optional** sub-field — legacy cards omit it and the validator (which keys on the `C`-profile top-level field) is unaffected. On a non-multimodal tool (Codex) the image-load instruction degrades to a harmless no-op and the structured fields carry the intent.

### Changed

- **The `/new` UI briefing loads the mockup IMAGE multimodally** (`framework/.claude/skills/new/references/implement.md` § "Design Reference"). The instruction now branches on the file type of `links.design`: an IMAGE (`.png`/`.jpg`/`.jpeg`/`.webp`/`.pdf`, or a `mockups/*.png` canonical path) is **`Read` directly** so it renders into the agent's context (pixels = fidelity target); an `.html` keeps the existing read-as-text behaviour. The briefing also now surfaces `component_bindings[].props`/`variant` (exact variant intent) and the card's responsive ACs (breakpoint behaviour a single-viewport image cannot show) as the structured complement to the image.
- **`mockup_analysis` schema enriched in-place with the four dimensions a flat affordance list + a single static image lose** (`framework/.claude/skills/prd/references/discovery-phase.md` § "Mockup analysis schema"). Each `screens[]` entry gains four OPTIONAL fields: `layout` (compact nesting sketch — the hierarchy), `component_props` (key props/variant legible per component — the reuse-vs-`reuse-variant` signal), `states` (richer `states_visible` mapping each state→the AC that mandates it; an untraced state is a Discovery gap by the same rule as an untraced affordance), and `responsive` (breakpoint behaviour). "Omit rather than guess" discipline extended to all four.
- **Component Reconciliation consumes the enrichment and persists it** (`framework/.claude/skills/prd/references/ui-design-phase.md` § "Component Reconciliation"). `REUSE+VARIANT` detection now compares `mockup_analysis.component_props` against the component's spec HEAD (props the spec doesn't cover ⇒ variant), and the persist step carries resolved props into `component_bindings[].props` and lifts `responsive` notes into the UI card's `acceptance_criteria` as testable ACs.

### Added

- **`component_bindings[].props` card sub-field** (`framework/agents/card-schema.md` § `component_bindings`) — the optional resolved props/variant each region must use, the exact-variant intent a static image leaves ambiguous; `ui-expert` implements it verbatim. Sub-field of the existing `C`-profile `component_bindings`, so no validator/matrix change.

## [4.72.0] - 2026-06-25

**Session-id provenance on PRDs and cards — trace an implementation bug back to the chat history that produced it.** When `/new` ships a card that turns out wrong, there was no link from the card back to *which Claude Code session* planned it (`/prd`) or implemented it (`/new`/`/new2`) — so you couldn't reopen the transcript to see the steps the session took (or skipped). Now every card carries an optional `provenance` block recording the session UUID for each lifecycle stage, and the PRD frontmatter records its authoring session. Resolve the chat history at `~/.claude/projects/*/<session-id>.jsonl` (the UUID is globally unique — no project slug needed). This turns "this feature is broken, why?" into a concrete pointer at the planning/implementation conversation, for improving the framework from real runs.

The enabling fact (verified against a live `env`): the session id IS available at runtime as **`CLAUDE_CODE_SESSION_ID`** — NOT `CLAUDE_SESSION_ID` (which does not exist). The capture discipline matters: the value is read by the **orchestrator** (`/prd`, `/new`, the `new2` skill) in its own context and passed down — never read inside a writer/commit subagent, whose own session id differs (`CLAUDE_CODE_CHILD_SESSION=1`).

**MINOR** — adds a traceability capability; **NOT a `baldart.config.yml` key** (it is a card-schema field + a native env var), so the schema-change propagation rule does NOT apply. The `provenance` field is **optional (`C`)** in every card profile — the validator ignores it and legacy cards never break (verified by running `validate-card-baseline.js` on cards with and without the block). Sub-fields are omitted when the env var is unset; an EPIC card (a tracker) carries `planning_session` only.

### Added

- **`provenance` card field** (`framework/agents/card-schema.md` — the schema SSOT) — `{ planning_session, implementation_session }`, both optional session UUIDs, added to the field-state matrix as `C`/`C`/`C` with a dedicated `§ provenance` section defining ownership-by-lifecycle and the `CLAUDE_CODE_SESSION_ID` resolution rule. Mirrored in the card template (`framework/.claude/skills/prd/assets/card-template.yml`) and a `planning_session` frontmatter line in the PRD template (`framework/.claude/skills/prd/assets/prd-template.md`).
- **`session_id` on the `/new` telemetry row** (`framework/.claude/skills/new/references/metrics.md`) — `skill-runs.jsonl` records now carry the run's session UUID so the metrics row also resolves to its transcript.

### Changed

- **`/prd` stamps the planning session** — `prd-card-writer` (`framework/.claude/agents/prd-card-writer.md`) sets `provenance.planning_session` from a new `PLANNING_SESSION_ID` field the `/prd` orchestrator passes (captured in the orchestrator's own context, `backlog-phase.md`), and `/prd` writes the same id into the PRD frontmatter (`prd-writing-phase.md`). The writer never reads the env var itself (subagent session ≠ conversation session).
- **`/new` + `/new2` stamp the implementation session** when a card reaches `status: DONE` — `/new` at `commit.md` Step 28 (`framework/.claude/skills/new/references/commit.md`); `/new2` via a new `sessionId` workflow arg captured by the skill and stamped by `new2.js` at the per-card commit step, the cross-card integration commit, and the skill's post-run deferred-card reconciliation — always merging into the existing `provenance` block, never overwriting `planning_session`.

### Fixed

- **`team-mode.md` used the non-existent `CLAUDE_SESSION_ID`** (`framework/.claude/skills/new/references/team-mode.md`) — the subagent-transcript locator (`ls ~/.claude/projects/*/${CLAUDE_SESSION_ID:-*}/subagents`) always fell through to the wildcard fallback because the variable name was wrong. Corrected to `CLAUDE_CODE_SESSION_ID` so it resolves the live session's `subagents/` dir directly.

## [4.71.0] - 2026-06-25

**Responsive-awareness in UI review — catch a viewport-scoped edit that regresses the *other* viewport.** When an agent works on one viewport (fix the mobile header, tweak the desktop table), nothing today catches it silently shifting the viewport it was *not* editing. The obvious fix — cloning the i18n ESLint gate into a `responsive-gate` that bans raw `matchMedia`/breakpoint literals — was **refuted** by an adversarial pass: that gate is blind to where responsiveness actually lives (Tailwind `md:` classes, CSS, CSS-in-JS — not `window.matchMedia`), so it is mayo-specific opinion, not a distributable mechanism. The real driver is served by **reactivating machinery that already exists and was inert**: `visual-fidelity-verifier` has had a `responsiveness-break` finding all along, but `/e2e-review` only ever screenshotted **one** viewport (1440px), so it could never fire.

**MINOR** — adds review capability; **no new config key** (rides on `features.has_e2e_review` + `features.has_design_system`), no schema-change propagation. The mobile viewport is fixed at 375px (the canonical small-screen audit width). Degrades to no-op on consumers without e2e-review. The per-route cost was **measured** on a real verifier run before shipping default-on: ~one extra verifier call per renderable route (~0.3% of a card run — the 375px screenshot is lighter than the 1440px one, and only renderable routes incur it).

### Changed

- **`/e2e-review` captures both viewports + runs a mobile responsive-integrity pass** (`framework/.claude/skills/e2e-review/SKILL.md`). Phase 3 now captures `<route-slug>.png` (desktop 1440×900, **unchanged path** — all existing consumers intact) AND `<route-slug>@mobile.png` (375×812), setting the viewport before navigating so SSR/responsive hooks render for that width. Phase 4 adds a second verifier invocation per route with `viewport_role: "mobile"` (step 2b), and the `skip`-level pre-filter no longer suppresses it (the mobile pass needs no mockup). Findings aggregate with `viewport_role` in the dedup key so a mobile `responsiveness-break` stays distinct and the Phase 5 gate / self-heal report show which viewport broke.
- **`visual-fidelity-verifier` gains a `viewport_role: "mobile"` responsive-integrity mode** (`framework/.claude/agents/visual-fidelity-verifier.md`). When set, it does NOT diff against the (desktop) mockup — the mobile layout reflows by design — and emits ONLY genuine breakage findings (`responsiveness-break`, `unreachable-action`, `layout-break`), conservatively. Reuses the existing taxonomy + output schema; no new finding codes.
- **`design-system-protocol.md` adds a "Responsive Scope Discipline" section** — the textual SSOT: (1) a viewport-scoped edit MUST NOT alter the other viewport's branch; (2) a structural/heavy component MUST NOT be swapped across viewports via `display:none`/`visibility`/`hidden`/`@media` CSS hiding (mounts both branches, runs both effect/query sets, ships dead DOM). Stated abstractly — concrete primitive names / breakpoints / any project-local gate live in the consumer's own files, never in the framework. Enforced as a **review check** (the CSS-hidden case is invisible to a JS AST rule), not a lint.
- **`code-reviewer` enforces it** (rule 9) — structural `display:none` swap = HIGH; viewport-scoped edit that mutates the other branch = MEDIUM. **`design-review.md`** checklist points to the same section. This is the surviving, generalized slice of the responsive work; everything stack-specific (a project's `eslint.responsive.config.mjs`, primitive names) stays a consumer-side project file, never in `.framework/`.

## [4.70.0] - 2026-06-25

**`/ds-edit` — the deliberate-edit twin of `/ds-new`, completing the single-element lifecycle.** `/ds-new` creates (its reuse-first step STOPS if the element exists); there was no symmetric guided entry to *change* an existing canonical element. "Update a component" decomposes into three cases — two were already covered: the **mechanical doc-resync** after a code change is automatic (serializer regenerates the HEAD from source, preserving agentic + prose; the BLOCKING `DS_COMPONENT_STALE` gate enforces it), and **agentic re-curation** is `doc-reviewer`'s domain. The gap was the **deliberate change with a decision** — add a variant/prop, a breaking change, or re-governing a closed-set role.

**MINOR** — adds a skill; no new config key (rides on `features.has_design_system`). A pure orchestrator over the same scripts as `/ds-new` — zero rule duplication, and the SAME canonical template (`framework/templates/component-spec.template.md`).

### Added

- **`/ds-edit` skill** (`framework/.claude/skills/ds-edit/SKILL.md`) — update ONE existing canonical component or token. Pre-flight locates the element (refuses + suggests `/ds-new` if it doesn't exist), then classifies the change via the edit decision tree: **resync** (spec catches up to changed source) / **extend** (new variant/prop — within the role contract, else it's a NEW role = `/ds-new`) / **breaking** (removed/renamed variant or prop → migration + spec Changelog) / **re-govern** (the edit changes `canonical_for`/`use_when`/closure). Delegates the source change to `ui-expert`, the agentic re-verify to `doc-reviewer`, and regenerates the spec via the serializer — which **preserves the existing prose + non-empty agentic fields** (`source_sha` anchor) and fills the canonical template only when a spec is missing, so an edited spec keeps the SAME standardized shape as a created one. Verifies via `ds-gate` + `DS_COMPONENT_STALE`. Codex-portable (deterministic spine zero-subagent; agentic/scaffold steps degrade to TODO). Token edits follow the Token Cascade (rename/removal is breaking — migrate the affected `token_bindings`).

### Changed

- **`/prd` routes a REUSE+VARIANT extension to `/ds-edit`** (`ui-design-phase.md`, new step 7b) — when the reconciliation gate finds a region maps to an existing component but needs a variant/prop not yet in its spec and the user confirms extending, route to `/ds-edit` (create-now) or a "modify Component X" card (defer) — symmetric to the NEW→`/ds-new` routing. A "variant" that is really a new role is a `NEW (governed)` case, not an extend.
- **`/new` (ui-expert Post-Intervention) follows the `/ds-edit` discipline for a MODIFIED primitive** (`implement.md`) — re-run the serializer to regenerate the HEAD from the new source (preserving agentic + prose), never hand-edit; a breaking change adds a spec Changelog entry. Registered in `design-system-protocol.md` (Consumers), `component-manifest-schema.md` (Producers), and `agents/index.md` routing. Propagates to `new2` (shared reference modules).

## [4.69.1] - 2026-06-24

**`/prd` and `/new` now actually USE `/ds-new` when a new STANDARDIZED component is detected** (v4.69.0 only wired it as an opt-in arm + left `/new`'s spec production unspecified). A new canonical element should be materialized through the one discipline, not hand-rolled per surface.

**PATCH** — wires the just-shipped skill into the two orchestrators; no new config key.

### Changed

- **`/prd` defaults to `/ds-new` create-now for a STANDARD** (`ui-design-phase.md`, NEW-governed branch) — when the user confirms a new component is `standard: yes` (canonical for its role, especially a `closed` family), `/prd` now invokes `/ds-new` to materialize the canonical element + its standardized spec + INDEX entry + Selection-Policy governance **before any card runs** (so `ds-gate` enforces it and every consumer reuses it), instead of scattering the governance lazily across `/new` cards. A genuine `scope: local` one-off stays bind-and-defer.
- **`/new` produces a new primitive's spec via the serializer, following the `/ds-new` discipline** (`implement.md`, ui-expert Post-Intervention) — a NEW canonical/standardized primitive's `components/<Name>.md` HEAD is **emitted by `extract-one.mjs` → `serialize-spec.mjs`** (filling `framework/templates/component-spec.template.md`), **never hand-written** (the serialization contract), with `canonical_for`/`use_when`/`selection_closed` set per governance when it is a standard. Closes the gap where "ship the spec" let ui-expert hand-author an untemplated, possibly invalid-YAML HEAD. Propagates to `new2` automatically (shared reference modules).

## [4.69.0] - 2026-06-24

**`/ds-new` — the fifth corner of the design-system discipline: guided creation of ONE canonical element + a canonical, ENFORCED doc template.** Discovery/reuse, prevention (`/prd`), enforcement (`ds-gate`), and bulk bootstrap (`/design-system-init`) all existed — but there was no user-invocable entry point to create a single new component or token *correctly* on-the-fly (research → create → document → register → govern → verify); `ui-expert` does it inside tasks, `/design-system-init` is bulk-only. Second gap the user named: the docs were **not standardized** — the rich prose template existed but was buried in `scripts/` and (critically) **not wired into the serializer** (`emitSpec` stubbed a one-liner on empty prose), so freshly-created specs were untemplated.

**MINOR** — adds a skill + a thin script wrapper; no new config key (rides on `features.has_design_system`). A pure orchestrator: every rule stays in its existing SSOT.

### Added

- **`/ds-new` skill** (`framework/.claude/skills/ds-new/SKILL.md`) — single-element guided creation, the on-the-fly twin of `/design-system-init`, also callable by `/prd` at its reconciliation gate's "NEW (governed)" branch. THIN orchestrator: reuse-first (Component Discovery Cascade + Selection-Policy-by-role + `signatureDuplicates`) → optional scaffold (delegated to `ui-expert`, opt-in; default documents an existing source) → extract HEAD (`extract-one.mjs`) → enrich agentic fields (`doc-reviewer`) → govern closures → document+register (serializer) → verify (`ds-gate`) → coherence report. Token path: DTCG node → `baldart tokens build` → bind. Gated on `features.has_design_system` (refuse + suggest `/design-system-init` when false). Codex-portable (deterministic spine zero-subagent; agentic steps degrade to TODO). No rule re-described — cross-pointers to `design-system-protocol.md` + `component-manifest-schema.md`.
- **`extract-one.mjs`** — single-component wrapper around `extract-manifest.mjs`'s now-exported `extractFile` (the bulk extractor was `--roots <dir>` only). One extraction implementation, zero-dep, Codex-portable; verified to produce byte-identical output to the bulk path.

### Changed

- **Canonical doc template promoted + ENFORCED.** Moved `component-spec.template.md` from `design-system-init/scripts/` to the visible **`framework/templates/component-spec.template.md`** (single SSOT, framework-internal, excluded from the consumer-copy loop) and enhanced it to explicitly encode the structural-retrieval contract (HEAD-first discovery, token IDs not values, tables>prose, cross-ref-over-repetition, one-owner-per-fact — cross-pointing `doc-writing-for-rag` + `component-manifest-schema.md` § Read contract). **Wired into the serializer**: `serialize-spec.mjs`'s `emitSpec()` now FILLS the template on empty prose (placeholders from `props`/`variants`/`token_bindings`, unfillable ones keep `<!-- TODO -->`), via a new `--prose-template` flag (default resolves to the promoted path). Non-empty prose is still preserved verbatim. One change → three beneficiaries: `/ds-new`, `/design-system-init` greenfield, `ui-expert` Post-Intervention regen — every newly-created spec now follows the standardized shape.
- **`extract-manifest.mjs`** — per-file extraction factored into exported `extractFile`; CLI guarded with `import.meta` so importing has no side effects.
- **`/prd` reconciliation** (`ui-design-phase.md`) gains an OPT-IN "create now" arm at the NEW-governed branch (invoke `/ds-new`); default stays bind-and-defer. `/ds-new` registered as the single-element create path in `design-system-protocol.md` (§ When a primitive is missing + Consumers), `component-manifest-schema.md` (Producers & consumers), and `agents/index.md` routing.

## [4.68.2] - 2026-06-24

**A change's OWN documentation now closes before merge — no more "post-merge, non-blocking" doc follow-ups for what the change itself touched.** Observed on a real `mayo` FEAT-0055 run that merged with two deferred doc items: a full-sync of the `/products` screen doc (the exact route it changed, beyond a stub) and a `must_rule` clarification on `PageLayout.md` (the component it modified). Both are *own-docs* — documenting what you just shipped — yet `/new` permitted deferring them (the DS Decision Matrix routed "needs judgement" to a follow-up, and Phase 6b force-marks cards DONE regardless). Leaving the doc of the thing you just changed stale at merge is the drift entry door the Post-Intervention philosophy exists to close.

**PATCH** — tightens the deferral classification of an existing gate; no new config key, no new agent/skill. Deliberately NOT a blanket "all docs block merge" (that regresses the follow-up mechanism by coupling ready code to genuinely-separate design decisions) — only *own-docs* block; adjacent/cross-cutting docs still defer to a tracked follow-up.

### Changed

- **New SSOT § "Documentation Closure"** in `agents/workflows.md` (twin of the AC Scope Closure Discipline, STRICTER) — a change's OWN docs must be SYNCED in the same PR before merge, no follow-up escape, a stub is not synced. "Own-docs" is **deterministic** (not the agent's discretion to relabel as adjacent): (1) the spec of any modified primitive, (2) the screen/route doc for a touched route, (3) the card's declared `documentation_impact`/`canonical_docs`. Anything outside those three is adjacent and still deferrable.
- **`design-system-protocol.md` Decision Matrix tightened** — a spec update that DOCUMENTS the behaviour THIS change introduced (a new/changed `must_rule`/prop/variant/slot/a11y hook) is **inline-BLOCKING, not deferrable**; only a genuinely SEPARATE design decision (a new variant's full semantics, breaking-change migration design, an unresolved a11y trade-off) remains the deferrable case.
- **`doc-reviewer` Own-Documentation Closure gate** — own-doc obligations are `synced` or a **BLOCKER** (`DOC_OWN_STUB` for a stub-only own-route doc, `DS_COMPONENT_STALE` for a modified-primitive spec), emitted as `BLOCKER` (never a `LOW`/`MED` note, never a silent "deferred"). Flows through the existing pre-merge severity gate (an unresolved BLOCKER blocks the merge — `final-review.md`) and the `doc` domain-override (never inline-bypassed). Shared agent + shared `new-final-review.js` → enforced in BOTH `/new` and `new2` with no separate wiring.

## [4.68.1] - 2026-06-24

**`/design-system-init` now actually ACTIVATES the closed-set gate it just populated — found on the first real `mayo` upgrade run.** The v4.68.0 upgrade correctly derived + confirmed the closed-set data (the two-header `page-header` family + a `range-slider` pair landed in the component HEADs with `canonical_for`/`selection_closed: true`), but `baldart ds-gate` stayed `no-policy` and never enforced. Root cause: step 6's serializer bundle merged "deterministic (step 3) + agentic (step 4)" but NOT the step-4b/4c selection-policy fields, so `emitSelectionPolicy` ran over components without `canonical_for` and generated an EMPTY `INDEX.md § "Selection Policy"` block (the HEADs got the fields via a side path; the gate reads the INDEX block). Data present, gate dormant.

**PATCH** — corrects the bootstrap write/verify flow; no schema/API change.

### Fixed

- **Selection-policy fields flow into the `--index` bundle** (`design-system-init/SKILL.md` step 6 + workflow step 5) — the serializer bundle MUST carry `canonical_for`/`use_when`/`selection_closed` for ALL components so `emitSelectionPolicy` generates the populated `INDEX.md § "Selection Policy"` block. Building the bundle from `extract-manifest` only / pre-4c agentic JSON leaves the block empty.
- **New BLOCKING verification step 6b/5b** — when any closure was confirmed, the skill now asserts `INDEX.md` contains `## Selection Policy` AND `baldart ds-gate` no longer reports `no-policy` (the gate reads the COMMITTED baseline, so the INDEX block must be committed) before declaring the closed-set bootstrapped. Closes the "data in HEADs but gate dormant" gap that would otherwise ship silently.

## [4.68.0] - 2026-06-24

**Closed-Set Selection Policy — the third semantic layer that makes the "third header" structurally impossible.** Found on real `mayo` PRDs: every PRD ignored the two existing canonical headers (`SectionHeader` depth-0, `DrilldownHeader` depth-1+) and wrote new ones, requiring an expensive consolidation epic (`FEAT-0045` + an ADR) and leaving `Header.tsx` deprecated-but-live in 3 call-sites. Root cause (confirmed by a BALDART trace + 40-source 2026 research): the registry has a token layer and a component layer, but **no layer asserts the *boundary* of a category** — so adding a third header violated no rule, only an undiscovered convention. The registry-first cascade was advisory at generation-time + reactive at review-time; nothing prevented the duplicate at build-time. The 2026 market confirms the gap: no vendor (Figma Code Connect, Storybook MCP, shadcn) hard-gates the reuse decision, and no deterministic tool detects a semantically-equivalent rewrite — but BALDART already ships the machine-readable manifest signature, so it is "one manifest-signature-diff away".

**MINOR** — adds a capability (the `baldart ds-gate` verb, three manifest fields, the `/prd` reconciliation gate); rides on `features.has_design_system` (NO new config key — the schema-change propagation rule does NOT apply), additive, no install change. Three slices.

### Added

- **Selection-policy manifest fields (Slice A)** — `canonical_for` (`"<family>@<context>"` roles a component is THE canonical answer for), `use_when` (the selection predicate, lifted from `must_rules`), and `selection_closed` (the role family is a closed set) join the AGENTIC HEAD in `component-manifest-schema.md`. The serializer (`serialize-spec.mjs`) emits them + aggregates a generated **§ "Selection Policy"** `families:` block into `INDEX.md` (a family is `closed` iff ≥1 member declares it) — the closure lives in the HEADs, the INDEX stays a pure derived router.
- **`baldart ds-gate` (Slice A — the deterministic build-time gate)** — BALDART-owned, the design-system analogue of the i18n gate. Reads the **committed** `INDEX.md` § Selection Policy (baseline closure) + the diff and BLOCKS (exit 1) on `DS_CLOSED_SET_VIOLATION`: a change that adds a new canonical component in a CLOSED family. Zero-false-positive (it checks declared `canonical_for` against baseline members). `DS_CLOSED_SET_SUSPECT` (name heuristic) and `DS_DUPLICATE_PRIMITIVE` (Slice C manifest signature-diff, Jaccard ≥0.7) are advisory warnings. Self-scoping + self-skipping (registry/policy absent → clean no-op); SKIPs, never silently passes. Logic in `src/utils/ds-reuse-gate.js`, verb in `src/commands/ds-gate.js`. Wired into `qa-sentinel`, `/new` Phase 2 (`implement.md` Step 8), `new2`, the Final review, and `code-reviewer`.
- **`/prd` Component Reconciliation gate (Slice B — the human-in-the-loop prevention)** — `ui-design-phase.md`'s passive "Component Registry Lookup" becomes an active *match-before-generate* gate: build a Mockup Element Inventory → classify each region by ROLE (closing the name-vs-need vocabulary gap) → ask ONLY on ambiguity (`REUSE+VARIANT` / `NEW`, never plain reuse) in the prohibition+alternative form → persist. A closed-family hit surfaces *"this becomes the standard?"*; a confirmed standard writes `canonical_for`/`selection_closed` + opens a migration follow-up card (never migrates existing views in-PRD).
- **`component_bindings` card field (Slice B)** — the resolved `region → component (reuse|reuse-variant|new)` map (`card-schema.md`) is the authority `ui-expert` implements against: the mockup becomes a visual reference, the bindings decide which component each region IS. Resolves the mockup-vs-registry contradiction in `implement.md` (the old "implementation MUST match the approved mockup 1:1" that let the approved-mockup pixels beat the dusty registry).

### Changed

- **`/design-system-init` derives the closed-set on upgrade — confidence-gated (Slice A)** — the bootstrap auto-fills only what is unambiguous (`canonical_for`/`use_when` lifted from each component's `must_rules`, e.g. *"MUST use only for depth 0"* → `page-header@depth-0`) and **interacts for the judgement calls**: it COLLECTS closure candidates (families that look like they partition all cases) + suspected duplicates (manifest signature-overlap ≥0.7, reusing `signatureDuplicates`; deprecated-but-used primitives) and confirms them with the user in ONE consolidated `AskUserQuestion` (step 4c — family-level, ~5–10 confirmations for a ~100-component registry, never per-component). `selection_closed` is written only for confirmed families; a "consolidate" answer opens a migration follow-up card. Unattended / Codex runs propose-only — never auto-close (closing a family without a human is the governance bypass the policy forbids). Idempotent: a settled family is not re-asked. The bootstrap-time twin of the `/prd` reconciliation gate.
- **Closed-set enforcement across the discipline (Slice B/C)** — the Component Discovery Cascade now queries by role first (`design-system-protocol.md` § Closed-Set Selection Policy, the textual SSOT); `ui-expert` (cascade step 6 + Implementation briefing), `ui-design` (closed-set query + bindings output), `code-reviewer` (rule 4 closed-set BLOCKER + Post-Intervention check (e) + the new drift code, UI-reinvention fixes route to `ui-expert`), and the weekly `ds-drift` routine (Selection-Policy integrity audit) all consume it. New canonical drift code `DS_CLOSED_SET_VIOLATION`.

## [4.67.1] - 2026-06-24

**Rule C is now i18n-aware — a presentational card that only changes labels no longer gets forced to `balanced`.** Found on a real `mayo` run: FEAT-0050 is a "PURAMENTE PRESENTAZIONALE" fidelity-drift epic, yet its child cards ran the full per-card review cluster (full-depth Codex + simplify) instead of the lighter UI path. Root cause: the i18n layer (v4.52.0) silently broke the v4.x UI-presentational carve-out in `prd-card-writer.md § Rule C`. The LIGHT branch (b) requires "≤3 files, every one a component/style file", and the SKIP rule requires "all files .md/.yml/CSS" — but with `features.has_i18n: true`, every label change forces an edit to **all** native locale files (`it.ts`/`en.ts`/`de.ts`/`es.ts`/`fr.ts`), which blows the ≤3 count and injects non-component `.ts` files. So a genuinely pixels-only card was structurally incapable of being classified `light`/`skip`.

**PATCH** — corrects a blind spot in existing Rule C logic; reuses `features.has_i18n` + `i18n.locales_root` (both already in the schema), no new config key, no install change. Schema-change propagation rule does NOT apply.

### Fixed

- **i18n locale-file transparency in Rule C** (`prd-card-writer.md`) — when `features.has_i18n: true`, locale files under `i18n.locales_root` are now TRANSPARENT to BOTH the LIGHT branch (b) and the SKIP rule: excluded from the file count AND the component/style test, classifying the card on its non-locale files only. A locale file is a flat string table whose correctness is owned by the i18n gate (anti-hardcoded lint) + `i18n-translator`, NOT the logic cluster — so it must not disqualify the presentational carve-out. Effect: a genuinely-presentational card (`.tsx`/`.css` + N locale files) now correctly lands on `light` (full→light Codex depth, doc-review deferred to the Final FULL gate), and a pure-cosmetic CSS/copy card that retouches locale strings now correctly lands on `skip`. The DS safety net is unchanged — a `light` card still runs its per-card Codex finder + `code-reviewer` rule 8 / registry-first cascade, plus the batch-wide Final FULL gate.
- **Why Codex stays per-card (not deferred to Final)** — deferring per-card Codex onto the Final gate for UI cards was considered and rejected: it reverts the deliberate v4.18.0 choice (at `light`, Codex is the SOLE deep finder) and re-opens the refuted `slimCodex` design (v4.56.1 — per-card runs PRE-fix/PRE-E2E, the Final runs POST-fix over the whole batch; they review different states). The sanctioned lever for "less Codex on a graphic card" is the `light` profile's reduced Codex depth, which this fix now lets presentational cards actually reach.

## [4.67.0] - 2026-06-23

**Token-binding CORRECTNESS check + PRD-Inventory-sourced enrichment — the survivor of an adversarial pass that killed a design-time "intent seed".** Goal: capture design intent ("when/how/which-tokens") so the implementing agent doesn't reverse-engineer the manifest's agentic fields from code. The proposed `component-intent.yml` seed (emitted by `/prd`) was refuted 3/3 — it duplicates the PRD's existing UI Element Inventory + the manifest (twin), goes stale via `/prd-add` REDO, survives-but-wrong (the v4.66.2 guard checks token EXISTENCE, not correctness-for-this-component), and as an on-disk prose file would be skipped/unowned/Codex-inert. This ships the deterministic survivors instead.

**MINOR** — adds a capability (usage check + intent-source wiring); no new config key, no install change. Schema-change propagation rule does NOT apply.

### Added

- **Token-binding correctness check (`serialize-spec.mjs --verify-usage`)** — closes the half v4.66.2 left open (existence ≠ correctness): a binding can exist in the DTCG yet be wrong for this component. `verifyBindingUsage(components, cwd)` (exported) reads each component's effective styling surface — its own `source` (+ sibling `.module.css`/`.css`) **plus the sources of the primitives it directly `composes`** (a wrapper presenting an overlay via a composed Popover legitimately binds `zIndex.overlay`) — and reports any `token_bindings` id absent from that surface as `unverifiedBindings`. **ADVISORY, never a drop/block** (a binding can be legitimately indirect — Tailwind classes, a non-sibling stylesheet — so a hard block would false-fire). Resolves source from the `source` field, never a guessed path. Validated on mayo: 96% of 603 bindings verified; the residual surfaces for review, exit 0.

### Changed

- **Intent-sourced enrichment (no new artifact)** — `doc-reviewer` + `/design-system-init` step 4 now fill the agentic HEAD fields (`purpose`/`a11y`/`must_rules`) from the PRD's existing **UI Element Inventory** (`Purpose`/`function_ref`/`Behavior`/`States`) when the component traces to a PRD (card `links.prd`/`links.design`), reconciled against the implemented source — capturing design-time intent instead of reverse-engineering it. Clean fallback to source-only when no PRD exists. `token_bindings` are now picked preferring ids the source actually uses.
- **Advisory propagated to the three Post-Intervention sites** — `design-system-protocol.md` (check 3), `code-reviewer.md` (Rule 8), and the `ds-drift` routine now flag a resolves-but-unused binding as a `DS_TOKENS_DRIFT` **advisory** (review-required: fix or justify as indirect), distinct from the hard existence drift. SSOT'd in `component-manifest-schema.md` (existence = HARD/drop; usage = ADVISORY/report). `/design-system-init` passes `--verify-usage` on every generation.

## [4.66.2] - 2026-06-23

**`token_bindings` are now validated against the real DTCG vocabulary — fixing dangling bindings found on a real upgrade.** After the v4.66.1 HEAD-serialization fix, a real `/design-system-init` upgrade still produced 26/586 `token_bindings` that didn't resolve in `.tokens.json` (`font.weight.medium` instead of `typography.fontWeight.medium`, `z.overlay` vs `zIndex.overlay`, `space.2` vs `space.sm`, …). Root cause: the agentic enricher (`doc-reviewer`) writes `token_bindings` as free text and guesses plausible-but-wrong names. A dangling binding id is `DS_TOKENS_DRIFT`. This closes the gap so it cannot be written, and instructs the enricher to pick from the actual token vocabulary.

**PATCH** — refines the v4.65.0/v4.66.1 layer; no config/install change.

### Fixed

- **Deterministic binding vocabulary guard** — `serialize-spec.mjs` gains `--tokens <.tokens.json>`: it flattens the DTCG to its real id set and **drops + reports** any `token_bindings` id that does not resolve (the `droppedBindings` count + the dropped list on stderr) — so a dangling binding can never be serialized. Exposed as `tokenIds()` + `validateBindings()`. Proven: `z.overlay`/`font.weight.medium` dropped, `zIndex.overlay`/`typography.fontWeight.medium`/`color.surface` kept.
- **Prevention at the source** — `/design-system-init` step 4 + `doc-reviewer` (HEAD owner) + `component-manifest-schema.md` now require `token_bindings` to be chosen from the real `${paths.design_tokens}` id list (exact ids, never invented names). The serializer invocation passes `--tokens ${paths.design_tokens}` so the guard runs every generation.

## [4.66.1] - 2026-06-23

**The component-manifest HEAD is now serialized DETERMINISTICALLY — fixing invalid-YAML HEADs found on a real `/design-system-init` upgrade.** A real upgrade run produced 74 per-component specs but **73/74 frontmatter HEADs failed to parse as YAML** — the v4.65.0 skill let the model hand-write the HEAD, so it emitted `props:{}` / `variants:[]` (no space after the colon) and unquoted/multiline `purpose`/`a11y` strings (embedded `:`, quotes, newlines). The HEADs were unparseable → "frontmatter-first" discovery silently fell back → the machine-readable benefit was lost. Same run also wired the DTCG token output to **overwrite a hand-authored `tokens.ts`** (semantic exports `uiTokens`/`categoryPalette` replaced by a generic `export const tokens` — would break every import). This release closes the model-in-the-loop-for-a-deterministic-artifact gap (the recurring lesson behind `setup-worktree.sh` / `merge-worktree.sh`).

**PATCH** — bugfix to the v4.65.0 layer; no new config key, no install change. Schema-change propagation rule does not apply.

### Fixed

- **Deterministic HEAD serializer** — new `framework/.claude/skills/design-system-init/scripts/serialize-spec.mjs` (pure Node, zero-dep): the model supplies field **data** (deterministic JSON ∪ agentic JSON), the script emits valid YAML (every string scalar via `JSON.stringify` = a valid YAML double-quoted scalar; empty collections `{}`/`[]`; block lists; flow-map props) **and the thin INDEX router**, then **self-validates every HEAD and fails closed** (never writes an unparseable HEAD). Proven on the exact DrilldownHeader content that previously failed (backticks, colons, quotes, newline in `purpose`) → parses clean. `/design-system-init` now routes all spec writes through it and **salvages** agentic values from existing malformed HEADs before re-serializing; the spec template marks the HEAD serializer-owned.
- **Token-output clobber guard (ENFORCED)** — `/design-system-init` must run an **export-preservation check** before pointing `design_tokens.outputs` at an existing token file: the output path may be a hand-authored file only if every existing `export` is reproduced; otherwise it writes to a NEW path and leaves the hand-authored SSOT untouched. The deferred "don't clobber in place" guard from v4.65.0 is now mandatory, not advisory. The config note must state the real direction (`.tokens.json` → output), never the reverse.
- **Invalid-HEAD detection (safety net)** — `baldart doctor` now parse-checks a sample of spec HEADs (not just "is there a `---`") and the `ds-manifest-nudge` reports `N/M sampled HEADs invalid YAML → re-run /design-system-init`; the `ds-drift` routine + `code-reviewer`/`doc-reviewer`/`component-manifest-schema.md` treat an unparseable HEAD as `DS_COMPONENT_STALE` drift. This is the check that would have caught the bug at ship time.

## [4.66.0] - 2026-06-23

**Server-side reviewer death is no longer a silent drop in the `/new` review workflows — from a real `new-final-review` run where `qa-sentinel` died on `API Error: 529 Overloaded`.** When a reviewer agent inside `new-final-review.js` / `new-card-review.js` was killed by a transient server-side error (529 Overloaded, 429, 5xx, rate-limit — the org-level limit is SHARED across every terminal/session), the worker pool's `try/catch → null` + `.filter(Boolean)` **silently discarded it**. For `qa-sentinel` this was worst: a dead gate-runner produced an **empty `gateTable`**, which the fan-in read as "no mechanical gates ran ⇒ nothing to block" — so the batch's lint/tsc/test/build/i18n merge gate was lost without any FAIL signal. The two workflows had NO transient retry and NO degraded-coverage signal at all; the canonical handling already existed in `new2.js` (`agentSafe` + `noteDegraded`) and in the prose SSOT (`references/team-mode.md` § transient-vs-genuine) but had **never been propagated** to the other two workflows.

**MINOR** — reliability hardening of two distributed dynamic workflows; additive, backward-compatible (the return shape gains an additive `summary.degradedReviewers`), no new config key, schema-change propagation rule does NOT apply. Claude-only (workflows are Claude-only).

### Fixed

- **`new-final-review.js` + `new-card-review.js` — transient-aware reviewer spawn (SSOT parity with `new2.js`).** Lifted the canonical `TRANSIENT`/`isTransient`/`agentSafe` helper (the same code `new2.js` already uses) into both workflows and routed every review thunk (codex / doc-reviewer / api-perf-cost-auditor / qa-sentinel / simplify / security) through a new `reviewSafe()` wrapper. A reviewer killed by a transient error is **retried in-workflow** (cap 3; honest about the lack of a reliable runtime sleep — this absorbs brief blips, it is NOT timed backoff). A sustained outage exhausts the cap and is **recorded**, never silently dropped.
- **A dead `qa-sentinel` now BLOCKS the merge instead of passing blind.** Both workflows inject a synthetic `{ gate: 'mechanical-gates (qa-sentinel)', status: 'FAIL', detail: … }` when qa-sentinel returns null after retries, so `summary.failingGates` (and the skill's F.5 / wave-boundary gate) treat the UNKNOWN mechanical gates as **non-PASS**. Verified deterministically: the synthetic FAIL flows into `failingGates` and does NOT read as a build PASS (so the O5(b) build-claim demotion stays suppressed). Also applied to `new-card-review.js`'s light-tier post-fix scoped qa-sentinel.
- **Degraded-coverage ledger.** Both workflows return an additive `summary.degradedReviewers: [{ reviewer, reason }]` listing reviewers that died after retries, and log it on completion — so the `/new` skill knows the batch ran with INCOMPLETE coverage rather than mistaking a thin `findings` set for "clean". `new2.js` was already correct (its authoritative mechanical gate is the owner-run Phase-2 `buildBlocked` path, transient-handled; the qa-sentinel reviewer death already triggers `noteDegraded`) — **left untouched** to avoid perturbing the A/B experiment.

### Changed

- **`references/final-review.md`** — documents the non-silent reviewer-death contract in the delegated-workflow return-shape section + the F.3 qa-sentinel row (dead qa ⇒ synthetic FAIL ⇒ merge blocked; `summary.degradedReviewers` ⇒ surface terse / AUTONOMOUS materializes a re-run follow-up). Framed as the **per-agent twin of the F.1.5 whole-workflow-killed recovery**.
- **`references/team-mode.md`** — the transient-vs-genuine § now cross-links the workflow twin, making the SSOT relationship bidirectional (same discipline, two surfaces: prose for orchestrator-spawned teammates, workflow JS for in-workflow fan-out).

## [4.65.0] - 2026-06-23

**Machine-readable component manifest + DTCG token SSOT — from a real slow-discovery diagnosis on a consumer (mayo).** UI component discovery was slow because the design-system `INDEX.md` had become a 33KB monolith read in full on every UI task, there was no token reference (tokens parsed from a 17KB `tokens.ts`), and 0/~135 components had a per-component spec — so every reuse lookup degraded to reading component source. The framework already prescribes per-component specs (owned by `doc-reviewer`, gated by `code-reviewer`, audited by `ds-drift`); this release **evolves that existing layer** (no twin) into a machine-readable, on-demand one, and gives tokens a real W3C DTCG source of truth.

**MINOR** — adds a CLI verb + a schema module + a util registry + config keys; additive, backward-compatible with existing prose specs, rides on `features.has_design_system` (NO new feature flag). The **schema-change propagation rule applies** to the new `paths.design_tokens` + `design_tokens:` block (template + configure + update detector + doctor + CHANGELOG).

### Added

- **`framework/agents/component-manifest-schema.md`** — SSOT for the machine-readable frontmatter HEAD on each `components/<Name>.md`: deterministic fields (from TS/git — `name`/`source`/`source_sha`/`props`/`variants`/`composes`/`category`/`status`) vs agentic fields (curated by `doc-reviewer` — `purpose`/`token_bindings`/`a11y`/`related`/`must_rules`). Defines the on-demand read contract (INDEX router → frontmatter-first, prose only when modifying), the `source_sha`-anchored regeneration contract, and the **transition-leniency** rule (prose-only specs and empty deterministic fields stay valid — the HEAD is an upgrade target, not a new floor). Mirrors `card-schema.md`. Registered in `agents/index.md`.
- **DTCG token SSOT + generator (zero npm dependency)** — `src/utils/token-emitters/` (REGISTRY pattern: `ts` + `css` emitters, extensible) + `src/utils/tokens-generator.js` (parse `.tokens.json`, resolve `{aliases}` with cycle guard, `build()`/`isStale()`) + the **`baldart tokens build`** CLI verb (`src/commands/tokens.js`). Generated outputs carry a `baldart-generated from .tokens.json` banner.
- **Deterministic, Codex-portable extractor** — `framework/.claude/skills/design-system-init/scripts/extract-manifest.mjs` (pure Node, no deps): emits the deterministic HEAD from React+TS sources with **no subagent** (validated on mayo's real 73 components: 100% source_sha, ~73% props).
- **`framework/docs/COMPONENT-MANIFEST-LAYER.md`** — authoritative layer design / lifecycle / invariants (parallels `I18N-LAYER.md` / `CODE-GRAPH-LAYER.md`).
- **Config keys** `paths.design_tokens` (DTCG source) + `design_tokens.outputs[]` (generation targets) — propagated through template, `configure` (autodetect + prompt, gated on `has_design_system`), `update` detector (nested block + array), and `doctor`.

### Changed

- **`/design-system-init`** — now bootstraps OR upgrades: a GREENFIELD mode and an UPGRADE mode (the mayo "prose-only / monolith / no DTCG" case) that regenerates the per-component HEAD from source, lifts prose into agentic fields, rebuilds the **thin INDEX router**, and bootstraps `.tokens.json` from the existing token file with a **round-trip verify**. No longer refuses on an existing registry. Spec template (`component-spec.template.md`) gains the frontmatter HEAD.
- **`design-system-protocol.md`** — canonical sources updated (INDEX = generated router; DTCG token SSOT inversion; `tokens-reference.md` demoted to a generated/narrative view); Authority Matrix made canonical in the protocol (out of the generated INDEX); Post-Intervention Check (b)/(c) require HEAD regeneration + DTCG edit-then-build; `DS_TOKENS_DRIFT` redefined as "generated output ≠ `baldart tokens build`".
- **Consumers wired** — `code-reviewer` Rule 8 (HEAD-first reads + transition leniency), `codebase-architect` Reuse Analysis (frontmatter-first discovery), `ui-expert` (read HEAD-first; write HEAD in the completion gate), `doc-reviewer` (owns the agentic HEAD fields; DTCG-aware token reconciliation), `context-primer` (UI fast-path), `ds-drift` routine (`source_sha` + DTCG output checks).
- **`baldart doctor`** — two backfills under `has_design_system`: `tokens-build` (regenerate drifted token outputs, autoOk) and `ds-manifest-nudge` (registry present but specs missing/prose-only → run `/design-system-init`) — the "default-on + nudge" behavior achieved without a new flag.

## [4.64.0] - 2026-06-23

**Off-context final-merge conflict resolution — from a real heavy-drift `/mw` on `/new`.** On a long epic batch the trunk had drifted 45 commits while the batch ran; the final `/mw` (invoked via `Skill()`, so it runs IN the orchestrator's context) hit mixed conflicts and resolved them **inline** — Read/grep/sed/python churn re-reading the giant end-of-batch cache, exactly when tokens cost the most (the `~20M-token` hand-merge regression `merge-cleanup.md` already measured). This release moves the merge off that context the right way, without repeating the v4.46.0/v4.53.0 "model-in-the-loop for plumbing" mistake.

**MINOR** — adds one agent + one SSOT script; opt-in, additive, runtime-detected, with the legacy `Skill(/mw)` path kept as the fallback. **No new `baldart.config.yml` key** → the schema-change propagation rule does NOT apply.

### Added

- **`scripts/merge-worktree.sh` — the deterministic SSOT for the whole merge** (the merge analogue of `setup-worktree.sh`), consumed identically by `/mw`, `/new` Phase 6, and `new2`. It runs the entire sequence — safety commit → rebase onto `origin/$TRUNK` → **deterministic** resolution of the ADDITIVE conflict classes (structured registries with `validate_structured_md`, metrics JSONL, generic docs/config) → toolchain-aware build (hard timeout, SKIP tier) → land via `git.merge_strategy` (`pr`/`local-push`) → local-trunk sync (markers written to the **status file**, not stdout) → worktree cleanup → registry removal. It writes a structured status file on **every** exit path (`success | code_conflict | build_fail | sync_needs_decision | error`) with a **provisional non-terminal** status written before any mutation, so a kill mid-rebase/mid-build never leaves a stale `success` (no false-pass). The ONE thing it never does is resolve a **code** conflict (irreducible judgment): it pauses the rebase, records `conflict_files`, and exits `code_conflict` for the resolver. "A deterministic script cannot fabricate or stall."
- **`merge-conflict-resolver` agent (Sonnet, effort medium)** — the judgment layer, spawned by `/new` Phase 6 / `/mw` **only** on `status:code_conflict`. It classifies each paused code/test hunk **additive** (distinct imports/decls/rows, colliding index → renumber → resolve, build-verified) vs **semantic** (same logic/value/signature → STOP, never guess), then re-invokes `merge-worktree.sh --continue` (looping per replayed commit) to land — all in a **fresh, isolated context** so the conflict churn never re-enters the orchestrator. It is the ONE bounded exception to `coder`'s "branches off-limits" rule (worktree-scoped, merge phase only) and returns COMPACT (a one-line `MERGED <sha>` / `STOP: <file>:<hunk> semantic` + `path:line`). The orchestrator runs an **anti-fabrication disk gate** on its return (no residual markers + `status:success`) — "came to rest ≠ done", never trusts the prose.

### Changed

- **`/new` Phase 6 launches the script as a BACKGROUND Bash and reads only the small status file** (`merge-cleanup.md`), instead of `Skill(/mw)` inline — so the orchestrator's in-context cost is one launch + one status read, with the conflict churn entirely off-context (deterministic in the script, judgment in the resolver subagent). The HARD "the orchestrator never hand-merges" rule stays (the script is a deterministic delegate, the resolver a subagent); `code_conflict` → spawn the resolver; `build_fail`/semantic-STOP/`error` → interactive `AskUserQuestion` / autonomous follow-up; **script absent OR `Task` tool unavailable (Codex / older subtree) → fall back to the legacy `Skill(/mw)` inline** (zero regression). Sync markers are transcribed from the status file into the tracker so Phase 6c's hygiene gate is unchanged.
- **`/mw` (worktree-manager skill) now consumes `merge-worktree.sh`** (new step 1b) the same way `/nw` consumes `setup-worktree.sh`; the inline steps 2–7 remain as the human-readable SSOT the script mirrors and the fallback for older subtrees.
- **`new2` merge phase runs the deterministic script** instead of the model-in-the-loop "/mw programmatic, run git yourself" prose — removing the model from the merge plumbing. Honoring its OPS/GIT role boundary (never edits source, F-030), a `code_conflict` is left+reported → tracked as a merge blocker → follow-up (behavior-preserving vs today's `/mw`-STOP); the full off-context resolver is wired on the classic `/new` path. `MERGE_SCHEMA` gains `status` + `conflictFiles`.
- **`new-session-audit.mjs` + `/new-audit`** gain a **merge-phase telemetry detector** (turns/cache_read between the merge launch and Phase 6b + `code_conflict_hit` YES/NO) — non-blocking, to measure ex-post that the off-context move reduced Phase 6 cost and to gauge real conflict frequency.

## [4.63.0] - 2026-06-23

**Fixes from the FEAT-0042 post-mortem — the first real `/new` run on v4.62.0.** v4.62.0 worked (that run billed −29% vs the FEAT-0041 baseline, 1 unplanned fix-coder vs 4, `/mw` and the durable tracker held), but the deep post-mortem surfaced a new class of bug and two real leaks — including one *introduced by* v4.62.0 itself.

**MINOR** — adds a `doctor` diagnostic + guidance/agent edits; no layout/install change, **no new `baldart.config.yml` key**.

### Fixed

- **CLI ↔ framework version-skew silently disabled the Codex broker teardown (the v4.62.0 leak)** — `merge-cleanup.md` / `new/SKILL.md` / `new2/SKILL.md` + `src/commands/doctor.js`. The framework prose ships via the git subtree (current), but `npx baldart` resolves the SEPARATELY-installed CLI; on a consumer whose global CLI predated v4.62.0, `npx baldart teardown-codex-broker` returned `unknown command`, which `2>/dev/null || true` masked as a no-op → the Codex broker + its Playwright/… MCP children survived every run (FEAT-0042: 5 brokers leaked). The teardown snippets now **capture stderr (no masking), detect `unknown command`, warn to `npm i -g baldart@latest`, fall back to a direct `pkill` of the broker by cwd, and assert the broker is actually gone**. `baldart doctor` gains a **CLI↔framework version-skew check** (warns when the installed CLI is older than the framework, so new `npx baldart` subcommands can't silently fail). The earlier `$MAIN`-non-persistence hypothesis was **refuted** by the post-mortem's adversarial verification — the run passed the correct absolute `--cwd`.
- **Fix-pass coder could edit the live MAIN repo (worktree-isolation gap)** — `coder.md` + `completeness.md`. The reactive AC-fix-pass spawn had no canonical worktree-isolation clause (it lived only in `implement.md`'s per-card brief), so an improvised thinner brief dropped it and a fix-coder Edited `<main>/src/lib/i18n/*.ts` on the MAIN repo, then reverted + re-applied (~3M wasted; the `tsc 2353` it triggered was a stale artifact, not a real defect). Now: `coder.md` STEP 0 carries an **intrinsic per-Edit path guard** (every Edit path must be worktree-rooted; `cd` is not enough; assert `.worktrees/` before editing shared/locale files) — applies to every spawn even when the brief omits it — and the fix-pass spawn brief must carry the verbatim isolation clause.
- **i18n locale-parity gate ↔ coder "do-not-translate" contradiction** — `implement.md` / `coder.md` / `i18n-protocol.md` / `prd-card-writer.md`. The `coder` writes the source locale only (by design), but a project's own locale-parity test (a `test`/`build` gate) fails on a source-only commit → `/new` spawned `i18n-translator` reactively (~2M round-trip). Resolved: a **planned per-card `i18n-translator` fill pass** runs after the coder and before the parity gate (when `has_i18n` + new source keys), a parity failure is explicitly an `i18n-translator` task (never a coder retry), and the ownership docs (coder STEP 9.6, the i18n cascade, the P3 DoD line) are reconciled so the coder stays source-only with no contradiction.

### Changed

- **`new-session-audit.mjs` detector false-positives fixed (audit tooling).** `B1_mw_delegation` now inspects the actual `Skill` tool-call inputs (not the raw transcript, which quotes `new:new`/`__phase6_merge__` as the merge-cleanup.md "don't do this" text → a false FAIL even when the `/mw` call was correct); `B3_nul_grep` matches only genuine output phrases (`Binary file … matches`, `graphify-out`), not the v4.62.0 guidance text; in-context-merge detection is scoped to Bash commands; and command/mode detection now selects the actual `/new` line (not a later `/model`).
- **Tracker-write cadence coarsened (O2/M3).** The v4.62.0 "lone tracker-Edit = violation" prose did not change behavior (a measured run still paid 38 standalone tracker-Edit turns ≈13.6M / 9%), so `SKILL.md` § Update rules now mandates tracker writes only at **card-state-changing boundaries**, always co-emitted with that action — recovery reconstructs the phase from commits + `phase_module_loaded` + backlog statuses, so per-phase writes are not needed.
- **markdownlint debt prevented at `/prd` authoring (L1/L2)** — `prd-writing-phase.md` gains a markdown-hygiene rule: escape a literal `|` in a prose/table cell (`\|` or a code span — the `isManagerOf||isOwner` incident), honor MD058, and run `markdownlint` on generated docs when configured, rather than shipping debt for `/new`'s doc-review to catch.

### Deferred (documented)

- **Codex sentinel-JSON prompt tax (~270k/batch)** — `new-card-review.js` already tolerates Codex's prose output (it fetches both the sentinel JSON and the prose tail), the gain is ~0.1%, and the Codex relay is a delicate, much-iterated subsystem; not worth a regression risk in this sweep.

## [4.62.0] - 2026-06-22

**Token-economy + reliability hardening for `/new` and `/prd`, distilled from a real `/new FEAT-0041 -full -auto` run that billed ~364M tokens (~95% cache_read).** A deep multi-agent post-mortem (with adversarial verification that refuted several plausible-but-wrong hypotheses — e.g. the workflow-`<result>`-payload-bloat theory was false by ~430×) traced the cost to a small set of root causes: in-context work that should have been off-context (a hallucinated `/mw` delegation and a killed-final-review recovery, ~36M combined), friction multipliers (a NUL-byte/GNU-grep trap ~8–12M, ~28 lone tracker-Edit turns ~11M), and `/prd` design gaps that pushed work downstream into `/new`'s expensive orchestrator context. Plus a proactive Codex-broker teardown so terminated reviews stop leaking MCP processes.

**MINOR** — adds one CLI verb (`baldart teardown-codex-broker`) + new util export + skill/agent guidance edits; no layout/install change, **no new `baldart.config.yml` key** (so the schema-change propagation rule does NOT apply). All `/new` paths without the `Workflow` tool, and all Codex-less installs, are unaffected.

### Added

- **`baldart teardown-codex-broker --cwd <path>`** (`src/commands/teardown-codex-broker.js` + `bin/baldart.js`) + **`tearDownCodexBroker({cwd, pluginScriptsDir})`** in `src/utils/codex-orphans.js` — a **proactive, cwd-scoped** teardown of the Codex `app-server` broker bound to a specific worktree, freeing its Playwright/Figma/… MCP children. The Codex plugin already ships a clean per-cwd broker shutdown but only fires it on Claude Code `SessionEnd`; a `/new` run that ends (or is killed mid-final-review) hours earlier leaves the broker + MCP subtree burning CPU. Tier 1 reuses the plugin's own `loadBrokerSession`→`sendBrokerShutdown`→`teardownBrokerSession` (graceful); Tier 2 (plugin lib absent / no `broker.json` but the process lingers) matches the broker by its `--cwd` token and SIGTERM-pgroup→SIGKILLs the subtree (reusing `listProcesses`/`collectTree`). Scoped strictly by cwd — never touches a broker for another worktree or the user's interactive session. Validated end-to-end against a real 50-minute-orphaned broker (killed 11 processes). Doctor's `reapOrphans` stays the reactive safety net.

### Fixed (`/new` — bugs surfaced by the post-mortem)

- **`/mw` Phase-6 merge now has a concrete `Skill()` invocation** (`merge-cleanup.md` + `worktree-manager/SKILL.md`). The prose-only "delegate to `/mw`" let the orchestrator hallucinate `Skill({skill:"new:new", args:"__phase6_merge__ …"})` → `Unknown skill` → it ran the **entire** Phase-6 merge inline in its ~590k-token context (~20M cache_read). Now: a literal `Skill({skill:"worktree-manager", args:"/mw worktreePath=… checksAlreadyPassed=true"})` + a HARD "never fall back to manual git" rule (interactive → ask; AUTONOMOUS → follow-up card).
- **Killed `new-final-review` workflow now recovers off-context** (`final-review.md` F.1.5 + `new-final-review.js`). When the workflow was KILLED mid-Fix-phase (no structured return, no notification), the orchestrator hand-parsed the journal (explicitly forbidden) and re-verified all findings + committed fixes in-context (~16M). New "workflow-killed recovery" sub-protocol: detect non-completion, delegate ONE off-context subagent to validate+commit the worktree's uncommitted Fix diffs, never re-verify in-context, never block on a human under AUTONOMOUS. The workflow now emits a **terminal completion sentinel** (its absence ⇒ killed) and the Fix output is documented as uncommitted-until-F.5.
- **`rg` over GNU `grep` for source searches** (`code-search-protocol.md` + `coder.md`). A single NUL byte in a large generated file silently flips GNU `grep` to binary mode (empty output + exit 1, indistinguishable from "no match"); three coders independently re-discovered one such byte in a `data.ts` (~8–12M). Now a HARD default: prefer `rg`; treat empty+exit-1 on an expected-match file as a binary-mode trap, never as "symbol absent".
- **Batch tracker relocated off volatile `/tmp`** (`SKILL.md` + `metrics.md`). The recovery-SSOT tracker lived in `/tmp` and was wiped during a pause, forcing an in-context rebuild. Now `$(git rev-parse --git-common-dir)/baldart/run/batch-tracker-<FIRST-CARD-ID>.md` — repo-local, session-stable, survives the worktree removal, never in `git status` (no partition/`.gitignore` change). Plus a degrade branch that reconstructs run state from commits + backlog statuses if the tracker is ever missing.

### Changed (`/new` — optimizations)

- **A turn whose only tool call is a tracker `Edit` is now an explicit VIOLATION** (`SKILL.md` § Context economy) — a measured run paid ~28 such standalone turns (~11M). Ride the tracker write along with the next real tool call.
- **Out-of-lane `preValidated` guard + build-claim reconciliation** in `new-final-review.js`. A narrow specialist (doc-reviewer / api-perf-cost-auditor) FP-checking a finding OUTSIDE its lane no longer short-circuits to VERIFIED — it routes to the correct domain specialist. And a VERIFIED BLOCKER/HIGH asserting a build/compile failure that the SAME review's build/tsc gate contradicts (PASS) is auto-demoted to NEEDS_MANUAL_CONFIRMATION instead of gating the merge on a phantom (the false-positive perf BLOCKER the post-mortem caught).
- **Codex broker teardown wired into the run-end finalizers** — `merge-cleanup.md` Phase 6c (cwd-scoped teardown FIRST, then the cumulative `reap-orphans`), `SKILL.md` terminal-hygiene (HALT-before-6c), `new2/SKILL.md` Step 5, and a `codex-gate.md` note that the broker is shared/warm and must NOT be torn down per-card.

### Changed (`/prd` — design gates that prevent downstream `/new` waste)

- **Consumer enumeration for shared multi-mode functions** (`discovery-phase.md` ISA dimension 9) — when the feature changes a shared list/query/selector/filter fn used in >1 mode/call-site, EVERY consumer is now enumerated as its own ISA touchpoint. This closes the gap that let an orders-wizard catalog picker keep showing hidden products (the exclusion was modeled only at the picker-mode call; the catalog-mode consumer page was unowned → surfaced as 2 HIGH only at the final review, ~28M).
- **Per-AC seam-file closure + `requirements`-at-write-time** (`card-schema.md` authoring invariants, `audit-phase.md` Codex audit, `validation-phase.md` BLOCKING gates d/e, `prd-card-writer.md`). Every AC's render/seam file (and every file named in requirements/AC/integration_points) must be in `files_likely_touched` (⊆ check); every ISA touchpoint must be owned by some card; and a derivable `requirements` block must be emitted at authoring time — so `/new` no longer pays MAY-EDIT stalls, unplanned fix-coders, or an in-context requirements back-fill.
- **i18n DoD parity gated on `features.has_i18n`** (`prd-card-writer.md`) — a card introducing user-facing strings now carries a DoD line requiring translation to ALL maintained locales (not just "registered"), shifting the reactive `i18n-translator` spawn author-time-left.

## [4.61.1] - 2026-06-22

**`/new` pre-flight no longer wrongly HALTs/skips well-specified `/prd` cards that merely omit the top-level `requirements` block.** Real-world friction: implementing a `/prd`-designed epic with `/new` repeatedly hit *"N cards lack the top-level `requirements` field (profile CHILD)"* even though those cards had complete `scope`, `scope_boundaries`, `acceptance_criteria`, and `business_rationale` — the requirements were fully derivable, and the user had to hand-back-fill them every time or risk the gate destroying the epic. Root cause: `card-schema.md` listed `requirements` as an **unconditional** HALT (non-derivable), so a derivable-but-omitted block forced an ask/skip. The `prd-card-writer` still mandates emitting `requirements`; this only adds the consumer-side safety net that already exists for `review_profile`/`owner_agent`.

**PATCH** — repairs an over-strict ingestion gate; no new capability, no new `baldart.config.yml` key, no install change. The validator and the matrix `R` classification are unchanged (the writer must still emit `requirements`; the back-fill makes it present before validation runs).

### Changed

- **`framework/agents/card-schema.md`** (SSOT consumer contract) — `requirements` reclassified from unconditional HALT to a **conditional** HALT member: it HALTs **only when its deriving material is itself absent** (`acceptance_criteria` empty OR `scope` empty — a genuinely thin card). When both are present and non-empty, a missing `requirements` block is recovered via a new BACK-FILL **sub-kind (b): faithful derivation** — a bounded restatement/decomposition of the existing AC + scope (faithful, never generative; never invents new scope).
- **`framework/.claude/skills/new/references/setup.md`** — pre-flight 1b-i drops `requirements` from the unconditional HALT list (now conditional on AC/scope being empty); 1b-ii gains the faithful-derivation back-fill (derive from AC + scope, persist to the card on disk in the main repo, log `[BACKFILL] … requirements=<N items, AC+scope-derived>`).
- **`framework/.claude/workflows/new2.js`** (parallel location, gate G4) — a card missing `requirements` with AC + scope present is now faithfully derived + written back (F-040 main-repo discipline) instead of excluded from the batch; only excluded when AC or scope is also absent.

## [4.61.0] - 2026-06-22

**UI design-quality critic — a separate, rigorous reviewer for *"is this good design?"*, distinct from the existing *"does it match the mockup?"* check.** Born from the observation that implemented UIs were often poor despite the full `ui-expert` + `/e2e-review` machinery. Root cause, confirmed by exploration: every existing UI verification checks **fidelity / conformance**, not **intrinsic quality** — `visual-fidelity-verifier` asks only whether the render matches the mockup (and never reads source, never judges aesthetics), `code-reviewer`'s design-system steward is mechanical (token/primitive compliance), and the only agent judging quality (`ui-expert`) is *also the implementer*, so it grades its own work. The generator-grades-itself antipattern is exactly what the critic-in-the-loop research (vision-guided iterative refinement; rubric-based LLM-as-judge with per-dimension calibration) identifies as the failure mode. This release adds the missing layer: a separate critic that judges the rendered output against a scientific rubric, in the existing bounded self-heal loop.

**MINOR** — adds one agent (`ui-quality-critic`) + one SSOT rubric section + `/e2e-review` Phase 4b wiring. **No new `baldart.config.yml` key** (gates reuse `features.has_e2e_review` + `features.has_design_system` + `features.e2e_review.fidelity_tolerance` + `max_self_heal_iterations`), so the schema-change propagation rule does NOT apply. Additive: with `features.has_e2e_review: false` (or no screenshots) the new phase is skipped and every path is byte-identical to v4.60.0.

### Added

- **`framework/.claude/agents/ui-quality-critic.md`** — new stateless multimodal agent (`model: opus`). Judges **intrinsic design quality** of a rendered UI screenshot against a fixed 10-dimension rubric (visual hierarchy, typographic rhythm, spacing rhythm, color/contrast harmony, density consistency, composition/balance, state & affordance completeness, motion quality, polish/AI-generic smell, brand coherence), each with explicit *excellent vs poor* calibration anchors. Emits a severity-tagged JSON report + per-dimension scores + region-anchored findings (`source: "design-quality"`). Inherits `visual-fidelity-verifier`'s anti-bias discipline (never reads source code), adds an explicit *never grades its own design* rule, grounds every numeric threshold in the existing Reference Tables (never invents numbers), and reuses the canonical severity taxonomy so its findings flow through the same gate.
- **`framework/agents/design-system-protocol.md` § "Design Quality Rubric"** — new SSOT section owning the 10 dimensions, their weights (sum 1.00), and the dimension→severity mapping. The per-dimension anchors live in the agent body; the dimension list/weights/mapping live here so they cannot drift across consumers. Not gated on `features.has_design_system` (quality is judged on every surface; thresholds ground in the project's `tokens-reference.md` when present, the Reference Tables otherwise).

### Changed

- **`framework/.claude/skills/e2e-review/SKILL.md`** — new **Phase 4b — Design Quality Critique**: invokes `ui-quality-critic` on every route with a screenshot (including no-mockup `skip` routes; no pixel-diff pre-filter — quality is not pre-filterable). Findings join Phase 5's `findings[]` under the same tolerance filter; Phase 5b's bounded self-heal loop now **routes `design-quality` (and UI-presentation `visual`) fixes to `ui-expert`** (the UI domain owner) and functional/structural fixes to `coder`. Report schema gains a `quality` block (`quality_score_avg`, `weakest_dimensions`). Orchestrator overview, Interaction, and Failure-Modes sections updated (added `quality_critic_protocol_violation` + non-converging-quality-loop diagnostics).
- **`framework/.claude/agents/visual-fidelity-verifier.md`** — added an explicit scope boundary vs `ui-quality-critic` (fidelity/conformance vs intrinsic quality) so the two never overlap.
- **`framework/.claude/agents/REGISTRY.md`**, **`README.md`** (agent count 29→30 + inventory entry), **`CLAUDE.md`** (agent count + new convention bullet) — kept in sync.

## [4.60.0] - 2026-06-21

**`/new` cost + autonomy release — born from a forensic post-mortem of a real `-full` epic run** (224M orchestrator cache_read over 543 turns, context grown to ~670k, 9 improvised on-context fix agents, Codex completing only 2/7 reviews). The dominant cost is mechanical — turn-count × monotonically-growing context, confirmed against Anthropic's current prompt-caching docs — and the back half of the run was the expensive half because the final-review fix cascade and the deploy phase ran ON the orchestrator context at ~670k/turn. Two adversarial review passes shaped and verified the change set (several proposed items were refuted as already-shipped or redundant; four blockers found in pre-release review were fixed before tagging).

### Added

- **`/new -auto` / `-auto-ship` — autonomous epic mode.** `/new <ID> -full -auto` runs the whole batch with ZERO `AskUserQuestion`: every decision gate resolves deterministically via the new **§ AUTONOMOUS RESOLUTION RULE** (SSOT in `SKILL.md`, cited by every gate site) — the recommended option where one exists, else a **no-drop follow-up card** (never a silent deferral, never a scope/AC reduction), with irreversible/outward actions governed by the **§ AUTONOMOUS SAFETY BOUNDARY**. Motivated by convenience AND cost: removing the human-gate idle waits keeps the prompt cache warm (the post-mortem run lost ~1.2–1.5M tokens re-warming a ~670k prefix across two idle gaps that exceeded the 1h subscription TTL). Reuses the deterministic-gate machinery already battle-tested in `new2`. `AUTONOMOUS` also activates from `BALDART_AUTONOMOUS`/`CI`/`GITHUB_ACTIONS` (unified codepath). Under `-auto`, team-mode is forced to sequential single-worktree (≈7× cheaper). **Safety:** bare `-auto` ships NOTHING outward (remote db:push / deploy / scheduled-functions / secret / DNS → deferred follow-up); `-auto-ship` + the new `git.auto_deploy` allowlist authorize ONLY the DoD-declared deploys, with `stack.schema_deploy_from_trunk_only` still a hard floor; force-push / reset-hard trunk and release/production merge are HARD-NEVER under any flag.
- **`git.auto_deploy`** config key — allowlist bounding what `-auto-ship` may auto-execute (default `[]` = nothing auto-deploys). Propagated end-to-end (template + configure prompt + update detector + doctor diagnostic).

### Changed

- **`framework/.claude/workflows/new-final-review.js`** — off-context Fix phase (OPT-IN, default OFF). The final review now applies its VERIFIED findings INSIDE the workflow (sequential security-reviewer → coder → doc-reviewer, off the orchestrator context) when the caller passes `applyFixes:true` + `editableFiles`. This removes the post-final-review fix cascade the post-mortem showed running on the orchestrator at 460–670k/turn (the dominant late-run cost). The classic `/new` opts in (final-review.md §F.1/§F.5); `new2` and any non-opting caller get the byte-identical review-only return (no regression).
- **`framework/.claude/workflows/new-card-review.js` + `new-final-review.js`** — Codex prose path. When Codex finishes but emits prose instead of the sentinel JSON (≈5/7 of the post-mortem run), the seeded path now runs a CHEAP transcription-only converter (no duplicate full re-hunt); its findings are NON-preValidated AND confidence-clamped (<80) so the existing Verify phase still validates each over real code (`codexEngine: 'codex (prose-converted)'`). Recovers the wasted second review pass while preserving the trust model.
- **`framework/.claude/workflows/new-card-review.js`** (+ the final-review Fix phase) — collision-safe `finding_id` reconciliation. A writer that echoes only the human-readable id tail no longer drops applied fixes into residual (the empty-coder re-spawn bug): `resolveId` maps a returned token to its canonical prefixed id by exact or UNIQUE colon-tail match, logging any unmatched. Also a light-tier test gate: after the Fix phase at light tier (where qa-sentinel was skipped), a diff-scoped qa-sentinel pass now catches a fix-introduced test regression via the existing gateTable gate (graceful SKIP when no diff-scoped runner).
- **`framework/.claude/skills/new/` references** — autonomous gate coverage + turn economy. Every `AskUserQuestion`/HALT gate across SKILL.md + `implement.md`, `setup.md`, `review-cycle.md`, `team-mode.md`, `completeness.md`, `merge-cleanup.md`, `commit.md`, `codex-gate.md`, `final-review.md`, `production-readiness.md` cites the AUTONOMOUS rule (so none silently blocks under `-auto` and none auto-approves a deferral); a sequential-path coverage assertion (review-cycle.md) catches a silently-skipped test-only card review; one concrete batching example added to the turn-economy HARD RULE.

**MINOR** — additive: two new flags (`-auto`/`-auto-ship`) + an opt-in workflow Fix phase + one new `baldart.config.yml` key (`git.auto_deploy`, schema-change propagation applied) + reference-prose gate coverage. No new agent/skill/command/routine; no install/layout change. With the flags/args absent and `applyFixes` not opted in, every path is byte-identical to v4.59.0 (`new2` unaffected).

## [4.59.0] - 2026-06-20

**Subagent return contracts — every orchestrator-facing agent now bounds its *final message* so a long body doesn't flood the spawning context.** Outcome of studying the [headroom](https://github.com/chopratejas/headroom) context-compression tool against BALDART's existing context-economy stack. The study's verdict: headroom is **not** adopted as a dependency — it is a heavyweight Python+Rust runtime (proxy/wrapper) that contradicts BALDART's "thin CLI + authored markdown, zero runtime" identity, is explicitly unsuitable for the ephemeral cloud sandboxes these sessions run in, and applies *lossy* compression to code/tool output (risky in a code-gen framework). Its techniques already have prompt-level analogues here: verbosity steering L0–L4 ≈ the Terse / Terse Ultra output styles, effort routing ≈ `effort-protocol`, AST/code compression ≈ the LSP + Graphify retrieval layers (which filter input *losslessly*).

The one real gap headroom surfaced is the **verbosity of subagent return messages**: when free-form analysts (`codebase-architect`, `doc-reviewer`, `plan-auditor`, `senior-researcher`) finish, their entire reasoning body lands verbatim in the orchestrator's context — and the orchestrator pays for it on *every* subsequent turn of a batch. A subagent's final message is the only text that costs the spawning context tokens (its intermediate reading/thinking is discarded), so the whole optimization reduces to one rule: **persist the long form, return the short form.** This is implemented natively — zero dependencies, portable across Claude + Codex, works inside sandboxes — as the input-side twin of the effort protocol: effort governs how *deeply* an agent reasons, the return contract governs how *compactly* it reports back.

The contract: COMPACT (default) = bounded headline/verdict + key points as `path:line` references (no reasoning dump, no pasted code) + `Report: <path>` when the long form is on disk; FULL narrative only when the user invoked the agent directly for prose. It reuses the existing pooled YAML findings schema (never a new one) and explicitly does **not** truncate substance — it cuts ceremony/narration/echo, not real findings. Agents already structured this way (`qa-sentinel` Return Protocol, `visual-fidelity-verifier` JSON, the verdict+pooled-YAML reviewers `code-reviewer` / `security-reviewer` / `api-perf-cost-auditor`) are the reference exemplars and were left untouched.

**MINOR** — adds a protocol module + a template snippet (additive capability); no new `baldart.config.yml` key, so the schema-change propagation rule does not apply (same class as `effort-protocol`). No install/layout change.

### Added

- **`framework/agents/return-contract-protocol.md`** — new SSOT protocol module: COMPACT vs FULL return modes, the persist-then-summarize rule, findings-schema reuse, honest limits, and the canonical `## Return Contract` body block. Modeled on `effort-protocol.md`.
- **`framework/templates/agent-return-contract.snippet.md`** — copy-paste starter for the `## Return Contract` block (twin of `skill-effort.snippet.md`), with placement guidance and the "skip agents already structured" carve-out.

### Changed

- **`framework/.claude/agents/{codebase-architect,doc-reviewer,plan-auditor,senior-researcher}.md`** — added an explicit `## Return Contract` block (tailored per agent: architect → file/symbol/pattern/risk map as `path:line`; doc-reviewer → condensed verdict + docs on disk; plan-auditor → verdict + pooled YAML, 10-section narrative is FULL-only; senior-researcher → §0–§11 report written to disk, return the Executive Summary + pointer).
- **`framework/agents/index.md`** — added the return-contract routing rule and a Modules-list entry.
- **`framework/.claude/agents/REGISTRY.md`** — added the return-contract convention note (exemplars + "paste the block for new analysis/audit/research agents").

## [4.58.0] - 2026-06-20

**`simplify` gains a 4th Reuse class — "a new third-party dependency reimplements a native/stdlib/platform capability" — closing the one genuine residual from the (otherwise refuted) ponytail analysis.** Follow-up to a user observation that `simplify` relates to the ponytail learnings. The accurate picture is the reverse: ponytail was **refuted, never integrated** (zero mentions anywhere in `framework/`), because `simplify` already covered its concepts in superior form. That analysis left exactly **one** genuinely uncovered item (recorded as a ~1-line residual): the reuse lens did not flag *pulling in a dependency* for something the language/stdlib/runtime/platform already does (e.g. a date/util library for what `Intl` / `Temporal` / built-ins provide). The v4.57.0 coder reflex touched only the adjacent *inline-reimplementation* (hand-rolling) case, not the new-dependency case.

The residual had been tagged "measure first" under the data-driven principle, but that principle targets **arbitrary numeric thresholds** ("N lines = inline"); this is a **qualitative** reuse lens (the established you-don't-need-lodash discipline), so it ships directly. Added consistently across every parallel location: `simplify` Agent 1 (a new dependency check + the new classification), Step 3 (aggregation), Step 4 (the fix: replace with the native primitive, remove the package, or flag when removal is non-trivial/shared), the delegated workflow's Reuse lens (`new-card-review.js`), and the coder's author-time reuse reflex (now two reflexes: no inline-reimpl + no native-duplicating dependency; a genuinely needed new dep stays an ADR decision). The simplify↔coder cross-pointer's Reuse example list is updated to match.

**MINOR** — additive review capability (a new finding class the `simplify` pass can surface) + the matching author-time reflex; no new agent/skill/command/routine/template, no new `baldart.config.yml` key (schema-change propagation rule does not apply), no install/layout change. The review contract (`runSimplify`, gating) is unchanged.

### Changed

- **`framework/.claude/skills/simplify/SKILL.md`** — Agent 1: new dependency-vs-native check (step 5) + new **"Dependency reimplements native"** reuse class; Step 3 aggregation gains it (reuse classes 1–4, then Quality/Efficiency renumbered 5/6); Step 4 gains its fix recipe; the author-time-mirror back-pointer's Reuse example list updated.
- **`framework/.claude/workflows/new-card-review.js`** — the delegated `simplifyPrompt` Reuse lens now also covers a new third-party dependency reimplementing a native/stdlib/platform capability (parity with the skill, so the `/new` delegated path applies the lens too).
- **`framework/.claude/agents/coder.md`** — `## Author-Time Simplicity Discipline` Reuse reflex extended from one to two reflexes: also do not add a dependency for a capability the stdlib/runtime/platform already provides (a needed new dep is an ADR decision, not a reflex).

## [4.57.0] - 2026-06-20

**`coder` now prevents the avoidable code smells at authoring time — an author-time subset of the `simplify` taxonomy applied WHILE writing — without removing the independent `simplify` review pass.** Triggered by a user question: "can we save ourselves the `simplify` agent by folding its job into the `coder` as prevention, avoiding the double pass?" An extensive review (codebase + scientific literature, then an adversarial pass) established that the "double pass = redundant" framing is a trap: the two passes are **complementary, not redundant**, and removing the independent one is the risky half.

The evidence cut both ways and shaped a **hybrid** outcome (the user's choice over full elimination):

- **For prevention (sound):** LLM-authored code carries measurably more code smells than human reference solutions (+63% average, Codex +85% — [arXiv 2510.03029](https://arxiv.org/abs/2510.03029)) and the more capable models bias toward over-building / bloated procedural logic ([AI-Generated Smells, arXiv 2605.02741](https://arxiv.org/html/2605.02741v1)). Shift-left says the cheapest smell is the one never written, and Self-Refine shows a *single* author-time critique step yields a real gain (~+13% on code — [arXiv 2303.17651](https://arxiv.org/abs/2303.17651)).
- **Against removing the reviewer (the risk):** without *external* feedback, intrinsic self-correction often *degrades* output ([LLMs Cannot Self-Correct Reasoning Yet, ICLR 2024, arXiv 2310.01798](https://arxiv.org/abs/2310.01798)), and an LLM is **not** reliably better at *discriminating* its own output than at generating it (54/56 experiments — [SELF-[IN]CORRECT, AAAI 2025, arXiv 2404.04298](https://arxiv.org/abs/2404.04298)). The author has a structural blind spot on the code it just wrote; the independent `simplify` pass supplies the external signal (component registry + grep/LSP/graph reverse-impact, fresh eyes, different context) that the in-flight coder cannot. This also matches BALDART's own recorded principles ("no self-judging adversarial"; the refuted "Lazy Ladder pre-gen in the coder").

So the fix adds **prevention in the coder** for the patterns that are pure author-time discipline (Quality + Efficiency — they need no fresh cross-codebase scan) and **keeps `simplify` untouched** as the independent net whose load-bearing value is cross-codebase **Reuse** detection. Net effect: the coder ships cleaner first drafts → the review pass finds and re-fixes less → fewer fix→re-review cycles, with the safety net intact.

Design properties that make this low-risk: (1) the discipline lives in the **agent** (`coder.md`), so it reaches every spawn path — `/new`, `new2`, `/bug`, standalone — with zero per-skill wiring; (2) **no change** to the review contract — `runSimplify`, `review-cycle.md`, `new-card-review.js`, `new2.js` are all untouched, so the independent pass still runs; (3) **anti-drift via bidirectional cross-pointers** (the same pattern already used for the Reference-Aliasing rule): the full taxonomy + detection/fix workflow stay SSOT in `simplify/SKILL.md` Step 2, the coder carries a terse "keep-in-sync" mirror of the author-time subset only; (4) it does **not** duplicate what the coder already does — Reuse and Design-System are already covered by `Before Creating New Components` and STEP 8, so the new section fills exactly the Quality+Efficiency gap.

**MINOR** — additive behaviour capability on the `coder` agent; no new agent/skill/command/routine/template, no new `baldart.config.yml` key (schema-change propagation rule does not apply), no install/layout change. README inventory counts unchanged.

### Changed

- **`framework/.claude/agents/coder.md`** — new `## Author-Time Simplicity Discipline (MUST)` section (between `Code Standards` and `Reference-Aliasing Mutation Hazards`): prevention-framed, terse one-line-per-pattern checklist mirroring the Quality + Efficiency author-time subset of `simplify/SKILL.md` Step 2, plus one Reuse reflex (no inline reimplementation of an existing utility/stdlib) and an explicit SSOT + keep-in-sync note. `Code Standards` and the closing "Keep it simple" line now point to it.
- **`framework/.claude/skills/simplify/SKILL.md`** — Step 2 gains a back-pointer to the coder mirror, declaring this pass the independent net (load-bearing on cross-codebase Reuse) and itself the SSOT for the taxonomy (keep the coder mirror in sync).

## [4.56.3] - 2026-06-20

**`/new` Codex relays made Haiku-proof — the JS now owns the success decision and deterministically recovers Codex's findings, so no relay model (whatever tier the runtime hands it) can drop them.** Direct follow-up to a user question — "why is the per-card Codex relay still running on Haiku?" Forensics on the real consumer run (reading `<session>/subagents/workflows/wf_*/agent-*.jsonl`) established two facts that recontextualise v4.56.2:

1. **`model: 'sonnet'` is a request, not a guarantee.** Across 19 real relay runs, the *same* `model:'sonnet'` line produced `claude-sonnet-4-6` on some runs and `claude-haiku-4-5` on others — non-deterministic, no `fallbackModel` configured, no settings pin. Claude Code silently downgrades a workflow subagent from Sonnet to Haiku under capacity pressure (the per-card relay, spawned in a repeated Discovery fan-out, saturates account-level Sonnet capacity in a long session; the lone Final relay rarely does). So v4.56.2's haiku→sonnet bump is best-effort only — it can't *guarantee* the relay isn't Haiku, and the whole "Sonnet fixes it" premise is unenforceable from a `.js` workflow (the runtime picks the model).

2. **Codex usually writes PROSE, not the sentinel JSON.** The real `/tmp/codexreview-wave-*.md` had **zero** `<<<FINDINGS_JSON>>>` sentinels — Codex ignored "emit ONLY the sentinel block" and wrote a `**Findings**` prose report ending in `Turn completed.`. So the sentinel-extraction path the relay was built around frequently extracts nothing, and the findings live in the prose. The v4.56.2 failure (empty `codexProse` → cold fallback ran *unseeded* → missed the BLOCKER) was the relay model failing to capture that prose, not failing a JSON pass-through.

Fix (chosen by the user over a runtime-config force, which can't ship and is blunt): **take the success decision out of the relay model and make recovery deterministic + JS-owned.** The relay now returns RAW command outputs (`marker`, `jsonBlock`, `proseTail`) and decides nothing. The JS fan-in parses `jsonBlock` itself (`null`=no JSON, `[]`=Codex ran clean, `[..]`=findings); when there's no parseable JSON and Codex didn't report ABSENT, a one-shot re-extraction agent (near-zero latitude — two fixed commands) re-runs BOTH the sentinel `awk` AND the `[codex]`-stripped prose `tail` over the on-disk file, so a relay that went off-script or returned an empty `proseTail` can no longer strand the cold fallback unseeded. Result: Codex's findings are recovered whether it emitted sentinels or prose, and whatever model the runtime assigned the relay. Verified deterministically against the real failing transcript (6/6 scenarios, incl. "Haiku returned empty proseTail" → 2304 chars of prose recovered → cold fallback seeded). `model: 'sonnet'` is kept as a best-effort request (it follows steps more reliably when capacity allows) but is no longer load-bearing.

`new2.js` is intentionally **not** touched: its per-card Codex path is a different design (a foreground `--wait` DRIVER that maps findings by model, already on Haiku by design, with its own `note:'codex-unavailable'` → code-reviewer fallback) and it is the experimental A/B variant — porting this would contaminate the comparison.

**PATCH** — robustness fix in the two workflow files that share the THIN-RELAY pattern (`new-card-review.js`, `new-final-review.js`); no install/layout change, no new config key, no prose-SSOT change (the relay's return schema is internal to the workflow). The Codex task prompt, the sentinel contract Codex is *asked* to honour, and the prose-salvage→code-reviewer trust model are all unchanged.

### Changed

- **`framework/.claude/workflows/new-card-review.js`** + **`framework/.claude/workflows/new-final-review.js`** — Codex relay restructured: `CODEX_SCHEMA` now returns raw `{marker, jsonBlock, proseTail}` (no `codexAvailable`/`findings` self-report); the relay prompt drives + returns verbatim stdout and decides nothing; the fan-in owns the parse via `parseFindingsArray` and adds a deterministic one-shot re-extraction (`EXTRACT_SCHEMA` `{jsonOut, proseOut}`, `reExtractPrompt`, `AWK_EXTRACT` SSOT) that recovers both the sentinel JSON and the prose from the on-disk file. `model: 'sonnet'` retained as best-effort with corrected rationale comments (runtime downgrades it under load; correctness no longer depends on the tier).

## [4.56.2] - 2026-06-19

**`/new` per-card Codex relay bumped haiku → sonnet — stops a real run from discarding a valid Codex review (incl. a BLOCKER) and falling back to a cold, weaker `code-reviewer`.** A user reported the symptom directly: in the per-wave Discovery phase Codex emitted a perfect sentinel-wrapped JSON array (3 findings, one a BLOCKER race condition on the idempotency key), yet the workflow still ran the cold `code-reviewer (fallback)` — which re-reviewed from scratch (~101k tok) and **missed the BLOCKER**.

Root cause: the per-wave Codex relay in `new-card-review.js` ran on `model: 'haiku'`. The relay is "thin" but is still a model in the loop that must follow a multi-branch decision tree AND faithfully pass a large JSON array through the structured schema. Haiku proved too weak on a real run — it went off-script (re-grepped source to "confirm" findings, explicitly forbidden by the prompt) and then failed the pass-through, returning `codexAvailable:false` with empty `codexProse`. The fan-in (`item.r.codexAvailable && Array.isArray(findings)`) therefore discarded every Codex finding and triggered the **plain** cold fallback (not even the codex-seeded one). The "a misfire degrades safely to the fallback" assumption baked into the old comment is false: the cold fallback is a *different model hunting from scratch*, not a backstop — it lost the BLOCKER.

`new-final-review.js` already learned this exact lesson (its relay is `model: 'sonnet'` "not haiku because this is the last cross-model gate before merge"); the per-wave relay was simply left behind. This brings it to parity. The relay prompt is low-volume plumbing, so the haiku→sonnet delta is negligible against a dropped BLOCKER plus a ~100k-token cold re-review. Matches the recorded lessons that a weak *model in the loop for pure plumbing* fabricates/abandons, and that the cheap-model-via-structured-output bet only holds if the model can actually emit the structure.

**PATCH** — single-line model change (+ rationale comment) in `framework/.claude/workflows/new-card-review.js`; no schema/prose/install/layout change, no new config key. The Codex prompt, awk extraction, sentinel contract, and prose-salvage path are all unchanged.

### Changed

- **`framework/.claude/workflows/new-card-review.js`** — per-wave Codex relay `model: 'haiku'` → `model: 'sonnet'`, with an expanded rationale comment explaining why the misfire-degrades-safely assumption is false (parity with `new-final-review.js`).

## [4.56.1] - 2026-06-19

**`/new` per-card review-cluster — corrected a prose-vs-implementation drift about Codex depth (doc-only, no behaviour change).** Triggered by a user question — "why does `/new` run a per-card review *and* a final review for a single card, when some agents overlap?". A full read of `new-card-review.js` + `new-final-review.js` + the orchestrating modules, followed by an adversarial refutation pass, established that the apparent double-review is **not** redundant: the per-card Codex/qa-sentinel run in the **Discovery phase (PRE-fix, PRE-E2E)**, while the final Codex/qa-sentinel run on the **post-fix, post-E2E** tree — the same temporal argument the framework already uses to keep `slimDoc:false` on the delegated path (`final-review.md`). A proposed `slimCodex` (drop the final Codex on N=1 when per-card Codex ran) was **refuted and NOT implemented**: it would review materially different code and reopen the v3.35.0 scope-reduction hole that v3.37.0 closed (Codex + qa-sentinel are the explicit never-slim exceptions to F-041).

The one genuine defect the review surfaced: `references/review-cycle.md` (Phase 2.5x delegation gate) claimed the Step-A detector "also tells the workflow's Codex pass the depth." It does not — `qaTier` is binary (`light|full`, gating `qa-sentinel` only), the workflow's Codex prompt is unconditionally a "deep code review" and never reads `qaTier`, and that is correct (per `codex-gate.md`, at `light` Codex stays the SOLE deep finder — only the breadth agents qa/api/doc are dropped). Fix: the prose now states Step-A drives `qaTier` only, the per-card Codex review is always full-depth regardless of profile, and the Codex stage takes no depth input.

**PATCH** — single-file prose correction in `framework/.claude/skills/new/references/review-cycle.md`; no workflow/agent/skill behaviour change, no install/layout change, no new config key.

### Changed

- **`framework/.claude/skills/new/references/review-cycle.md`** — Phase 2.5x `qaTier` bullet: removed the false "Step-A tells the Codex pass the depth" claim; clarified that Step-A drives `qaTier` (qa-sentinel inclusion) only and the per-card Codex pass is always a full deep review.

## [4.56.0] - 2026-06-19

**Hardening from a real `/new` session post-mortem (mayo, FEAT-0035-followup batch): three protocol-violation classes the framework *allowed*, a recurring Codex waste, and a context-economy win.** The orchestrator recovered and merged all 5 cards, but only by improvising around failures the framework should have prevented. Root cause across the board: the framework trusted **prose discipline** where it needed **enforced constraints**.

1. **`codebase-architect` locked to read-only.** In the incident a read-only "arch baseline" agent — which had full tool access (no `tools:` allowlist) including the `Task` tool — decided to "complete the work" and **spawned two `general-purpose` sub-agents that implemented and committed two cards** (`[FEAT-0035]` prefix, b3e925b6a/45496634d), bypassing the entire coder→gate→review→backlog pipeline (one card's i18n AC was left undone → forced cleanup re-spawns). Fix: `framework/.claude/agents/codebase-architect.md` gains a `tools:` allowlist that **omits `Task`** (mechanically cannot spawn implementers) + a binding **Role boundary** block (mirrors `code-reviewer.md`): read-only analyst, no implement/commit/spawn, the only writes are its own agent-memory; a prompt carrying line-level fix detail is treated as malformed → return the map + flag it.

2. **Worktree isolation clause for every spawned subagent.** A `doc-reviewer` told to work "in the worktree (cd there)" edited docs via **absolute main-repo paths** (`/…/mayo/docs/…`) instead of the worktree copy — corrupting a checkout shared live with other terminals. `cd` does not stop an absolute path. Fix: a new **SSOT clause in `framework/AGENTS.md` § Worktree isolation** (every Edit/Write MUST resolve under the worktree root; never the main repo; the merge/finalizer step is the only exception), a belt-and-suspenders qualifier in `doc-reviewer.md`, and the clause injected at every spawn site (`implement.md` MISSION BRIEFING, `team-mode.md` coders, `review-cycle.md` Phase-3 doc-reviewer, `final-review.md` F.5 writers, `new2.js` `cardBrief`, `new-card-review.js` Doc stage).

3. **Arch-baseline prompt de-contamination.** The incident's baseline prompt was an implementation order ("remove ONLY lines 90-91", "Fix: use X"). A new **baseline content contract** in `implement.md` step 3 + `team-mode.md` Step A bars line-level fix directives / "implement" framing / AC-as-task from the architect prompt — it requests a read-only MAP only.

4. **Codex prose-only salvage-seed (recurring from v4.54.1).** v4.54.1's best-effort prompt hardening did NOT stop GPT-5.x emitting a prose review without the `<<<FINDINGS_JSON>>>` sentinel; it recurred (1 of 3 workflow runs), discarding 3 real findings (2 HIGH) and paying a full cold `code-reviewer` re-review (~107k tok). The deterministic fix (a JSON output-schema) exists in the companion engine but only on its `adversarial-review` path, not the `task` CLI surface we call — an **upstream** request, not shippable here. So: when Codex finishes prose-only, the relay (still dumb — never parses prose into findings) now returns the cleaned prose tail as `codexProse`, and the **fan-in seeds the fallback `code-reviewer`** with it as UNVERIFIED leads to verify against the real diff. Cross-model finds are recovered, not discarded; engine telemetry gains `code-reviewer (codex-seeded)`. In `new-card-review.js` + `new-final-review.js` (parallel locations). Validated deterministically on real prose/sentinel/not-found fixtures.

5. **Robustness.** `final-review.md` now tells the orchestrator to read the workflow's **structured return value** (final shape `{ codexEngine, findings, noActionFindings, gateTable, summary }` — distinct from `new-card-review`'s `perCard` shape) instead of hand-parsing the raw run-journal (the incident's stale `codexEngine: None` mis-read); `codexEngine` is telemetry and never gates F.5. A **test-filter precision** guard in `prd-card-writer.md` (authoring) + `coder.md` (runtime): a loose `npm test -- products` silently ran the wrong suite (matched `products-identifiers-search` not `product-picker-guard`) — prefer explicit spec paths + an `expect` that asserts the right suite ran; never report a validation passed on a no-match.

6. **Per-wave doc-review moved INTO `new-card-review.js` (context/cache economy).** doc-review was the last review-cluster piece still running at orchestrator level (its analysis + Read/Edit re-entered the orchestrator context every group). It now runs as a new **`phase('Doc')`** stage inside the workflow (off-context): doc-reviewer in audit-and-apply mode over the post-code-fix wave diff, **relevance-gated** (skips a pure code-only wave), edits docs UNCOMMITTED in the worktree (the skill stages them at Phase 4 via `perCard.docFixesApplied`). The "must see FINAL code" invariant is preserved by the **batch-Final doc-reviewer** (post-E2E backstop). The inline Phase-3 path stays as the Codex/no-Workflow fallback (prose SSOT unchanged).

**MINOR** (new enforcement + a new workflow stage; no agent/skill/command removed, no install/layout change, no new `baldart.config.yml` key — schema-propagation does not apply). A separate follow-up sweeps `tools:` allowlists across the other read-only/analyst agents.

### Changed

- **`framework/.claude/agents/codebase-architect.md`** — `tools:` allowlist (no `Task`) + Role boundary block.
- **`framework/AGENTS.md`** — new `### Worktree isolation` SSOT subsection under Git Workflow.
- **`framework/.claude/agents/doc-reviewer.md`** — worktree-rooted `docs/` qualifier.
- **`framework/.claude/skills/new/references/{implement,team-mode,review-cycle,final-review}.md`** — isolation clause at each spawn site; baseline content contract (implement/team-mode); delegation-gate update so doc-review is owned by the workflow when delegated (review-cycle); structured-return parse note (final-review).
- **`framework/.claude/workflows/new-card-review.js`** — Codex prose salvage (`codexProse` schema field + relay capture + seeded fallback); new `phase('Doc')` stage + `perCard.docFixesApplied` + `summary.docApplied`.
- **`framework/.claude/workflows/new-final-review.js`** — Codex prose salvage (parity).
- **`framework/.claude/workflows/new2.js`** — worktree-isolation line in `cardBrief`.
- **`framework/.claude/agents/{prd-card-writer,coder}.md`** — test-filter precision guard (authoring + runtime).

## [4.55.1] - 2026-06-19

**`baldart doctor`'s toolchain fresh-write paths now also wire the i18n pre-commit command (v4.55.0 follow-up).** v4.55.0 wired the guarded `i18n` lefthook command only through `configure`. The doctor backfill actions that write a FRESH `lefthook.yml` — `toolchain-install` (when lefthook resolves after a devDep install) and `toolchain-init-config` (restore a missing default config) — did not pass `i18nCommand`, so a `lefthook.yml` born from a doctor backfill on a `has_toolchain` + `has_i18n` project would be Biome-only. Both now compute the resolved (trailing-` .`-stripped) i18n command from `state.i18nEnabled` and pass it through, exactly like `configure`. Fresh-write only (`initConfig` still skips an existing file); the `i18n-precommit-snippet` WARN still covers an already-present Biome-only `lefthook.yml` (no duplication — the snippet check requires the file to exist).

**PATCH** (closes a fresh-write gap in doctor; no behaviour change for projects without `has_i18n`; no new agent/skill/command/key, no install/layout change).

### Fixed

- **`src/commands/doctor.js`** — `toolchain-install` + `toolchain-init-config` actions pass `i18nCommand` to `ToolchainInstaller.install` / `initConfigs` when `has_i18n`.

## [4.55.0] - 2026-06-19

**The i18n anti-hardcoded gate now also runs as a guarded lefthook pre-commit command — closing the direct-commit hole for projects that have BOTH the toolchain and i18n enabled.** Until now the gate ran only inside the BALDART workflows (`/new`, `/new2`, `/qa`, qa-sentinel, routine), so a *direct* `git commit` (a human, or an agent editing outside `/new`) bypassed it. When `features.has_toolchain` AND `features.has_i18n` are both true, a fresh `lefthook.yml` now gains an `i18n` pre-commit command that runs the standalone anti-hardcoded gate over `{staged_files}`. It is **guarded** (`sh -c 'if [ -f eslint.i18n.config.mjs ] || …; then exec <cmd> "$@"; fi'`): the gate's non-zero exit propagates (a hardcoded JSX string fails the commit), but when the gate config is absent it is a clean no-op (exit 0) — never a spurious block. Bypassable with `git commit --no-verify` (a strong default, not a wall); covers JSX text + allowlisted attributes (`jsx-only`); the non-JSX backstop stays with `code-reviewer`; consumers without the toolchain keep the `/new`/qa-sentinel gates as enforcement.

**Scope discipline (adversarial-reviewed).** The original idea also included a deterministic *registry-membership* pre-commit check (every `t('key')` must be in the registry). A 3-reviewer adversarial pass **rejected it** and it is NOT shipped: `I18N_REGISTRY_DRIFT` is already triple-covered (coder STEP 9 + code-reviewer + doc-reviewer); the `extract_command` activation gate was a fake proxy (the enumerator is a regex, and `extract_command` is auto-populated only for i18next/lingui/react-intl → off for the mainstream next-intl/vue/custom); the static-key regex would false-block correct commits (`useTranslation('ns')` + `t('key')`, overlay-overridable conventions); and it would break inside `/new` git worktrees (`.claude/skills/` is untracked there). The pre-commit hook was likewise scoped down — it rides the **existing** toolchain lefthook (no new installer, no new config key, no husky branch, no installing lefthook on a non-toolchain project, no js-yaml round-trip of user files).

**MINOR** (extends the toolchain lefthook's fresh-config behaviour, gated on existing flags; no new agent/skill/command, no new `baldart.config.yml` key, no install/layout change).

### Added / Changed

- **`src/utils/toolchain-adapters/lefthook.js`** — `initConfig(cwd, opts)` accepts `opts.i18nCommand`; a fresh `lefthook.yml` gains the guarded `i18n` pre-commit command alongside `biome`. Non-destructive behaviour unchanged (writes only when absent).
- **`src/utils/toolchain-installer.js`** — `install()` / `initConfigs()` forward `opts.i18nCommand` to the adapters.
- **`src/commands/configure.js`** — the toolchain block resolves `i18nCommand` from `features.has_i18n` (the diff-scoped gate command, trailing ` .` stripped) and passes it through.
- **`src/commands/doctor.js`** — WARN-only backfill: when `has_toolchain` + `has_i18n` and `lefthook.yml` has no i18n command (toolchain predating v4.55.0), surface the exact snippet to paste. doctor never mutates a user-owned `lefthook.yml`.
- **`framework/agents/i18n-protocol.md`** — documents the pre-commit layer (gated on `has_toolchain`, `--no-verify` caveat) and records why the registry-membership check was rejected.

## [4.54.1] - 2026-06-19

**Two bugs in the Codex relay of the `/new` review workflows — found live as a ~20-minute "freeze".** During a real `/new FEAT-0037` the Codex step of `new-card-review` appeared stuck; on-disk diagnosis (`/tmp/codexreview-wave-FEAT-0037*.md` + `ps`) showed **Codex had actually finished 9 minutes earlier** (full review, zero actionable findings) but the relay couldn't detect it, because of two independent, compounding bugs:

1. **`$$` placeholder in the `/tmp` filenames.** The workflows built `/tmp/codextask-${cardId}-$$.txt` / `/tmp/codexreview-wave-${cardId}-$$.md`. In a JS template string `$$` is the literal two characters — the `Write` tool keeps it verbatim, but **bash expands `$$` to the PID of each subshell**, so the launch, the poll, and the extract Bash calls each referenced a *different* path that never agreed. Observed: the first launch `cat`-ed a PID-named task file that didn't exist → Codex ran with an empty task ("Provide a prompt…", 68 bytes); the relay retried with clean names (Codex then ran for real), but the poll loop kept `grep`-ing a `-$$.md` file with a fresh PID each iteration → never terminal.
2. **Codex emits prose, not the sentinel block.** Even with the right file, Codex (GPT-5.x via the companion) wrote its findings as a prose report (`**Findings** No actionable… **Cleared Concerns** …`) instead of between `<<<FINDINGS_JSON>>>`/`<<<END_FINDINGS_JSON>>>`, so the relay's only terminal condition (`grep END_FINDINGS_JSON`) never matched → it spun to the 10-minute window before falling back to `code-reviewer`.

Nothing broke downstream (the fallback is wired), but every occurrence cost ~12 wasted minutes + a redundant full `code-reviewer` re-run. Fix:
- **Stable per-wave/batch filenames (no `$$`)** — `cardId`/`firstCardId` is already unique per wave/batch and the files are `/tmp` scratch truncated by `>` on launch, so the launch/poll/extract calls now agree on one literal path.
- **Robust terminal detection** — the relay also polls for the companion's reliable `Turn completed.` marker. New return logic: a non-empty awk extraction ⇒ `codexAvailable:true` + parsed findings; `Turn completed.`/`CODEX_NOT_FOUND` seen **but** the awk output empty (Codex finished with prose, no machine-readable block) ⇒ `codexAvailable:false` + **exit immediately** (route to the `code-reviewer` fallback — never parse findings out of prose, which would risk dropping a BLOCKER); only a full 10-minute window with no marker ⇒ false. So a prose-only completion now falls back at Codex-done time, not +10 min.
- **Task-prompt hardening (best-effort)** — Codex is told its FINAL message must be ONLY the sentinel block (no `Findings`/`Cleared Concerns`/`Validation` prose; cleared concerns go as `requires_action:false` entries inside the JSON), reducing how often the prose-fallback path is hit.

Verified by parsing both workflows (`node -c`), confirming zero `$$` remain, and running the exact `awk` + marker logic against the real incident fixture (prose-only → empty awk + `Turn completed.` → `false`/fallback) and a synthetic sentinel file (→ valid JSON, `true`). Scope is the two review workflows only; `new2.js`'s Codex usage is a different shape (a `$FILE` var, no sentinel-poll dependency; a synchronous `--wait` per-card call) and inherits the fix at the Final via its delegation to `new-final-review`. **PATCH** (relay bugfix; no new agent/skill/command, no `baldart.config.yml` key, no install/layout change).

### Fixed

- **`framework/.claude/workflows/new-card-review.js`** + **`new-final-review.js`** — Codex relay: (1) dropped the `$$` placeholder from the `/tmp` task/review filenames (stable per-wave/batch names); (2) added `Turn completed.` as a poll terminal + a deterministic 3-way return rule (non-empty JSON ⇒ available; finished-without-block ⇒ unavailable, immediate fallback, never prose-parse; full-window-no-marker ⇒ unavailable); (3) task prompt now demands the FINAL message be ONLY the sentinel block.

## [4.54.0] - 2026-06-19

**The i18n anti-hardcoded gate is now DIFF-SCOPED in every per-change context, and the two UI-authoring surfaces (`ui-expert`, `ui-design`) finally honor the i18n layer.** Surfaced by a real `/new FEAT-0037` run on a consumer (`mayo`) that is *partially* i18n-adopted: the coder had to manually triage 25 eslint-i18n errors + 4 parity failures and discovered they were **pre-existing baseline debt in files the card never touched** (`products/edit`, `products/new`, `sw-register`, …). Root cause: the i18n gate ran **whole-repo** (`npx eslint --config eslint.i18n.config.mjs .`) in every per-change consumer, while the *normal* lint gate was already diff-scoped — so on any mid-adoption project every card failed the i18n gate on unrelated baseline debt. The coder did the right thing (distrusted its own report, verified independently) but should never have had to. (The parity failure and the `Cannot find module '@/lib/i18n'` tsc note in that run were project-side, not framework — handled soundly by the coder.)

1. **Diff-scoped per-change gate.** `/new` per-card (Phase 2 step 8), the `new2` Phase-2 gate, the Final review batch gate, and `qa-sentinel` now strip the trailing ` .` full-sweep target from the resolved command and lint **only the change's changed `*.{ts,tsx,js,jsx,mjs,cjs}` files** (the card's diff per-card, the batch's diff at the Final review). Pre-existing debt in untouched files no longer fails an unrelated change; no changed JS/TS files → SKIP. A diff-scoped failure is therefore unambiguously a string the change introduced — `implement.md` step 9 now says so (no more "is this mine?" triage). The string transform (`${GATE% .}` + changed-file append) was verified on flat/legacy/path-less commands and end-to-end on a real git repo (baseline file in HEAD but outside the diff is correctly excluded from the eslint args).
2. **`i18n.lint_command` semantics unchanged** = the canonical **full-sweep** command. The full-sweep consumers (`/i18n` audit, `/i18n-adopt` migration, `i18n-align` routine) keep running it whole-repo by design (they reconcile the entire codebase). No config key change → the schema-change propagation rule does not apply.
3. **`ui-expert` + `ui-design` honor i18n.** `ui-expert` (which WRITES UI code) gains a `coder.md` STEP-9 mirror gated on `features.has_i18n` — no hardcoded user-facing strings, registry stub population, `paths.i18n_registry` in its Project Context, two new i18n red flags. `ui-design` (which produces throwaway mockups) gets a lighter note: mockup copy is realistic placeholder, the UI Element Inventory must list user-facing strings for the implementer to externalize, no registry population.

**MINOR** (corrects gate scoping behaviour + extends the capability of existing agents/skills; no new agent/skill/command, no `baldart.config.yml` key, no install/layout change).

### Fixed / Changed

- **`framework/.claude/skills/new/references/implement.md`** — Phase 2 step 8 i18n gate replaced with the diff-scoped recipe (compute changed JS/TS files via `git diff --name-only "$TRUNK...HEAD"`, strip the trailing ` .`, lint only those files, SKIP when none); step 9 states a diff-scoped failure is always the card's own string.
- **`framework/.claude/workflows/new2.js`** + **`new-final-review.js`** — Phase-2 / Final batch i18n gate prompts now instruct diff-scoping to the card's / batch's changed files.
- **`framework/.claude/agents/qa-sentinel.md`** — i18n gate prose: diff-scope to the current review scope (card diff per-card, batch diff at Final), consistent with the normal lint gate; whole-repo reserved for the full-sweep consumers.
- **`framework/agents/i18n-protocol.md`** — SSOT: new "Scope — diff-scoped in per-change contexts, whole-repo only in full-sweep contexts" paragraph.
- **`framework/templates/baldart.config.template.yml`** + **`src/utils/i18n-gate.js`** — doc/comment clarifying `lint_command` is the full-sweep command and the per-change gates diff-scope it (no functional change).
- **`framework/.claude/agents/ui-expert.md`** — `paths.i18n_registry` + `features.has_i18n` gating in Project Context; new "i18n — No Hardcoded Strings" BLOCKING section mirroring `coder.md` STEP 9; two i18n entries in the Internationalization red flags.
- **`framework/.claude/skills/ui-design/SKILL.md`** — `features.has_i18n` note in Project Context; Step G inventory must list user-facing strings (placeholder copy, no registry population).

## [4.53.8] - 2026-06-19

**Five `/new` refinements from a deep cost/correctness pass over the same `/new FEAT-0035` epic run that produced v4.53.4–6** (re-reading the orchestrator + all subagent + workflow transcripts of session `e502c890`). The v4.53.5 gate prose was present in this very run yet still not executed — the analysis isolated *why*, plus four other concrete wastes. Recurring theme: in two places the orchestrator ignored a discipline that already existed — so these make existing rules **executable/forcing**, they do not add new policy.

1. **Empty-result gate: the "read the rested agent's LAST event" step is now executable for background teammates.** v4.53.5 told the orchestrator to classify rate-limit-death vs genuine-empty-result by reading the dead teammate's last event — but a background teammate killed mid-flight delivers *nothing* to the orchestrator, so "no report + no diff" is the identical observable for both causes. In this run the orchestrator therefore classified the two rate-limit deaths (03, 07) as *generic empty-result* and re-fired both within 16s (same burst), the exact anti-pattern v4.53.5 forbids. The gate now spells out HOW: the failure signature lives ONLY in the teammate's on-disk transcript (`<session-subagents-dir>/agent-*.jsonl`, matched by label via the sibling `*.meta.json`) — locate + `tail` it, look for `isApiErrorMessage`/`Rate limited`/`Server is temporarily limiting requests`/`(not your usage limit)`, and STATE the classification + evidence before choosing the remedy. (`new2` is unaffected — its F-019 workflow runtime surfaces transient errors to the script directly, no disk dig needed.)
2. **No more SendMessage-to-fetch-a-report.** Same gate locator is reused for report retrieval: when a teammate rests without its final message captured, `SendMessage`-ing it forces a full-context **resume** (~10–15M tokens reloaded just to re-emit a report) — observed twice on the L0 coders in this run, then ignored. Step C now forbids it: the completion-report is already the teammate's last transcript message; read it from disk for free.
3. **Per-wave fix-agent scope discipline.** The wave-3 "apply ALL verified findings" coder ran away to 215 turns / ~32M tokens (the single most expensive agent in the run) by re-surveying the architecture. `new-card-review.js` `applyFixPass` brief now mandates going DIRECTLY to each finding's cited `evidence` file:line, finding-by-finding, no re-mapping the codebase (a verified finding is a precise instruction, not a research prompt).
4. **grep binary-heuristic false-empty added to the grep-verification discipline.** The orchestrator burned ~6 turns when plain `grep` treated an 8000-line accented `data.ts` as *binary* and returned 0 matches for present tokens. The discipline (completeness.md) gains a third failure mode: prefer the already-mandated `rg`; if you must use `grep`, `grep -a`; a 0-match from plain `grep` on a known-modified file is almost always this artifact.
5. **Codex companion path resolved once per run, passed via workflow args.** Both review workflows spawned a Haiku resolver agent per invocation (4× in this run). They now accept `args.codexScriptPath` (resolved once by the orchestrator, threaded at all three delegation sites) and **fall back to the in-workflow resolver when absent** — fully backward-compatible.

**PATCH** (refinements to `/new` orchestration + one optional, fallback-guarded workflow arg; no new agent/skill/command, no `baldart.config.yml` key, no install/layout change).

### Fixed / Changed

- **`framework/.claude/skills/new/references/team-mode.md`** — empty-result gate: concrete on-disk transcript locator + rate-limit signatures + mandate to state the classification (Int 1 of the pass); Step C forbids `SendMessage`-to-fetch-report and points at the same locator (Int 2); Step-D delegation args gain `codexScriptPath` with a resolve-once glob (Int 5).
- **`framework/.claude/skills/new/references/completeness.md`** — Grep-verification discipline: "Two failure modes" → "Three", adds the binary-heuristic false-empty bullet (`rg`/`grep -a`).
- **`framework/.claude/workflows/new-card-review.js`** — `fixBrief` gains a SCOPE DISCIPLINE clause (go to the cited file:line, no re-survey); accepts `args.codexScriptPath` (caller-provided) with Haiku-resolver fallback.
- **`framework/.claude/workflows/new-final-review.js`** — accepts `args.codexScriptPath` with Haiku-resolver fallback.
- **`framework/.claude/skills/new/references/review-cycle.md`** + **`final-review.md`** — the two remaining `Workflow({...})` delegation sites pass `codexScriptPath` (resolve-once, omit-if-empty).

## [4.53.7] - 2026-06-19

**`setup-worktree.sh` no longer turns an empty `toolchain.commands.build` into a spurious `baseline: fail`.** Observed on a real `/new FEAT-0037` run in a consumer (`mayo`) whose `baldart.config.yml` declares `build: ''` (single-quoted empty scalar — the project's build is not a separate toolchain gate; per `toolchain-protocol.md` an empty command must fall back to the default `npm run build`). The deterministic SSOT worktree script (v4.53.0) read that value with a sed that stripped only DOUBLE quotes (`\"?([^\"#]*)\"?`), so `build: ''` came through as the **literal 2-char string `''`** — non-empty — so the `[ -n "$TC_BUILD" ]` fallback never fired and the script ran `bash -c "''"`, an empty *quoted* command → `bash: : command not found` (exit 127) → `baseline: fail`. tsc + biome had actually passed and the worktree was green on disk; the failure was entirely the unresolved empty scalar. Fix: (1) `_tc()` now resolves a YAML scalar to its true value — `build:`, `build: ''`, `build: ""`, `build: ~`, `build: null` all collapse to empty so the default fallback fires (strips one layer of matching single OR double quotes + YAML null markers + inline `# comment`); (2) the script now tracks build provenance and implements the protocol's **SKIP tier** — the *fallback* `npm run build` on a project with no `build` npm script is a no-build library → SKIP (baseline stays pass), never a spurious fail, while a *configured* (non-empty) build command is always run as a real gate. Verified by executing the resolver against every empty-scalar form + real values, the fallback/SKIP decision on a has-build vs no-build `package.json`, and reproducing the exact `bash: : command not found` exit-127 from the old sed. The prose inline fallbacks (`setup.md`, `worktree-manager/SKILL.md`, `new2.js`) are model-driven and never had the sed bug — no parallel fix needed. **PATCH** (deterministic-script bugfix; correctly-configured consumers see no behaviour change; no new agent/skill/command/config key, no install change).

### Fixed

- **`framework/.claude/skills/worktree-manager/scripts/setup-worktree.sh`** — `_tc()` rewritten to dequote single/double quotes + treat `~`/`null` as empty (was stripping double quotes only, so `build: ''` resolved to the literal `''` and skipped the fallback → `bash -c "''"` → `bash: : command not found` → spurious `baseline: fail`). Build resolution now records `BUILD_FROM_CONFIG` and the baseline build step adds the toolchain-protocol SKIP tier: the fallback `npm run build` on a no-`build`-script project is SKIPPED (baseline pass), not failed; a configured command is always a real gate.

## [4.53.6] - 2026-06-18

**`/new` now caps concurrency at 3 agents per wave — the API rate-limit guardrail, end to end.** Follow-up to the v4.53.5 root-cause finding (parallel coders killed by org-level rate-limiting): the org rate limit is SHARED across every terminal/session on the account, so a wide fan-out saturates the shared pool and the server throttles, killing agents mid-work and forcing full re-spawns (paying twice). Fix puts a hard ceiling of **3 concurrent agents** at every fan-out point in a `/new` epic: (1) `prd-card-writer` caps each `execution_strategy.groups[]` at ≤3 cards — a wider topological layer is split into sequential sub-groups of ≤3 (always safe: same-layer cards are independent by construction, so serializing them is correctness-neutral; the per-card `parallel_group` keeps its true logical layer, only the execution schedule is capped), so the **plan/cards reflect the rule**; (2) team-mode runtime never spawns more than 3 coders per wave and sub-batches a wider group sequentially (protects legacy epics + manual runs — the **guarantee**); (3) the per-wave and final review workflows cap their finder/verify fan-out at 3 via a rolling worker pool (not chunked barriers — those would block fast finders behind the slow background Codex poll). The cap is **3** (the user's choice — prefer never clogging over re-paying), a fixed operational guardrail like the existing retry caps (NOT a `baldart.config.yml` key). Honest limits: BALDART cannot see agents in OTHER terminals — the cap bounds one run; total cross-terminal parallelism is still the user's to manage. `new2` inherits the prd-card-writer group cap (same cards) and the final-review cap (shared workflow); its own per-card scheduling is not separately capped here (follow-up if it sees use). **PATCH** (operational guardrail on existing orchestration; no new agent/skill/command/config key, no install change).

### Fixed

- **`framework/.claude/agents/prd-card-writer.md`** — Parallel Group Computation step 7: cap each execution group at ≤3 cards (split wider independent layers into sequential sub-groups); example + `max_concurrent_agents: 3` added; per-card `parallel_group` unchanged (logical layer).
- **`framework/.claude/skills/new/references/team-mode.md`** — Step B gains a MANDATORY "Concurrency cap — MAX 3 coders concurrently" (sub-batch a wider group sequentially; cross-terminal caveat); Per-Group Execution header notes the cap.
- **`framework/.claude/workflows/new-card-review.js`** + **`new-final-review.js`** — new `MAX_PARALLEL=3` + `parallelCapped()` rolling worker pool; the Discovery/Review finder fan-out and the Verify fan-out run through it (≤3 concurrent agents, order-preserving, null-on-throw, no head-of-line blocking behind the Codex poll).

## [4.53.5] - 2026-06-18

**The team-mode empty-result gate now distinguishes a rate-limit death from a genuine fabrication — and re-spawns transient failures STAGGERED, not in the same parallel burst.** Root-cause finding (from reading the actual subagent transcripts of the v4.53.4 incident — `~/.claude/projects/<mayo>/<session>/subagents/agent-*.jsonl`): the two `/new FEAT-0035` L2 coders that "came to rest" with no work were **killed mid-flight by API rate-limiting** — both transcripts end on the identical `API Error: Server is temporarily limiting requests (not your usage limit) · Rate limited`, after the agent had done real work (file reads, a written plan) but before any edit. It was NOT model fabrication or a "decided nothing to do" stall (the v4.53.4 framing): a wide parallel wave (3 coders + orchestrator + other agents) saturated the API, the background teammate runtime does not auto-resume a rate-limited teammate, so it rested with no report + no diff, and the full re-spawn re-did the lost reads (the "paid twice" cost). Fix: the gate now reads the rested agent's LAST event — a transient API/rate-limit/overload error ⇒ a **transient infra failure** (re-spawn STAGGERED with backoff; never re-fire several transient-failed agents in the same parallel burst — it re-saturates the API; narrow the wave's fan-out width if rate limits recur; does not consume the genuine-failure Step-B budget), vs a CLEAN rest with no error ⇒ the genuine empty-result/fabrication path (one Step-B re-spawn → `AskUserQuestion`). Honest boundary: the cheaper real fix — *resuming* a rate-limited teammate instead of killing+re-spawning it — is Claude Code **runtime** behaviour, not BALDART's to fix in prose; BALDART mitigates (correct classification + staggered backoff + fan-out narrowing). **PATCH** (refines the v4.53.4 team-mode gate; no code/CLI/workflow change, no `baldart.config.yml` key).

### Fixed

- **`framework/.claude/skills/new/references/team-mode.md`** — the Step-C empty-result gate gains a CAUSE classification: transient infra failure (rate-limit/overload — re-spawn staggered with backoff, narrow the fan-out if recurring, no budget consumed) vs genuine empty-result/fabrication (clean rest, no error — the existing one-shot Step-B → `AskUserQuestion`). Corrects the v4.53.4 framing that attributed the silent rest to model fabrication.

## [4.53.4] - 2026-06-18

**Team-mode now has a MANDATORY empty-result / fabrication gate — a parallel coder that rests without doing the work no longer slips past on luck.** Observed on a real `/new FEAT-0035` L2 wave: of three coders spawned in parallel (03, 06, 07), **03 and 07 "came to rest" with NO completion report and zero file changes** (the empty-result failure pattern — same model-in-the-loop class as the v4.53.0 worktree-setup fabrication/stall), while 06 worked and reported normally. The orchestrator caught it only because it *improvised* a `git status` and noticed only 06's files had changed — the team-mode protocol (`team-mode.md` § Step C) merely said "read the completion report from the agent's output", assuming a report exists, and its failure branch only handled an agent that *reports* `status: failed`. So detection of a silent idle was model-discretion, not protocol: a less diligent run would have carried the empty work into the D pipeline and only hit it late/expensively at the D.3a AC-Closure or the Final build. Fix: Step C gains an explicit gate — for EACH agent, before logging it `done`, verify on disk that (a) it emitted the Step-7 completion report and (b) its `files_changed` actually changed in the worktree (`git status --porcelain` ∩ the card's File Ownership Map, using the v4.53.2 grep-verification discipline). No report on rest = empty-result; claimed-but-absent files = fabrication; both → the existing one-shot Step-B re-spawn (re-briefed to write + disk-confirm before reporting), second failure → `AskUserQuestion` skip/abandon. A legitimate zero-diff no-op (AC pre-satisfied) is the rare explicit exception, confirmed from the report's `evidence`. The coder brief (Step 7) is reinforced symmetrically: a `done` report is invalid with an empty diff. Cannot make coders not-idle (they do real work; the internal cause isn't recoverable from the trace — the agents left no report), but the DETECTION is now deterministic + mandatory instead of lucky. **PATCH** (protocol-hardening prose for team mode; no code/behaviour change in CLI or workflows, no `baldart.config.yml` key).

### Fixed

- **`framework/.claude/skills/new/references/team-mode.md`** — Step C gains a MANDATORY "Empty-result / fabrication gate": a teammate resting without a completion report, or with `files_changed` not present on disk, is an empty-result/fabrication failure → one Step-B re-spawn (disk-verify re-brief) → `AskUserQuestion` on a second failure; never accept "came to rest" as "done" without on-disk evidence. Step-7 completion-report mandate reinforced (a `done` report is invalid with an empty diff).

## [4.53.3] - 2026-06-18

**The per-wave/Final Codex review wrapper is now a thin relay on a cheaper model, not an Opus agent that re-investigates Codex's findings.** Observed on a real `/new FEAT-0035` run: the `codex` discovery agent burned ~127k tokens / 26 tool calls — it launched the Codex companion (correct) but then **re-grepped/sed'd the source to "mechanically confirm" Codex's findings** before returning, despite the design already trusting them (`preValidated`, FP-checked by Codex). Root cause: the wrapper consumed Codex's **freeform `task` output** (prose), which forced a smart+expensive model to parse 6 fields per finding out of prose — and it over-reached into self-investigation. Fix: instruct Codex to emit its findings as strict JSON between explicit `<<<FINDINGS_JSON>>>` / `<<<END_FINDINGS_JSON>>>` sentinels (verified on two real companion runs: output round-trips cleanly through `JSON.parse` with all our fields incl. `domain` + `requires_action`, and the sentinels make extraction deterministic despite the companion's `[codex]` trace lines + the truncated "Assistant message captured:" echo — the truncated copy lacks the closing sentinel so the last-complete-pair awk skips it). The wrapper is now a **pure relay**: write the task to a temp file (no shell-quote breakage — same discipline as v4.53.2), launch Codex in the background, poll for the END sentinel, extract via a fixed awk, `JSON.parse`, return — **no re-grep, no re-verify**. Model dropped from the inherited Opus to **`haiku` per-wave** (a poll misfire degrades safely to the `code-reviewer` fallback, and the Final re-runs Codex batch-wide) and **`sonnet` at the Final** (last cross-model gate before merge — no downstream Codex backstop). The review *mechanism* (Codex `task` + `--cwd` worktree + the changed-file list) is unchanged from production — only the output format, the wrapper's scope, and the model tier change. **PATCH** (cost/economy refinement of existing workflows; no new capability, no schema/config key, no install change).

### Fixed

- **`framework/.claude/workflows/new-card-review.js`** — `codexPrompt` rewritten as a thin relay (write task → background launch → poll END sentinel → deterministic awk extract → `JSON.parse` → return; explicitly forbids self-investigation); Codex instructed to emit sentinel-delimited strict JSON findings; the `codex` discovery thunk runs on `model: 'haiku'`.
- **`framework/.claude/workflows/new-final-review.js`** — same relay rewrite + sentinel-JSON contract; the Final `codex` thunk runs on `model: 'sonnet'` (last cross-model gate, no downstream backstop).

## [4.53.2] - 2026-06-18

**`/new` completeness/AC verification no longer treats a quote-broken grep as proof an AC is unimplemented.** Diagnosed from a real `/new FEAT-0035` team-mode L1 run: the orchestrator spot-checked AC-3 by grepping `data/products.ts` for the `mode:'catalog'|'picker'` token, the pattern's embedded single-quotes + pipe broke the shell/regex parse inside the `echo "=== … ===" && grep` block → **0 matches even though `mode` was fully implemented** (`ListProductsMode` type + `mode??'catalog'` + picker branch all present) → a false "AC-3 not implemented" alarm. It self-recovered (the orchestrator noticed the grep contradicted the coder's completion report and re-grepped correctly), but a false "AC absent" can otherwise trigger a needless gap-fix coder re-spawn — the same waste class as v4.53.1. Two compounding causes, both addressed: (1) grep-as-AC-evidence over **code-shaped tokens** is brittle (the verification phase is prose-mandated to "grep changed files for the implementation"), and (2) `0 matches` was conflated with "absent". Added a **Grep-verification discipline** to the completeness SSOT: search literals with `rg -F` (fixed-string, one fully-single-quoted pattern per call — never folded into an `echo "===…" && grep` composition), and a `0-match` that CONTRADICTS the coder's structured completion report is a likely broken grep → re-confirm by Reading the cited `file:line` BEFORE flagging or re-spawning (never re-spawn a coder on a 0-match alone). **PATCH** (verification-guidance hygiene to existing prose; no code/behaviour change in the workflows, no `baldart.config.yml` key).

### Fixed

- **`framework/.claude/skills/new/references/completeness.md`** — Phase 2.5 gains a "Grep-verification discipline" MUST note governing every grep check in the phase (step 0 binary-outcome branches, step 4 symbol verification, the `data_fields` checks): `rg -F` for code-shaped literals, one pattern per call; `0 matches` ≠ absent; a 0-match contradicting the coder report → re-confirm at the cited `file:line` before classifying Missing/Partial or spawning a gap-fix agent.
- **`framework/.claude/skills/new/references/team-mode.md`** — the post-completion (D.1+) on-disk spot-check step now cites that discipline, so a quote-broken AC/symbol grep in team mode cannot trigger a needless fix-coder.

## [4.53.1] - 2026-06-18

**Review workflows no longer spawn fixers for verified-but-no-action findings, and a `finding_id` collision that turned applied fixes into false residuals is closed.** Diagnosed from a real `/new FEAT-0035` team-mode run where the per-wave review cluster spawned `security-reviewer` (Sonnet) + `coder` (Opus) for ~110k tokens to apply ~3 LOW fixes and **no-op ~8 non-actionable findings** — finders (simplify / security / codex / api-perf / doc) emit *verified observations that need NO change* (`minimal_fix_direction` = "No fix required" / "acceptable as-is" / "verified non-issue") as findings; these survive the false-positive check (they are TRUE, not false positives), are `preValidated` → skip Verify → classified `VERIFIED` → flood the fix partitions. The system conflated **VERIFIED with actionable**. Fix (two fail-safe layers): the `FINDING` schema gains an optional `requires_action` flag, every finder prompt is told to set it `false` (or simply not emit) a no-change observation, and a deterministic `isActionable()` gate — honoured for MED/LOW only, a **BLOCKER/HIGH is always actionable** so a cross-wave merge-ordering HIGH still surfaces as residual — segregates no-action findings into a recorded/counted bucket (`summary.noAction`, `decision=skipped reason=no-action`) that **never reaches a writer and is never silently dropped**. A regex backstop on the fix direction covers a finder that omits the flag. **Second defect (root cause of an observed empty-coder re-spawn):** finders number `finding_id` (`F###`) independently, so a security `F003` and a migration `F003` collide; the Fix-phase bookkeeping (`appliedIds`/`unresolvedIds`/`codeResidual`/`bucket`) keys on `finding_id` with flat Sets, so one finder's `unresolved` id dragged another finder's *applied* same-id fix into residual → the skill then re-spawned a coder that found nothing to do. Fixed by making `finding_id` globally unique at the fan-in (prefix `source` + index). Verified with a deterministic test that extracts the real gate from source and runs it over the actual FEAT-0035 findings (4 no-action LOW suppressed incl. a mid-prose "no fix required", 3 real LOW + the cross-wave HIGH kept) plus a collision regression (bug reproduces pre-fix, gone post-fix). Applied to **both** review workflows + their prose-SSOT (so the inline fallback benefits) + (opt-in) the finder agents' system prompts. **PATCH** (economy/quality refinement of existing workflows; no new agent/skill/command, no install-layout change, no `baldart.config.yml` key).

### Fixed

- **`framework/.claude/workflows/new-card-review.js`** — `FINDING` schema gains `requires_action`; finder prompts (codex / simplify / security / code-reviewer fallback) instructed to flag no-action observations; new `NO_ACTION_RE` + `isNoAction()` gate (MED/LOW only) segregates a `noAction` partition excluded from `securityFix`/`actionable`/`docResidual` and logged; `finding_id` made globally unique at the fan-in (fixes the cross-finder `appliedIds`/`unresolvedIds`/`codeResidual` contamination); coder fix-pass prompt clarifies `applied` = edit made even if the build stays red for out-of-scope reasons, `unresolved` = the edit itself could not be made; `summary.noAction` added.
- **`framework/.claude/workflows/new-final-review.js`** — same `requires_action` schema field + finder-prompt instruction (codex / doc / api-perf / code-reviewer fallback) + `finding_id` uniqueness at the fan-in; no-action VERIFIED MED/LOW findings segregated out of the returned `findings` into a new `noActionFindings` field + `summary.noAction`, so the skill's F.5 writer partition never routes them.
- **`framework/.claude/skills/new/references/review-cycle.md`** / **`final-review.md`** — drift-note + Phase-2.55 step-4 / F.4 classification + F.5 partition updated to document the actionability gate and the `finding_id` uniqueness (the inline fallback path mirrors the workflow). `codex-gate.md` unchanged (its per-card gate only auto-fixes BLOCKER/HIGH, which are never no-action).
- **`framework/.claude/agents/code-reviewer.md`**, **`api-perf-cost-auditor.md`**, **`security-reviewer.md`**, **`doc-reviewer.md`** — an Actionability clause after the FP challenge: a verified-safe / already-accurate observation is a cleared concern, not an actionable finding (set `requires_action:false` or omit; MED/LOW only) — defense-in-depth that also benefits standalone `/codexreview`, `/prd`, `/bug`.
- **`framework/.claude/workflows/new2.js`** / **`new2-resolve.js`** — unchanged (justified): they route only BLOCKER/HIGH findings, which the severity invariant always keeps actionable, so no-action MED/LOW never reach them; they group by file/area, not `finding_id`.

## [4.53.0] - 2026-06-18

**Worktree setup for `/new` / `/nw` / `new2` is now a deterministic SSOT script — no model in the loop for pure plumbing.** Diagnosed from a real `/new FEAT-0036 -full` run where the Phase-0 worktree-setup subagent **created the worktree but went idle before `npm install`** (node_modules missing, no registry entry), forcing the orchestrator to burn turns probing → messaging the agent → waiting → cleaning the half-built orphan → falling back to inline `/nw` (which reloads the ~1200-line skill body into the orchestrator — the exact cost the delegation was meant to avoid). Same root cause as the earlier haiku *fabrication* (a well-formed `baseline: pass` with nothing on disk, 2/2): a **model in the loop for mechanical work** — it either fabricates the expected return or stalls mid-execution. The fix removes the model: the entire build sequence (worktree add → sync untracked cards → install → copy `stack.env_files` → allocate port → write registry → baseline) now lives **once** in `framework/.claude/skills/worktree-manager/scripts/setup-worktree.sh` and is invoked identically by all three callers as a background `Bash`. A deterministic script cannot fabricate or stall: it does the work or fails honestly with a log + a structured manifest. The orchestrator's disk-verification gate (setup.md step 6a) and git-authoritative idempotency pre-check (step 4a2) are unchanged — defence-in-depth survives the redesign. Non-interactive HARD (stdin closed, `CI=1`, build under `timeout 600`); port allocation is atomic under a `.worktrees/.wt-setup.lock` mutex (mirrors `allocate-id.sh`) so concurrent `/nw` runs never collide; the registry is written via `node` (atomic tmp+rename, no jq dependency); the script never self-repairs a failing baseline (role boundary — the coder does, downstream). Verified end-to-end on a throwaway repo: happy path (worktree + env+PORT + card-sync + `buildVerified:true` + manifest), collision (loud fail, exit 3), build-fail (`baseline:fail`, port held, `buildVerified:false`), and second-worktree port allocation (3001+3002 reserved → 3003). Opt-in-with-fallback: a pre-this-release subtree without the script falls back to inline `/nw`. **MINOR** (new shipped SSOT asset + behaviour change to existing skills; backward-compatible, no install-layout or command change, no `baldart.config.yml` key — `stack.env_files` / `toolchain.commands.*` are reused).

### Added

- **`framework/.claude/skills/worktree-manager/scripts/setup-worktree.sh`** — the single source of truth for code-worktree creation + baseline, consumed by `/nw`, `/new` (setup.md step 4), and `new2`. Resolves config (`$MAIN`/`$TRUNK`/`stack.env_files`/`toolchain.commands.*`/port range) from args with `baldart.config.yml` fallback; writes a stable manifest block (`status`/`error`/`worktree_path`/`branch`/`port`/`created_at`/`baseline`/`baseline_log`) on every exit path.

### Changed

- **`framework/.claude/skills/new/references/setup.md`** — Phase-0 step 4 replaces the Sonnet-subagent-via-Skill delegation (+ its `/nw`-inline fallback chain) with a background launch of `setup-worktree.sh`; step 5 barrier + step 6a disk gate now read the manifest; fallback collapses to `script → inline /nw → HALT` (honest/deterministic failures are not fallback-eligible). `SKILL.md` worktree `Created:` note updated.
- **`framework/.claude/skills/worktree-manager/SKILL.md`** — `/nw` steps 3-6 (create/isolate/baseline/registry) collapsed to an invocation of `setup-worktree.sh` + a prose description of what it does; steps 7-8 read `worktree_path`/`port` from the manifest instead of in-process shell vars.
- **`framework/.claude/workflows/new2.js`** — the Pre-flight agent's worktree bullet now runs `setup-worktree.sh` instead of hand-rolling the sequence; it still independently verifies the worktree on disk for the (structurally weaker) E2.5 evidence gate, and falls back to the inline sequence if the script is absent.

## [4.52.4] - 2026-06-18

**SSOT hygiene: `/cont` (context-primer) no longer restates the code-search tier hierarchy inline.** Follow-up to v4.52.3. The prompt template `context-primer` hands to `codebase-architect` re-stated the "prefer LSP/graph over grep for symbols" guidance inline (step 3a) instead of citing the protocol modules — so after v4.52.3 added the anti-flail rule to `code-search-protocol.md` / `code-graph-protocol.md` / the `codebase-architect` system prompt, that inline copy was the one spot that didn't carry it. Functionally harmless (the spawned `codebase-architect` already governs via its own system prompt + the protocols), but a latent drift surface. Trimmed step 3a to a citation of both protocol modules ("follow their tier order, budgets, and anti-flail rule — don't restate them here"), keeping context-primer's own task shaping (canonical-router-first, agent memory, backlog, git, verification) intact. Verified no other skill restates the hierarchy (`lsp-bootstrap`'s line is a user-facing confirmation, not a protocol copy). **PATCH** (doc/guidance hygiene; no behaviour change — the search behaviour was already governed by the agent + protocols; no `baldart.config.yml` key).

### Changed

- **`framework/.claude/skills/context-primer/SKILL.md`** — step 3a of the SEARCH STRATEGY block now cites `code-search-protocol.md` + `code-graph-protocol.md` for tier order / budgets / anti-flail instead of restating the LSP/graph-over-grep hierarchy inline.

## [4.52.3] - 2026-06-18

**`codebase-architect`: structural tier (graph/LSP) is the PRIMARY source for symbol queries, with a hard anti-flail grep cap.** Diagnosed from a real `/prd` discovery run: a foreground `codebase-architect` verification burned its whole turn budget (56 tool uses, 88.1k tokens) grepping ~10× for one function (`createDraft`) inside an 8000-line re-export barrel, **terminated mid-investigation before producing its `totals:` report** (the truncation `/prd` then mis-recovered by resuming the completed agent in the background, breaking the foreground/blocking discovery contract). Root cause: the agent reached for the graph (`graphify explain`, which located the symbol instantly) only *after* the grep flail — because the protocols framed graph/LSP as "escalation" (i.e. *after* grep) and excluded the graph from single-symbol queries entirely, and nothing capped per-symbol grep. Fixed by three surgical guidance edits: (1) `codebase-architect` step 3 now routes symbol location/relation queries to the structural tier **first** when its flag is on (grep is the fallback), with an **anti-flail heuristic** — don't grind grep to locate one symbol; after ~2 misses (or on a barrel / very large file) switch tools: to `graphify explain`/`query` or LSP `go-to-definition` when a flag is on, or — **when neither flag is on (the default config)** — read the file's exports/headings directly instead of full-text-grepping it again; (2) `code-search-protocol.md` adds the same flag-aware per-symbol anti-flail rule to **Budget** + a fallback rule (LSP off & grep can't find a definition → graph is a valid locator); (3) `code-graph-protocol.md` narrows the "graph adds nothing for single-symbol" exclusion — the graph IS a valid fallback for *locating* a definition when grep/LSP miss on a barrel. The downstream `/prd` post-condition gap (how to react to an incomplete foreground grounding return) is intentionally **not** fixed here — this release attacks the root cause (budget burn), not the symptom. The "prefer LSP/graph over grep" guidance already existed and was ignored on the incident run; the reframe + graph-as-locator fix are net-positive regardless, but the behavioural claim is unverified — the real test is the next run. **PATCH** (behaviour-correcting guidance to existing agent + protocol modules; no new agent/skill/capability, no `baldart.config.yml` key).

### Fixed

- **`framework/.claude/agents/codebase-architect.md`** — step 3 reframed: structural tier first for symbol location/relation queries when `has_lsp_layer` / `has_code_graph` is on; grep is the fallback; new flag-aware anti-flail heuristic (after ~2 grep misses, escalate to `graphify explain`/`query` or LSP if a flag is on, else read exports/headings directly — never grind grep on a barrel).
- **`framework/agents/code-search-protocol.md`** — **Budget** gains a flag-aware per-symbol anti-flail rule (with a no-flag fallback to reading exports/headings); **Fallback rules** gain "LSP off & grep can't locate a definition → code graph is a valid locator".
- **`framework/agents/code-graph-protocol.md`** — the single-symbol exclusion is narrowed to *reference enumeration*; `graphify explain`/`query` is now documented as a valid fallback for *locating* a symbol's definition on barrel/large files, with a matching query-type table row.

## [4.52.2] - 2026-06-18

**Anti-hardcoded gate: kill the false positives — the generated ESLint config now allowlists only user-facing JSX attributes.** v4.52.0's `mode: 'jsx-only'` flagged *every* JSX attribute string, so on a real component-heavy codebase it drowned the signal in technical props — `variant="inset"`, `size="lg"`, `href="/login"`, `padding="md"` were all reported as "hardcoded user-facing strings". Measured on a production app: **3260 raw findings, mostly false positives.** Fixed by adding `'jsx-attributes': { include: ['placeholder','title','alt','label','aria-label'] }` to the generated `eslint.i18n.config.mjs` / `.eslintrc.i18n.json`, so attribute checks are restricted to the genuinely user-facing ones while JSX text nodes are still fully checked. Same codebase after the fix: **~1500 findings, 0 technical-attribute false positives** — every remaining finding is a real hardcoded label. Verified against `eslint-plugin-i18next`'s `shouldSkip` include-allowlist semantics. **PATCH** (tightens the generated gate config to be usable on real codebases; no new key, no behaviour change beyond fewer false positives).

### Fixed

- **`src/utils/i18n-gate.js`** — both the flat (ESLint 9) and legacy (ESLint 8) generated configs now carry the `jsx-attributes.include` user-facing allowlist; `framework/agents/i18n-protocol.md` cascade updated to describe the precise coverage (JSX text + 5 user-facing attributes; technical props excluded).

## [4.52.1] - 2026-06-18

**Hotfix: `baldart configure` (and the `configure-drift` doctor action) crashed with `detectedI18nFramework is not defined` whenever the i18n block ran.** The v4.52.0 i18n config block lives in `interactivePrompts()`, but referenced two helpers (`detectedI18nFramework`, `detectLocalesRoot`) that are scoped to the separate `detect()` function — a `ReferenceError` the moment `features.has_i18n` was true (so the framework / locales-root prompt defaults threw, taking down configure and the doctor drift action). Fixed by passing the detected values through the `detected` object (`detected.i18n.{framework,locales_root}` — real snake_case keys that also seed `merged.i18n.*` via the existing merge, no junk serialized). Also **broadened i18n framework autodetection** to `next-i18next` / `react-i18next` (→ i18next) and `@nuxtjs/i18n` / `@nuxt/i18n` (→ vue-i18n), so the `has_i18n` flag defaults to Yes on those stacks instead of being silently skipped. **PATCH** (bugfix to a v4.52.0 CLI path; no behaviour change beyond fixing the crash + wider detection; no new `baldart.config.yml` key).

### Fixed

- **`src/commands/configure.js`** — i18n config block no longer references `detect()`-scoped helpers from `interactivePrompts()`; detected framework + locales root flow through `detected.i18n`. Autodetection covers `next-i18next`, `react-i18next`, `@nuxtjs/i18n`, `@nuxt/i18n`.

## [4.52.0] - 2026-06-18

**BALDART becomes opinionated about internationalization: a new opt-in `features.has_i18n` layer makes consumers develop multi-language by default — zero hardcoded labels + a stack-agnostic context registry that makes LLM translation context-aware instead of blind.** The differentiator vs ordinary i18n is the **per-key context**: a stack-agnostic YAML registry (`paths.i18n_registry`, default `docs/i18n/registry.yml`) holds `source` + a *hyper-brief usage context* + `domain` for every label, so the translator (the new `i18n-translator` agent, or a delegated bulk backend) sees *what the label is for and where it appears* — the 2026 best-practice fix for "the void of context", the #1 translation-quality threat. **Translations live in the stack's NATIVE locale files** (i18next / next-intl / react-intl / lingui / vue-i18n); BALDART wraps the consumer's framework + best-in-class OSS rather than reimplementing i18n. The anti-hardcoded gate is a **deterministic, BALDART-owned standalone ESLint config** (`eslint.i18n.config.mjs` running `eslint-plugin-i18next` `no-literal-string` — the most-adopted at ~550k weekly downloads; framework-agnostic so it covers next-intl/vue too) written + installed by `configure`/`doctor` (`src/utils/i18n-gate.js`). It is a **dedicated ESLint run independent of the project's main lint** (so it works even on a **Biome-only toolchain**, which does not execute ESLint plugins) and **never touches the user's config**. The gate runs via `i18n.lint_command` in `qa-sentinel` (`/new`, `/qa`), the `new2` Phase-2 gate, and the `i18n-align` routine; `coder` STEP 9 and `code-reviewer` enforce the same rule at the agent level (real layers, not hollow — they do not assume the linter is wired). Key extraction reuses the stack's extractor (`i18next-cli`/`lingui`/`formatjs`, detection-only). Drift is contained in **3 tiers** mirroring the design-system discipline: per-task (`coder` STEP 9 populates the registry stub), per-merge (`doc-reviewer` curates context/orphans in `/new`), weekly (`i18n-align` routine translates the maintained `i18n.target_languages`, lints, and commits **directly to the trunk — no PR**, since translation files are low-risk). The always-on HARD rule ("never hardcode a user-facing string") lives in `framework/AGENTS.md` § 4.b gated on the flag. Strict specialization throughout: `coder` externalizes + populates, `doc-reviewer` curates, `i18n-translator` (`model: sonnet`, `effort: low`, bounded flag-not-guess) translates — none crosses lanes. **MINOR** (additive, opt-in: new agent + two new skills (`/i18n`, `/i18n-adopt`) + new routine + new flag/`i18n:` block + `paths.i18n_registry`; no removed surface). Schema-change propagation rule applies (`features.has_i18n` + `paths.i18n_registry` + `i18n.*` through template + configure + update detector + doctor + CHANGELOG).

### Added

- **`framework/agents/i18n-protocol.md`** — the textual SSOT: authority model (code/native files own keys+translations; registry owns context/domain), registry schema, naming convention + ICU/no-concatenation, the BLOCKING anti-hardcoded cascade, drift codes (`I18N_REGISTRY_DRIFT` / `I18N_KEY_ORPHANED` / `I18N_CONTEXT_MISSING`), and strict ownership. Registered in `framework/agents/index.md`.
- **`framework/.claude/agents/i18n-translator.md`** — new context-aware translator agent (`model: sonnet`, `effort: low`): translates into native locale files using registry context, preserves ICU/placeholders, and flags anomalies (`needs-attention`) instead of guessing. Added to `REGISTRY.md`.
- **`framework/.claude/skills/i18n/SKILL.md`** — new `/i18n` skill (portable Claude + Codex): `audit` (registry ↔ `t()` ↔ native files drift report) and `translate` (spawns `i18n-translator` with review gate). Gated on `features.has_i18n`.
- **`framework/.claude/skills/i18n-adopt/SKILL.md`** — new `/i18n-adopt` skill (`effort: medium` — the orchestration is light; the spawned `coder` keeps its Opus model since it mutates production code, `i18n-translator`/`doc-reviewer` are Sonnet): one-shot migration that adopts the i18n layer on an EXISTING codebase — the gate linter finds every hardcoded user-facing string, `coder` replaces each with `t()` + registers key/context (STEP 9), `doc-reviewer` curates, `i18n-translator` translates the target languages, the gate + build verify. Full-auto on a dedicated branch `chore/i18n-adopt-*` (no mid-run pauses), idempotent + resumable via `.baldart/i18n-adopt/progress.json`, never auto-merges. The i18n analogue of `/design-system-init`.
- **`framework/routines/i18n-align.routine.yml`** — new 8th routine (weekly): translates the maintained target languages, lints, commits direct to trunk. Registered in `framework/routines/index.yml`. Optional (skips when feature off / no target languages).
- **`framework/docs/I18N-LAYER.md`** — operator guide (why / moving parts / lifecycle / invariants / fallback).
- **`baldart.config.template.yml`** — `features.has_i18n`, `paths.i18n_registry`, and the `i18n:` block (`framework`, `source_language`, `target_languages`, `locales_root`, `extract_command`, `bulk_translator`, `lint_command`).
- **`src/utils/i18n-gate.js`** — the standalone anti-hardcoded gate util: writes `eslint.i18n.config.mjs` (only when absent), installs `eslint` + `eslint-plugin-i18next` + `typescript-eslint`, and resolves the gate command (explicit `i18n.lint_command` → standalone config by convention → SKIP).

### Changed

- **`src/commands/configure.js`** — autodetects the i18n framework from `package.json`, prompts the `has_i18n` flag + the `i18n:` block (source/target languages, locales root), populates `extract_command` defaults, sets `i18n.lint_command`, offers the standalone gate setup (`I18nGate.setup`), and points to `npx baldart routines install i18n-align`.
- **`src/commands/update.js`** — update detector diffs the `i18n:` block (`missingI18n`) alongside `graph:` / `toolchain:`.
- **`src/commands/doctor.js`** — detection-only i18n backfill: `i18n-gate-setup` (installs the gate deps + writes `eslint.i18n.config.mjs`) and `i18n-registry-missing` (registry absent); never installs the consumer's i18n framework.
- **`framework/.claude/agents/qa-sentinel.md`** + **`framework/.claude/workflows/new2.js`** — run the i18n anti-hardcoded gate (resolve `i18n.lint_command` → standalone config → SKIP) in addition to the normal lint gate when `features.has_i18n: true`.
- **`framework/.claude/agents/coder.md`** — STEP 9 (i18n registry population, gated on `features.has_i18n`): externalize strings + write the registry stub; enforcement of "no hardcode" is delegated to the linter (no prompt duplication).
- **`framework/.claude/agents/code-reviewer.md`** — i18n anti-hardcoded backstop (syntactically scoped HIGH finding for the residue the linter can't see).
- **`framework/.claude/agents/doc-reviewer.md`** — Internationalization Scope: owns registry curation (context completeness, dedup, orphan/missing), does not translate.
- **`framework/AGENTS.md`** — one gated MUST in § 4.b (no hardcoded user-facing strings when `has_i18n: true`); gating table row.
- **`framework/.claude/skills/prd/SKILL.md`** — HARD RULE 20 (cards with UI copy declare label keys + context).
- **`framework/.claude/skills/new/references/implement.md`** — coder mission briefing gains an i18n-constraint section (ambient flag, no orchestrator handoff).
- **`README.md` / `CLAUDE.md`** — counts (29 agents, 34 skills, 8 routines), inventory, and a new feature section / convention bullet.

## [4.51.0] - 2026-06-17

**A second, more aggressive output style — `Terse Ultra` (caveman "ultra") — ships alongside `Terse` as an opt-in maximum-compression profile.** Research into the [caveman](https://github.com/JuliusBrussee/caveman) tiers (lite → full → ultra → wenyan) + its issues/reviews established the frontier: the community consensus is that `full` (≈ our `Terse`) is the sweet spot, `ultra` buys more compression at a small-but-real cost (it "occasionally drops important caveats" and is "a bit much when you actually need to understand what went wrong"), and `wenyan` (classical Chinese, −80/90%) is out of scope for a general dev framework. So `Terse Ultra` is shipped as a SECOND style, not a replacement: fragments, prose-word abbreviations (db/auth/config/req/res/fn/impl — **prose words only**), arrows for causality (`X → Y`), with HARDENED guardrails — code symbols / function & API identifiers / file paths / `path:line` / error strings / CLI commands / commit keywords stay **verbatim even in ultra**, and the Auto-clarity carve-out is **extended to the final answer/explanation** (per caveman issue #510) on top of the security / irreversible-action / multi-step / ambiguity cases. `baldart configure` still offers only `Terse` (the recommended default); `Terse Ultra` is discovered and activated manually via `/config` → Output Style. Distribution is free — the v4.50.0 merge picks up any new `.md` under `framework/.claude/output-styles/` automatically. **MINOR** (additive payload — a new selectable style; **no new `baldart.config.yml` key** — `outputStyle` is a native Claude Code settings key; no removed surface).

### Added

- **`framework/.claude/output-styles/terse-ultra.md`** — the `Terse Ultra` output style (frontmatter `name: Terse Ultra`): caveman *ultra* prose compression with a never-abbreviate-code guardrail and an Auto-clarity carve-out extended to the final answer. Self-documents that it's the aggressive tier with a small readability/nuance cost and that `Terse` is preferred for everyday work.

## [4.50.0] - 2026-06-17

**BALDART ships a `Terse` Claude Code output style as a latent, opt-in deliverable — installed with every Claude-enabled consumer, activated only by the user.** Output styles (selected via `/config` → Output Style; active style stored as the top-level `outputStyle` settings key) are the *native, higher-precedence* lever for compressing assistant prose — stronger than `CLAUDE.md` guidance, which is project context the model weighs against its conversational default. The new `framework/.claude/output-styles/terse.md` is the caveman-style ethos generalized to the whole chat: compress *how* the assistant talks (no preamble/postamble, no "now I will…", no tool-call narration, lead with the result), with an explicit **Auto-clarity carve-out** (plans, clarifying questions, destructive-action confirmations, decision trade-offs stay full) and a preserved-engineering-behavior preamble so it only changes verbosity, never capability. Distribution reuses the v4.14.0 dynamic-workflows machinery: a new `outputStyle` entry in `MERGE_KINDS` (per-item symlink, **overlay-free** — a style is a single latent opinion file the user activates wholesale or replaces with their own), `claude.js` exposes `outputStylesDir()`, `codex.js` returns `null` (**Claude-only** — Codex has no output styles). Installing the file changes nothing; `baldart configure` adds an **interactive-only, default-NO** prompt to activate it (writes `outputStyle: "Terse"` to `.claude/settings.local.json` — personal, gitignored, never forced on the team). **MINOR** (additive payload type + CLI capability; **no new `baldart.config.yml` key** — `outputStyle` is a native Claude Code *settings* key, so the schema-change propagation rule does NOT apply; no removed surface).

### Added

- **`framework/.claude/output-styles/terse.md`** — the `Terse` output style (frontmatter `name: Terse`): prose-only terseness with an Auto-clarity carve-out, preserving all software-engineering behavior.
- **`src/utils/symlinks.js`** — `outputStyle` kind in `MERGE_KINDS` + `mergeOutputStyles()` wrapper, wired into `createAllSymlinks()` and the install-validation reconcile loop. `_mergeBulkDir` generalized with an optional `srcDir` field (defaults to `<kind>s`) so the hyphenated `output-styles/` source dir Claude Code requires resolves correctly (the kind name `outputStyle` would otherwise pluralize to the wrong `outputStyles/`).
- **`src/utils/tool-adapters/{claude,codex}.js`** — `outputStylesDir()` (`.claude/output-styles` for Claude; `null` for Codex).
- **`src/commands/configure.js`** — `maybeActivateTerseOutputStyle()`: interactive-only, default-NO prompt to activate the style by writing `outputStyle: "Terse"` to `.claude/settings.local.json`; no-ops when the file is absent (Codex-only / older framework) or already active. New NEXT STEPS line pointing to `/config` → Output Style.

## [4.49.1] - 2026-06-17

**The four `/new` dynamic workflows shed their paragraph-length `meta.description` — the last prose that re-entered the orchestrator context on every workflow completion.** The `Workflow` tool contract specifies `description` as a *one-line* field shown in the permission dialog, but `new-card-review` / `new-final-review` / `new2` / `new2-resolve` each carried a full-paragraph rationale (~150–230 words) that the runtime echoes back into the orchestrator prefix on the completion line — pure prose with zero runtime value (the orchestrator already knows what it invoked). Each is collapsed to a single line; the full semantics were already duplicated in the script's own args-contract comment block (source code, never injected into context) plus the `references/*.md` SSOT, so nothing is lost. Same family as the v4.49.0 prose-leak closures, applied to the workflow layer. **PATCH** (no behaviour change — descriptions are display/permission-dialog text only; **no new `baldart.config.yml` key**).

### Changed

- **`framework/.claude/workflows/{new-card-review,new-final-review,new2,new2-resolve}.js`** — `meta.description` collapsed from a full paragraph to one line each (e.g. `'Per-wave review+fix cluster for /new (off-context).'`). A two-line comment above each marks why it must stay a one-liner (re-enters orchestrator context on completion) and points to the args-contract comment block + reference module that hold the full semantics.

## [4.49.0] - 2026-06-17

**Caveman-style terse contracts close the last prose-leak the v4.47.0 turn-economy left out of scope: `codebase-architect` gets an opt-in `OUTPUT=terse` grounding mode, and `/new`'s orchestrator is forced near-silent.** A study of the [caveman](https://github.com/JuliusBrussee/caveman) skill (compress *how the agent talks*, not *what it builds* — output tokens only, reasoning untouched) plus a **prose-leak surface map of BALDART's own fleet** found the surface already covered almost everywhere — finders emit YAML/schema (`code-reviewer`, auditors, `plan-auditor`, `prd-card-writer`, `coder` completion-report), `senior-researcher` is file-resident, `context-primer` is already capped at a 30–50-line XML — with ONE genuinely-uncovered prose exploder: `codebase-architect`, whose budget cap is *input*-side only (`codebase-architect.md`), whose "Communication Style" explicitly licenses narrative prose, and whose return is outside the v4.47.0 rule by design (`new/SKILL.md` § "Context economy": *"It does NOT change any agent's return contract"*). The earlier ponytail-style "cap all finder prose" idea was **refuted** by a 3-skeptic adversarial pass (finders already capped + data-driven; new2 already isolates subagent output; driver is turn-count) — this release ships only the two data-validated survivors. A follow-up **return-consumption audit** of the whole fleet confirmed the floor is already reached everywhere else — finders/auditors return a parsed schema or override-to-`## FINDINGS`, and `qa-sentinel`/`doc-reviewer`/`senior-researcher`/`prd-card-writer` are file-resident with a thin manifest return; the only remaining lever was wiring the new `OUTPUT=terse` flag into the remaining machine-to-machine `codebase-architect` call-sites (`context-primer`, `/prd` discovery dimension-resolution + ISA spawns). The near-silent rule is scoped to `/new` only — `/prd` and `/bug` are conversational, where the prose IS the deliverable (caveman's own Auto-Clarity principle: never compress what a human reads to decide). `/new` is the opposite: not conversational — the user keeps the orchestrator attached only so it can surface a genuine decision, so its running narration serves no one. **MINOR** (additive agent capability + skill behaviour; **no new `baldart.config.yml` key** ⇒ schema-change propagation rule N/A; no removed surface).

### Changed

- **`framework/.claude/agents/codebase-architect.md`** — new **"Output mode: terse grounding (opt-in)"** section: when the invoking prompt carries `OUTPUT=terse`, the architect emits the structured substance only (the `## Reuse Analysis` table + `## Canonical Evidence` block unchanged, affected code as `path:line — symbol — pattern/role` rows, high-risk paths one-line each, a closing `totals:`), drops the high-level-overview/​"why" narrative, keeps code identifiers verbatim, and keeps an **Auto-clarity carve-out** (`⚠` one-liner for a genuine ambiguity/risk a row can't carry). The existing "Communication Style" is retitled *(interactive / explanatory invocations only)* — without the flag the narrative path is unchanged.
- **`framework/.claude/skills/{new/references/implement.md` (Phase 1 step 3), `bug/SKILL.md` (Phase 0 step 3), `context-primer/SKILL.md` (loader spawn template), `prd/references/discovery-phase.md` (dimension-resolution spawn + ISA spawn)}** — every machine-to-machine `codebase-architect` grounding spawn now passes `OUTPUT=terse`. The `/new` baseline persisted verbatim at step 5b stays lean while keeping everything a lean `/codexreview` needs (paths, signatures, patterns, high-risk paths).
- **`framework/.claude/skills/new/SKILL.md`** — new **HARD RULE in § "Context economy": "orchestrator is near-silent / maximally cryptic"** — the orchestrator emits **nothing user-facing during the run** (no progress, no "now I will…", no per-phase recap, no agent-activity description, no tracker restatement, no tool-call narration) and speaks in exactly **three** cases: (1) decision gates — `AskUserQuestion` + security/​irreversible-action confirmations, full clarity, never compressed; (2) genuine blockers / items needing user action, terse and exact; (3) one minimal end-of-batch result block (one line per card + actionable residue). Governs the orchestrator's *prose*, not any agent return contract.
- **`framework/.claude/agents/plan-auditor.md`** — the FULL-mode `## OUTPUT FORMAT` 10-section block is compressed in place (~78 → ~22 lines): same 10 sections + every load-bearing semantic (verdict definitions, `Priority ≥ 16 → BLOCK`, the three distinct pre-mortem root-cause classes), with the verbose per-section sub-bullet templates collapsed to one-line specs. This is dormant weight on the common QUICK-mode spawns (which explicitly skip it), trimmed from every spawn's system prompt. Compressed **in place** rather than extracted to an on-demand reference — agents don't Read external format files (that's a skill-only pattern), so extraction would introduce a cross-context path dependency not worth ~1k tokens. (Return-consumption audit verdict: this is a per-spawn *subagent input* cost, not the orchestrator's replayed-context multiplier — a deliberately small, low-risk trim, the only one of the dormant-format set worth touching.)
- **`framework/.claude/skills/new/references/final-review.md`** (Step F.6 item 4) — the per-card + batch **13-field summary report** is collapsed to the minimal result block (`CARD-ID: merged @<sha> | failed: … | deferred → …` one line per card + actionable residue only). All process telemetry (files changed, test/build/lint status, fix cycles, review counts, QA verdict, commit/merge hashes, cleanup, knowledge-sync) is no longer rendered — it lives in the tracker + `skill-runs.jsonl`. The Next Steps & Launch Command (4b) and Production Readiness manual steps stay (they are actionable residue).

## [4.48.0] - 2026-06-16

**`new2` is relaxed: a post-batch, interactive-only escape hatch hands genuinely-blocked hard cases to `/new`'s real human gate — without re-introducing rigidity, a twin, or breaking the A/B.** `new2` runs the batch autonomously, so a card the deterministic policy can't salvage becomes a tracked follow-up + is left `IN_PROGRESS` — the "edge case that wants human intelligence" the user flagged. A 3-lens **adversarial review before implementation** killed the obvious fix (a Step-5 `AskUserQuestion` that lets the skill implement/resolve the residual): (1) **correctness** — the workflow auto-merges and removes the worktree *before* Step 5, so a skill-side fix has no worktree, lands unreviewed code on trunk, and bypasses the F-029 DONE gate; (2) **value** — on the only run with a full `deferral_breakdown` (14 residuals) a post-batch fix changes the outcome of *zero* of them (the batch already merged); only a mid-batch checkpoint could salvage-before-merge, and the data (n=1) doesn't justify breaking `new2`'s background/no-poll contract; (3) **prior-art** — it would twin `/new`'s Phase 2.5b AC-Closure gate and muddy the autonomous-vs-`/new` A/B. What survived all three: the sound escalation is **not to re-implement a gate but to invoke `/new` on the already-materialised follow-up** — `/new` owns the real worktree+review+AC-Closure+F-029+merge pipeline. **MINOR** (additive skill behaviour, interactive-only, autonomous-mode-safe; **no new `baldart.config.yml` key** ⇒ schema-change propagation rule N/A; no removed surface).

### Changed

- **`framework/.claude/skills/new2/SKILL.md`** — new **Step 3b "Escape-hatch escalation"** in the skill's post-batch reconciliation: in INTERACTIVE mode only (skipped when `BALDART_AUTONOMOUS`/`CI`/`GITHUB_ACTIONS` is set), after follow-ups are materialised on disk (offline-safe ordering preserved — the offer is additive over an already-safe ledger) and any `degraded` resume has converged, it presents **one batched `AskUserQuestion`** offering to run `/new` on the **code-actionable** hard-case follow-ups (`deferralClass ∈ {unresolved, out-of-ownership, scope-expansion}`; `owner-gated`/`not-a-code-defect`/`policy-deferred-ac` excluded — `/new` can't perform infra steps). "Sì" invokes `/new <followup-id …>` via the Skill tool (which closes them through its own gates — the skill never marks DONE itself, never re-implements the gate). The **ZERO-ASK CONTRACT** banner is rewritten to scope it precisely to *the workflow during the batch* (the skill may interact pre-launch AND post-batch, interactive-only). New `escape_hatch` telemetry field. Honest limitation documented: post-batch, so it gives the human gate on the follow-up but does NOT salvage a card before its merge (that would need a mid-batch checkpoint — out of scope by design).

## [4.47.0] - 2026-06-16

**`/new`'s orchestrator context economy is re-aimed at its real driver — turn count — and the user-visible Progress Bar + native Task spine are removed.** Telemetry of two real 8-card batches (FEAT-0028/0029 on a consumer) showed the orchestrator paying ~285M `cache_read`: 613 turns each replaying a ~490k-token accumulated context (growing toward ~800k), so total cost ≈ **turn count × accumulated context**. A 3-lens **adversarial review before implementation** refuted the obvious diagnoses: the static prefix is only ~77k (not the ~225k first assumed — context is ~86% *accumulated*, not static); narration prose is only ~7% of the fuel; and the existing § "Context economy" (bulk-content-inline) rule targets a channel that totals only ~119k cumulatively. The measurement reviewer surfaced the actual missed lever — **0 of 274 tool turns batched any calls, and ~55% of turns carried no tool call at all** — and the correctness + prior-art reviewers established that delegating bookkeeping out of the orchestrator is a previously-trodden trap (v4.15.0 reverted a Write-from-memory tracker flush; the tracker is the recovery SSOT; `card_status: DONE` needs orchestrator-side disk re-read; the weak-subagent fabrication precedent applies). What survived: (1) a new turn-economy HARD RULE (batch independent tool calls; no narration-only turns; never poll/wait), and (2) since the Progress Bar + Task spine are pure *mirrors* of the internal tracker (recovery reads the tracker, never the spine), removing them is correctness-safe and eliminates ~45 dedicated visibility turns (~8% of a batch's `cache_read`, guaranteed, not batching-dependent). **MINOR** (skill behaviour change; **no new `baldart.config.yml` key** ⇒ schema-change propagation rule N/A; no removed agent/command/skill/routine).

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — replaced the `## Progress Visibility (MANDATORY)` section (native Task spine + transition Progress Bar + Phase→ledger mapping) with `## State surface — the tracker only`: the internal `/tmp/batch-tracker-*.md` is now the SINGLE state surface (it was already the recovery SSOT; the spine/bar were never read by recovery). Added a second HARD RULE to § "Context economy" — **"turn count is the multiplier"** — instructing the orchestrator to (1) batch independent tool calls into one message, (2) never emit a narration-only turn, (3) never poll/wait on background work. The § "Context Tracking" mirror bullet and the § "Routing" core-invariant list are reconciled.
- **`framework/.claude/skills/new/references/{setup,implement,commit,team-mode,final-review,merge-cleanup}.md`** — every `**→ Visibility**: TaskUpdate … / emit a Progress Bar` step is removed; the underlying *state* writes (card `IN_PROGRESS` claim at implement.md 2b, `card_status: DONE (verified)` at commit.md 29 / team D.6) are **preserved** (they are correctness, not visibility) and now explicitly note the tracker is the only surface. Pre-flight no longer creates a Task spine (setup.md step 2b).

## [4.46.0] - 2026-06-15

**Worktree env-file copy is unified onto one stack-agnostic config key, `stack.env_files` — closing the v4.42.0-deferred divergence (`/nw` copied `.env.local`+`.env`; new2's pre-flight copied `env/.env.local/.env.example/supabase/.temp`) WITHOUT the superset the adversarial pass had refuted.** A worktree is a fresh checkout, so the gitignored env artifacts a build needs must be copied from main — but the SET was hard-coded and divergent across paths. This release makes the set a single SSOT list (`stack.env_files`, default `['.env.local', '.env']`) read identically by `worktree-manager` (`/nw`), `/new`, and `new2`. A 3-skeptic **adversarial review before implementation** shaped the design: it (1) confirmed `framework/agents/runbook.md:36` (`cp .env.example .env`) is a documentation-template onboarding idiom — a DIFFERENT context — and **excluded** it; (2) refuted folding `supabase/.temp` into the copy set (it carries the remote project-ref → every copying worktree auto-links to the shared remote, the exact footgun `stack.schema_deploy_from_trunk_only` exists to prevent, and worse unattended in `/new` than in manual `/nw`; it is not even a build input); and (3) found the bash bug class fixed below. Crucially, the layer is **stack-agnostic**: the generic skill/installer/template never name Supabase or any stack — copying a tool's local-state directory is a per-project opt-in the user adds to their own `stack.env_files` / overlay, never a framework default. **MINOR** (additive: a new `baldart.config.yml` key, propagated end-to-end per the schema-change rule; no removed surface).

### Added

- **`stack.env_files` config key** (`framework/templates/baldart.config.template.yml`) — list of gitignored env artifacts (files OR dirs) copied into each worktree, default `['.env.local', '.env']`. Files are copied with `cp` (dereferences symlinks — never `cp -P`, which would dangle a relative env symlink inside `.worktrees/`), directories with `cp -r` (mirrored: stale files removed on a `/nw` resume). A missing FILE is WARNed at copy time (never `exit 1` — that would strand `/new`'s programmatic path; the build gate is the real enforcement); nothing-copied escalates to a loud WARN. Replaces the old `cp … 2>/dev/null || true` that silently swallowed a missing critical env → cryptic later build failure.

### Changed

- **`framework/.claude/skills/worktree-manager/SKILL.md` — `/nw` step 4 copy loop reads `stack.env_files`** (file/dir-aware, WARN-not-fatal). The dev-server PORT is now written to the first FILE actually copied (not `ENV_FILES[0]`, which could be a directory — `echo >> dir` errors), falling back to `.env.local`. The env-sync staleness check (`/mw` step 2) compares the NEWEST mtime among all `stack.env_files` (portable `stat -f %m || stat -c %Y`, recursive for dirs) instead of only `.env.local`; the `/lw` status line, the port-reuse rule, the docs-mode "forbidden" list, and the Project-Context header are reconciled to the list. The skill stays **stack-agnostic** — it iterates the list and never names a stack.
- **`framework/.claude/workflows/new2.js` + `framework/.claude/skills/new/references/setup.md`** — the pre-flight worktree briefing stops hard-coding `env/.env.local/.env.example/supabase/.temp` and copies `stack.env_files` instead (new2 interpolates the list resolved from `args.config`). `.env.example` is dropped (tracked → already in the checkout).
- **`src/commands/configure.js`** — autodetects `stack.env_files` from the universal gitignored env conventions (`.env.local`/`.env`) only (stack-agnostic — no per-database special-casing), a comma-separated interactive prompt, and a summary-box line.
- **`src/commands/update.js`** — the `missingStack` detector now also catches top-level ARRAY stack keys (`stack.env_files` is `typeof 'object'`, so it needs `Array.isArray`); sub-object keys (`charting`/`animation`/`testing`) stay excluded (`Array.isArray({}) === false`). Without this the new key would silently never be asked on update.
- **`framework/docs/PROJECT-CONFIGURATION.md`** — documents `stack.env_files` (§4.4) as a stack-agnostic list.

## [4.45.0] - 2026-06-15

**The main-repo-root (`$MAIN`) resolution across `worktree-manager` + `/prd` + `/new` is unified onto one correct, `separate-git-dir`-safe canonical — fixing two latent bugs while *refuting* the obvious "unify on `--git-common-dir`" fix the deferred plan proposed.** v4.42.0 deferred the `$MAIN` cleanup after an adversarial pass flagged it as a likely trap; this release does it properly. The root resolution lived in ~5 forms across three genuine execution contexts (main-checkout cwd, worktree cwd, the `allocate-id.sh` startup), and the deferred plan wanted to collapse them onto `git rev-parse --git-common-dir` + `/..` (the recipe already in `allocate-id.sh` `resolve_main()`). A 3-skeptic adversarial review **before implementation** demolished that plan with git experiments: (1) `--git-common-dir` + parent returns the *parent of the git dir*, which is the WRONG directory under `git init --separate-git-dir` (the git dir lives outside the working tree) — so the "canonical" recipe was itself fragile, and `resolve_main()` only worked because BALDART repos use in-tree `.git`; (2) the alleged `prd/SKILL.md` Step-1 bug ("`--show-toplevel` from inside a worktree") does **not** exist — Step 1 runs only at fresh kickoff on the main checkout, and resume reads the persisted `main_path`, so `--show-toplevel` is correct there *and* is the only form that survives `separate-git-dir`; (3) `git worktree list` head is also not `separate-git-dir`-safe (returns the git dir). The corrected design keeps the task's **context classification** but fixes the **primitive**: every site resolves the root through `--show-toplevel` (the true working-tree root in all cases), using `git -C .. rev-parse --show-toplevel` from a worktree (its parent `.worktrees/` lives inside the main repo). `resolve_main()` is rewritten as the canonical reference (detects a linked worktree via `git-dir != git-common-dir`, walks one level up) — **byte-identical output to the old form for in-tree repos** (verified), correct for `separate-git-dir`, and fails cleanly (`return 1`) instead of emitting a garbage path. Two real fragilities fixed along the way: the `/mw` `$MAIN` fallback (`--git-common-dir` + `/..` under `git -C`, relative-base hazard + `separate-git-dir` break) and the `3b` card-sync (`--show-superproject-working-tree || pwd`, which returned the SUPERPROJECT for a real submodule and only worked otherwise by a cwd accident). All recipes dogfooded across normal + `separate-git-dir` repos in every cwd context. **MINOR** (hardening + bug fixes across skills/agents; no removed surface, **no new `baldart.config.yml` key** ⇒ schema-change propagation rule N/A).

### Fixed

- **`framework/.claude/skills/worktree-manager/scripts/allocate-id.sh` — `resolve_main()` rewritten `--show-toplevel`-based.** Now resolves the working-tree root (correct under `git init --separate-git-dir`, where `--git-common-dir` + parent pointed at the wrong directory), detects a linked worktree (`git-dir != git-common-dir`) and walks one level up. Identical output to the prior form for in-tree repos; returns non-zero instead of a garbage path when not in a git repo.
- **`framework/.claude/skills/worktree-manager/SKILL.md` — `/mw` step 4c `$MAIN` fallback** no longer derives the root via `git -C "$WORKTREE_PATH" rev-parse --git-common-dir` + `/..` (parent-of-git-dir, wrong under a separate git dir; appending `/..` to a possibly-relative `--git-common-dir` under `git -C` resolves against the wrong base). It `cd`s into the worktree then `git -C .. rev-parse --show-toplevel`, with a loud guard replacing the silent failure path.
- **`framework/.claude/skills/worktree-manager/SKILL.md` — `3b. Sync untracked cards`** drops the `git -C "$WORKTREE_PATH" --show-superproject-working-tree || pwd` form (returned the *superproject* for a real git submodule, where the cards do not live; only resolved otherwise by relying on cwd). At step 3b cwd is still the main checkout, so it now uses `git rev-parse --show-toplevel` — correct for normal, submodule, and `separate-git-dir` repos.

### Changed

- **`framework/.claude/skills/worktree-manager/SKILL.md` — new `## Resolving the main repo root` section** documenting the three execution contexts, the correct primitive per context, and an explicit "do NOT use `--git-common-dir` + `/..`" rule (codifying the adversarial finding so the next maintainer doesn't re-attempt the refuted unification). The `/nw` step-1, nw-docs step-0, and `/nw` env-copy sites gain cross-ref comments; env-copy keeps `git -C .. --show-toplevel` but replaces its silent `|| echo '../..'` fallback with a loud guard.
- **`framework/.claude/skills/prd/references/validation-phase.md` + `framework/.claude/agents/{prd,prd-card-writer}.md`** — the legacy `$MAIN` fallback and the descriptive `$MAIN` comments switch from `--git-common-dir` to the canonical `--show-toplevel` resolution (and point at the new section). `prd/SKILL.md` Step 1 is deliberately **unchanged** — `--show-toplevel` there is correct (refuted "bug").
- **`framework/.claude/skills/new/references/setup.md` — `$MAIN` resolution corrected**: from a worktree-invocation, resolve via the worktree's parent toplevel rather than `--show-superproject-working-tree` (a submodule primitive) or `git worktree list` (not `separate-git-dir`-safe).

## [4.44.0] - 2026-06-15

**The toolchain layer (v4.41.0) now reaches the *execution-side* gates too: the worktree baseline/merge builds and the classic `/new` reference suite stop hard-coding `npx tsc`/`npx eslint`/`npm run build` and run the consumer's configured `toolchain.commands.*` instead.** v4.41.0 wired the *review* gates (`qa-sentinel`, `coder`, the `/new`+`/new2` review workflows) through `toolchain.commands.*`, but a dozen gates that actually run a build/lint/typecheck were still hard-coded — so a Biome/Vitest consumer's worktree baseline silently ran `eslint`, and the classic `/new` per-card + final gates ran `npm run build` regardless of config. This release makes those gates toolchain-aware while preserving each site's existing **severity** policy (a configured command failing is the same STOP/continue verdict the default produced — the protocol's no-fallback rule governs only *which* command runs, never what to do with its exit code). The mechanism follows the **execution context**: a markdown **prose note** where a full model writes the gate command (the `/new` references, `/bug`, `/simplify`), and an inline **shell resolver** in `worktree-manager` — which has no `args.config` and whose baseline runs in a weak/background subagent, so it resolves `toolchain.commands.*` from `baldart.config.yml` on disk the same way it already resolves `git.trunk_branch`, rather than relying on a prose note a weak model could skip. The design was adversarially reviewed before implementation (3 skeptics), which corrected the initial plan on three points: (1) the `/mw` best-effort `npm run test 2>/dev/null || true` keeps its `|| true` swallow (never a STOP trigger); (2) the `/mw` changed-files lint scope is preserved as the default (a configured whole-tree command like `biome check .` runs by the consumer's contract); (3) two missed twins — `/bug` and `/simplify` verify steps — were added to scope. **Declared debt (intentionally untouched):** `npm install` stays hard-coded (there is no `toolchain.commands.install` key — adding one would be a 5-layer schema change; keeping this MINOR), `markdownlint` is not in the curated map, and the descriptive `(tsc+lint+build)` labels in `new2.js`/`setup.md` are not load-bearing (the `new2` projectBrief already injects the verbatim instruction — refuted by review). **MINOR** (additive: more gates honor the existing `features.has_toolchain` + `toolchain.commands.*`; no removed surface, **no new `baldart.config.yml` key** ⇒ schema-change propagation rule N/A).

### Changed

- **`framework/.claude/skills/worktree-manager/SKILL.md` — gates toolchain-aware via an inline shell resolver.** New `## Toolchain-aware gates` convention + a Project-Context dependency on `features.has_toolchain` + `toolchain.commands.{typecheck,lint,build,test}`. Each gate block (`/nw` step 5 baseline, `/mw` step 2 pre-merge, step 4b rebase build, step 5 post-merge) inlines a `_tc()` resolver that reads the config from disk (mirroring the existing `git.trunk_branch`/`git.merge_strategy` grep idiom) and runs the configured command via `eval "${VAR:-<default>}"`, falling back to the hard-coded default when unset/flag-off. Three carve-outs documented: `npm install` is not a gate key, severity is unchanged, and the `/mw` best-effort `test` line keeps its `|| true` swallow.
- **`framework/.claude/skills/new/SKILL.md` — new `## Toolchain gates` core invariant** (cited as `§ "Toolchain gates"` by the reference modules, same pattern as `§ "Context economy"`): when `has_toolchain`, the mechanical gates use `toolchain.commands.<gate>` verbatim, a non-zero configured command is a real FAIL (no masking fallback), and `npm install`/`markdownlint` are not toolchain keys.
- **`framework/.claude/skills/new/references/{implement,review-cycle,completeness,final-review,commit,team-mode}.md` — gate sites cite `§ "Toolchain gates"`** so the classic `/new` per-card (Phase 1 verify, Phase 2.55/2.5x re-runs, Phase 3 doc re-verify, completeness gap re-run, commit pre-check), team-mode (per-card coder briefing, group build, group simplify re-run), and final-review build all resolve `toolchain.commands.*` verbatim when the flag is on.
- **`framework/.claude/skills/{bug,simplify}/SKILL.md` — standalone verify steps toolchain-aware** (the two twins surfaced by the adversarial scope review): `/bug` PHASE 5 regression checks and `/simplify` Step 5 verify gain a self-contained Toolchain note + a Project-Context dependency on `features.has_toolchain` + `toolchain.commands.{lint,typecheck,test}`.
- **`framework/agents/toolchain-protocol.md` — Consumers section updated** to record the execution-side gates and the two resolution mechanisms (prose note vs on-disk shell resolver), and to restate the `npm install`/`markdownlint` carve-outs.

## [4.43.1] - 2026-06-15

**The npm publish workflow is now race-safe: pushing several `v*.*.*` tags close together can no longer leave `latest` on a lower version.** Publishing v4.41.0/4.42.0/4.43.0 together exposed the bug — the three tag-push workflow runs executed in **parallel**, and `npm publish` (with no explicit dist-tag) points `latest` at whichever version finishes **last chronologically**, not the highest. Real outcome: `latest` landed on **4.42.0**, so `npx baldart` installed a non-latest version (fixed by hand with `npm dist-tag add baldart@4.43.0 latest`). This release closes the race two ways: (1) a workflow-level **`concurrency` group** (`group: publish-npm`, `cancel-in-progress: false`) serializes all publish runs so they never race on the dist-tag and a publish is never cancelled; (2) a new final step **reconciles `latest` to the highest published version** — it computes the max stable semver across `npm view baldart versions` (unioned with the just-published `VERSION` to survive registry propagation lag, numeric per-segment compare so `4.10.0 > 4.9.0`) and repoints `latest` only when it has drifted. Because the runs are serialized, the last run always leaves `latest` on the maximum regardless of push order. The existing "Verify tag matches VERSION" guard is untouched. **PATCH** (CI/release-machinery bugfix, no change to installed surface or `baldart.config.yml` ⇒ schema-change propagation rule N/A).

### Fixed

- **`.github/workflows/publish-npm.yml` — race-safe `latest` dist-tag.** Added a `concurrency` group to serialize parallel tag-push publish runs, and a `Reconcile 'latest' dist-tag to highest published version` step that repoints `latest` at `max(npm view baldart versions ∪ VERSION)` via a self-contained numeric semver comparator (no `semver` dependency), only when `latest` drifts. Fixes `latest` landing on a lower version when multiple tags are pushed close together.

### Changed

- **`MAINTAINING.md` — release protocol note** that pushing multiple tags at once is now safe (since v4.43.1) and why.

## [4.43.0] - 2026-06-15

**The classic `/new` team-mode director stops bloating its own context by legitimately running the review cluster inline — the delegation gate is hardened to the same no-discretion enforcement the sequential path already carries.** A real run analysis (the same FEAT-0028/0029 investigation that produced v4.42.0) found the dominant orchestrator-context driver in team mode is NOT the review token volume (that already runs out-of-context in the `new-card-review` workflow) but the **director's turn count**: the team-mode delegation gate `D.1.6` was labelled "opt-in, additive" with a soft `IF available → delegate`, while the sequential path (`review-cycle.md` Phase 2.5x) carries a strict **MECHANICAL · NO discretion · MUST delegate · GATE VIOLATION if inline** clause. That asymmetry let a director legitimately run the chattiest sub-steps (the per-card Codex loop + per-card Simplify fan-out) inline, re-reading a growing prefix every turn. This release ports the strong clause verbatim into `D.1.6`, plus two pure-distillation fixes (batch the D.1.5 per-card diffs to `/tmp` and read only the compact summary; the same diffs-to-disk discipline on the sequential path for symmetry). Also lands the one **safe** survivor of the deferred worktree-prose cleanup: the `/nw` code-mode now **auto-appends** `.worktrees/` to `.gitignore` (idempotent + WARN) instead of leaving a programmatic `/new` run to create a git-tracked worktree. An adversarial review refuted the rest of that cleanup and a separate review-gate proposal (per-card AC-conformance gate + a second relevance-gating layer) as net-negative/regressive — both deferred. **MINOR** (FIX A changes which path runs by default in team mode — an observable behavior change; no removed surface, no new `baldart.config.yml` key ⇒ schema-change propagation rule N/A).

### Changed

- **`framework/.claude/skills/new/references/team-mode.md` — `D.1.6` delegation gate hardened** (`v4.34.0 → enforcement hardened`). The "opt-in, additive" framing is replaced with the verbatim no-discretion clause from `review-cycle.md` Phase 2.5x ("MECHANICAL, BINARY decision — you have NO discretion … the factors that feel like reasons to run inline are NOT inputs to this gate"), plus the twin telemetry log `D.1.6: GATE VIOLATION — ran inline despite Workflow available + new-card-review.js linked + non-trivial group`. Forces the cluster out of the director's context whenever the `Workflow` tool + `new-card-review.js` are present and the group is non-trivial.
- **`framework/.claude/skills/new/references/team-mode.md` — D.1.5 diff batching (context economy).** The per-card effective-profile computation is now an explicit ONE-bash pass that writes each card's diff to `/tmp/diff-<CARD-ID>.txt` and derives `source|non-source` + doc-touch + the Step-A trigger count there; only the compact per-card summary + `--name-only` lists enter the director's context, never the raw diffs.
- **`framework/.claude/skills/new/references/review-cycle.md` — sequential-path symmetry.** Phase 2.5x input-building writes the card's diff to `/tmp/diff-<CARD-ID>.txt` once and derives both the Step-A grep and `scopeFiles` from it (same diffs-to-disk discipline as team-mode, so the two paths stay symmetric).
- **`framework/.claude/skills/worktree-manager/SKILL.md` — `/nw` code-mode `.gitignore` auto-heal.** When `.worktrees/` is not ignored, the code-mode path now auto-appends it (idempotent) and WARNs, instead of either silently creating a git-tracked worktree or hard-aborting (`exit 1`) on a non-attended `/new` run. (nw-docs keeps its interactive `exit 1` — that path is human-driven.)

## [4.42.0] - 2026-06-15

**The `/new` worktree-setup can no longer pass on a fabricated baseline: the orchestrator now verifies the worktree on disk instead of trusting the subagent's self-report, and the setup subagent moves off `haiku`.** A real `/new` run reproduced **2/2** a silent failure — the background **worktree-setup** subagent (on `haiku`) returned a well-formed block reporting `baseline: pass` in ~6s **with no worktree on disk**: it pattern-matched the expected output instead of running the multi-step `/nw` skill, and the orchestrator trusted it because the baseline gate only ever checked the *returned* field, never the disk. This release closes both halves: (1) the worktree-setup subagent moves **`haiku → sonnet`** (running `/nw` via the Skill tool is a sustained tool-execution chain a too-weak model fabricates rather than executes); (2) a new **worktree integrity gate** (`setup.md §6a`) verifies the worktree with **orchestrator Bash** — `git worktree list --porcelain` + `test -d` + branch + `node_modules` — which a subagent cannot fabricate, and routes any failure to a **non-circular fallback chain** (`subagent → inline /nw → HALT`, no loop, with a build `timeout`). `new2`'s pre-flight gets the parallel mitigation (it cannot run Bash, so it returns non-falsifiable evidence the workflow string-matches — explicitly declared structurally weaker). Separately, the card-baseline validator is now **dependency-free**: it parsed cards with `js-yaml`, absent from the framework payload, so `require('js-yaml')` failed in consumers and silently disabled card-baseline validation (1b-iii) — replaced with a node-core `parseCardYaml`. **MINOR** (additive integrity gate + a behavior change to the worktree subagent model + a dependency-removing bugfix; no removed surface, the `/nw` `{path,branch,port}` contract is unchanged, and no new `baldart.config.yml` key ⇒ schema-change propagation rule N/A).

### Added

- **`framework/.claude/skills/new/references/setup.md` §6a — worktree integrity gate (BLOCKING, orchestrator-Bash).** On `baseline: pass`, the orchestrator verifies the worktree on disk *itself* (`git -C "$MAIN" worktree list --porcelain` lists the path, `test -d`, branch matches, `node_modules` present; registry `buildVerified` + a build-artifact dir as best-effort corroboration) — **never** a `verified` field returned by a subagent (which is exactly as fabricable as the block). Any of checks 1-4 failing routes to the §4d fallback. Honest-failure signals (`baseline: fail`/`timeout`) are trusted and STOP first (recreating cannot fix a broken or hung build).
- **`framework/scripts/validate-card-baseline.js` — `parseCardYaml()`**, a node-core reader for the card-YAML subset (`|`/`>` block scalars, nested maps, scalar/map lists, inline flow collections, number fidelity for `group.sequence` which drives epic detection). Replaces the `js-yaml` dependency. Exported and guarded by a new **Check C** in `scripts/check-card-baseline.js`.

### Changed

- **`framework/.claude/skills/new/references/setup.md` §4b — worktree-setup subagent `model: "haiku" → "sonnet"`** (rationale rewritten: a sustained Skill-tool execution chain, not one-shot plumbing; its return is verified by the §6a disk gate regardless). §4d fallback broadened from "empty / 0-tool-uses" to "did not produce a VERIFIED worktree", made a genuinely **different executor** (inline `/nw`, never a re-spawn or a frozen script) with an explicit cap (`subagent → inline → HALT`; transient-vs-deterministic classification), and the briefing wraps the build in `timeout 600` → new `baseline: timeout`.
- **`framework/.claude/agents/REGISTRY.md`** — the haiku plumbing carve-out drops the worktree-setup sub-clause (now sonnet); only the file-scoped revert agent remains a sanctioned haiku use.
- **`framework/.claude/workflows/new2.js`** — `PREFLIGHT_SCHEMA` gains `worktreeVerified` + `worktreeEvidence` (literal `worktree list --porcelain` / `ls` / baseline-log tail) and `baseline: 'timeout'`; a new **E2.5 worktree integrity gate** string-matches that evidence (the workflow JS cannot run Bash, so it is declared structurally weaker than classic `/new`'s orchestrator-Bash gate), and a timeout is batch-fatal.
- **`framework/templates/ci/check-card-baseline.yml`** — drops the now-unnecessary `npm i js-yaml` step (the validator is dependency-free).

### Fixed

- **Card-baseline validation (1b-iii) silently skipped in every consumer lacking `js-yaml`.** The shipped validator `require`-d a package not in the framework payload, so `/new` / `/new2` pre-flight could not run it and degraded to skip-with-note. Now node-core — it runs everywhere.
- **`/new` worktree-setup accepting a fabricated `baseline: pass`** (reproduced 2/2 on a real consumer) — closed by the §6a integrity gate + the sonnet model + the broadened fallback above.

## [4.41.0] - 2026-06-15

**BALDART becomes opinionated about the *tools*, not just the workflow: a new curated toolchain layer installs best-in-class JS/TS dev tools at first install and makes the agents actually use them.** Until now BALDART shipped agents and workflows but was agnostic about linters/formatters/test-runners — every quality gate hard-coded `eslint`/`tsc`/`jest`. This release adds an opt-in toolchain layer (`features.has_toolchain`) that, on a JS/TS project, PRESELECTS and installs a curated set as devDependencies — **Biome** (format + lint + import organizer), **Vitest**, **tsc**, **Lefthook** (pre-commit) — then records literal gate commands in `toolchain.commands.*` that the gate flows (`/new`, `/new2`, `/qa`, `qa-sentinel`, `coder`) run verbatim instead of guessing. The design is the fourth member of the install-adapter family (alongside routine-/tool-/lsp-adapters) and inherits their invariants: **opinionated but askable** (default Y, opt-out), **non-destructive** (configs written only when absent; existing ESLint/Prettier/Jest/husky are detected and a migration is only ever PROPOSED, never automatic — `.husky/` is never overwritten), **never silent in CI** (`--non-interactive` writes the flag only; `baldart doctor` backfills), and **silent fallback** (an unset command degrades to the project-standard default; the layer is invisible to non-JS projects and consumers with their own toolchain). **MINOR** (additive capability + new `features.has_toolchain` + `toolchain.*` config keys, propagated end-to-end per the schema-change propagation rule; backwards-compatible — flag defaults `false`, every gate falls back to today's behavior when unset).

### Added

- **`src/utils/toolchain-adapters/`** — new adapter family (same dispatcher pattern as `lsp-adapters/`). Curated installers `biome.js` / `vitest.js` / `tsc.js` / `lefthook.js` (each: `installCommand`/`verifyCommand`/`commands()`/`initConfig`/`static replaces`/`static detect`) + incumbent **detectors** `eslint.js` / `prettier.js` / `jest.js` / `husky.js` (detection-only, drive the migration proposal). `index.js` exposes `REGISTRY` + `INCUMBENTS` + `detectAll` + `detectIncumbents`.
- **`src/utils/toolchain-installer.js`** — orchestrator (modeled on `graphify-installer.js`): `recommend()` (clean installs), `migrations()` (incumbent-blocked tools → manual migration), `install()` (devDeps; writes a tool's config BEFORE the package's own postinstall so Lefthook registers our Biome pre-commit, not its commented-out example), `initConfigs()` (non-destructive), `commandsFor()`, `activeTools()`, `certify()`. Never throws.
- **`framework/agents/toolchain-protocol.md`** — runtime protocol: per-gate resolution (explicit `toolchain.commands.*` → project-standard fallback → skip), the command map, and the rule that a configured command which FAILS is a real failure (fallback applies only to *unset* commands). Routed from `framework/agents/index.md`.
- **`framework/docs/TOOLCHAIN-LAYER.md`** — operator guide (plumbing, lifecycle, edge cases, how to add a language/tool).
- **`framework/.claude/skills/toolchain-bootstrap/SKILL.md`** — explicit install path (`/toolchain-bootstrap`), CI-safe; mirrors `/lsp-bootstrap`.
- **`framework/templates/baldart.config.template.yml`** — `features.has_toolchain` + the `toolchain:` block (`installed_tools`, `commands.{lint,format,typecheck,test,test_related,build,audit}`, `auto_verify`).

### Changed

- **`src/commands/configure.js`** — autodetects the prompt default (JS/TS project AND no incumbent → Y), and on enable runs the install/migrate/write/certify flow (clean set preselected; incumbents → migration proposal, never automatic).
- **`src/commands/update.js`** — the schema-drift detector now diffs the `toolchain:` block including its nested `commands.*` map (same nested-block contract as `graph:`), so a future gate key surfaces for pre-existing consumers.
- **`src/commands/doctor.js`** — backfill actions `toolchain-install` (devDeps missing) and `toolchain-init-config` (default config missing), gated on `features.has_toolchain`, never blocking.
- **`src/utils/tool-currency.js`** — `_toolchainRecords` reports curated devDeps behind their npm `latest` (installed version read from `node_modules`, not a global binary); `autoUpgradable:false` (devDep upgrades touch `package.json` — surfaced as a command, user-run). Honest `unknown` offline.
- **Consumers wired to read `toolchain.commands.*` with silent fallback** — `framework/.claude/agents/qa-sentinel.md`, `framework/.claude/agents/coder.md`, `framework/.claude/commands/qa.md`, `framework/.claude/workflows/new-card-review.js` (post-fix re-verify + qa gate brief), `framework/.claude/workflows/new2.js` (pre-flight baseline config facts).

## [4.40.0] - 2026-06-15

**The classic `/new` now slims its single-card final review the same way `new2` already does — closing an asymmetry that made an N=1 batch re-run review finders it had already run.** `new-final-review.js` has carried an F-041 single-card slim since v4.17.x (keep the unique-value cross-model Codex pass + the qa-sentinel merge gate; drop the duplicate Claude breadth finders that already ran per-card), but it only ever fired for **`new2`**, which passes `singleCard`. The classic `/new` final-review delegation (`final-review.md` Step F.1.5) **never passed the flag**, so `slim` was always false and a one-card `/new` batch ran the full breadth set in the final review even though the per-card pass had already covered the same files — a duplicate cross-model Codex + redundant doc review on every single-card run. The user's framing ("for one card, skip the per-card review") was the wrong lever (it would drop Simplify, break the fail-fast ordering that runs review *before* E2E + doc-sync, and lose the early security pass — and was already litigated: v3.35.0 introduced an N=1 final-review skip, v3.37.0 reverted it precisely because the per-card pass can run shallow under `light`). The right lever, already chosen for `new2`, is the inverse: keep the per-card pass, slim the *final*. This release extends that slim to classic `/new` — but **per-finder, coverage-gated**, because classic `/new` differs from `new2`: its `doc-reviewer` runs in the skill's Phase 3 (and can defer to Final under `light`), and its `api-perf-cost-auditor` is deferred-to-final *by design* (never runs per-card). **MINOR** (changes the N=1 merge-gate composition; no new surface, no `baldart.config.yml` key — `singleCard`/`slimDoc`/`slimApi` are workflow args ⇒ schema-propagation rule N/A).

### Changed

- **`framework/.claude/workflows/new-final-review.js`** — the single `slim = a.singleCard` flag is decoupled into **per-finder** `slimDoc` / `slimApi`, each defaulting to `a.singleCard` when absent (so callers passing only `singleCard` — i.e. `new2` — are byte-for-byte unchanged, and a caller passing neither gets an unconditional FULL pass — a safe default that never silently over-skips). `doc-reviewer` is gated on `!slimDoc`, `api-perf-cost-auditor` on `!slimApi && hasApiDataFiles`; Codex + `qa-sentinel` always run. Per-finder skip logs replace the single coupled log line.
- **`framework/.claude/skills/new/references/final-review.md`** — Step F.1.5 now passes `singleCard: cardPaths.length === 1`, `slimDoc: cardPaths.length === 1 && <Phase-3 doc-review ran, not deferred>`, and `slimApi: false` (api-perf never runs per-card in classic `/new` — the final is its only run, so it is never slimmed). The v3.37.0 "FULL gate" invariant prose is reconciled to describe the coverage-gated slim (Codex + qa-sentinel are the unconditional safety gate; breadth finders slim per-finder for N=1 only, gated on per-card coverage — the coverage-gate is the backstop the v3.35.0 blanket skip lacked). The inline Step F.3 fallback (SSOT the workflow mirrors) and the fan-out completion barrier ("ALL THREE" → "every launched Task") carry the same rule, so the no-`Workflow`-tool path stays coherent.
- **`framework/docs/WORKFLOWS.md`** — the `new-final-review` row's single-card description corrected from "skips the duplicate doc/api reviewers" to the accurate per-finder, coverage-gated behavior (`new2` drops both; classic `/new` drops only `doc-reviewer` and keeps `api-perf-cost-auditor`).

## [4.39.0] - 2026-06-15

**`/new`, `/new2`, and `/prd` now auto-reap orphaned Codex MCP servers at their workspace-hygiene finalizers — the v4.37.0 doctor reaper, made automatic.** v4.37.0 added an on-demand reaper to `baldart doctor`, but the leak compounds *per skill run*: every batch's Codex finder calls (`/new`/`new2` per-card review + final review, `/prd` discovery-completeness + plan audit) drive `codex app-server`, whose detached broker spawns the `~/.codex/config.toml` MCP servers (Playwright, …) as children that orphan to init (ppid 1) when the broker dies and keep burning CPU. Waiting for a manual `baldart doctor` let them accumulate between runs. Now each batch ends by sweeping them. A new focused, non-interactive CLI command (`baldart reap-orphans`) is the SSOT the three finalizers call; it shares the v4.37.0 `codex-orphans.js` detection/reaping logic and the same hard safety invariant — it reaps ONLY orphaned MCP servers (ppid 1 ⇒ broker dead ⇒ stdio broken), and NEVER kills a live `codex app-server` broker (a shared, detached runtime that may still serve the user's interactive session). Because an MCP child of a still-warm broker is not yet orphaned, this is a cumulative orphan sweep (catches this run's debris once its broker dies, plus any prior runs'), not a per-run broker teardown. **MINOR** (new CLI command + skill-finalizer wiring; backwards-compatible — non-blocking hygiene step, no-op when nothing is orphaned, no install/layout change, no `baldart.config.yml` key ⇒ schema-propagation rule N/A).

### Added

- **`src/commands/reap-orphans.js`** + **`bin/baldart.js`** — new `baldart reap-orphans` command: detects orphaned MCP servers (ppid 1 + MCP signature) and reaps each process tree via syscall, then prints a one-line summary. `--dry-run` reports without killing; `--json` emits a machine-readable result (`schema:"baldart.reap-orphans/1"`). Always exits 0 (hygiene, never a blocker). Reuses `src/utils/codex-orphans.js` (the v4.37.0 SSOT); live `codex app-server` brokers are detected and reported but never killed.

### Changed

- **`framework/.claude/skills/new/references/merge-cleanup.md`** (Phase 6c, new step 5b), **`framework/.claude/skills/prd/references/validation-phase.md`** (Step 7.5, new non-blocking closer), **`framework/.claude/skills/new2/SKILL.md`** (Step 5, new item 6) — each workspace-hygiene finalizer now runs `npx baldart reap-orphans` as a NON-BLOCKING step and folds its summary into the phase log. `new2` runs it in the main context after the workflow returns (the workflow sandbox cannot run Bash). All three carry the explicit "reaps orphans only, never the broker" note so a future maintainer does not escalate it into a broker kill.

## [4.38.0] - 2026-06-15

**`baldart doctor` now checks whether the external tools BALDART installs are out of date upstream, and offers a one-command upgrade.** BALDART installs external tools into consumer machines but **pins none of them** — `pipx install graphifyy`, `npm install -g typescript-language-server`, … all grab "latest" at install time, and neither pipx nor npm ever auto-upgrades. So a consumer who installed months ago is frozen on whatever version they got, and never receives upstream security/correctness fixes; the `add`/`update`/`configure` flows can't help because they only run at install time. Concrete trigger: Graphify shipped `0.8.37` (SSRF guard thread-safety + prompt-injection mitigation + a macOS NFC/NFD re-extraction loop fix), `0.8.38` (`calls` edge-direction + JS/TS default import/export + tsconfig `paths` correctness) and `0.8.39` (a `graphify affected` `KeyError` crash fix — a command BALDART agents actually invoke) — all invisible to a frozen install. This release adds the continuous currency check that was missing: the tool-dependency analogue of the `baldart` CLI's own `UpdateNotifier`. **MINOR** (new doctor diagnostic + self-heal action; backwards-compatible — network-gated and skipped under `--offline`, zero output when every tool is current, no install/layout change, no `baldart.config.yml` key ⇒ schema-propagation rule N/A).

### Added

- **`src/utils/tool-currency.js`** — external-tool currency prober. `probeAll()` returns one record per managed tool BALDART can query against its upstream registry, comparing the installed version to the latest published one. Probes today: **Graphify** (`graphifyy` on PyPI, gated on `features.has_code_graph` + binary present) and **LSP language servers** (gated on `features.has_lsp_layer`; `npm-global` adapters queried via the npm registry, `system` adapters — gopls, rust-analyzer — reported `unknown` with their canonical `@latest` reinstall command rather than a probe we can't cheaply run). Invariants mirror the rest of doctor's probes: **best-effort** (a failed probe is omitted, never throws, never blocks), **honest** (`status: 'outdated'` only with BOTH versions known and installed < latest — `unknown`/`current` are never surfaced as a nag), **offline-aware** (the caller skips it entirely under `--offline`).
- **`src/utils/http.js`** — dependency-free `getJson(url, {timeoutMs})` over `https` core (no reliance on the Node 18 `fetch` flag), resolving `null` on any failure. Used by the currency layer to read registry metadata.
- **`src/utils/semver-lite.js`** — tolerant dependency-free version comparison (`cmpVersions`) for heterogeneous registry strings (`0.8.39`, `v1.2`, `4.3.3`, `1.2.3-rc1`); returns `0` (no nag) on unparseable input. Shared by `graphify-installer` and `tool-currency`.
- **`src/utils/graphify-installer.js`** — `checkUpgrade()` (installed CLI vs latest `graphifyy` on PyPI; best-effort, async, never nags on uncertainty) and `upgrade()` (`pipx upgrade graphifyy`, falling back to `pip --user --upgrade`; never throws).
- **`src/commands/doctor.js`** — new probe (`state.toolCurrency`, network-gated on `!--offline`) and one non-blocking upgrade action per tool **confirmed** behind upstream (`tool-upgrade:<tool>`, `autoOk: false` — upgrading a system-level tool warrants explicit intent). Auto-upgradable tools (pipx/npm) run the upgrade in place; non-auto ones (system toolchains) print the canonical `@latest` command. `unknown`/`current` tools produce no action — no nag without proof.

## [4.37.0] - 2026-06-15

**`baldart doctor` now detects and reaps orphaned MCP-server processes left behind by BALDART's Codex calls.** A real machine hit ~100% CPU from ~45 orphaned `@playwright/mcp` processes (plus stray `obsidian-mcp-server` instances), all children of OpenAI Codex CLI sessions that had since died. Root cause traced through the Codex companion plugin: every BALDART Codex finder call (`/new`, `new2`, `/codexreview`, the cron review engine) drives `codex app-server` via `codex-companion.mjs`, which attaches to a **shared, `detached + unref'd` broker** (`broker-lifecycle.mjs`). That broker spawns every MCP server declared in the user's `~/.codex/config.toml` (Playwright, Figma, …) as its own children; when the broker dies the OS reparents those MCP servers to init (ppid 1) and they keep running — an `@playwright/mcp` can peg a core for days. The leak compounds across sessions. We cannot suppress the MCP spawn per-call (the companion attaches to a broker it does not control, and exposes no shutdown verb), so the fix is a **safe reaper** owned by the doctor. **MINOR** (new doctor diagnostic + self-heal action; backwards-compatible — zero output on a clean machine, no install/layout change, not a `baldart.config.yml` key ⇒ schema-propagation rule N/A).

### Added

- **`src/utils/codex-orphans.js`** — orphaned-MCP-server detector + reaper. `detectOrphans()` snapshots `ps -axo` and returns MCP servers that are **orphaned (ppid 1) AND match an MCP-server command signature** (`@playwright/mcp`, `playwright-mcp`, `*-mcp-server`, `@modelcontextprotocol/*`, `obsidian-mcp`, npx `*-mcp@*`). `reapOrphans()` kills each orphan's full process tree (so a Playwright MCP's browser children go too) via `process.kill(pid, 'SIGKILL')` — a direct syscall, immune to sandboxed shells that silently swallow multi-arg `kill`/for-loops. **Safety invariant**: ppid 1 means the parent is dead, so an MCP server's stdio pipe is broken and the process is unreconnectable dead weight — safe to reap. The `codex app-server` broker is deliberately NOT reaped: it is `detached + unref'd` by design, so a *live, in-use* shared runtime also shows ppid 1 and ppid 1 cannot tell a leaked broker from a healthy one. Broker processes are detected for visibility only. Fully fail-safe (Windows / any error → "no orphans"); no age threshold (an orphan is dead weight at any age).
- **`src/commands/doctor.js`** — new probe (`state.mcpOrphans`), diagnostic line (`Codex MCP leak — N orphaned MCP server(s) running`, shown only when present so a clean machine prints nothing), and self-heal action `reap-mcp-orphans` (`autoOk: false` — killing processes warrants explicit intent; re-detects against a fresh snapshot at run time so it never acts on a stale list).

## [4.36.0] - 2026-06-13

**`/new` security-domain fixes are now applied by `security-reviewer`, not `coder` — the v4.26.1 canonical writer map, finally propagated from `new2` to `/new`.** Auditing the `new2` lessons for guards/logic missing on `/new` surfaced one real gap (the others — args-string guard, JS router clamp, no-self-judge + specialist-owned lane, relevance-gated fan-out — were already present on `/new`). `new2-resolve.js` routes security fixes to `security-reviewer` (`fixerAgent = {doc:'doc-reviewer', ui:'ui-expert', security:'security-reviewer'}[domain] || 'coder'`), but the canonical writer map was never propagated to `/new`'s SSOT: the `Domain-Override Domains` table (SKILL.md) and every fix-routing site still sent `security` → `coder`. A coder applying a one-line RLS/permission/auth fix lacks the security-invariant contract that lives in `security-reviewer`'s system prompt — the same class of error as "wrong agent for the card", and a direct violation of the user's standing strict-specialization principle. **MINOR** (changes which agent applies security fixes across `/new`; backwards-compatible — `migration` stays `coder`, no install/layout change, no `baldart.config.yml` key ⇒ schema-propagation rule N/A).

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — `Domain-Override Domains` table: `security` owning agent `coder` → **`security-reviewer`** (write mode), plus a new "Why `security` is owned by `security-reviewer`" rationale mirroring the `doc` one. The sequential-mode overview line aligned too.
- **`framework/.claude/skills/new/references/review-cycle.md`**, **`final-review.md`**, **`team-mode.md`**, **`codex-gate.md`** — every security-fix routing site (Phase 2.55 Domain-Override delegation, the delegated-workflow residual routing, the Final FULL merge-blocking partition, the Phase 3.7 codex fix sub-loop, and the doc-drift→bug security path) now routes `security` → `security-reviewer` and runs it before the `coder` pass. `migration` stays `coder`.
- **`framework/.claude/workflows/new-card-review.js`** — the Fix phase no longer folds security into the single coder pass. It partitions `VERIFIED` findings into a `security-reviewer` pass (domain `security`) and a `coder` pass (`code`/`perf`/`migration`/`test`/`simplify`), run sequentially (security first) over the disjoint-by-ownership editable set so shared-file edits never conflict; a FAIL in either pass fails the wave. `new-final-review.js` needed no change (it is read-only — the calling skill applies fixes — and its `domainVerifier` already routes security verification to `security-reviewer`).
- **`framework/.claude/agents/security-reviewer.md`** — new "Dual mode — review vs. apply" Behavior Rule: by default it audits and proposes (read-only), but when invoked as the security domain writer (by `/new`/`new2`/the codex fix loop) it APPLIES the remediation directly via Edit/Write and re-verifies — security fixes are owned by it, never deferred to a coder.

## [4.35.1] - 2026-06-13

**`/new` workflow delegation no longer degrades to a silent no-op when `args` arrives as a JSON string.** A live `/new FEAT-0027 -full` team-mode run delegated its per-wave review cluster to the `new-card-review` workflow and got back a degenerate result (`cards:0`, 0 agents, ~24ms) — the orchestrator correctly fell back to the inline cluster, but the delegation (the single biggest context-economy win in team mode) was wasted on every wave. Root cause: the `Workflow` tool sometimes serializes a structured `args` object to a JSON **string**; `new-card-review.js` and `new-final-review.js` read `args.cards` / `args.reviewScopeFiles` directly, so a string `args` left those `undefined` → empty scope → the early-return guard fired. The `new2` family (`new2.js`, `new2-resolve.js`) had already been hardened against exactly this (`F-001/F-004` parse-or-default guard), but the fix was **never propagated** to the two `/new` workflows — a parallel-location miss. **PATCH** (bugfix to shipped workflow payload, no behaviour change to install, no config key ⇒ schema-propagation rule N/A).

### Fixed

- **`framework/.claude/workflows/new-card-review.js`**, **`framework/.claude/workflows/new-final-review.js`** — added the same defensive `if (typeof a === 'string') { try { a = JSON.parse(a) } catch (_) { a = {} } }` guard already present in `new2.js`/`new2-resolve.js`. All four workflows now tolerate `args` delivered as a JSON string, so `/new`'s delegated review cluster and Final Review fan-out run as intended instead of no-op'ing into the inline fallback.

## [4.35.0] - 2026-06-13

**Card-baseline standardization — every backlog card, any prefix/origin, conforms to one profile-aware SSOT; `/new` normalizes foreign cards at ingestion.** A real `CHORE-0007` (consumer repo `mayo`) reached `/new` **without `review_profile`** (and without `scope`/`scope_boundaries`/`canonical_docs`): it was hand/ad-hoc authored after a graph-align finding, never by the canonical writer. `/new` and `/new2` consume cards **type-blind** — they scale per-card review depth on `review_profile` and run the same pipeline regardless of prefix — so an off-baseline card silently degrades the pipeline. Root cause: the baseline was scattered across `card-template.yml` + `prd-card-writer`'s Required-Fields + Rule C with **no single SSOT and no validator**; `prd-card-writer` only documented the `/prd` epic+children flow (its standalone single-card mode, used by `new2-resolve`, was undocumented); and three writers diverged — `/prd`/`new2`/`new2-resolve` emit the full baseline, but the `/new` AC-deferral stub (`completeness.md`) and `/issue-review` (`issue-review.md`) wrote partial cards.

The fix is **profile-aware** by design (an adversarial review caught that a flat "every field required, non-empty" rule would wrongly reject valid cards — an epic legitimately has `owner_agent: ""` and `files_likely_touched: []`; a standalone CHORE legitimately omits `scope_boundaries`/`links.prd`). A new SSOT module defines per-profile (epic/child/standalone) field states (REQUIRED / may-be-empty / conditional). Every creation path now delegates to `prd-card-writer` (the canonical writer), and `/new`/`/new2` ingestion **normalizes** a non-conformant card: HALT only on truly non-derivable fields (`scope`/`requirements`/`acceptance_criteria`/`files_likely_touched`), **back-fill + persist** the derivable ones (`review_profile` via Rule C, `owner_agent`→`coder`) to the card on disk in the **main repo** (check-and-skip idempotent; committed with the `merge-cleanup.md` Phase 6b COMMIT_LOCK discipline), WARN on the rest. A shipped validator (profile-aware) + CI gate would have failed `CHORE-0007` at authoring time. **MINOR** (new protocol module + new agent capability + new validator/CI gate; backwards-compatible — conformant cards still validate, back-fill only *adds* fields; `review_profile`/baseline are not `baldart.config.yml` keys ⇒ schema-propagation rule N/A).

### Added

- **`framework/agents/card-schema.md`** — Atomic Card Baseline Schema: the universal, profile-aware (epic/child/standalone) field-state matrix every backlog card satisfies, plus the consumer HALT/BACK-FILL/WARN contract. Cites Rule A (`owner_agent` enum, REGISTRY.md) and Rule C (`review_profile` criteria) — never copies them. This is the missing peer of the `agents/index.md` module manifest ("if you create/mutate a card, read card-schema.md"). 23rd protocol module.
- **`framework/scripts/validate-card-baseline.js`** — profile-aware validator. Parses the field-state matrix from `card-schema.md` and the enums from `REGISTRY.md` (no embedded copy → no drift), detects the card profile, and asserts each field's state + enum validity. Exit 1 + per-field errors. Module API (`validateCard`/`detectProfile`/`loadSchema`/`loadEnums`) for reuse.
- **`scripts/check-card-baseline.js`** + new step in **`.github/workflows/check-reference-integrity.yml`** — CI gate: self-tests the validator against fixtures (valid epic/child/standalone + a broken CHORE-0007 shape + invalid enums) and asserts `card-template.yml` ↔ `card-schema.md` do not drift.
- **`framework/templates/ci/check-card-baseline.yml`** — opt-in consumer CI template: fails a PR when a hand-authored `backlog/*.yml` drifts off the baseline (catches manual cards at PR time, not only at `/new` time).

### Changed

- **`framework/.claude/agents/prd-card-writer.md`** — new § "Standalone Single-Card Mode" (documents the no-epic single-card path used by `/new`/`new2`/`new2-resolve`/`/issue-review`: relaxes epic+children and conditional fields, but still emits the full STANDALONE baseline — *standalone ≠ minimal*). "Required Fields Per Card" now cites `card-schema.md` instead of being the implicit SSOT; `description:` frontmatter acknowledges the standalone invocation.
- **`framework/.claude/skills/new/references/setup.md`** — pre-flight step 1b restructured into **1b-i** (HALT on non-derivable fields, incl. `scope`; `scope_boundaries` is NOT a HALT field), **1b-ii** (back-fill `review_profile`/`owner_agent`, persist to `$MAIN`, check-and-skip, dedicated `[BACKFILL]` commit), **1b-iii** (run the profile-aware validator).
- **`framework/.claude/skills/new/SKILL.md`** — QA Profile Selector simplified to a pure read (the compute-from-Rule-C fallback is consolidated into setup.md 1b-ii, which persists the value — no second copy of the criteria).
- **`framework/.claude/skills/new/references/completeness.md`** — AC-deferral follow-up now delegates to `prd-card-writer` Standalone Mode (full baseline) instead of writing a 3-field stub; minimal stub only on total agent outage (self-heals on first ingestion).
- **`framework/.claude/commands/issue-review.md`** — backlog-card creation delegates to `prd-card-writer` (standalone for one-off issues, epic+children for new features) instead of hand-filling the template.
- **`framework/.claude/skills/new2/SKILL.md`** — offline reconciliation cites `card-schema.md` + the standalone mode; notes the crisis stub self-heals via the `/new` back-fill.
- **`framework/agents/index.md`**, **`framework/.claude/skills/prd/assets/card-template.yml`** — routing rule + template header pointer to `card-schema.md` as the field SSOT.

## [4.34.2] - 2026-06-13

**`/new` workflow-delegation gates hardened to BINARY/imperative — close the "ran inline by judgment" escape hatch.** A real `/new` run (mayo, single-card CHORE-0007, BALDART 4.34.1 with both workflows linked + the `Workflow` tool available) ran **both** review-delegation gates inline instead of delegating: the orchestrator invented escape criteria absent from the gate ("single-card → marginal context-economy benefit", "memory note on workflow degeneration") and mis-cited the `feedback_dynamic_workflows_fit` guidance, which actually names the Final Review fan-out as the canonical workflow fit. The degeneration reports it conflated concern the `new2`/`new2-resolve` whole-batch host — a different workflow. Net effect: the run validated only the inline fallback and gave **zero** signal on the `new-card-review` (v4.34.0) path. Both gates were already worded as `IF tool+script → delegate / ELSE inline`, but the prose left enough softness for the model to insert discretion. **PATCH** (gate-prose hardening; no behavior change to the install layout or the workflows themselves — the delegated path was always the intended one).

### Changed

- **`framework/.claude/skills/new/references/review-cycle.md`** — Phase 2.5x branch reframed as a "MECHANICAL, BINARY decision" with an explicit **no-discretion** rule: the soft factors (N=1, small diff, marginal benefit, "a workflow once degenerated", "inline is safer") are named and declared **NOT inputs to the gate**; the `new2` degeneration conflation is corrected inline; running inline while the conditions held is now defined as a **gate violation** with a mandated telemetry log line.
- **`framework/.claude/skills/new/references/final-review.md`** — Step F.1.5 branch given the same imperative treatment (MUST-delegate, no-discretion rule, `new2`-conflation correction, gate-violation log line).

## [4.34.1] - 2026-06-12

**`reference-integrity` CI gate: recognize the plumbing carve-out.** v4.33.2 introduced a legitimate `subagent_type: general-purpose` dispatch for the mechanical file-revert agent (the REGISTRY.md plumbing carve-out — mechanical git/file ops, never code authoring), but the CI `reference-integrity` gate's R11 anti-fallback rule banned `general-purpose` **unconditionally**, so `main` went red on the v4.33.2 / v4.34.0 tags (the npm publish workflow is independent and succeeded — both versions are live). The gate now allows `subagent_type: general-purpose` **only** on a dispatch line explicitly marked `plumbing carve-out`; every other `general-purpose` dispatch stays banned (R11 intact, intentionally narrow so it can't be abused as a generic fallback). **PATCH** (CI gate fix; no behavior change to `/new`).

### Changed

- **`scripts/check-reference-integrity.js`** — Check A allows `subagent_type: general-purpose` when the dispatch line carries the literal `plumbing carve-out` marker; otherwise the R11 ban stands (with an error message pointing to the marker + REGISTRY.md).

## [4.34.0] - 2026-06-12

**`/new` per-card review cluster extracted into the `new-card-review` dynamic workflow — kills the dominant context-cost driver on long epics.** On a long epic (10-12 cards) `/new` accumulated context monotonically: every card ran the full review fan-out (Simplify ×3 + Codex/code-review + qa + the verify/FP pass) *inside the orchestrator*, and each step re-paid the whole growing prefix (cumulative `cache_read`). The driver is **turn-count × prefix**, not subagent output volume. Following the validated `new-final-review` pattern (read-only fan-out hosted in a workflow whose context is discarded), the review cluster now runs OUTSIDE the orchestrator and returns only a compact result.

A **single workflow parametrized on the number of cards** covers both modes: sequential delegates `cards:[1]`; team-mode delegates `cards:[the whole wave]`, so it runs **once per wave, not once per card** — the key win, since the team-mode per-card Codex (D.4b) + Simplify (D.3b) loops were the chattiest sub-steps. The workflow fans out the finders per card (Simplify + cross-model Codex with `code-reviewer` fallback + qa-sentinel at group-max tier + security-reviewer on high-risk only), each specialist FP-checking its own findings; then **one `coder` applies all VERIFIED code/perf/security/simplify findings in a single pass** (files disjoint by ownership) and re-verifies lint/tsc/build. Only `{perCard:{fixesApplied,residual}, gateTable, summary}` re-enters the orchestrator — the minority `residual` (doc, needs-manual, scope-expanding, unconverged) is resolved by the skill with the right specialist or a user gate.

**Boundaries (by design):** E2E (Phase 2.6 / D.3c — human-gated + nests a skill) and doc-review (Phase 3 / D.2+D.4a — write-mode, must see final code) stay in the skill; doc runs **after E2E on final code** in the delegated path. `/codexreview` (Phase 3.7 / D.4b — a skill that nests sub-agents) cannot run inside a workflow, so the workflow launches the Codex **binary** directly via an agent (haiku preflight + background poll), exactly like `new-final-review`; the **Final FULL gate** remains the cross-card depth net. api-perf stays deferred to the Final. **Opt-in + additive**: delegates only when the `Workflow` tool is available AND the script is linked, else the inline prose (SSOT) runs unchanged — Codex/cross-tool consumers are unaffected. **MINOR** (new capability, fallback inline, no breaking change; no `baldart.config.yml` key ⇒ schema-propagation rule N/A). Validation deferred to telemetry on the first real run (per `/new` convention).

### Added

- **`framework/.claude/workflows/new-card-review.js`** — the per-wave review+fix workflow (1..N cards). Modeled on `new-final-review.js` (shared schemas, deterministic Codex preflight haiku + background poll, specialist-owned FP-check) + Simplify/security finders + the single-coder batch-fix pass.

### Changed

- **`framework/.claude/skills/new/references/review-cycle.md`** — new "Phase 2.5x — Review-cluster workflow delegation gate" (replica of final-review.md F.1.5): delegates Phase 2.55+3.5+3.7 to `new-card-review` (`cards:[this card]`), consumes the typed `residual`, then proceeds E2E→Doc→Commit; else inline SSOT. `IS_TRIVIAL` never delegates.
- **`framework/.claude/skills/new/references/team-mode.md`** — new "D.1.6 — Review-cluster workflow delegation gate": delegates the group's D.3b+D.4+D.4b in ONE call per wave (`cards:[the wave]`), runs D.3a (AC-closure) before it and D.3c (E2E) + post-E2E doc + D.5/D.6 after; else inline.
- **`framework/.claude/skills/new/references/codex-gate.md`** — Phase 3.7 carve-out: SKIP when the cluster was delegated (the workflow already ran the Codex pass).
- **`framework/.claude/skills/new/SKILL.md`** — routing table notes the `new-card-review` delegation for the 2.55+3.5+3.7 cluster (parity with the final-review row).
- **`framework/docs/WORKFLOWS.md`** — documents the `new-card-review` workflow.

## [4.33.2] - 2026-06-12

**`/new` runs its two purely-mechanical spawns on haiku instead of opus.** Audit of every agent `/new` spawns (not `new2`) found that the spawns inherit each agent's frontmatter default and none ever override `model:` — correct for the reasoning agents (`coder`/`ui-expert` = opus; the review/analysis agents = sonnet), but the **background worktree-setup subagent** that runs `/nw` is a `general-purpose` agent with no `model:`, so it inherited the **session model (opus)** to do pure git plumbing (create worktree, install deps, allocate port, write registry). Same shape for the **file-scoped revert agent** (was spawned as `coder`/opus to do a mechanical revert). Both now run on **haiku** — cheaper and faster on the background barrier, with no quality loss because neither does any reasoning or code authoring. `qa-sentinel` deliberately stays on sonnet (failure interpretation + test-tier selection). `/mw` is invoked inline by the orchestrator (not a subagent) so it is unaffected. No `baldart.config.yml` key involved (model selection is native) ⇒ schema-change propagation rule does not apply. **PATCH** (cost/perf, no new capability, no breaking change).

### Changed

- **`framework/.claude/skills/new/references/setup.md`** — the §4b background worktree-setup subagent (`general-purpose`, runs `/nw`) now spawns with `model: "haiku"`, with a rationale noting it is deterministic plumbing.
- **`framework/.claude/skills/new/references/implement.md`** — the Phase 2.4 unauthorized-file revert is now a `general-purpose` + `model: "haiku"` agent with an explicit ROLE BOUNDARY (mechanical file-scoped revert, never code authoring), replacing the prior `coder`/opus revert spawn.
- **`framework/.claude/agents/REGISTRY.md`** — the Model Selection Matrix "never haiku" rule is scoped to **named specialist agents**, with an explicit **plumbing carve-out** sanctioning `model: "haiku"` for mechanical `general-purpose` plumbing (the two `/new` uses above); `qa-sentinel`'s stay-on-sonnet rationale made explicit.

## [4.33.1] - 2026-06-12

**Discoverability follow-up to v4.33.0** — the new `stack.schema_deploy_from_trunk_only` gate guard was enforced in the `/new` reference modules + `new2.js`, but the **"Reads from `baldart.config.yml`"** manifest line at the top of each skill body still omitted it. Surface it so the key is discoverable from the skill's config-reads declaration (parity with `stack.database` / `stack.deployment`). Manifest-only — **no behavior change**. **PATCH**.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — add `stack.schema_deploy_from_trunk_only` to the "Reads from" gate-guards list.
- **`framework/.claude/skills/new2/SKILL.md`** — make the `stack.*` read explicit about `stack.schema_deploy_from_trunk_only` (the trunk-only schema-deploy gate guard).

## [4.33.0] - 2026-06-12

**New opt-in `stack.schema_deploy_from_trunk_only` safety invariant — `/new` + `new2` never push a remote schema/RLS/index/migration deploy from a non-trunk branch or worktree.** Before this release, the Production Readiness checklist (`/new` Phase 7 / `new2` Production phase) and the front-loaded Migration Gate would auto-execute a remote schema deploy (`supabase db push`, `prisma migrate deploy`, `firebase deploy --only firestore:rules|firestore:indexes|storage`, CDK/Terraform apply, any SQL migrator) gated only on `stack.database`/`stack.deployment` — with no guard against running it from an unmerged feature branch / git worktree. On projects that develop multiple cards in parallel worktrees against a single shared remote datastore, that desyncs migration history and produces "code-ahead-of-schema" outages (a real consumer hit a `42703 undefined_column` production outage this way).

The new key is **stack-agnostic and opt-in**, a no-op for projects that don't set it:
- `false` (default) / unset → current behavior preserved exactly (unset → the skill asks; `baldart configure` defaults it to `true` only when it autodetects parallel worktree usage).
- `true` → a remote schema/RLS/index/migration deploy is auto-executed **only** when `git rev-parse --abbrev-ref HEAD == git.trunk_branch`. From any non-trunk branch / worktree it is **never** auto-run: the migration stays committed and the remote deploy is deferred ("happens from trunk after merge") and routed to the existing **`db-migration-deploy`** residual class (dedup-collapsed onto one external action).

**The guard enforces only the branch invariant** — the actual deploy command, env loader, and gate UX stay delegated to the project overlay (`.baldart/overlays/new.md`); core does not move the command in. Datastore-agnostic by design (any persistence layer, not Supabase-specific). The principle is documented as a MUST-rule in `production-readiness.md` and the generic `deployment-protocol.md`. **MINOR** (new opt-in config key + capability; default false ⇒ zero behavior change for existing consumers). Schema-change propagation rule applied end-to-end (template + configure prompt/autodetect + update detector + skills + docs + CHANGELOG).

### Added

- **`framework/templates/baldart.config.template.yml`** — new `stack.schema_deploy_from_trunk_only: false` key with documentation.
- **`src/commands/configure.js`** — `detectWorktreeUsage()` probe (`git worktree list`); the autodetected default and an interactive confirm prompt for the new key (pre-checked when parallel worktrees are detected).

### Changed

- **`framework/.claude/skills/new/references/production-readiness.md`** — new "Schema-deploy-from-trunk-only guard" subsection + a MUST-rule in § Rules: when the flag is on, gate every remote schema/RLS/index/migration auto-deploy on `current branch == git.trunk_branch`, else defer as a `db-migration-deploy` residual.
- **`framework/.claude/skills/new/references/setup.md`** — Migration Gate (Phase 0 step 1b) honors the same guard: a command modality (front-loaded apply) runs only when `$MAIN` is on trunk; otherwise it degrades to the end-of-batch owner-gated deferral.
- **`framework/.claude/workflows/new2.js`** — reads `stack.schema_deploy_from_trunk_only`; the Production phase agent gates schema deploys on `HEAD == TRUNK` and returns `schemaDeploysDeferred`, which the workflow converts into `db-migration-deploy` residuals (owner-gated, collapsed by `ownerGatedActionKey`).
- **`framework/.claude/skills/new2/SKILL.md`** — Migration Gate Step 3.5 step 6 honors the off-trunk guard (degrade to deferred when `$MAIN` is not on trunk).
- **`src/commands/update.js`** — the `missingStack` schema-drift detector now surfaces boolean scalar `stack.*` keys (not just strings), so `schema_deploy_from_trunk_only` is offered to pre-v4.33.0 consumers on update.
- **`framework/agents/deployment-protocol.md`** — new datastore-agnostic "Schema deploys run from trunk only" best-practice subsection.
- **`framework/docs/PROJECT-CONFIGURATION.md`** — documents the new `stack.schema_deploy_from_trunk_only` key in § 4.4.

## [4.32.0] - 2026-06-11

**`new2`: `deep` is now relevance-gated too — it no longer fires all 5 reviewers on every card regardless of content.** Diagnosing a real batch (the FEAT-0023 supplier-associated-users epic — permissions/RLS/migrations, so 7 of 11 cards are legitimately `review_profile: deep` per `prd-card-writer` Rule C) exposed an **asymmetry**, not a misclassification: the per-card review matrix relevance-gated `balanced` (a specialist runs only if its domain is evidenced by `scopeFiles ∪ MAY-EDIT`) but left `deep` on an **unconditional 5-way fan-out** (`FULL_FANOUT` = code-reviewer + doc-reviewer + qa-sentinel + api-perf-cost-auditor + security-reviewer). So a `deep` card with no doc surface still paid for a doc-reviewer, and one with no API/data surface still paid for api-perf-cost-auditor — every time. This is exactly the "every card gets a full review regardless of what's inside" symptom.

The fix unifies the review MODEL: **the profile controls review DEPTH, the surface controls review BREADTH.**
- `deep` and `balanced` now share the same relevance-gated reviewer SET (extracted into `relevanceGated()`): `touchesCode` → code-reviewer + qa-sentinel; `touchesDocs` → doc-reviewer (the core finder on a doc-only card); `touchesApiData` → api-perf-cost-auditor; `securityRelevant` → security-reviewer.
- `deep`'s extra DEPTH is unchanged and is carried downstream, NOT by the fan-out set: the full QA suite + Codex-full posture reach each spawned reviewer via `cardBrief`'s "Review profile" line, and the Phase-1 architect audit still keys on `reviewProfile === 'deep'` (`needAudit`). So the reviewers that DO run review at full depth — only the ones whose domain the card never touches are dropped.
- **Fail-safe preserved**: `noEvidence` (empty `scopeFiles ∪ MAY-EDIT`) still falls back to `FULL_FANOUT` for BOTH balanced and deep — absence of evidence ≠ evidence of absence. `light`/`skip` unchanged. Batch-level coverage is unchanged: the Final Review's batch-wide doc + api/perf passes remain the safety nets, and `securityRelevant` (which a `deep` card almost always trips) keeps security-reviewer on genuinely sensitive cards.

Not touched: `prd-card-writer` Rule C (it classifies correctly — `deep` for migration/RLS/permission/schema/HIGH-integration cards, `light` for pure-UI, `skip` for epic/doc — verified against the FEAT-0023 cards) and `/new`'s interactive prose path (which already scales per-card by profile + relevance — api-perf is deferred to the Final Review F.3, qa deferred at balanced, doc deferred at light-no-doc — so the unconditional-deep fan-out was unique to `new2.js`). **MINOR** (review-behavior change on the EXPERIMENTAL `new2` surface only; no config key — the schema-change propagation rule does not apply; no change to `/new`).

### Changed

- **`framework/.claude/workflows/new2.js`** — per-card review matrix (B7): the relevance gate now applies to `deep` as well as `balanced` via a shared `relevanceGated()` helper; the `reviewProfile === 'deep'` arm of the `FULL_FANOUT` branch is removed, leaving only `noEvidence` as the conservative full-fan-out fail-safe. The `review-matrix` ledger row's `→conservative-full` annotation now fires for any `noEvidence` profile (was `balanced`-only). Comment block updated to document the depth-vs-breadth model.

## [4.31.1] - 2026-06-11

**`new2`: fix v4.31.0's dedup coverage + drop empty residuals + honest A/B cost — verified against the real FEAT-0022 telemetry/report (not the narrative).** Reading the actual run's `skill-runs.jsonl` + workflow report (instead of the prior turn's recollection) exposed three things v4.31.0 missed:
- **The owner-gated dedup was too narrow** — it filtered by `deferralClass ∈ {owner-gated, not-a-code-defect}`, but the run's FIVE identical `db:push` residuals on card -01 carried kinds `ac-unmet`/`blocker`/`policy-deferred-ac`; the two `policy-deferred-ac` (an F-016 class that is ALWAYS an external/infra action) slipped through. Fix: dedup on `EXTERNAL_DEFERRAL = {owner-gated, not-a-code-defect, policy-deferred-ac}`. A second defect caught by a self-test before publishing: the action key split those 5 across `db-migration-deploy` and `migration:<file>` (the filename branch ran first), collapsing to 2 not 1 — **one `db:push` pushes all pending migrations, so the deploy intent must win over a bare filename**; reordered the key so every db-migration-deploy residual maps to one key. Verified on the real residuals: 7 db:push residuals (5 on -01, 1 on -02, 1 batch-level `F001`) → 2 (one per real card; `F001` dropped), matching the hand reconciliation.
- **Empty out-of-scope residuals** — the run emitted two `out-of-scope` residuals with no file/line/evidence (`":"`), which would mint contentless follow-up cards. `resolve()` now skips an out-of-scope finding whose composed evidence is empty after stripping `:`/space separators.
- **A/B cost telemetry was ~8× under** — the workflow reported `total_tokens=476k` (`budget.spent()`, output-only) / `agent_count=30` while the harness saw ~4.15M subagent tokens / 64 agents, because the workflow counters exclude the subagents spawned inside the nested workflows (`new2-resolve`, `new-final-review`). Since `total_tokens` was non-null, the skill's "backfill only if null" rule accepted the wrong figure — defeating new2's whole purpose (A/B on context economy). SKILL.md Step 5.5 now mandates **always** reading the real transcript `usage` via the `-stats` script as the headline cost, labelling the workflow figures as partial.

**PATCH** (corrects v4.31.0's dedup coverage + noise filter + telemetry honesty on the EXPERIMENTAL `new2` surface; no new capability, no config key). Meta: same lesson as the feature itself — the v4.31.0 dedup was written from the prior turn's recollection; reading the raw telemetry/report found the gap. Validate fixes against the data, not the memory of the data.

### Changed

- **`framework/.claude/workflows/new2.js`** — `dedupOwnerGatedResiduals` now keys on `EXTERNAL_DEFERRAL` (adds `policy-deferred-ac`); `ownerGatedActionKey` reorders the db-migration-deploy branch ahead of the filename branch (one key per deploy action); `resolve()` skips empty `out-of-scope` residuals.
- **`framework/.claude/skills/new2/SKILL.md`** — Step 5.5 cost recording: always backfill real transcript `usage` (the workflow's `total_tokens`/`agent_count` are partial — output-only + exclude nested workflows); keep them labelled `*_workflow`.

## [4.31.0] - 2026-06-11

**`new2`: the residual ledger self-corrects before returning — no more duplicate-of-done follow-ups, no more N defers for one external action.** A real `new2` run (FEAT-0022 epic, 3 cards) surfaced two over-report classes in the offline-safe residual ledger that the skill was absorbing **by hand** every run: (1) **4 of 8 follow-ups were false-open** — scope-expansion residuals deferred early in the batch but satisfied LATER by another card's commit / a final-review fix, which only `integrateCrossCard()` retracts; a residual closed by any *other* in-batch path stayed falsely-open (and left a best-effort, uncommitted follow-up YAML in the worktree). (2) **3 follow-ups for ONE physical action** — one migration's remote `db:push` re-raised per-card AND batch-wide by the final review → three near-identical owner-gated cards. Both were caught only by the skill's manual per-residual disk grep + consolidation (load-bearing, repeated every run). This release moves that work into the workflow, once, deterministically:
- **Ledger self-correction (`reconcileLedgerAgainstHead`)** — before returning, re-verify every still-pending code/doc residual (`scope-expansion`/`out-of-scope`/`unresolved`/`file-diff-violation`/`merge-artifact-skipped`/`out-of-ownership`, `materialized` true OR false) against the worktree HEAD via one read-only agent, and retract only the ones it can back with a `file:line` proof. **Conservative by design** — default KEEP: a false-open residual is recoverable downstream (the skill re-checks disk), but a wrong retract silently drops real work (F-029). Owner-gated / not-a-code-defect / policy-deferred are EXTERNAL actions a commit cannot close → never auto-retracted here. Telemetry `ledger_reconciled`.
- **Owner-gated dedup (`dedupOwnerGatedResiduals`)** — collapse owner-gated / not-a-code-defect residuals that share one action key (migration filename · `db:push`/`db:check-sync` · deploy · secret · DNS), keeping **one per distinct real card** (the skill marks each card DONE only after ITS follow-up exists) and dropping only batch-level duplicates (residual `card` is a finding id → no DONE-linkage); a batch-level residual with no matching per-card entry is a genuinely-new action and is kept. Telemetry `owner_gated_deduped`.

The skill's Step 5.1 disk reconciliation **still runs** — it is now a safety net over a pre-cleaned ledger (the self-correction is conservative; F-040 worktree-not-merged still applies), not the sole defence. Same recurring shape as prior `new2` fixes: the splice existed in ONE location (`integrateCrossCard`), the resolution-detection was needed in ALL paths that close a residual. **MINOR** (additive capability + observability on the EXPERIMENTAL `new2` surface; no behavior regression — the guards/policies are unchanged and the retract is proof-gated + conservative; no `baldart.config.yml` key, so the schema-change propagation rule does not apply; no change to `/new`).

### Added

- **`framework/.claude/workflows/new2.js`** — `reconcileLedgerAgainstHead()` (new `Reconcile` phase, conservative proof-gated retract of already-satisfied residuals) + `dedupOwnerGatedResiduals()` (collapse duplicate external actions), both run right before `finalReturn`. New telemetry fields `ledger_reconciled` + `owner_gated_deduped`; new counters wired through `buildTelemetry`.
- **`framework/.claude/skills/new2/SKILL.md`** — documents that `residuals[]` now arrives pre-cleaned (Step 5.1 reframed as a safety net) and records the two new telemetry fields in the A/B step.

## [4.30.1] - 2026-06-11

**`new2`: stop spending opus on mechanical ops steps — explicit per-step model overrides.** Three `general-purpose` agents had no `model:` override, so they inherited the session's main-loop model (opus) for work that needs none. The Merge step is a deterministic OPS/GIT executor (git merge + YAML status reconciliation + grep-based epic closure + leave-and-report hygiene gates) whose correctness-critical checks (F-029 forcedDone guard, F-040 deferred guard) are enforced in JS AFTER it returns — independent of the agent's reasoning → **sonnet**. The per-card Codex review agent is a pure DRIVER (runs the companion, strips `[codex]` traces, maps findings) — the review intelligence is Codex, run externally → **haiku**. The post-merge Production Readiness checklist is non-blocking report-not-execute → **sonnet**. The Pre-flight agent (DAG + ownership map + idempotency — it grounds the whole batch) intentionally **stays opus**. **PATCH** (cost optimization on the EXPERIMENTAL `new2` surface; no behavior change — the deterministic guards/policies are unchanged; no config key).

### Changed

- **`framework/.claude/workflows/new2.js`** — `merge` → `model: 'sonnet'`; `review:<card>:codex` driver → `model: 'haiku'`; `production-readiness` → `model: 'sonnet'`. `preflight` unchanged (opus).

## [4.30.0] - 2026-06-11

**`new2`: Cross-Card Integration Pass — implement residuals deferred by a per-card ownership artifact in-batch, instead of leaving them as follow-ups to manage.** A residual deferred `out-of-ownership` is often a *process* artifact: during the per-card pipeline each card may only touch its own MAY-EDIT, so a fix that needs another card's files is deferred. But `new2` is strictly sequential — once every card is committed that boundary no longer binds. A new pass, run BEFORE the final review, re-runs `resolve()` for those residuals with ownership widened to `card MAY-EDIT ∪ remedy files` (domain-routed — never a raw coder), but ONLY when the remedy lands inside the batch union (`within()`); a remedy outside the batch is a future card's work and stays a follow-up. Transient `outage` residuals are retried with the card's own ownership. Everything else (`owner-gated` / `not-a-code-defect` / `baseline-not-reached` / new-AC `scope-expansion`) still defers — a coder physically cannot do those, or it would be silent scope creep. Integrated fixes are committed (one `[integrate]` commit) so the final review covers them, and a fully-freed card is marked DONE. **MINOR** (additive capability on the EXPERIMENTAL `new2` surface; scope per the user decision = out-of-ownership-within-batch + outage only; no `baldart.config.yml` key, so the schema-change propagation rule does not apply; no change to `/new`).

### Added

- **`framework/.claude/workflows/new2.js`** — `integrateCrossCard()` (new `Integrate` phase between `Implement` and `Final`): collects `out-of-ownership`/`outage` residuals, batch-union `within()` gate on remedy files, surgical ownership widening, domain-routed `resolve()` re-run, one integration commit + YAML DONE-marking for fully-freed cards, ledger cleanup. New telemetry `cross_card_integrated`; the residual ledger now carries `domain` + `remedyFiles`, and `cardMayEdit` records per-card MAY-EDIT for the union.
- **`framework/.claude/workflows/new2-resolve.js`** — `noFollowup` arg: when the integration pass re-runs an already-tracked residual, a failed retry returns the verdict WITHOUT minting a duplicate follow-up card (the original residual stays tracked). `materialiseFollowup()` now propagates `remedyFiles` so the caller can decide whether a remedy is in-batch.
- **`framework/.claude/skills/new2/SKILL.md`** — the A/B record step keeps `cross_card_integrated` alongside `deferral_breakdown` (how many deferrals were genuinely undeferrable vs absorbed in-batch).

## [4.29.1] - 2026-06-11

**`new2`: `deferral_breakdown` telemetry — see WHY residuals became follow-ups, per class.** Follow-up cards in `new2` are not "useless deferred work": an in-scope fixable anomaly is fixed directly by `resolve()`'s domain fixer in-batch; a follow-up is created ONLY for a residual that is structurally undeferrable (`out-of-ownership` would edit another card's files · `owner-gated`/`not-a-code-defect`/`baseline-not-reached` aren't code a coder can apply · `unresolved` already failed fixer+judge+tier-2 · `scope-expansion` with a new AC needs a PRD decision · `outage`). To diagnose a run with *many* follow-ups, the telemetry now counts residuals per class so a skewed breakdown points at the real root cause (many `out-of-ownership` → PRD MAY-EDIT too narrow · many `unresolved` → fixes genuinely hard · many `scope-expansion` → cards under-specified upstream) — data first, before any change to the deferral logic. **PATCH** (observability only on the EXPERIMENTAL `new2` surface; diagnosis-only — nothing is auto-implemented from a residual; no behavior change, no config key).

### Added

- **`framework/.claude/workflows/new2.js`** — `telemetry.deferral_breakdown` (per-class residual counts, derived from the `deferralClass`/`kind` already carried on each residual) + a "Ripartizione per classe" line in the human report's Residui section.
- **`framework/.claude/skills/new2/SKILL.md`** — the A/B record step now keeps `deferral_breakdown` and names it the data to consult BEFORE proposing any deferral-logic change.

## [4.29.0] - 2026-06-11

**Final review (F.4): each domain specialist owns its lane — no cross-domain `code-reviewer` re-judge, no self-judge.** The verification step classified findings by spawning `code-reviewer` over EVERY low-confidence finding. That was wrong on two counts: (a) a `doc` finding from `doc-reviewer` (or an `api`/`perf` finding from `api-perf-cost-auditor`) was re-validated by `code-reviewer` — the WRONG specialist judging another domain (a `code-reviewer` judging prose); (b) when Codex is unavailable the `code-reviewer` *fallback* produces the findings and `code-reviewer` then re-judged its OWN findings — self-judging with no model diversity (the same waste removed from the resolve pass in v4.27.1–2). Now every domain specialist FP-checks its OWN findings in the finding pass (doc-reviewer, api-perf-cost-auditor, and the Codex/`code-reviewer`-fallback code engine), so surviving findings arrive already validated; the residual `confidence < 80` path is routed to the finding's DOMAIN specialist (doc→doc-reviewer, api/perf→api-perf-cost-auditor, security/migration→security-reviewer, test→qa-sentinel, else code-reviewer), and when that specialist is the originating finder the finding is surfaced as `NEEDS_MANUAL_CONFIRMATION` rather than re-judged. Applies to BOTH `/new` (inline F.4 prose, the SSOT) and `new2` (the `new-final-review` workflow). **MINOR** (behavioral refinement of the final-review classification rule; the returned `{findings, classification, summary}` contract is unchanged; no `baldart.config.yml` key, so the schema-change propagation rule does not apply).

### Changed

- **`framework/.claude/skills/new/references/final-review.md` Step F.4 step 9** (SSOT) — replaced "Claude agent findings with confidence < 80 → cross-validate by spawning code-reviewer" with the specialist-ownership rule: specialists self-FP-check their own domain, residual unresolved findings route to the domain specialist, finder===verifier ⇒ `NEEDS_MANUAL_CONFIRMATION`.
- **`framework/.claude/workflows/new-final-review.js`** — `doc-reviewer`/`api-perf-cost-auditor` prompts now mandate an in-domain false-positive check; their findings (and the `code-reviewer` fallback's) are marked `preValidated` so they skip the re-judge; `verifyFinding()` uses a new `domainVerifier()` router (never a hardcoded `code-reviewer`) and short-circuits a finder===verifier case to `NEEDS_MANUAL_CONFIRMATION`. `meta.description` + Verify phase detail updated.

## [4.28.1] - 2026-06-11

**Doc: list the `new.migration-example.md` overlay example in the overlays README.** Completes the v4.28.0 Migration Gate docs — the example overlay file shipped in v4.28.0 was not yet listed in the `framework/templates/overlays/README.md` example table. **PATCH** (doc-only).

### Changed

- **`framework/templates/overlays/README.md`** — added the `new.migration-example.md` row (Migration-Gate apply modalities for `/new` Phase 0 step 1b · `new2` Step 3.5).

## [4.28.0] - 2026-06-11

**`new2`: front-loaded Migration Gate — declared DB migrations are applied BEFORE the batch, so downstream cards verify against a live schema.** The workflow runs autonomously and cannot apply an owner-gated DB migration mid-run, so a declared migration was deferred to end-of-batch — and every downstream card was built/verified against a schema not yet live (`validation_commands` / QA / E2E / DB-generated `tsc` types fail falsely → those cards cascade into deferral). The new gate runs in the **skill** (the main loop, which can interact — the workflow's zero-ask contract is untouched), mirroring `/new` Phase 0 step 1b: it finds the batch epic's `migration_plan` (`required: true`), verifies artifacts on disk, assembles apply modalities (epic `apply_modalities` + `.baldart/overlays/new.md` `## Migration modalities` + project memory + built-in `Già applicata`/`Abort`), asks ONE pre-launch question, executes the chosen modality in `$MAIN`, and passes a `migration` manifest into the workflow. **MINOR** (additive capability on the EXPERIMENTAL `new2` surface; the gate is a silent no-op when no migration is declared and degrades to today's end-of-batch deferral when artifacts are missing — no regression; the `migration` arg is the skill→workflow contract only, NOT a `baldart.config.yml` key, so the schema-change propagation rule does not apply).

### Added

- **`framework/.claude/skills/new2/SKILL.md` Step 3.5 — Migration Gate.** BLOCKING only when a migration is declared; otherwise a silent no-op. Emits the `migration = { status: 'none'|'applied'|'skipped'|'degraded', modality?, summary?, artifacts?, affects_cards? }` manifest passed to the workflow (Step 4 args block). The final-record step also logs `migration_gate: <status>` so the A/B comparison stays honest about when a migration was front-loaded (a pre-launch interaction, NOT a mid-batch question — the zero-ask-during-batch invariant holds).
- **`framework/.claude/workflows/new2.js` — migration manifest consumption.** Reads `a.migration` (args contract documented), and when `status === 'applied'` injects an EXCEPTION to the F-016 owner-gated rule in the preflight (the schema-apply / db-push AC of the affected cards must NOT be classified as a policy-deferred AC — it is satisfied) and a `MIGRATION LIVE` note into each affected card's brief (run REAL validation against the live schema, do not defer or build against an absent schema). A migration NOT covered by the manifest still follows the existing end-of-batch owner-gated deferral.

## [4.27.2] - 2026-06-11

**`new2` resolve: kill the second self-judging adversarial pass for `security` too.** v4.27.1 removed the wasteful adversarial doc review (doc-reviewer judging doc-reviewer). An audit of every fixer→judge pair in the resolution pass found `security` is the exact same structural case: the fixer is `security-reviewer` (protected domain) and so is the judge, so the mandatory cross-check was `security-reviewer` judging `security-reviewer` — no cross-model diversity, the same waste. The skip is now driven by a structural guard, `selfJudges = (fixerAgent === judgeAgent)`, instead of a per-domain allowlist, so it covers doc + security today and can't silently regress if the fixer/judge map changes. Every domain where fixer ≠ judge (ui/perf/test/code/migration) keeps the mandatory adversarial cross-check unchanged. **PATCH** (cost/latency optimization on the EXPERIMENTAL `new2` surface; no config key, no change to `/new`).

### Changed

- **`framework/.claude/workflows/new2-resolve.js`** — introduced `selfJudges`; both `judgeVerify()` and the terminal-judge ratification branch short-circuit when `selfJudges` is true (was `domain === 'doc'`). `meta.description` updated to describe the structural fixer===judge rule.

## [4.27.1] - 2026-06-11

**`new2` resolve: never run an adversarial doc review.** When a resolution pass fixed a `doc`-domain finding, the fixer was `doc-reviewer` (which applies *and* self-verifies the fix in one pass) — but `judgeVerify()` then spawned `doc-reviewer` *again* as the mandatory adversarial judge, plus the terminal-judge ratification used it a third time. That is `doc-reviewer` judging `doc-reviewer`: zero cross-model diversity, pure waste of tokens and time. The doc domain now trusts the single reviewer-writer pass — the adversarial judge and the terminal-judge ratification are both skipped for `doc` only (every other domain keeps the mandatory cross-check unchanged, since their fixer and judge are genuinely independent specialists). **PATCH** (cost/latency optimization on the EXPERIMENTAL `new2` surface; no config key, no change to `/new`, no behavior change for code/ui/security/perf/test resolutions).

### Changed

- **`framework/.claude/workflows/new2-resolve.js`** — `judgeVerify()` short-circuits for `domain === 'doc'` (accepts the first verified attempt without re-spawning), and the terminal-judge branch ratifies a doc terminal verdict directly instead of re-running `doc-reviewer`. `meta.description` updated to document the doc exception.

## [4.27.0] - 2026-06-11

**`/prd`: the Obsidian back-reference is now two-phase — the spec note is linked the moment the PRD starts, not only at merge.** The end-of-run snippet (v4.22.0) is gated on a successful merge, so a PRD that is paused or abandoned mid-flow (e.g. parked at the UI-design step, worktree never merged) left its origin note empty — even though detection already happened at Step 1. Observed in the wild: a `realtime-order-collaboration` PRD stuck at `ui-design` never wrote back, while a completed sibling PRD did. **MINOR** (additive capability on `/prd`; reuses the existing slug-keyed markers + the snippet `status:` lifecycle field; no `baldart.config.yml` key, so the schema-change propagation rule does not apply).

### Added

- **`framework/.claude/skills/prd/references/obsidian-backref.md`** — new shared SSOT for the back-reference: the marker contract (`<!-- BALDART:prd:<slug> START/END -->`), the idempotent read-replace-or-append write procedure, the skip conditions (no note / `path: none` / unwritable → log, never fatal), and the two write modes. Both call sites cite it, so the snippet shape can't drift between them.
- **Provisional write at detection (discovery-phase Step 1 point 2).** Right after the Obsidian note is detected, `/prd` writes a lightweight `status: drafting` block (slug + deterministic branch `prd/<slug>` + planned PRD path + trunk) into the note and sets state `## Obsidian Spec Note / status: provisional-written`. The note reflects "a PRD was started here" immediately.

### Changed

- **Final write upgrades the provisional block in place (validation-phase Step 7.6).** Same slug-keyed markers, so the post-merge full payload (epic + card IDs, merge SHA/PR, `resources:`, `status: planned`) replaces the `drafting` block — no duplicates. Step 7.6 is now a lean delegation to `obsidian-backref.md § Mode: FINAL` instead of an inlined copy of the snippet.
- **Snippet `status:` enum tightened to `drafting | planned | done`** (was `planned | in-progress | done`) — a clean ladder: PRD in stesura → mergiato/pianificato → in produzione.
- **State `## Obsidian Spec Note / status` enum** gains `provisional-written`: `none → detected → provisional-written → back-reference-written`.
- **Coupling fix:** `prd-writing-phase.md` Step 4 point 2 now adds the PRD Canonical Sources row when the note status is anything other than `none` (was hard-coded to `== detected`), so the provisional-write status change doesn't suppress the traceability row.
- **`SKILL.md` HARD RULE 19** rewritten for the two-phase lifecycle; flow-table gains a `1.2` row and annotates `7.6` as the FINAL upgrade.

## [4.26.1] - 2026-06-11

**`new2`: specialization integrity — a full audit of every agent spawn; no role mixing anywhere.** User principle: code is written ONLY by `coder`, UI only by `ui-expert`, each agent does one thing. The audit of all 16 spawn sites found two genuine violations and three under-specified plumbing roles. **PATCH** (role-integrity fixes on the EXPERIMENTAL `new2` surface; no config key, no change to `/new`).

### Fixed

- **`new2.js` E2 — the ops pre-flight agent no longer repairs the baseline.** "baseline FAILS → fix once" had a general-purpose git agent editing source code. Now: the pre-flight returns `baseline:'fail'` + an actionable log (explicit role boundary: it never edits source/doc files), and the WORKFLOW spawns the **coder** specialist for one bounded repair attempt — verified by deterministically re-running the baseline gates (no claim to trust), `E2-baseline: FIXED-BY-CODER` ledger row; still failing → batch-fatal as before. Zero extra spawns in the happy path.
- **`new2.js` — the Codex driver never reviews.** Its runtime fallback was "perform the review yourself with the code-reviewer lens" — a general-purpose agent doing code review. Now the driver returns `note:'codex-unavailable'` with empty findings, and the workflow spawns the REAL `code-reviewer` (same `stdReview` call as the matrix), ledgered as `review-codex: FALLBACK`.
- **`new2-resolve.js` — judge map completed per domain.** `doc` fixes are now judged by **doc-reviewer** (code-reviewer judging prose was cross-domain) and `test` fixes by **qa-sentinel**; `security`/`migration` → security-reviewer and `perf` → api-perf-cost-auditor unchanged; `ui`/`code` stay with code-reviewer (the judge verifies a CODE change — its charter, incl. DS rule 8 for UI). Fixer and judge of the same type remain two independent adversarial instances.

### Changed

- **`new2.js` — explicit ROLE BOUNDARY lines on every plumbing agent** (pre-flight, commit, merge, production-readiness): they never edit source/doc files; commit/merge touch ONLY card YAML status fields + registry rows; a content-level merge conflict is leave+report, never hand-resolved; production-readiness executes stack-matched commands but reports (never edits) code/config changes. The full writer map is now: code/perf/migration/test fixes → `coder`; UI → `ui-expert`; security fixes → `security-reviewer`; docs → `doc-reviewer`; backlog YAML → `prd-card-writer`; bookkeeping (status/registry) → mechanical commit/merge agents.

## [4.26.0] - 2026-06-11

**`new2`: Phase 1 decomposed into specialist agents — the owner implements, it no longer explores.** The old per-card pipeline gave the owner agent one mega-prompt ("you ARE claim + architect + plan-auditor + owner"), so the owner absorbed the whole codebase exploration into its own context and reached the actual coding with a degraded window. Phase 1 now runs as dedicated specialists with file handoff (`/tmp`), per the "ognuno fa una cosa" principle. Verified premise: nested subagent spawning does NOT exist in Claude Code (official docs: "Subagents cannot spawn other subagents" + 2 empirical probes), so the decomposition lives at the WORKFLOW level — the JS is the orchestrator, exactly what dynamic workflows are for. **MINOR** (pipeline capability on the EXPERIMENTAL `new2` surface only; no config key, no change to `/new`).

### Changed

- **`framework/.claude/workflows/new2.js` — per-card `codebase-architect` specialist (B7).** Context retrieval is now a dedicated `codebase-architect` spawn that writes the COMPLETE untruncated baseline to `/tmp/arch-baseline-<CARD>.md` (the owner and reviewers Read the file; the structured return stays minimal). Bonus over the old pre-flight snapshot: the per-card architect sees prior in-batch commits — card N's baseline reflects what cards 1..N-1 changed. The pre-flight no longer writes baselines (generalist relieved of specialist work; `archBaselinePaths` removed from its contract). **Fail-safe**: architect crash/not-ok degrades to the old inline behavior (the owner explores itself) — never blocks the card; transient outage re-queues it.
- **`new2.js` — `plan-auditor` demoted to a deterministic DRIFT gate.** `/prd` already validates every card at creation time, so an unconditional re-audit per card was duplicate work. plan-auditor now runs ONLY when execution-time drift is plausible: (a) a prior in-batch commit touched the card's declared surface (`filesLikelyTouched` ∩ prior `filesChanged`, computed in JS), (b) the architect found declared paths missing (stale card — factual `missingPaths` check, no judgement), or (c) `review_profile: deep` (Rule C escalation). Its corrections amend the implementation BRIEFING (never the backlog YAML). A fresh single-card batch with no drift evidence skips it at zero cost. Crash → non-blocking skip, ledgered.
- **`new2.js` — epic guard moved to JS (zero spawns for trackers).** The pre-flight returns `cardGraph[].isEpic` (implement.md §6b rule); an epic card now short-circuits in JS before ANY spawn — the old path burned a full owner-agent spawn just to learn the card was a tracker. The impl agent's own epic flag stays as backstop.
- **`new2.js` telemetry — `per_card[].phase1`** records `{architect: done|inline-fallback, audit: pass|fixes|skipped-no-drift|skipped-error|skipped-no-baseline}` plus `phase1-architect`/`phase1-audit` ledger rows, so the A/B accounting can price the decomposition.

## [4.25.0] - 2026-06-11

**`new2`: deterministic per-card review matrix — `balanced` no longer spawns every specialist on every card.** The old fan-out ran code-reviewer + doc-reviewer + qa-sentinel + api-perf-cost-auditor (+ security-reviewer) unconditionally at `balanced`: an api-perf pass on a doc-only card, a doc pass on a pure-code card — 1–3 wasted spawns per card. Each specialist now runs IFF its domain is evidenced by the card's actual surface (`scopeFiles ∪ MAY-EDIT`), computed deterministically in JS (no agent judgement) and audited via a `review-matrix` ledger row per card. **MINOR** (review-behavior capability on the EXPERIMENTAL `new2` surface only; no config key, no change to `/new` — its interactive profiles are untouched; the schema-change propagation rule does not apply).

### Changed

- **`framework/.claude/workflows/new2.js` — relevance-gated `balanced` fan-out.** `touchesCode` → code-reviewer + qa-sentinel; `touchesDocs` (`.md`/`.mdx` or under `paths.docs_dir|references_dir|wiki_dir|prd_dir|design_system`) → doc-reviewer (on a doc-ONLY card it IS the core finder — code-reviewer and qa-sentinel are skipped); `touchesApiData` (superset of the final review's `hasApiDataFiles` regex) → api-perf-cost-auditor; the v4.24.2 `securityRelevant` gate → security-reviewer. Surface = `scopeFiles ∪ MAY-EDIT`, so a DoD-mandated doc the implementer FORGOT to edit still triggers doc-reviewer (it's in MAY-EDIT). **Fail-safes**: `deep` keeps the unconditional full fan-out (Rule C escalation is respected); an empty surface (no evidence) falls back to the full fan-out — absence of evidence is not evidence of absence. `light`/`skip` unchanged.
- **`new2.js` Phase Final — the F-041 `singleCard`/slim skip is now coverage-gated.** The final review's doc pass is skipped for a single-card batch ONLY when doc-reviewer ACTUALLY ran per-card (`reviewersRun`); otherwise it runs as the missing-doc-update safety net — so gating a per-card specialist off never leaves its domain unreviewed at batch level. This also closes a **latent F-041 hole**: under `light`, doc-reviewer never ran per-card, yet the slim skip still suppressed the final doc pass for single-card batches.
- **`new2.js` telemetry — `per_card[].reviewers`** records the gated matrix per card, so the A/B spawn accounting can measure exactly what the gating saves.

## [4.24.3] - 2026-06-10

**`new2`: two follow-ups from the v4.24.2 logic review — epic closure no longer misses deferred cards, and the owner-gated classification reaches the final review.** Both are "parallel location" gaps of earlier fixes (the v4.17.2 meta-lesson): v4.22.1's epic-closure and v4.24.1's owner-gated deferral each missed one site. **PATCH** (bug-fix to the EXPERIMENTAL `new2` surface only; no config key, no change to `/new`).

### Fixed

- **`new2/SKILL.md` (Step 5.2) — epic-closure re-check after deferred-DONE.** The merge agent's epic-closure (Phase 6b step 5e) runs BEFORE the skill marks deferred cards DONE post-run — so an epic whose last open child was a deferred card stayed TODO forever (nothing re-checked it). The skill now re-runs the all-children-DONE check for the parents of the cards it just marked DONE, closing the epic in the same reconciliation commit.
- **`new2.js` (Phase Final) — owner-gated deferrals no longer block the merge at final review.** The `merge-blocker`/`qa-fail` resolves checked only `.status`: a finding whose sole remedy is an external/infra step (e.g. the same pending remote `db:push` a reviewer re-raises batch-wide) set `mergeBlocked` and stranded a complete batch. The `deferralClass` check from the per-card loops (v4.24.1/v4.24.2) now applies here too: `owner-gated`/`not-a-code-defect` → follow-up tracked + `DEFERRED-OWNER-GATED` ledger row, merge proceeds; a genuine unresolved code defect still blocks.

## [4.24.2] - 2026-06-10

**`new2`: holistic logic review — deterministic owner_agent routing, a merge gate that no longer strands the batch, `deferralClass` end-to-end, and −N agent spawns per run.** A full logic review of `new2.js`/`new2-resolve.js`/`SKILL.md` against `/new`'s reference modules found four correctness defects and five sources of wasted spawns. The headline routing bug: the pre-flight's `ownerAgent` was passed RAW as `agentType` — the G25 "unknown→coder" rule lived only in the prompt, so any freeform value (`claude`, `backend`, a typo) was a PERMANENT spawn error → card `failed` → (combined with the merge-gate bug) the whole batch unmerged. **PATCH** (bug-fix/hardening of the EXPERIMENTAL `new2` surface only; **no `baldart.config.yml` key**, **no change to `/new`** — the schema-change propagation rule does not apply).

### Fixed

- **`new2.js` (A1) — deterministic owner_agent clamp in JS, not prompt.** After the pre-flight, `cardGraph[].ownerAgent` is clamped through the exact `/new` router table (`implement.md` §6b): `coder`/`ui-expert` pass through; `plan`/`visual-designer`/`motion-expert`/unknown/missing degrade to `coder`, each with a `[ROUTER]` ledger row (audit trail). The RAW value is kept for the security-relevance heuristic. Fixes the "wrong agent for the card" failure class at the code level.
- **`new2.js` (A2) — the merge integrity gate matches its own comment.** `followup`/`blocked` cards (rolled back + tracked in the offline-safe residual ledger) no longer count as "incomplete": one failed card used to strand every committed card in an orphaned worktree with no resume path (the run wasn't `degraded`, so the skill never resumed). `failed` (crash) counts as complete ONLY when its cleanup verified the worktree clean.
- **`new2.js` + `new2/SKILL.md` (A3) — `deferralClass` end-to-end (v4.24.1 covered only the review-blocks loop).** The ac-unmet loop now records WHY each deferral happened; `residuals[]` carry `deferralClass` and committed cards carry `deferredClasses[]`. The skill marks a deferred card DONE post-run ONLY when every class is `owner-gated`/`not-a-code-defect`/`policy-deferred-ac`; an `unresolved` class (an AC the workflow tried and failed to implement) keeps the card IN_PROGRESS and is surfaced explicitly — a genuinely unmet DoD can no longer be auto-DONE'd with a follow-up as a fig leaf (F-029).
- **`new2.js` (A4) — crashed cards clean up after themselves.** `crashResult` now runs the rollback (whole-worktree scope — safe: all committed work lives in HEAD), so a crashed card's dirty files no longer poison the next card's E4 reconcile with a misleading `out-of-ownership` revert.
- **`new2.js` (A5) — `${paths.backlog_dir}` instead of a hardcoded `backlog/*.yml`** in the epic-closure merge prompt (consumer portability).

### Changed

- **`new2.js` (B1) — per-card idempotency probe removed (−N Haiku spawns/run).** `cardGraph[].alreadyCommitted` is now probed once by the pre-flight (git-authoritative: commit in `trunk..HEAD` + validations green + no open follow-up); on a fresh worktree it costs nothing, and resume stays covered by the journal cache.
- **`new2.js` (B2) — `security-reviewer` fan-out only when the card's files intersect `high_risk_modules`** (the bare `highRisk.length` fired it on every card of every project that configures the list); the owner-agent and brief-token triggers are unchanged.
- **`new2.js` (B3) — review blocks and scope-expansion findings batched per kind+domain → ONE `resolve()` per group** via the existing `findings[]` contract (the same F-007 batching the final review already used). One-by-one routing was the dominant per-card cost driver (each resolve = fixer + judge + possible Tier-2 fan-out).
- **`new2.js` (B4) — `executionMode`/`groups` removed from the pre-flight contract**: the DAG scheduler is strictly sequential (single worktree) and nothing ever read them; telemetry now reports the honest constant. Real parallelism is a future release, after A/B data.
- **`new2.js` (B5/C) — dead code removed**: the never-populated `lessons` mechanism (every cardBrief promised batch lessons that were always `(none)`), the unreachable `resolve()` `'fatal'` branch (`new2-resolve` never returns it — v4.17.2 G1 rule), the always-empty per-card `telemetry` stub (per_card now reports `deferred`/`deferredClasses`/gate count), and the now-unused `batchFatal` flag.

## [4.24.1] - 2026-06-10

**`new2`: an owner-gated gate no longer destroys a completed card (silent work-loss → commit + defer).** A real `/new2` run on a schema-change card produced **zero output** despite 52 min of work: the card's only obstacle was an *owner-gated* step (`db:check-sync` needs an approved remote `db:push`). That step was correctly `policy-deferred` up front — but the **same** condition was *also* re-raised by a reviewer as a fresh `MIGRATION_NOT_DEPLOYED` blocker, which (via the review-block branch's `s !== 'resolved' → cardBlocked`) triggered `rollbackCard`'s `git clean -fd` and **erased the completed migration**. Compounding it: the `E4-file-diff` gate logged `AUTO-REVERTED` while reverting *nothing* (leaving the card's DoD-mandated ADR/ER-doc edits orphaned in the worktree — its MAY-EDIT map was narrower than the DoD), and the residual follow-up was written *inside the worktree* and marked `materialized:true` without disk proof, so it vanished when the batch didn't merge. Root cause confirmed via the gate ledger + on-disk state + two rounds of adversarial review (the obvious "E4 reverted it" diagnosis was **wrong** — E4 was a no-op; `rollbackCard` was the eraser). **PATCH** (bug-fix to the EXPERIMENTAL `new2` skill + its workflows; **no `baldart.config.yml` key** and **no change to shared `/new` prose** — the DONE-deferral is handled entirely inside `new2`'s own merge prompt + skill, so `/new` interactive is provably unaffected; the schema-change propagation rule does not apply).

### Fixed

- **`framework/.claude/workflows/new2.js` + `new2-resolve.js` — classification-based card-block (the primary fix, F-040).** `resolve()` now propagates a structured `deferralClass` from `new2-resolve`'s terminal-judge. A review-block that resolves to an **owner-gated / not-a-code-defect** deferral no longer sets `cardBlocked` → the card's *complete* code is committed (it is **not** rolled back); only a genuine unresolved **code** defect (or `out-of-ownership`/`baseline`/`outage`) still blocks + rolls back. A deliberately-broken migration (`db:reset` failing) is still classified code-defect → blocks, so this is not an over-match escape hatch.
- **`new2.js` — `E4-file-diff` is honest.** The old `AUTO-REVERTED` log reverted nothing. The owner agent now reconciles out-of-ownership edits itself (per `implement.md §11b`) and reports `revertedOutOfOwnership`; the gate logs `REVERTED` / `FLAGGED` to match reality, and an unresolved violation becomes a tracked residual (never silent, never orphaned).
- **`new2.js` pre-flight — ownership map ⊇ DoD (root cause of the E4 false positive).** A card's MAY-EDIT now = `files_likely_touched` ∪ paths **named explicitly** in its `acceptance_criteria`/`definition_of_done` (the ADR/data-model/ER doc a schema-change must touch), so editing a DoD-mandated doc is no longer a file-diff violation.
- **`new2.js` + `new2/SKILL.md` — DONE deferred to the skill, gated on the follow-up existing on disk (closes the F-029 false-DONE the review surfaced).** A card carrying an open owner-gated/policy-deferred AC commits its code but stays **NON-DONE**; the merge agent is told to leave those cards non-DONE; the **skill** marks them DONE post-run **only after** verifying the deferral's follow-up exists on disk in the main repo (fail-loud otherwise — never DONE with a dropped requirement).
- **`new2-resolve.js` + `new2/SKILL.md` — follow-ups are reliable.** Follow-up materialisation is best-effort inside the workflow (it rides the merge if the batch merges); the **skill is the SSOT**, verifying every residual against the **main-repo** disk and creating any missing follow-up there — so a non-merged batch never loses one. `materialized` is now advisory only.
- **`new2.js` (Fix G) — AC-deferral dedup is text-drift-proof.** The policy-deferred-AC key is now scoped to the AC *number* (`acSig`), so a deferred AC is no longer re-routed to `resolve()` a second time when the pre-flight and implement agents word the AC text slightly differently.
- **`new2/SKILL.md` — telemetry reconciled against disk.** Before recording, the skill verifies each `committed` card actually has a commit on trunk and never presents progress the disk does not show; adds `cards_deferred_done_pending` so the A/B record distinguishes "code landed" from "DONE".

## [4.24.0] - 2026-06-10

**Atomic backlog-ID allocator — no FEAT/BUG collisions across parallel worktrees.** When several `/prd` (or `/new`/`new2` follow-up) sessions run in parallel on sibling worktrees, each branched from the same trunk, the old `max(^id: FEAT-) + 1` scan made them all land on the **same next integer**: the other session's card was in flight on an unmerged sibling branch, invisible to both the local backlog and the trunk merge-base — so two epics both became `FEAT-0024` and conflicted at rebase/merge. The `git fetch` + merge-base scan only ever covered *already-merged* IDs, never in-flight ones. A new allocator anchors a lock + per-prefix high-water mark in `$MAIN/.worktrees/` (the shared coordination point every worktree already reaches via `git rev-parse --git-common-dir`, gitignored like `registry.json`), so a reservation is atomic across every worktree **on the same machine**. The high-water mark bumped under the lock is the correctness anchor; `max()` against the real backlog + sibling-worktree backlogs + reservations log + trunk merge-base makes it **self-healing**. **MINOR** (additive capability on the `worktree-manager` skill; opt-in — callers fall back to the inline merge-base scan + `[ID-RACE-RISK]` note when the script is absent, so older installs and cross-machine cloud agents are unaffected. **No `baldart.config.yml` key** — the allocator reuses `paths.backlog_dir` + the gitignored `.worktrees/` convention, so the schema-change propagation rule does not apply).

### Added

- **`framework/.claude/skills/worktree-manager/scripts/allocate-id.sh`** — prefix-parametric (`FEAT`/`BUG`/`UI`/`DOC`/`PERF`/…) atomic ID allocator. `reserve <PREFIX> <slug>` prints the next free integer zero-padded to 4 digits; `release <worktree-path>` prunes a finished worktree's reservations. Cross-process mutex via atomic `mkdir` (stale-stolen after 30s via the lock dir's own mtime — a directory, not a git ref, so it never touches the shared `refs/stash` that `git stash` in worktrees is forbidden over). The slow `git fetch` runs **before** the lock, keeping the critical section local-FS-only (sub-second). Reads `paths.backlog_dir` + `git.trunk_branch` from `baldart.config.yml` and resolves `$MAIN` at runtime — no hardcoded project facts (passes the framework-edit-gate). Same-machine scope; the trunk merge-base scan is the best-effort cross-machine guard.

### Changed

- **`framework/.claude/skills/worktree-manager/SKILL.md`** — new section **"ID Allocation"** documenting the allocator, the shared `.id-alloc.lock/` / `.id-hwm-<PREFIX>` / `id-reservations.jsonl` files, the same-machine scope + opt-in fallback contract, and the gap-tolerant monotonic-counter design. `release` is wired into the cleanup paths of `/mw` (step 7), `mw-docs` (step 6), and `/cw` (step 4).
- **`framework/.claude/agents/prd-card-writer.md`** — § "FEAT-XXXX numbering" + Pre-Generation Checklist item 1 now reserve the integer via the allocator (primary path), with the trunk fetch + merge-base scan + `[ID-RACE-RISK]` note as the documented fallback when the script is absent. Applies to any prefix the writer mints.
- **`framework/.claude/agents/prd.md`** — § 4.1 NAMING CONVENTIONS prefers the allocator (any prefix) over the plain backlog scan, with the same documented fallback.

## [4.23.0] - 2026-06-09

**Functional Traceability Gate — no orphan UI affordances from mockups.** Handed-off mockups (Claude Design / Figma / the internal generators) routinely add interactive-looking chrome — buttons, icons, menu entries — that no requirement backed. Implementing a mockup **1:1 for fidelity** materialised that chrome into real markup + dead handlers + unused icon imports, and the cruft propagated forever. A new **unconditional** gate (independent of `features.has_design_system`) now requires every interactive or iconographic artifact to trace to a function before it becomes code: each artifact is classified **backed** (implement) / **decorative** (static-render, `aria-hidden`) / **orphan** (drop or escalate — never silent). The oracle is the PRD **UI Element Inventory `function_ref`** allowlist → card AC → orphan. Enforced at three boundaries: PRD mockup-intake (orphans become BLOCKING Discovery items), `ui-expert` implementation-time (BLOCKING pre-work), and `code-reviewer` / `/design-review` per-merge (`UI_ORPHAN_AFFORDANCE` HIGH finding). The `visual-fidelity-verifier` gains a `resolved_orphans` carve-out so a deliberately-dropped orphan is *expected-absent*, not a `component-missing` false positive. **MINOR** (additive discipline across existing UI surfaces; no new agent/skill/command and **no `baldart.config.yml` key** — the gate is registry-independent and reads its allowlist from the PRD inventory, so the schema-change propagation rule does not apply).

### Added

- **`framework/agents/design-system-protocol.md`** — new SSOT section **"Functional Traceability Gate (BLOCKING — unconditional)"**: the rule, the oracle hierarchy (UI Element Inventory `function_ref` → card AC → orphan), the backed/decorative/orphan classification + outcome matrix, the drop-or-escalate contract for orphans, where it runs (design / implementation / review), and the canonical `UI_ORPHAN_AFFORDANCE` finding code. The Gating section gains an explicit carve-out: this is the one part of the module **not** gated on `features.has_design_system`.

### Changed

- **`framework/.claude/agents/ui-expert.md`** — BLOCKING step `0b` (Functional Traceability Gate) added to "When Designing New Interfaces"; check `3b` added to "When Reviewing UI Code"; new red-flag category **Functional traceability (`UI_ORPHAN_AFFORDANCE`)**.
- **`framework/.claude/agents/code-reviewer.md`** — new section **"Functional Traceability (MANDATORY for UI work — unconditional)"**: live affordances tracing to nothing are `UI_ORPHAN_AFFORDANCE` HIGH findings that block merge; positioned as the per-merge net behind `ui-expert`'s implementation-time gate.
- **`framework/.claude/agents/visual-fidelity-verifier.md`** — input contract gains optional `resolved_orphans[]`; `component-missing` / `interactive-state-missing` carve-out so a dropped/decorative orphan is expected-absent rather than a fidelity finding.
- **`framework/.claude/skills/e2e-review/SKILL.md`** — populates `resolved_orphans` in the verifier payload from the card's UI Element Inventory (`function_ref: decorative` or gate-resolved drops).
- **`framework/.claude/skills/prd/references/discovery-phase.md`** — mockup-intake `mockup_analysis.screens[]` schema gains an `affordances[]` block (`label` / `traces_to` / `classification`); new **Affordance discipline (Functional Traceability)** making every `orphan` a BLOCKING Discovery item resolved via Scope Expansion / `/prd-add` / downgrade / drop.
- **`framework/.claude/skills/prd/references/ui-design-phase.md`** — Step 3d gates (Full + Hybrid) now require every interactive/iconographic inventory row to carry a `function_ref` with no unresolved `orphan`.
- **`framework/.claude/skills/prd/assets/claude-design-handoff-prompt.template.md`** — § 3 gains a functional-traceability constraint (only ship controls tied to a listed function; mark decorative explicitly); § 7 gains a per-element traceability map + orphan-to-decide list requirement.
- **`framework/.claude/skills/ui-design/references/inventory.md`** — UI Element Inventory table gains a **Function ref** column + the rule binding it to the gate.
- **`framework/.claude/commands/design-review.md`** — new step 6 (Functional Traceability check) emitting `UI_ORPHAN_AFFORDANCE` HIGH→major findings.

## [4.22.1] - 2026-06-09

**Epic-parent closure at end of `/new` + `/new2`.** When the last child of an epic is implemented, the epic/parent card now gets marked `DONE` automatically instead of being left stuck at `TODO`. The Phase 6b post-merge reconciliation gate only ever touched the **batch cards** (the children); the epic/parent card (`group.is_epic: true`) is excluded from the batch by the epic guard, so it never got reconciled — final-review only *reported* "all children done" without writing the epic's own status. **PATCH** (strengthens the existing Phase 6b reconciliation gate; no new agent/skill/command/config key — `/new2` inherits it because its merge agent follows `merge-cleanup.md` as SSOT).

### Changed

- **`framework/.claude/skills/new/references/merge-cleanup.md`** — Phase 6b gains **step 5e (Epic-parent closure)**: for each distinct `group.parent` of the batch (plus any epic card in the batch itself), if **every** child is `DONE` (checked across the whole backlog, not just the batch subset), force-update the epic card to `status: DONE` + `completed_date` + note, folded into the same reconciliation commit. Leaves the epic untouched if any child is still open. The step-6 tracker block now records `Epics closed: …`.
- **`framework/.claude/skills/new/references/final-review.md`** — the "all cards in the epic are DONE" Next-Steps message now states the epic card itself is marked `DONE` (by Phase 6b step 5e), with a note that a still-`TODO` epic under all-DONE children is a Phase 6b failure to flag.
- **`framework/.claude/workflows/new2.js`** — the merge agent's deterministic policies gain an explicit **EPIC CLOSURE** rule (all-children-DONE gate, distinct from the F-029 forced-DONE prohibition since the epic is a tracker), `MERGE_SCHEMA` gains `epicsClosed[]`, and a new `epic-closure` ledger entry records which epics were closed.

## [4.22.0] - 2026-06-09

**`/prd` Obsidian back-reference.** When a PRD is kicked off from an Obsidian note (the documental starting point), the same note now receives, at the end of the run, an idempotent snippet indexing how the feature was developed and where it landed — PRD path, epic + card IDs, branch, merge commit/PR, and the PRD's Canonical Sources — so the user can later query the note ("come è stata sviluppata la feature X, dov'è finita"). **MINOR** (additive capability on the `prd` skill; no new `baldart.config.yml` key — the note path is runtime-detected and state-persisted, so the schema-change propagation rule does not apply).

### Added

- **`framework/.claude/skills/prd/references/validation-phase.md`** — new **Step 7.6 — Obsidian back-reference** (runs after the merge, when the canonical paths/IDs/SHA are final). Writes a marker-delimited section (`<!-- BALDART:prd:<slug> START/END -->`) holding a fenced `yaml` payload + human-readable links back into the spec note. **Idempotent** (re-runs / `/prd-add` replace the block in place), **non-blocking** (no note, unresolved vault path, or unwritable file → skip with a log line). The note lives in the user's vault, outside the worktree and outside git — a plain filesystem write, never staged.

### Changed

- **`framework/.claude/skills/prd/references/discovery-phase.md`** — Step 1 point 2 Obsidian detection broadened beyond `[[wikilink]]` to also accept `obsidian://open?…` URIs and pasted `.md` paths, and now resolves + persists the note's **absolute `path:`** (vault root from `paths.obsidian_vault` / memory / one-time ask) so Step 7.6 can write back without re-deriving.
- **`framework/.claude/skills/prd/assets/state-template.md`** — new structured `## Obsidian Spec Note` section (`wikilink:` / `path:` / `status:`).
- **`framework/.claude/skills/prd/SKILL.md`** — new **HARD RULE 19** (Obsidian back-reference) + Step 7.6 row in the flow table.
- **`framework/.claude/skills/prd/references/prd-writing-phase.md`** — Canonical Sources row now reads the structured `wikilink:` field and cross-links Step 7.6.

## [4.21.1] - 2026-06-09

**Nightly reconciliation of the canonical API registries (errors.md / schemas.md).** Makes the nightly `doc-review` run explicitly keep `${paths.api_errors}` and `${paths.api_schemas}` in sync with the day's code changes, instead of relying on the generic doc↔code drift step. **PATCH** (strengthens existing routine/agent behaviour; no new agent/skill/command/config key).

### Changed

- **`framework/routines/doc-review.routine.yml`** — new step 3b: when `features.has_api_docs: true`, reconcile the canonical `errors.md` (stable error codes added/renamed/removed → apply or flag `DOC_GAP`) and `schemas.md` (Zod/TS shape changes → update or flag `SCHEMA_DRIFT`) against the last-24h commits. Skips silently when the paths are unset.
- **`framework/.claude/agents/doc-reviewer.md`** — new "Canonical API registries — reconciliation (nightly-owned)" section: the agent now explicitly OWNS the code→doc direction for the two registries (the `doc-writing-for-rag` skill owns the doc-authoring direction). Ties together the existing `SCHEMA_DRIFT` + error-code checks.

## [4.21.0] - 2026-06-09

**Code knowledge graph layer (Graphify) + Karpathy wiki rebind + nightly doc-alignment command.** Introduce [Graphify](https://github.com/safishamsi/graphify) as the new retrieval engine replacing the removed RAG (v4.20.0): a single language-agnostic tool (tree-sitter, 28 langs, local/offline, native Leiden communities; pip `graphifyy`, CLI `graphify` + `graphify-mcp`). Opt-in, runtime-detected, silent fallback to LSP→Grep→Git. **MINOR** (additive; new agent + 2 skills + 1 routine + 1 protocol module; new config key `features.has_code_graph` + `graph.*` propagated end-to-end per the schema-change rule).

### Added

- **`features.has_code_graph` + `graph:` block** (`framework/templates/baldart.config.template.yml`) — `graph.{output_dir, auto_rebuild_hook, register_mcp}` (scalars only). Propagated: configure autodetect + prompt + interactive install (`src/commands/configure.js`), update schema-drift detector **extended to diff nested top-level blocks** like `graph:` (`src/commands/update.js`), doctor diagnostic + actions (`src/commands/doctor.js`).
- **`src/utils/graphify-installer.js`** — thin wrapper over Graphify's NATIVE commands (no per-language adapter registry): `detect/install/verify/buildGraph/installGitHook/installAlwaysOn/registerMcp/isStale`. Composes `pipx install graphifyy`, `graphify update` (offline build), `graphify hook install` (auto-rebuild), `graphify <tool> install` (always-on), `graphify-mcp` (MCP).
- **`framework/agents/code-graph-protocol.md`** — new protocol module: when to prefer the graph (structural/relational) vs LSP (symbol) vs Grep (text) vs Git; CLI-first, MCP optional; budget; silent fallback. Wired into `code-search-protocol.md`, `index.md`, `codebase-architect`, and the exploration skills (`context-primer`, `bug`, `prd`, `new`, `simplify`) — all gated on `has_code_graph`.
- **Karpathy wiki rebind** — the LLM-wiki auto-learning loop (removed with RAG in v4.20.0) is re-activated on Graphify's native `GRAPH_REPORT.md` (god nodes / Leiden communities / suggested questions → synthesis candidates), offline. Updated `llm-wiki-methodology.md`, `wiki-curator.md`, `wiki-review.routine.yml`, `capture/SKILL.md`. Gated; dormant when `has_code_graph: false` (no regression).
- **Nightly doc-alignment command** — new agent `doc-graph-aligner` (rebuild graph → cross-check graph vs docs: uncovered core code, stale docs, registry drift → report + synthesis candidates; propose-only), routine `doc-graph-align.routine.yml` (+ `routines/index.yml`, `optional: true`), interactive skill `/graph-align`, and install skill `/graphify-bootstrap`.
- **`framework/docs/CODE-GRAPH-LAYER.md`** — operator guide (lifecycle, schema, native commands, invariants, CI behavior). MCP-INTEGRATION.md + PROJECT-CONFIGURATION.md updated.

### Notes

- **CI install is deliberate**: a Python tool is never installed silently in `--yes`/non-interactive runs — `configure` writes the flag, `baldart doctor` (`graph-install`/`graph-rebuild`/`graph-hook`) and `/graphify-bootstrap` backfill.
- **`graphify-out/`** added to `.gitignore` (generated artifact). **`package.json`** version is unrelated drift (tracked separately from `VERSION`).

## [4.20.0] - 2026-06-09

**Rimozione del motore RAG (strato A); wiki overlay preservato dormiente.** Il layer RAG (`doc-rag` MCP, `search_docs`, loop auto-learning con tap-point `wiki_log`) non era attivo in nessun consumer ed era un costo di istruzioni morto in ~20 file tra agenti, skill e moduli protocollo. Rimosso completamente; il retrieval del codice si riduce a **LSP → Grep → Git**. Il wiki overlay (loop di Karpathy: `/capture`, `wiki-curator`, `paths.wiki_dir`, `has_wiki_overlay`) resta nel framework ma **dormiente**, scollegato dal motore — engine-agnostic e riutilizzabile in futuro su un nuovo motore di retrieval. **MINOR** (nessun agente/skill/comando rimosso, nessun cambio di layout d'install; rimossa la sola chiave config `paths.wiki_log`, non usata da alcun consumer).

### Removed

- **Motore RAG (strato A) da tutti gli agenti, skill e moduli protocollo.** `mcp__doc-rag__search_docs` / `search_docs` / `mode:"hybrid"`, i verdetti `rag_telemetry` (`useful/weak/empty/fallback_degraded`), il tier "RAG hybrid" in `code-search-protocol.md`, l'Investigation Protocol "RAG-first" di `codebase-architect`, il tap-point di instrumentazione `wiki_log`, e la community detection GraphRAG (`search_synthesis`, `mode:"drift"`). File toccati: `framework/agents/{code-search-protocol,llm-wiki-methodology,index}.md`; `framework/.claude/agents/{codebase-architect,doc-reviewer,wiki-curator,code-reviewer,security-reviewer,senior-researcher,plan-auditor,qa-sentinel,coder,api-perf-cost-auditor,prd-card-writer}.md`; `framework/.claude/skills/{context-primer,bug,capture,prd}/` + `prd/references/discovery-phase.md`; `framework/.claude/skills/issue-review/SKILL.md`; `framework/.claude/commands/codexreview.md`.
- **Chiave config `paths.wiki_log`** da `framework/templates/baldart.config.template.yml` + autodetect in `src/commands/configure.js`. `paths.wiki_dir` e `features.has_wiki_overlay` restano (infrastruttura wiki viva).
- **Sezione "Documentation retrieval (RAG)"** da `framework/docs/MCP-INTEGRATION.md`; riferimenti doc-rag Required/optional aggiornati.

### Changed

- **Retrieval hierarchy → LSP → Grep → Git** (era RAG → LSP → Grep → Git) in `code-search-protocol.md` e nell'Investigation Protocol di `codebase-architect` (ora "Canonical router FIRST" via `ssot-registry.md`).
- **`wiki-curator` ridotto** a drift / anchor / frontmatter validation + synthesis-candidate da ADR/PRD (engine-agnostic); rimosse § "RAG Instrumentation" e § "Graph Community Synthesis".
- **Routine `wiki-review`** ripulita dallo scan dei verdetti RAG; resta safety-net strutturale. `routines/index.yml` e README riformulati ("maintenance" invece di "auto-learning loop").
- **`/capture`** scrive ancora la synthesis page; il logging passa da helper `wiki_log` a bullet markdown manuale in `${paths.wiki_dir}/log.md`.

### Backup

- **`docs/archive/rag-layer/`** — backup integrale recuperabile: `README.md` (cos'era / perché / come ripristinare), `snapshot/` (copie pre-modifica dei file toccati), `rebind-to-graphify.md` (come ricablare il loop di Karpathy su un code knowledge graph con community detection nativa). `doc-writing-for-rag` (skill di densità documentale) **non** rimossa — è engine-agnostic.

## [4.19.0] - 2026-06-08

**`/prd`: cross-check di completezza della discovery con Codex (shift-left).** Aggiunge un singolo passaggio cross-model **al termine della discovery**, prima del design — il punto a più alto ritorno per Codex in `/prd`. **MINOR** (additivo/opt-in: scatta solo se il companion Codex è risolto a runtime; nessuna chiave `baldart.config.yml` → schema-propagation N/A).

### Added

- **Discovery Completeness Cross-Check (Codex) — `framework/.claude/skills/prd/references/discovery-phase.md`.** Dopo che i gate di uscita esistenti passano (comprehension ≥99% OR user-proceed, mockup-gap risolti, research risolta) e PRIMA di flippare lo status a `ui-design`, un passaggio Codex (`--wait`, pattern di 6.6d) rilegge lo state file di discovery (dimensioni + risposte + Discovery Log) + il brief di kickoff e cerca i **buchi che la formula 99% non può vedere**: domande mai poste, ambiguità latenti, NFR/error-path/edge mancanti, touchpoint d'integrazione. Razionale: la formula misura la copertura delle dimensioni *note*; un modello diverso intercetta i blind spot del questionatore, e catturare un gap di requisiti qui (pre-design) costa molto meno che a Step 6.6d (post-PRD/post-card). Per ogni gap: o diventa la prossima domanda di discovery (one-per-turn) o uno skip esplicito nel Discovery Log (stessa disciplina dei mockup-gap). **Un solo passaggio** (niente loop). **Rispetta l'intent esplicito**: se l'utente ha forzato l'uscita ("proceed"), il check è advisory-only. **Fallback**: Codex assente → skip con log (gate additivo; nessun fallback same-model Claude — annullerebbe la diversità di modello che è l'intero valore del check). Menzione nella flow table di `SKILL.md`.

## [4.18.0] - 2026-06-08

**Codex come finder a profilo `light` (`/new` + `new2`) + engine Codex opt-in nelle routine cron.** Sposta il carico di review da Claude alla licenza Codex dove la review è read-only e genera finding, mantenendo la diversità di modello (Codex trova, l'agent Claude verifica). **MINOR** (additivo/opt-in con fallback; nessuna rottura install, **nessuna chiave `baldart.config.yml`** → schema-propagation N/A — `review_engine` è schema routine-payload).

### Changed

- **Profilo `light` → Codex è il finder per-card (inverte la scelta v3.38.0).** Prima `light` faceva `code-reviewer` (finder) + FP-gate e deferiva Codex al solo Final FULL gate. Ora `/codexreview` a `light` lancia **Codex come unico finder** + lo Step 3 FP-gate (`code-reviewer` + `codebase-architect` validano): la lettura full-diff esce da Claude, Claude conserva solo la validazione. `full` invariato (entrambi i finder). **Fallback obbligatorio**: Codex non disponibile → `light` degrada a `code-reviewer` finder (comportamento pre-v4.18.0), la review non salta mai. L'invariante BUG-0530 ne esce rafforzata (Codex per-card su ogni card non-triviale). File: `framework/.claude/commands/codexreview.md` (Step -0.5/Step 2/invariant/Method-log), `framework/.claude/skills/new/references/codex-gate.md` (razionale Phase 3.7).
- **Team-mode allineato: i `LIGHT_CARDS` non skippano più D.4b.** La partizione LIGHT/FULL esisteva per tenere Codex fuori da `light`; ora ogni card non-triviale gira `/codexreview` a D.4b (light → Codex-light, full → Codex-full). La D.2 group pass diventa **doc-reviewer only** (il code-reviewer di gruppo sui light cards è rimosso — la loro review si sposta a D.4b). Coverage-assertion e clausola "REVIEW DEPTH SCALES" aggiornate. File: `framework/.claude/skills/new/references/team-mode.md`, `framework/.claude/skills/new/SKILL.md`, `framework/.claude/agents/prd-card-writer.md` (Rule C — semantica Codex-`light`).
- **`new2.js`: mirror del Codex finder a `light`.** Il companion Codex è ora risolto in pre-flight per **ogni** batch (prima solo multi-card per il cross-card), con il path propagato in `sharedCtx`. A `light` il fan-out per-card usa un reviewer **`codex`** (un agent general-purpose che lancia `codex-companion.mjs --wait` sul diff) quando il companion è risolto, con fallback `code-reviewer`. Il FP-gate è preservato a valle dal judge avversariale di `new2-resolve` (F-015). Nuovo campo schema `codexScriptPath`; ledger `codex-companion`; telemetria `codex_resolved` ora probed per ogni batch.

### Added

- **`review_engine: codex` opt-in nelle routine di review (solo backend cron).** `code-review.routine.yml` accetta `review_engine: claude|codex` (default `claude`) + un `prompt_codex:` per la review olistica single-pass. L'adapter `cron` risolve il companion Codex e lo invoca (`--wait`), con **guard `CODEX_NOT_FOUND` → fallback a `claude --agent`**; il check `ANTHROPIC_API_KEY` è ammorbidito a warning sul ramo Codex (Codex usa `~/.codex`). I backend `github-actions` e `claude-code-cloud` **non** supportano Codex (no plugin/auth in CI, schema RemoteTrigger ignoto) → emettono l'artefatto Claude + un warning esplicito. Codex non spawna specialist → la variante Codex perde la profondità security-reviewer/api-perf (documentato; usa `claude` per quella). File: `framework/routines/code-review.routine.yml`, `src/utils/routine-adapters/{cron,github-actions,claude-code-cloud}.js`, README.

## [4.17.2] - 2026-06-08

**`new2` hardening: 7 lacune latenti chiuse da una revisione adversariale dell'intero workflow.** Dopo i fix di v4.17.1 (nati da run reali), una revisione olistica di `new2` (orchestratore + resolve + final-review + skill) ha isolato 7 lacune non ancora emerse nei run ma reali alla lettura del codice — validate di persona scartando i falsi positivi (l'integrity gate di merge è corretto; lo skill "prose-only" è by-design). **PATCH** (hardening dei workflow `new2`/`new2-resolve` + skill/doc; nessuna nuova capability/layout, **nessuna chiave `baldart.config.yml`** → schema-propagation N/A).

### Fixed

- **G1 — `new2-resolve.js`: rimossi gli handler self-healing orfani (`agent-crash`, `baseline-fail`).** `new2.js` non vi instradava MAI (un crash → `crashResult` marca `failed` + residual; un baseline fail → `batchFatal`). Erano codice morto irraggiungibile: `runCard` va in crash PRIMA della fase commit, quindi un eventuale `resolved` non potrebbe comunque committare/mergiare. Il path residual (lo skill materializza la follow-up) è la SSOT; commento esplicativo in `crashResult`.
- **G2 — `new2.js`: detection Codex G3 cross-card deterministica + skip mono-card.** Il check Codex del pre-flight usava ancora il **soft self-report** che v4.17.1 aveva corretto solo nella final review (stesso falso negativo, per batch multi-card). Ora il bullet G3 risolve PRIMA il glob del companion (esistenza decisa dal glob, non dal giudizio dell'agent), gira Codex in **background con poll**, e vieta un `CODEX_NOT_FOUND` prematuro; su batch **mono-card** il check è saltato del tutto (niente cross-card). Nuovo `codexResolved` nello schema pre-flight + ledger `G3-codex` + campo telemetria `codex_resolved` → un falso negativo è ora visibile.
- **G3 — `new2.js`: niente silent drop dei finding merge-artifact.** Il filtro F-035 scartava i finding "esiste già nel worktree" con solo un ledger SKIP; un regex lasco poteva far sparire un BLOCKER reale. Ora ogni finding scartato diventa anche un `residuals` (`merge-artifact-skipped`, `materialized:false`) → tracciato, nulla perso.
- **G4 — `new2/SKILL.md`: materializzazione follow-up offline-safe via `prd-card-writer`.** Il fallback dello skill scriveva uno stub minimale, incoerente col fix #2 (F-039, workflow → prd-card-writer). Ora lo Step 5.1 delega a `prd-card-writer` (card-template, Rule C, owner_agent, traceability), con stub minimale solo come ultima risorsa in outage totale.
- **G5 — `new2.js`: routing `security-reviewer` non solo da `scopeFiles`.** Una card di sicurezza con file non-matchanti saltava la security review quando `high_risk_modules` non è configurato. Ora il trigger considera anche l'`owner_agent` della card e il suo brief (titolo/requirements), oltre allo scope.
- **G6 — classificatore di dominio doc allineato.** `isDocDomain` (new2.js) e la branch `doc` di `normDomain` (new2-resolve.js) ora portano un commento di sincronizzazione reciproca (stessi token `doc|wiki|ssot|readme`) per prevenire drift.
- **G7 — sync documentazione.** `WORKFLOWS.md` (righe `new-final-review`/`new2`/`new2-resolve`) e `new2/SKILL.md` allineati al comportamento v4.17.1/v4.17.2 (doc routing, prd-card-writer followup, Codex deterministico, final review snella mono-card, rimozione handler orfani).

## [4.17.1] - 2026-06-08

**`new2` post-refactor: tre falle di routing/detection emerse al secondo run reale (card FEAT-0019-followup).** Dopo l'hardening v4.17.0, gli agent principali sono correttamente specializzati, ma il *resolve* e il *final review* mostravano tre errori di instradamento. **PATCH** (bugfix di orchestrazione nei workflow `new2-resolve`/`new-final-review`; nessuna nuova capability, nessun cambio layout, **nessuna chiave `baldart.config.yml`** → schema-propagation N/A).

> **Why.** Tutti e tre i bug condividono la stessa radice: una decisione che dovrebbe essere **deterministica** (quale agent per quale dominio / esiste Codex) era affidata o a una mappa incompleta o al self-report soft di un subagent. Il dominio freeform `documentation` non matchava la chiave `doc`, il companion Codex (installato e funzionante) veniva dichiarato assente per un timeout di una Bash-call sincrona, e le follow-up card nascevano da un Haiku general-purpose anziché dal loro owner.

### Fixed

- **`framework/.claude/workflows/new2-resolve.js` — routing per dominio robusto (F-038/F-024).** Le review agent emettono `domain` come **stringa libera** (`documentation`, `frontend`, …); la mappa fixer keyava solo su `doc`/`ui`, così `documentation` cadeva sul costoso `coder` (Opus) per un finding di pura documentazione. Aggiunto `normDomain()` che normalizza ai bucket di routing (`doc|ui|security|migration|perf|test|code`) **prima** di scegliere fixer e judge, e mappa fixer estesa (`doc→doc-reviewer`, `ui→ui-expert`, `security→security-reviewer`, resto→`coder`) — reviewer-owns-its-domain anche in fase di *apply*, non solo di *judge*.
- **`framework/.claude/workflows/new2-resolve.js` — follow-up card via `prd-card-writer` (F-039).** La materializzazione della follow-up usava `general-purpose` + Haiku scrivendo uno stub YAML minimale; ora delega all'agent **`prd-card-writer`** (owner del `card-template`, Rule C per `review_profile`, `owner_agent`, traceability) — le card di follow-up rispettano le stesse regole SSOT di quelle native.
- **`framework/.claude/workflows/new-final-review.js` — detection Codex deterministica + background (F-040).** Codex risultava "non disponibile" pur essendo installato e funzionante: il workflow delegava l'intera invocazione a un subagent generico che la eseguiva **in modo sincrono** dentro una Bash-call e auto-dichiarava `codexAvailable` — un run reale di Codex (minuti) superava il timeout della call → falso negativo → fallback a `code-reviewer`. Ora un **pre-flight deterministico** (un Haiku risolve SOLO il glob del companion e ritorna il path: la decisione di esistenza è tolta al review agent), poi Codex gira in **background con poll su file** e finestra 10 min (allineato a `references/final-review.md` F.3/F.4); `codexAvailable:false` è ammesso solo per `CODEX_NOT_FOUND` reale o file vuoto a fine timeout. Un **unico** fallback deterministico a `code-reviewer` copre sia "companion assente" sia "run non completato".
- **`new2.js` + `new-final-review.js` — final review snella per batch mono-card (F-041).** Un batch da 1 sola card eseguiva la batch final review **per intero** (`doc-reviewer` + `api-perf-cost-auditor` + `qa-sentinel`), **duplicando** la review per-card (Phase 3) che aveva già girato gli stessi agent sugli stessi file — e senza alcun conflitto cross-card da trovare (~300k token ripetuti su un run reale da 1,28M). Ora `new2.js` passa `singleCard: committed.length === 1`; con quel flag la final review tiene **solo** il passaggio cross-model Codex (valore unico: modello diverso → bug diversi) + i gate `qa-sentinel` (letti dall'integrity gate di merge), e salta i duplicati Claude. Batch multi-card invariati.
- **`new2.js` + `new2-resolve.js` — fan-out Tier-2 inutile su finding doc (F-038/F-033/F-036).** Un singolo merge-blocker doc-domain già **risolto al Tier-1 e confermato dal judge** (`best:1`) faceva comunque partire **3 `doc-reviewer` paralleli (~150k token)** — con rischio di follow-up spuria per un problema già chiuso. Tre cause composte, tutte corrette: **(A)** la resolve riceveva come MAY-EDIT i file di **codice** della card (`reviewScopeFiles`), così il fixer doc editava `docs/` "fuori ownership" — ora `domainMayEdit()` dà ai finding doc la loro territory reale (`docs_dir`/`references_dir`/`wiki_dir`/`prd_dir`); **(B)** la guardia anti-fabbricazione del judge usava `.every()` (rigetto se *anche un solo* file confermato cade fuori MAY-EDIT) — ma il judge lista naturalmente i file adiacenti (il codice che una doc descrive); ora `someInScope()` (`.some()`) rigetta solo un verdetto che non conferma **nessun** file di proprietà (phantom puro); **(C)** il fan-out a 3 angoli (`root-cause`/`call-site`/`reuse-utility`) è code-bug-shaped → per domini deterministici non-code (doc/test) ora un **singolo** retry mirato invece di 3.

## [4.17.0] - 2026-06-07

**`new2` hardening single-wave: 36 finding dal primo run reale risolti in un colpo.** Il primo run reale di `new2` (batch FEAT-0021, 9 card) sotto un outage 529 prolungato è costato **7.4M token / 100 agent / 5h08m producendo solo 6/9 card reali**, con 2 stati `DONE` falsi nel backlog, codice security RLS-bypass non revisionato finito su `develop` via un "safety commit", e 1 follow-up persa. Una raccolta diagnostica end-to-end (36 finding verificati su codice + telemetria reale) ha isolato cause **strutturali**; questo rilascio le risolve tutte in un'unica wave, riportando il costo atteso a ~1.5-2.5M e — prioritario — azzerando le violazioni di integrità/sicurezza sotto degradazione. **MINOR** (cambio sostanziale di capacità del workflow esperimentale `new2`; nessuna rimozione/cambio layout d'install; **nessuna nuova chiave `baldart.config.yml`** → schema-propagation N/A).

> **Why.** `new2` è un A/B di `/new` per misurare l'economia di contesto del workflow-hosting. Il run reale ha mostrato che, oltre al costo, l'autonomia zero-ask degradava in modo **non sicuro** sotto outage (DONE falsi, codice non revisionato su trunk, follow-up persa) — esattamente il recovery che `/new` (gated, main-loop) avrebbe fermato. Decisioni utente vincolanti: *Decision-auto, infra-deferred* (niente `BALDART_AUTONOMOUS=1`; azioni infra owner-gated sempre policy-deferred) e *agenti specializzati come `/new`* (no general-purpose dove conta la qualità).

### Changed
- **`framework/.claude/workflows/new2.js` — scheduler a DAG + resilienza + integrità.** Il for-loop sequenziale e il ramo team-mode sono sostituiti da uno **scheduler dependency-gated** (semantica single-worktree): una card parte solo quando TUTTE le dipendenze in-batch sono `committed`; i dipendenti di una dep fallita vanno `blocked` (niente resolve doomed, F-017/F-021); fallimento transitorio → `pending` + re-invocazione (cap); outage sostenuto → **degrada pulito** (`degraded:true` + worktree clean + return-as-ledger, F-012/F-019/F-020). Implement usa l'**owner_agent** della card; review è un **fan-out specializzato** sfrondato dal `review_profile` (F-024/F-025); commit è **Haiku** con `git status` reconcile, mai `git add -A` (F-023). Skip-completed idempotente su commit-sha+receipt+no-open-followup, NON sul flag `DONE` (F-026). Merge **integrity-gated**: mai force-DONE (F-029), mai commit di codice non revisionato (F-030), mai merge di batch incompleto/degradato (F-031). Telemetria veritiera: `total_tokens` via `budget.spent()`, `agent_count`, `degraded`, `cards_real_done` vs `cards_force_done`=0 (F-032). Parse-guard `args` (F-001).
- **`framework/.claude/workflows/new2-resolve.js` — short-circuit + anti-fabbricazione + batching.** Verdetto **terminale** che salta il multi-attempt quando il problema è impossibile-by-definition (`out-of-ownership` verificato in JS contro MAY-EDIT; altri reason ratificati dal judge → niente escape-hatch, F-008). **Judge avversariale OBBLIGATORIO** su ogni claim `verified`, con **cross-check JS** dei file confermati indipendentemente dal judge (defeats fabricated success, beccata 2× nel run reale — F-015/F-033). Lista **`findings` batchata** (1 resolve per fix-area, F-007/F-035/F-037). Fixer/judge **per dominio** (doc→doc-reviewer, ui→ui-expert, security→security-reviewer judge, perf→api-perf-cost-auditor judge — F-024). `outOfScopeFindings` non più persi (F-022). Follow-up su Haiku, **offline-safe** (deferita al skill se nessun agente può scrivere, F-020). Parse-guard `args` (F-004).
- **`framework/.claude/skills/new2/SKILL.md` — reconcile + resume + telemetria reale.** Step 4 passa `ts` (il workflow non ha clock). Step 5: materializza le follow-up dei `residuals` non ancora su disco (offline-safe — il skill ha il filesystem → chiude "nothing dropped" sotto outage, F-020), **rilancia** con `resumeFromRunId` se `degraded` (idempotente via skip-completed), registra telemetria veritiera (`wall_clock_s`, `followups_on_disk` contando i file — F-032).
- **`framework/.claude/agents/prd-card-writer.md` — Rule A.1 (AC↔ownership consistency).** Fix a monte del difetto ricorrente "AC fuori dal MAY-EDIT" (F-016): ogni AC dev'essere soddisfabile editando i file della propria card.
- **`framework/docs/WORKFLOWS.md`** — righe `new2`/`new2-resolve` aggiornate al comportamento v4.17.0.

## [4.16.2] - 2026-06-06

**`baldart update` non si blocca più all'infinito su un edit `.framework/` ormai assorbito a monte.** Un edit locale a un file framework (es. `agents/ui-expert.md`) committato con un subject che *non* contiene la parola "baldart" sfuggiva alla classificazione `chore-wrapper` (rumore) e veniva marcato `custom-overlay-able`. Quel commit resta per sempre ancestor di `HEAD` ma mai di `FETCH_HEAD`, quindi `git log FETCH_HEAD..HEAD -- .framework` lo ripesca a **ogni rilascio successivo** — e il flusso seamless si fermava sul gate `divergence-capture-incomplete` perché l'overlay (additivo) non riproduceva byte-a-byte il file `.framework`, anche quando il contenuto era già identico all'upstream e la personalizzazione viveva correttamente nell'overlay. Risultato: lo stesso commit "fantasma" rifaceva muro a ogni `update`, senza via d'uscita via overlay (riconciliare un overlay additivo non può mai far passare il round-trip stretto). **PATCH** (bugfix del detector di divergenza + reset-safety degli overlay additivi; nessun cambio di capability, nessuna chiave `baldart.config.yml`).

> **Why.** Il check di fedeltà capture+verify pretendeva `merge(base_upstream, overlay) == file_.framework`. Per un overlay `extend` (additivo) questo è **strutturalmente falso** quando `.framework` è pristine-upstream: il merge aggiunge sezioni, il file no → "stale" eterno. E il classifier non distingueva un edit *vivo* (da preservare) da uno *assorbito* (contenuto già a monte → un re-sync non perde nulla). Due correzioni complementari chiudono la classe di falso positivo per tutti i consumer.

### Fixed

- **`src/utils/git.js` — categoria `absorbed` nel classifier di divergenza.** `classifyDivergence` ora demota a rumore (`absorbed`) ogni commit `custom-*` i cui file `.framework/` toccati sono **byte-identici all'upstream (`FETCH_HEAD`)** — confronto di blob-SHA, prefisso `.framework/` strippato per mappare il path upstream; conservativo (file fuori da `.framework/`, mancante a monte, o SHA diverso → resta drift reale). `aggregateDivergenceClass` include `absorbed` tra le `noiseCategories`, così un edit ormai assorbito non spinge più la classe verso `overlay-able`/`mixed`: la batch torna `all-noise` → auto-resolve pulito.
- **`src/utils/overlay-capture.js` — reset-safety per overlay additivi (`captureAndVerify`).** Quando esiste già un overlay agent/command, il file `.framework` viene riconosciuto come **reset-safe** (nessun edit inline da perdere) se è pristine: (a) byte-identico alla base upstream verso cui si farebbe reset, **oppure** (b) corrisponde alla base con cui l'overlay è stato autorato (`base_file_sha` in frontmatter) — pristine anche dopo che l'upstream è andato avanti. In quel caso è `captured` (non più `stale`), invece di pretendere il round-trip stretto che un overlay additivo non può mai soddisfare. Il caso *genuinamente* stale (edit inline reale che l'overlay non riproduce) resta correttamente un blocker.
- **Test.** `classify-divergence.test.js`: 3 fixture per `absorbed` (da solo / con merge plumbing → `all-noise`; con overlay-able genuino → `overlay-able`). `overlay-capture.test.js`: 2 fixture per la reset-safety (overlay additivo + `.framework` == base → `captured`; upstream spostato ma `.framework` pristine vs `base_file_sha` → `captured`).

## [4.16.1] - 2026-06-05

**Review di allineamento `new` ↔ `new2`: 9 fix di fedeltà sui workflow `new2`/`new2-resolve` emersi da una review adversariale a 3 reviewer paralleli.** La v4.16.0 ha introdotto `new2` come variante workflow-hosted di `/new`; una review di allineamento (gate-coverage · child-workflow wiring · control-flow) ha refutato l'implementazione su difetti reali — di semantica/allineamento, non di esecuzione (il wiring tra i 3 workflow era contract-fedele). **PATCH** (bugfix di allineamento del nuovo `new2`, nessun cambio di capability, nessuna chiave `baldart.config.yml`).

> **Why.** `new2` deve essere un A/B equo di `/new`: ogni gate va mappato con policy *fedele* alla semantica di `/new` (meno l'AskUserQuestion), e tutti i path/child-workflow vanno cablati. La review ha trovato punti dove la policy deterministica divergeva silenziosamente dalla SSOT o dove un segnale veniva letto ma non agito.

### Fixed

- **`framework/.claude/workflows/new2.js`**:
  - **Final review `NEEDS_MANUAL_CONFIRMATION` non più ignorato** — il filtro blocker includeva solo `VERIFIED` mentre il commento dichiarava "treat as VERIFIED": ora include anche `NEEDS_MANUAL_CONFIRMATION` (tentato il fix, mai scartato in silenzio → ripristina l'invariante zero-drop di `final-review.md` F.4).
  - **`status:'fatal'` di `new2-resolve` ora onorato** — `resolve()` setta `batchFatal` e sospende final review + merge, invece di degradarlo a `followup`.
  - **Phase 7 Production Readiness aggiunta** (post-merge, non-bloccante) — ripristina la full-parity dichiarata (auto-esegue i deploy stack-matched, riporta gli item manuali).
  - **Card crashata non più persa** — il catch del driver passa da `resolve('agent-crash')` (retry specialista, mai bypass su security/migration) e garantisce un follow-up tracciato (path `agent-crash` prima irraggiungibile).
  - **G3 cross-card / G5 depends-on rafforzati** — il prompt pre-flight ora *applica* le azioni cross-card (force-sequential/warning in executionMode/groups), non le logga soltanto; G5 garantisce l'esclusione **transitiva** dei dipendenti.
  - **Minori**: `fileDiffViolation` (E4) ora loggato nel gate ledger; `hasApiDataFiles` non più sempre-true (rimosso il disgiunto che rendeva irraggiungibile lo skip dell'api-perf-auditor); i resolve di `merge-blocker`/`qa-fail` passano `mayEditPaths` (confine ownership sui fix batch-wide).
- **`framework/.claude/workflows/new2-resolve.js`**: obbligo esplicito di **re-run E2E (Phase 2.6)** quando un fix tocca file UI (app/components/`*.css`/`*.scss`), allineato a `codex-gate.md` step 6.

## [4.16.0] - 2026-06-05

**Nuova skill sperimentale `new2` + i workflow `new2`/`new2-resolve`: l'INTERA batch di `/new` può girare come un singolo dynamic workflow autonomo in background, per fare A/B testing sul context economy.** `/new` è una skill che gira nel main loop, quindi l'output di ogni subagent (architect, coder, reviewer, Codex, QA, E2E) gonfia il contesto dell'orchestrator — il driver di costo che si sposta d'asse a ogni fix (v4.12→v4.15). I dynamic workflow spostano l'orchestrazione in un runtime **isolato**: i risultati intermedi vivono in variabili di script e nel main loop torna **solo il report finale**. `new2` è la versione workflow-hosted di `/new`, **parallela e additiva** (NON sostituisce `/new`), pensata per misurare il delta di `usage` reale a parità di semantica. **MINOR** (nuova skill + 2 workflow, opt-in, Claude-only; nessun breaking change, nessuna chiave `baldart.config.yml` — il gating è runtime via presenza del tool `Workflow`, come v4.14.0).

> **Il nodo: i workflow non possono chiedere input mid-run** ("For sign-off between stages, run each stage as its own workflow"). `/new` ha ~25 gate `AskUserQuestion` bloccanti. Invece di spezzare in stage interattivi, `new2` **rende deterministico ogni gate**: mappati tutti → zero `AskUserQuestion` → la batch intera è un solo workflow autonomo. Ogni gate collassa in due archetipi — **(A) auto-risolvi col default seamless** (stash/ff-pull/retry/revert) o **(B) fail → self-healing `new2-resolve`**. Le operazioni distruttive/outward (`reset --hard`, force-push, drop stash) non si auto-eseguono MAI → degradano a "lascia intatto + segnala". Decisioni utente recepite in design: **auto-merge completo** (G24, l'unica azione outward autonoma), **card in fail → si auto-risolve senza lasciare nulla indietro** (non skip, non halt), **E2E `error` di tooling → procedi con flag**, **full parity** (sequential + team-mode coder paralleli via `parallel()`).

> **`new2-resolve` — "nulla lasciato indietro".** Su decisione esplicita dell'utente, un gate bloccante non salta né ferma la batch: innesca un workflow di risoluzione self-healing con 8 `kind` su 3 assi — **fail/blocco** (`ac-unmet`/`blocker`/`qa-fail`/`e2e-blocked`/`merge-blocker`), **scope-expansion** (un finding legittimo che allarga lo scope oltre la AC: confine deterministico ancorato a file-ownership map + Domain-Override, niente soglie arbitrarie → integra-entro-confine-sicuro **o** follow-up), e **edge trasversali** (`agent-crash` re-instradato al giusto specialista mai bypassando security/migration; `baseline-fail` = unico terminale batch-fatale). Strategia a tier: fix mirato → multi-attempt giudicato (`parallel()` + judge avversariale) → re-verifica → **terminale = follow-up card auto-materializzata** (l'opzione sancita dalla Scope Closure Discipline, NON un deferral silenzioso). La decisione umana si sposta da ~25 interruzioni mid-run a una sola review post-run della coda follow-up.

> **A/B validity.** Per attribuire il delta SOLO al cambio di host (skill main-loop → workflow background) e non a drift semantico, `new2` riusa gli **stessi agenti** con gli **stessi briefing**: i prompt-agente del workflow **citano i moduli reference di `/new`** (`args.refModulesBase`) come SSOT semantica, e `new2.js` codifica solo la *shape di orchestrazione* + le policy di gate. La telemetria Phase-8 esce con discriminatore `variant: "new2"` su `skill-runs.jsonl`, confrontabile con i run `/new`. Limiti documentati (è un esperimento): Claude-only, recovery solo intra-sessione, ogni op git/bash diventa azione di un agent (lo script non ha fs/shell). Distribuzione: i `.js` sono raccolti automaticamente dal per-item symlink esistente (`MERGE_KINDS.workflow`, overlay:false, Codex `workflowsDir()=null`) — nessuna modifica all'installer; lo `framework-edit-gate` li scansiona by-path → tutti i fact via `args`.

### Added

- **`framework/.claude/skills/new2/SKILL.md`** — skill sottile di ingresso (`effort: high`): degradation gate (`Workflow` tool assente / `.claude/workflows/new2.js` non linkato → "usa /new" + HALT), parse args (stessa grammatica di `/new`: range, `-full`, `-stats`, `effort=`), feature gate (`has_backlog`), raccolta fact da `baldart.config.yml`, `Workflow({ name:'new2', args })`, e presentazione del report + telemetria `variant=new2`. Resta ~1 schermata nel main loop → costo context trascurabile.
- **`framework/.claude/workflows/new2.js`** — orchestratore della batch: Pre-flight (un ops-agent esegue Phase 0 + worktree + cross-card con le policy deterministiche G1–G5/E2), pipeline per-card sequential **o** team-mode (`parallel()` per wave), Final review (delega a `new-final-review`), auto-merge (G24 + riconciliazione 6b/6c conservativa G19–G23). Ogni gate bloccante / finding di scope → `workflow('new2-resolve', …)`. Gate ledger + report + telemetria nel return.
- **`framework/.claude/workflows/new2-resolve.js`** — workflow di risoluzione self-healing (8 `kind`, 3 assi; tier fix→judge→verify→follow-up). `agent()`/`parallel()` only (figlio, nessun nesting).

### Changed

- **`framework/docs/WORKFLOWS.md`** — tabella "What ships today": aggiunte le righe `new2` e `new2-resolve` + nota "esperimento, non rimpiazzo".
- **`README.md`** — conteggio skill 29 → 30; `new2` aggiunto alla categoria Workflow con nota EXPERIMENTAL/Claude-only.

## [4.15.0] - 2026-06-04

**Il pre-flight di `/new` smette di gonfiare il contesto dell'orchestrator: il setup worktree + baseline build e il check Codex cross-card girano in background fuori dal prefisso cached, e le scritture del tracker di pre-flight si fondono in un singolo flush.** Forensics su una run reale (`/new FEAT-0020 -full`, 8 card, team mode) — estrazione del campo `usage` per turno dal transcript — ha mostrato **~150k token di occupazione al solo pre-flight, prima di toccare una card**, con un `cache_read` cumulativo di ~10M token (il prefisso ri-letto a ogni turno × ~70 turni). La diagnosi ovvia («è il bloat della prosa della skill») è stata **refutata dai numeri**: dopo lo split di v4.14.0, la prosa skill-controllata in quei 150k è solo ~18%; il floor statico ambiente (schemi tool MCP + CLAUDE.md + memoria) è ~42% e non è toccabile da `/new`; il vero driver residuo è il **numero di turni dell'orchestrator nel pre-flight** (~70) — ognuno ri-bolla il prefisso — alimentato da tre sprechi concreti. **MINOR** (cambio di orchestrazione capability-preserving di `/new`: stessi gate, stessi output, stesso install-contract; nessuna chiave `baldart.config.yml`, nessuna modifica CLI).

> **Why.** L'orchestrator di `/new` è long-lived: tutto ciò che entra nel suo contesto ci resta per l'intera run e viene ri-letto (a tariffa cache-read) a ogni turno. Tre cose lo gonfiavano al pre-flight senza alcun bisogno di restare nell'orchestrator: (1) l'invocazione inline di `/nw` caricava il corpo da **~1200 righe** di `worktree-manager` nel prefisso cached + camminava `npm install`/`tsc`/`lint`/`build` per ~15-20 turni foreground; (2) il trace `[codex]` del check cross-card rientrava nell'orchestrator perché il result-handling faceva `tail`/`cat` del file grezzo (pieno di righe di comando-trace) invece del solo verdict; (3) il tracker veniva creato e poi editato ~5 volte in modo incrementale durante un pre-flight che è idempotente e non ha bisogno di persistenza mid-flight. La cura sposta (1) e (2) in **contesti usa-e-getta** (subagente background + Bash background) di cui l'orchestrator trattiene solo il verdict strutturato, e collassa (3) in **un solo flush**. Il pattern è additivo agli invarianti già presenti (§ "Context economy" del core, regola sul canale inline) e non tocca nessun gate BLOCKING né il contratto di recovery per-fase dell'esecuzione card. Lezione metodologica confermata: dopo ogni fix il driver dominante può spostarsi d'asse — qui da "prosa nel prefisso" (v4.13/4.14) a "turni di pre-flight × prefisso crescente".

> **Onestà sulla magnitudine (post code-review + DUE giri di adversarial review).** Il risparmio vero è il **sollievo del prefisso dell'orchestrator** (corpo di worktree-manager fuori dal prefisso cached + ~15-30 turni foreground collassati in background ops): i token del build NON spariscono — vengono spesi nel subagente, sullo stesso pool di sessione — ma smettono di moltiplicarsi contro il prefisso crescente dell'orchestrator a ogni turno. La review avversariale ha **refutato** quattro punti rivelatisi difetti reali (poi fixati): (1) l'assunzione "un subagente può invocare una Skill" era **non provata e senza fallback** → **fallback inline a `/nw`** se il subagente non ritorna il blocco; (2) la barriera a due op poteva **risvegliarsi sul primo completamento** e leggere un `$AUDIT_FILE` a metà → **wait-for-ALL esplicito**; (3) la delega worktree apriva un **recovery-gap** (un worktree di codice NON è ri-creabile — `/nw` fa fail-loud su collisione). Il primo rimedio (pre-check su `registry.json`) è stato **a sua volta refutato dal secondo giro**: worktree-manager scrive l'entry registry SOLO a fine `/nw`, *dopo* il build, quindi per tutta la finestra di barriera il worktree esiste su disco ma NON è in registry → il check era cieco proprio nello scenario da coprire. Fix corretto: **rilevamento git-autoritativo** (`git worktree list` sul branch deterministico) con resume-se-completo / reset-e-ricrea-se-orfano (a pre-flight non c'è lavoro card da perdere); (4) il flush proponeva un `Write`-da-memoria che dopo una compaction avrebbe **droppato i campi Phase-0** (`$MAIN`/`$TRUNK`) → ora solo `Edit` chirurgici, mai overwrite; più una clausola nel **Context recovery protocol** del core che codifica il rientro in Pre-flight step 4. Per contro l'adversarial è stata essa stessa **refutata dai dati** sul filtro `[codex]`: lo static-grep del `.mjs` diceva "prefisso mai emesso", ma il file audit reale ha **63 righe `[codex]` su 73** (il wrapper Codex di Claude Code le emette nel redirect) → il filtro è corretto e tenuto.

### Changed

- **`framework/.claude/skills/new/references/setup.md`** — Pre-flight (once) step 4-6 riscritti:
  - **Step 4 (Lever #1)**: il setup worktree non è più invocato inline; l'orchestrator spawna **un subagente background** (`general-purpose`, **`mode: bypassPermissions`**, `run_in_background`, `name: worktree-setup-<FIRST-CARD-ID>`) che invoca `/nw` programmatic e ritorna SOLO un blocco strutturato (`worktree_path`/`branch`/`port`/`created_at`/`baseline`/`baseline_log`). Il corpo da ~1200 righe di `worktree-manager` + i log install/build vivono e muoiono nel contesto del subagente. **Guardie (post-review, 2 giri)**: step **4a2** pre-check **git-autoritativo** (`git worktree list` sul branch deterministico — NON il registry, che è scritto solo a fine `/nw` ed è cieco durante il build) → worktree completo: resume; orfano mid-build: reset (`worktree remove --force` + `branch -D`) e ricrea; step **4d** fallback inline a `/nw` se il subagente non ritorna il blocco (capability Skill-da-subagente non assunta, mai strandare la barriera); campo `created_at` nel blocco (e letto dal registry sul path di resume) per non rompere `cycle_time_mins`; flush con soli `Edit` chirurgici (mai `Write`-da-memoria, che dopo compaction droppa i campi Phase-0).
  - **Step 5 (barrier)**: l'orchestrator lancia 3d (Codex) + 4 (worktree subagent) insieme e **chiude il turno** su una barriera **wait-for-ALL** (attende OGNI op lanciata, non la prima — altrimenti leggerebbe un `$AUDIT_FILE` a metà; compaction mid-barrier rientra in pre-flight, resa sicura dal 4a2) — niente più ~30 turni foreground di walk-through (no `sleep`/`echo`, ri-invocazione automatica a completamento).
  - **Step 6 (Lever #3)**: al resume, gate sul `baseline: fail` (STOP), verdict Codex distillato, e **un singolo Edit** che scrive l'intero blocco di pre-flight (`## File Ownership Map` + `## Execution Mode` + `## Worktree` + `## Cross-Card Conflicts (Codex)`) — niente Edit incrementali per sub-step (pre-flight idempotente; la persistenza incrementale resta per l'esecuzione card).
  - **Step 3d "Result handling" (Lever #2)**: aggiunta la **verdict-extraction discipline** — il file `$AUDIT_FILE` contiene il trace `[codex]` davanti al verdict; vietato `cat`/`tail`/`head` grezzo, si legge solo via filtro `grep -vE '^\[codex\]' "$AUDIT_FILE" | tail -n 60` e si logga il finding distillato, mai il trace.
- **`framework/.claude/skills/new/references/team-mode.md`** — Team Mode Pre-flight: le sezioni `## Team Mode` + `## Parallel Groups` confluiscono nel **flush unico** di setup.md step 6c (niente Edit separato); il wave layout è risolto prima del lancio dei background ops, così è in-context per il flush.
- **`framework/.claude/skills/new/SKILL.md`** — § "Context economy": nuovo **punto 6** (delega il setup body-heavy a un contesto usa-e-getta — background subagent/Bash, l'orchestrator tiene solo il verdict; batch delle scritture tracker di pre-flight) + chiusura della regola estesa a «what the orchestrator chooses to retain».

## [4.14.1] - 2026-06-04

**Il gate Simplify del Team Mode di `/new` (D.3b) ora skippa solo le card davvero trivial (`TRIVIAL_CARDS`), non tutte quelle con `review_profile == skip` — chiude una divergenza col path sequential (Phase 2.55 `IS_TRIVIAL`) e con la stessa D.1.5 di team-mode.** La pipeline per-card di `/new` vive in due modalità (sequential e team) che ri-enunciano alcuni gate; due predicati gemelli erano davvero divergenti. **Drift 1 (bug, fixato)**: il sequential Phase 2.55 ri-conferma `IS_TRIVIAL` sul diff ACTUAL committato (tutte e 3 le condizioni, incl. il check non-source) → una card `review_profile=skip` con un file SORGENTE nel diff reale NON è trivial → ESEGUE Simplify; il team D.3b invece skippava sul solo `review_profile == skip`, ignorando il diff reale → la saltava. Aggravante: team-mode era incoerente con sé stesso — a D.1.5 calcola GIÀ `TRIVIAL_CARDS` col check non-source completo (e la sua nota SSOT dichiara «D.3b already skipped for trivial»), ma il gate D.3b gateava sul set più lasco. Blast radius: solo qualità (la card prendeva comunque D.2 code-review + il Final FULL gate), ma in team mode si perdeva la difesa contro la deviazione del coder che il sequential ha. **Drift 2 (intenzionale, documentato)**: lo scope per-GRUPPO della QA gate (D.4) vs per-CARD del sequential (21b) — stesso predicato (`deep`/risk → FULL, altrimenti defer al Final), granularità per necessità del modello a wave (qa-sentinel gira una volta sul diff combinato del gruppo); è il lato deliberatamente più conservativo (una card rischiosa → tutto il gruppo FULL), nessun gap di correttezza → comportamento invariato, aggiunta solo una frase di chiarimento. **PATCH** (bugfix di un falso-skip in team mode; nessuna capability, nessuna chiave `baldart.config.yml`).

> **Why.** I gate ri-enunciati nei due path di `/new` devono restare predicato-identici, o le due modalità divergono silenziosamente. Qui il fix è un puro allineamento: D.3b passa dal set lasco (`review_profile == skip`) al set già calcolato a monte a D.1.5 (`TRIVIAL_CARDS` = `review_profile==skip` AND 0 Step-A triggers AND non-source diff), che è esattamente l'`IS_TRIVIAL` del sequential 2.55. Nessuna nuova logica, nessun ricalcolo: si consuma il set che D.1.5 dichiarava già di pilotare. Il Drift 2 si lascia perché la granularità per-gruppo è una necessità del modello a wave (un solo qa-sentinel sul diff combinato), non una svista, ed è il lato più sicuro — ma ora è documentato come intenzionale per non farlo rifixare in futuro.

### Changed

- **`framework/.claude/skills/new/references/team-mode.md`** — gate D.3b (Simplify) riscritto da «SKIP per `review_profile == skip`» a «SKIP per card in `TRIVIAL_CARDS`» (il set già computato a D.1.5), con log-line allineata al sequential (`simplify: SKIPPED (trivial — non-source diff)`); una card `skip` con file sorgente nel diff ora esegue Simplify come nel sequential 2.55. D.4 (QA gate) — aggiunta una frase che documenta lo scope per-gruppo come INTENZIONALE (stesso predicato del sequential 21b, valutato a livello di gruppo perché qa-sentinel gira una volta sul diff combinato; lato deliberatamente più conservativo). Nessun cambio di logica su D.4.

## [4.14.0] - 2026-06-04

**Nuovo tipo di payload `.claude/workflows/` (dynamic workflow di Claude Code) + il primo workflow `new-final-review`: la Final Review F.2–F.4 di `/new` può girare come script di orchestrazione deterministico, opt-in e con fallback inline.** I dynamic workflow (research-preview di Claude Code) spostano l'orchestrazione fan-out fuori dal context-window dell'orchestrator, nel codice. Una review avversariale ha **refutato** l'idea di farne il motore *per-card* di `/new` (rompe la portabilità Codex — sono Claude-only; dipende da una preview come breaking change; duplica la logica di gate col team-mode; introduce attrito-permessi e regressione di recovery proprio dove `/new` è più intrecciata con worktree/git/stato). Il bersaglio scelto è invece la parte dove i rischi spariscono e il valore resta: la **Final Review F.1–F.6** è **read-only** (gli agenti leggono solo il diff committato), ha **definizione SSOT unica** condivisa da sequential e team mode (→ nessun twin), e ha l'**adversarial-verify già nativo** (FP-check di Codex in F.3, cross-validation `code-reviewer` a confidence<80 in F.4). **MINOR** (nuova capability additiva: nuovo payload type + un workflow; nessun breaking change, nessuna chiave `baldart.config.yml` — il gating è runtime via presenza del tool `Workflow`; i consumer senza workflow o Codex-only eseguono la Final Review inline come prima).

> **Why.** La Final Review è l'unica fase di `/new` con una definizione singola condivisa da entrambe le modalità e senza ownership di stato git o gate-utente: estrarla in un workflow non duplica nulla e non regredisce nulla. Il confine è netto — il workflow incapsula **F.2 (baseline) + F.3 (fan-out Codex ‖ doc-reviewer ‖ api-perf-cost-auditor ‖ qa-sentinel) + F.4 (consolidamento + verifica)** e ritorna i finding già classificati (`VERIFIED`/`NEEDS_MANUAL_CONFIRMATION`, `FALSE_POSITIVE` scartati) + la gate table; **F.1 (scope), F.5 (apply fix + build + `AskUserQuestion`) e F.6 (wrap-up) restano nella skill**, perché un workflow non può eseguire git né chiedere all'utente. La prosa di `final-review.md` resta SSOT e fallback sempre-disponibile; il workflow ne mirrora solo la *forma di orchestrazione* (i suoi agent-brief citano il modulo per la *semantica* di review). Distribuzione: symlink per-item come skill/agent/command, ma **senza overlay** (l'overlay-merger è per `.md`; un `.js` non ha sezioni markdown) e **solo per Claude** (`codex.js` ritorna `workflowsDir()=null` → mai linkato in un albero Codex).

### Added

- **`framework/.claude/workflows/new-final-review.js`** (nuovo) — workflow `meta.name: 'new-final-review'`, phases `Baseline`/`Review`/`Verify`. Fan-out F.2–F.4 con fallback Codex→`code-reviewer` come branch JS, schema sui finding (`VERIFIED`/`FALSE_POSITIVE`/`NEEDS_MANUAL_CONFIRMATION`), `api-perf-cost-auditor` saltato via branch quando non ci sono file API/data. Tutto da `args` (portabile, anti-contaminazione).
- **Payload type `.claude/workflows/`** — `src/utils/tool-adapters/claude.js` espone `workflowsDir()='.claude/workflows'`; `codex.js` `workflowsDir()=null` (Claude-only). `src/utils/symlinks.js`: `_mergeBulkDir` generalizzato con mappa per-kind (`MERGE_KINDS`: estensione + overlay on/off), nuovo `mergeWorkflows()`, chiamato in `createAllSymlinks()` e nel reconcile di `update.js`; `verifySymlinks()` esteso a `workflows`. `BALDART_MANAGED_PATTERNS` (`update.js`) include `workflows`. `doctor.js`: probe + repair advisory dei workflow non linkati (Claude-only, non bloccante).
- **`framework/docs/WORKFLOWS.md`** (nuovo) — cos'è il payload `.claude/workflows/`, il modello opt-in + degradazione, il perché Claude-only e no-overlay.

### Changed

- **`framework/.claude/skills/new/references/final-review.md`** — nuovo **Step F.1.5 (workflow delegation gate)**: se il tool `Workflow` è disponibile e `new-final-review.js` è linkato → delega F.2–F.4 al workflow e porta i finding/gate-table a F.5; altrimenti inline come prima (prosa = SSOT/fallback). Drift-note per i maintainer. `SKILL.md` § Routing annota la delega.

**Il SKILL.md di `/new` scende da 61k a ~10k token (−82%) — il dettaglio passo-passo di ogni fase è uscito dal corpo skill (prefisso di sistema, ri-letto a ogni turno) in moduli `references/*.md` caricati on-demand.** Diagnosi data-driven su una run reale (`/new` su ~4 card → 480k token di contesto): forensics sul transcript dell'orchestrator → il **floor statico è 152k token PRIMA che parta una card**, e il colpevole #1 è il SKILL.md monolitico da **2766 righe / 61k token**, residente nel prefisso di sistema cached e ri-validato a ogni turno (362 turni nella run osservata). Una prima ipotesi ovvia (comprimere la *prosa* degli agent, stile "caveman") è stata **misurata e scartata**: la prosa è ~8k = 2% del totale. La cura replica 1:1 il pattern già in prod su `/prd` (core snello + reference modules + routing table): il core tiene solo gli invarianti cross-fase (Context Tracking, Progress Visibility, § Context economy, Agent Routing, QA Profile, Trivial fast-lane, Risk-signal detector, Fix Application Log) + una **§ Routing** che mappa fase→modulo→skip; ogni fase carica il suo modulo via Read quando ci entra. Contenuto spostato **verbatim** (zero cambi comportamentali; le stesse fasi/gate, nello stesso ordine). **MINOR** (refactor capability-neutral + nuovo contratto interno di module-loading on-demand; install-contract e invocazione invariati, consumer auto-aggiornati via directory-symlink; nessuna chiave `baldart.config.yml`, nessuna modifica CLI).

> **Why.** Il corpo di una skill vive nel prefisso di sistema: a ogni turno l'harness lo ri-legge dalla cache, quindi un SKILL.md da 61k si paga 362 volte su una run lunga e occupa il tetto-contesto per l'intera durata. Spostare il dettaglio per-fase nella conversazione (via Read on-demand) lo rende **compaction-eligible** (il system prompt non lo è) e fa sì che le card trivial/light non carichino mai i moduli delle fasi che skippano. Caveat onesto: nel caso peggiore (no compaction, ogni fase eseguita una volta) il risparmio è il solo shrink del prefisso (~50k × ogni turno); nel caso tipico è maggiore. La disciplina è preservata da una HARD RULE ("leggi il modulo PRIMA di eseguire la fase") + un campo tracker `phase_module_loaded:` che la § Context recovery ri-legge dopo una compaction. Il Level 2 (card-runner sub-agent per collassare i ~90 turni/card in contesto isolato) resta valutato a parte — i docs ufficiali confermano l'isolamento del contesto subagent, superando la refutazione v4.12.0, ma è fuori da questa release.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — riscritto come **core ~543 righe / ~10k token** (da 2766 / 61k). Tiene: frontmatter + meta-rules BLOCKING, § Context Tracking (+ nuovo campo `phase_module_loaded:`), § Progress Visibility, § Context economy, **nuova § Routing** (mappa fase→modulo→skip + HARD RULE "Read il modulo prima della fase"), § Parallelism rules, § Context recovery protocol (+ step di ri-Read del modulo post-compaction), § Agent Routing, § QA Profile Selector, § Trivial-card fast-lane, **nuova § Risk-signal detector** (il vecchio Phase 3.7 Step A, rilocato nel core perché usato sia dal Trivial-classifier a Phase 1 sia da Phase 3.7), § Fix Application Log + Domain-Override.

### Added

- **`framework/.claude/skills/new/references/*.md`** (11 moduli on-demand, contenuto verbatim dal SKILL.md originale): `setup.md` (Phase 0 + Pre-flight), `implement.md` (Phase 1-2), `completeness.md` (Phase 2.5/2.5b), `review-cycle.md` (Phase 2.55/2.6/3/3.5), `codex-gate.md` (Phase 3.7, con Step A che ora rimanda al § Risk-signal detector del core), `commit.md` (Phase 4-5), `final-review.md` (Final F.1-F.6), `merge-cleanup.md` (Phase 6/6b/6c), `production-readiness.md` (Phase 7), `team-mode.md` (Team Mode + Step D), `metrics.md` (Phase 8). Arrivano ai consumer (Claude **e** Codex) via il directory-symlink esistente, senza modifiche CLI.

## [4.12.1] - 2026-06-04

**I gate Phase 0 di `/new` (e il finalizer `/mw`) non scattano più sugli artefatti BALDART-managed lasciati dirty da una reinstall/update — `.baldart/generated/`, `.baldart/state.json`, `.baldart/skill-conflicts.json`.** Sintomo riportato dall'utente: *ogni* volta dopo un aggiornamento, `/new` Phase 0 segnalava «stato non pulito» (N file dirty `.baldart/generated` + `state.json`) e «commit ahead di origin», girando all'utente la domanda su come gestire lo stato della main repo. La verifica del sorgente ha confermato la radice: i due gate Phase 0 sapevano auto-ignorare **solo** la telemetria sotto `${paths.metrics}` (dirty-tree gate) e i commit sotto `.framework/` o con subject `baldart update/add` (divergence gate) — ma `baldart update`/`add`/`push` scrivono una **terza** classe di rumore sotto `.baldart/` (i file reali del merge overlay in `generated/`, il ledger `state.json`, il log `skill-conflicts.json`) che nessuno dei due riconosceva, così finiva in `$USER_DIRTY` / `GENUINE_AHEAD` e bloccava. Stessa classe di bug del fix v4.2.2 (telemetria), per gli artefatti `.baldart/` che quel fix non copriva. Cura: **classificazione chirurgica, non rilassamento** — `.baldart/overlays/` resta deliberatamente FUORI dal set ignorato (un overlay editato è lavoro genuino e *deve* gateare). **PATCH** (bugfix di un falso positivo nei gate; nessuna nuova capability, nessuna chiave `baldart.config.yml`).

> **Why.** Un gate di igiene workspace deve distinguere il *lavoro feature in corso* dall'*output del framework stesso*. La v4.2.2 aveva insegnato ai gate a ignorare la telemetria che `/new` scrive di suo; questa fix estende la stessa logica all'output che `baldart update`/`add` lascia dirty per costruzione (rigenera `generated/*` e bumpa `state.json` a ogni sync, committati solo al reconcile successivo). La distinzione critica è `generated|state|skill-conflicts` (framework-managed → ignora) **vs** `overlays/` (user-owned → gatea): un partizionamento di pattern, mai un rilassamento del gate. Verificato con bash reale che le tre forme (porcelain partition, file-list reclassification, `case` glob del finalizer) trattino l'overlay edit come genuino in ogni punto.

### Fixed

- **`framework/.claude/skills/new/SKILL.md`** — Phase 0 step 3 (dirty-tree gate) e Phase 6c step 3 (clean-tree assertion) droppano da `$USER_DIRTY`/`$POST_DIRTY` anche `.baldart/(generated/|state.json|skill-conflicts.json)`; Phase 0 step 4 (divergence reclassification loop) accetta come framework-management i commit che toccano *solo* `.framework/` **e/o** quei `.baldart/` artifacts. Label tracker `framework-telemetry-only` → `framework-managed-only`.
- **`framework/.claude/skills/worktree-manager/SKILL.md`** — il finalizer `/mw` (sync main repo) classifica gli stessi `.baldart/` artifacts come OWNED (auto-reconcile lossless, come il metrics log) invece che FOREIGN, evitando il falso `[SYNC-NEEDS-DECISION]` che `/new` Phase 6c convertirebbe in un gate utente.

## [4.12.0] - 2026-06-04

**`/new` non gonfia più il context dell'orchestrator su long-run — il bulk content (log dei gate, `git diff`, file interi, baseline) non scorre più nel canale inline dell'orchestrator: vive su disco e viaggia per *path*.** Sintomo riportato dall'utente: su batch lunghi il consumo si impennava per il context bloat. Una prima ipotesi (i *finding* delle review che tornano inline + un card-runner sub-agent isolato + un "context purge" da rinforzare) è stata **sottoposta a review avversariale e refutata**: (1) l'orchestrator DEVE leggere il contenuto dei finding ai punti-decisione (domain-override partition, AC-Closure, `NEEDS_MANUAL_CONFIRMATION`, merge/dedup del Final gate) → un contratto "verdict-only" che li nasconde romperebbe la skill e i consumer condivisi (`/prd`, `/codexreview`); (2) un card-runner sub-agent non può ospitare i gate utente né garantire il nested-spawn dei reviewer; (3) un "compaction-lossless checkpoint" non regge (`/tmp` è single-point-of-failure, l'auto-compact non è controllabile da una skill). La verifica sul codice ha trovato che **il vero driver del bloat è il bulk content reso inline nei tool-result dell'orchestrator stesso**, non i finding: output dei gate `lint`/`tsc`/`test`/`build` (resi inline da `| tee`, ~40-50%), `git diff` completo passato inline ai simplify agent (~25-30%), Read di file interi nel completeness check, e baseline concatenati nel Final F.3. **MINOR** (cambia il comportamento di esecuzione di `/new`; nessuna nuova capability, nessun nuovo agente, nessuna chiave `baldart.config.yml`, nessun contratto di return degli agent toccato — la schema-propagation rule non si applica).

> **Why.** L'orchestrator di `/new` è un unico contesto long-lived su tutto il batch: ogni byte che rende inline (Bash tool-result, prompt che compone, file che Read-a) resta in finestra per l'intero run, e nessun "purge" via prompt recupera token. La fix attacca la radice — **tenere il bulk fuori dal canale inline** — senza toccare ciò che l'orchestrator può leggere ai punti-decisione (i finding restano file-resident e leggibili on-demand). Onestà sui limiti: questo riduce il *rate di accumulo* (stima ~40-50% sul footprint dominante), non azzera la crescita del context; il bound vero richiederebbe isolamento per-card a livello di harness, fuori dalla portata di una skill.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — nuova sezione SSOT **`## Context economy (MANDATORY)`** (HARD RULE in 5 punti: gate-output discipline, diff-to-disk, findings file-resident, pass-paths-not-concatenations, targeted Reads). Applicata ai siti che la violavano:
  - **Gate bash (Phase 2 step 8)** — `cmd 2>&1 | tee /tmp/x.txt` → `cmd > /tmp/x.txt 2>&1; echo "<gate>:$?"`: l'output completo resta su disco, inline solo l'exit code; su fallimento si legge un estratto *bounded* (`tail -n 30`), mai il log intero. Step 9 passa al fix agent il **path** del log, non il blob.
  - **Tutti i re-run di gate foreground** (gap-fix Phase 2.5, Phase 2.55 step 5, Phase 3 step 16, pre-commit Phase 4, Final F.5, team D.1/D.3b) — pointer alla disciplina (redirect a disco, mai inline).
  - **`git diff` (Phase 2.55 step 2-3)** — `git diff > /tmp/diff-<CARD>.txt`; i 3 simplify agent ricevono il **path** e leggono il diff da soli, niente «full diff» nei prompt. Team D.3b allineato.
  - **Blocchi Codex (Phase 0 cross-card audit, Final F.3)** — `2>&1 | tee "$FILE"` → `> "$FILE" 2>&1` (niente copia inline; il file si riempie comunque in streaming per il polling F.4, commento aggiornato).
  - **Final F.3 baseline** — `${ARCH_BASELINE}` (contenuti concatenati) → `${ARCH_BASELINE_PATHS}` (lista di path che Codex Read-a on-demand); F.2 non concatena più i per-card baseline.
  - **Completeness check (Phase 2.5 step 4)** — Read *mirati* (grep con `offset`/`limit` stretti attorno al `file:line` del completion report) invece di Read di file interi.
  - **Progress Bar** — blocco markdown completo solo alle boundary pesanti (card/wave/gate-decision/STOP); il movimento di fase intra-card si mostra via Task spine sub-state, non ri-emettendo la tabella a ogni fase.
  - **Phase 5 step 31** — il "CONTEXT PURGE" è ri-etichettato come *reference-hygiene* (non token-reclamation, che un prompt non può ottenere); il leverage reale è upstream (la § Context economy tiene il bulk fuori dal context fin dall'inizio).

## [4.11.0] - 2026-06-04

**Gli overlay non costringono più a riconciliare a ogni `baldart update` — la VERSION non è più incisa nell'identità dei file generati né usata come segnale di drift quando il base file non è cambiato.** Sintomo riportato dall'utente: ogni update richiedeva una riconciliazione. La verifica del sorgente ha trovato **due trigger indipendenti con la stessa radice** — la VERSION del framework veniva trattata come parte dell'identità di overlay/file-generati, ma viene bumpata a *ogni* release a prescindere dal fatto che lo specifico base file sia cambiato. **(A) churn dei file agent/command generati:** `buildMarker()` stampava `base_version=<VERSION>` nel commento `<!-- baldart-generated -->`; a ogni update la VERSION cambiava → la riga del marker cambiava anche su contenuto byte-identico → git vedeva un diff → scattava il commit `post-update reconcile` *a ogni update*. Il campo era puramente informativo (`readMarker` lo leggeva ma nessuno gate-ava la rigenerazione su di esso). **(B) falso-drift degli skill overlay:** `doctor` e `status` confrontavano il pin `base_skill_version` *direttamente* con la VERSION installata (ignorando `base_file_sha`), così ogni bump faceva apparire ogni overlay come "drifted" anche con base SKILL.md identico — mentre `overlay drift` lo faceva già correttamente SHA-first. **MINOR** (cambia il formato del marker generato; nessuna chiave `baldart.config.yml` — la fix riusa `base_file_sha` già presente).

> **Why.** Il drift deve segnalare *modifiche reali di contenuto*, non bump di versione. La radice comune era usare la VERSION come proxy di "il base è cambiato": un proxy sempre falso-positivo dato che VERSION sale a ogni release. La fix rende ogni layer **content-addressed** (SHA) e relega la VERSION a metadato informativo: (A) il marker porta solo `base_sha` + `overlay_sha` — entrambi cambiano *solo* su modifica reale → i file generati tornano byte-stabili tra update di sola-versione (il commit di reconcile scatta solo quando base o overlay cambiano davvero); (B) `doctor`/`status` diventano SHA-first come `overlay drift`, col confronto di versione come fallback solo quando manca `base_file_sha`. In più il pin `base_<kind>_version` viene rinfrescato al merge **solo quando lo SHA combacia** (base invariato) — così resta veritiero senza mai mascherare drift reale: se il base è cambiato lo SHA differisce e il pin stale resta un segnale legittimo da riconciliare. Retrocompatibile in lettura: `readMarker` tollera ancora i marker legacy con `base_version=` (i consumer pre-4.11 subiscono un diff one-shot al primo update, poi stabilità).

### Changed

- **`src/utils/overlay-merger.js`** — `buildMarker()` non stampa più `base_version` (identità content-addressed: solo `base_sha` + `overlay_sha`); `readMarker()` rende `base_version` opzionale nella regex (legge sia il formato nuovo sia il legacy, `baseVersion: null` quando assente); call-site di `mergeOverlay` semplificato. Nuovo helper `refreshOverlayVersionPin(overlayPath, kind, frameworkVersion, currentBaseSha)` — rinfresca il pin di versione **solo** quando `base_file_sha` combacia col base corrente (no-op su mismatch → drift preservato; no-op su already-current; fail-safe su I/O). `baseVersionTargeted` ora usa `frameworkVersion` come fallback significativo invece di `'unknown'`.
- **`src/commands/doctor.js`** — `overlayDrift` SHA-first: confronta `base_file_sha` con lo SHA del base SKILL.md corrente; fallback al confronto di versione solo quando `base_file_sha` manca (o il base è illeggibile). Esclude i file `*.example.md`.
- **`src/commands/status.js`** — display drift SHA-first allineato al contratto di `overlay drift` (mostra `base changed: sha X → Y` su drift reale, `base unchanged` quando lo SHA combacia; fallback di versione solo senza `base_file_sha`).
- **`src/utils/symlinks.js`** — `_generateOverlayedFile` calcola `baseSha` una volta e, dopo il backfill di `base_file_sha`, chiama `refreshOverlayVersionPin` per tenere truthful il pin informativo quando il base è invariato.

### Added

- **`src/utils/__tests__/overlay-merger.test.js`** — 7 test: stabilità byte-identica del merge tra bump di sola-versione (fix Trigger A), assenza di `base_version` nel marker generato, lettura back-compat del marker legacy, e i tre rami di `refreshOverlayVersionPin` (refresh su sha-match, no-op su sha-mismatch con drift preservato, no-op su already-current).

### Fixed

- **`framework/.claude/skills/new/SKILL.md`** (Phase 3.7 Step A, card-scoped diff) — risoluzione della backlog card per nome file resa **case-insensitive** (`find -name` → `-iname`). Era l'unico punto in tutte le skill che risolve una card by-name, e `-name` è case-sensitive: i file su disco hanno prefisso maiuscolo (`FEAT-`/`BUG-`/`CHORE-`), ma se `CARD_ID` arriva con case diverso (es. `/new feat-0019`) `find` tornava **vuoto in silenzio** → `CARD_FILE=""` → la card veniva trattata come inesistente (stessa famiglia del case-collision docs/PRD noto). Stesso blocco: il placeholder `${paths.backlog_dir:-../../backlog}` non era bash valido (il punto nel nome variabile → `bad substitution` se eseguito verbatim) — ora emesso già risolto in una variabile `BACKLOG_DIR` (fallback `backlog`).
- **`framework/.claude/skills/worktree-manager/SKILL.md`** (`/nw` step 3b, sync card non-tracciate) — stesso fix di normalizzazione: il token `${paths.backlog_dir:-backlog}` letterale (bash non valido) sostituito da una variabile `BACKLOG_DIR` risolta, allineando il fallback (`backlog`) tra le due skill.

## [4.10.0] - 2026-06-04

**`/new` ora ha una todo list persistente e una progress bar — durante un batch sai sempre quale fase gira, quale wave del team mode è in volo, e quale gate è risolto o skippato (con motivo).** Finora `/new` era un orchestratore autonomo cieco verso l'utente: l'unica traccia di stato era il tracker file interno `/tmp/batch-tracker-<FIRST-CARD-ID>.md`, che l'utente **non vede**. Risultato: in un run lungo (15 fasi per card, 50+ gate, sequential o team mode a wave) non avevi modo di capire dove fosse il processo. La fix replica il doppio meccanismo che dà a `/prd` la sua todo list amata, adattato alla struttura `wave → card → fase/gate`: (1) un **Task spine nativo** (TaskCreate/TaskUpdate) — il pannello todo sempre visibile della UI — tenuto *coarse*, un task per card etichettato per wave (`Wave 1 · FEAT-0502 — <title>`) più i framing task `Pre-flight`/`Final review`/`Merge & cleanup`; (2) una **Progress Bar markdown** emessa **alle transizioni** (cambio fase/wave/card, decisione gate, ogni `AskUserQuestion`/STOP — non a ogni messaggio, per non fare rumore nei tratti autonomi) con tabella card×wave + un **Gate ledger** che mostra ogni gate `✅ risolto` / `⏭️ skippato (+motivo)`. **MINOR** (capability aggiunta alla skill `/new`; nessuna chiave `baldart.config.yml` — usa solo i Task tool nativi e l'`execution_strategy` già esistente).

> **Why.** Il valore è la *visibilità*, ma il vincolo è non trasformare il ledger in una licenza di skip. Per costruzione: ogni entry `⏭️` deve citare un motivo **verbatim** dalle skip-reason già enumerate nelle Gate table (`IS_TRIVIAL`, `review_profile=light`, `backend-only diff`, `features.has_e2e_review:false`, `balanced→Final`, `holistic_audit provenance`); un `⏭️ (time budget)` / `(to save tokens)` è esso stesso una violazione di protocollo per la clausola top-of-file "NO PHASE SKIP FOR PERCEIVED TIME". Il ledger rende i skip **auditabili**, non li autorizza. Il tracker interno resta la SSOT di recovery; il Task spine + Progress Bar sono il suo specchio live — esattamente il pattern `/prd` state-file ↔ progress bar.

### Added

- **`framework/.claude/skills/new/SKILL.md`** — nuova sezione autoritativa **"Progress Visibility (MANDATORY)"** (subito dopo "Context Tracking"): definisce una volta sola (A) il Task spine nativo per-card etichettato per wave con la tabella delle transizioni di stato (`Pre-flight`/card/`Final review`/`Merge & cleanup`), (B) il format e i trigger della Progress Bar a transizione con il Gate ledger e la regola dei motivi-verbatim, (C) il mapping fase→cella del ledger (2.5b/2.55/2.6/3/3.5/3.7 + i gate di setup). Reminder a una riga `→ Visibility` nei punti di transizione già esistenti — Pre-flight (creazione spine dopo la risoluzione di modalità/wave), Phase 1 step 2b (claim → card `in_progress`), Phase 4 step 29 (DONE → card `completed`), team mode Step B/D.6/E (confini di wave) e Step F.1/Phase 6 (transizioni di batch) — senza duplicare logica in ogni fase. Aggiunta una "Update rule" del tracker che lega ogni transizione interna ai due surface visibili.

## [4.9.0] - 2026-06-04

**`prd-card-writer` Rule C apre `light` alle card UI-presentazionali — la metà `-ui` di una coppia split non paga più una review Codex piena su un diff di pixel.** Finora il pavimento era assoluto: *ogni* card `feature`/`enhancement` era `balanced` minimo, mai `light`. La regola nasceva giusta (non sotto-revisionare feature che *sembrano* piccole ma nascondono logica), ma non distingueva una UI card che *contiene* logica da quella genuinamente presentazionale — tipicamente la metà `-ui` di una coppia `logic/ui` già splittata dalla atomicity rule (es. `…-submit-ui`, la cui logica vive nella card sorella `…-submit-logic`). Risultato: su un epic come FEAT-0019, card puramente compositive venivano marcate `balanced` → review Codex-`full` per-card su un diff di soli componenti. Nuovo **ramo LIGHT (b)**, recintato e deterministico: scatta SOLO se TUTTE valgono — card pure-UI (`areas == [ui]`, o `owner_agent: ui-expert` quando `areas` è assente) · ≤3 file, tutti componenti/stili (zero data-fetching / state / API / server) · **nessun path sotto `paths.components_primitives`** (compone primitive esistenti, non ne crea/modifica) · nessun trigger DEEP. **MINOR** (nuova capability di routing in Rule C; nessuna chiave `baldart.config.yml` — si appoggia su `areas`/`owner_agent`/`paths.components_primitives` già esistenti).

> **Why.** Il taglio è sicuro perché **la rete DS non dipende dal `review_profile`**: una card `light` mantiene la sua code-review per-card/group con `code-reviewer` rule 8 + la cascade registry-first (gated su `features.has_design_system`), ed è coperta dalla **Final FULL gate** cross-batch — si rimanda solo il Codex-*adversarial* per-card. Quindi la DS-drift (hex/shadow/radius hardcoded = HIGH) viene comunque beccata. Il ramo è **stretto per costruzione**: ogni indizio di logica/state/data-binding, un primitive nuovo/modificato, o `areas` oltre `[ui]` → la card resta `balanced` (la guida resta "in dubbio, `deep`"). Coerente col principio "data-driven sopra threshold arbitrarie": niente soglia a righe, solo segnali strutturali autorati in PRD. Validato su FEAT-0019: dei 3 card `-ui`, solo `05b` (2 file componenti, logica su 05a, compone `BottomActionBar`) qualifica; `06b` (4 file + validazione UI/View Transitions) e `07` (5 file, tocca i view-model + render permission-gated) restano correttamente `balanced`.

### Changed

- **`framework/.claude/agents/prd-card-writer.md`** — Rule C: riga LIGHT estesa col ramo (b) UI-presentational (predicato recintato a 4 condizioni AND); paragrafo "Critical" riscritto — il pavimento `balanced` per `feature`/`enhancement` ora ha **l'unica eccezione** del ramo (b), col razionale (rete DS profilo-indipendente) e i disqualifier espliciti. SSOT invariato altrove: la fallback table del `/new` continua a puntare a Rule C, il routing escalation-only `light`→`full` resta intatto (una UI card `light` è sempre promuovibile da un trigger Step-A).

## [4.8.0] - 2026-06-03

**`baldart update` diventa seamless: un solo comando, zero iterazioni — aggiorna e preserva gli overlay.** Un transcript reale (`v4.2.1 → v4.3.0`) ha richiesto **3 invocazioni fallite + git surgery manuale** per un'operazione il cui lavoro è sempre lo stesso. Le cause erano bug di design del CLI, ora rimossi: (1) **asimmetria stash** — `--yes` auto-stasciava il tree sporco ma il path `--reset` lo *rifiutava*, buttando via l'autorizzazione dell'utente e costringendo a stash/drop/mv-a-/tmp/pop a mano; (2) **hard-stop su divergenza già sicura** — il CLI calcolava `overlayCoversTouched` (sa che gli edit sono già negli overlay, lo scrive: "--reset is safe") ma sotto `--yes` *rifiutava* invece di procedere; (3) **flag-maze** — per sbloccarsi serviva `--reset --yes --i-know`, combo da indovinare; (4) **finalizer assente sul path reset** — i generated-agents + `state.json` restavano uncommitted; (5) **preview morta** — `git log HEAD..FETCH_HEAD -- .framework/` è sempre vuoto col subtree-squash. Ora `npx baldart update --yes` fa tutto in un colpo: auto-stash (motore unico, `git stash push -u` senza pathspec → niente più fallimenti su dir `.gitignore`d), auto-risoluzione della divergenza (overlay-covered → reinstall pulito che riapplica gli overlay; **uncovered → auto-cattura in overlay popolati dal diff + VERIFICA che il merge riproduca l'edit** prima del reinstall), riconcilia + committa gli artefatti BALDART-managed (finalizer anche sul path reset), e ripristina lo stash. **MINOR** (capability aggiunta, comandi invariati e retrocompatibili; nessuna chiave `baldart.config.yml`).

> **Why.** L'intent dell'utente è vincolante: *"voglio che tutto il processo sia seamless, niente iterazioni — le cose da fare sono sempre le stesse: aggiornare e preservare gli overlay"* (coerente col principio "rispetta l'intent esplicito vs advisor"). L'unica collisione possibile tra "seamless" e "non perdere dati" è un edit a un file framework **non** coperto da overlay: la scelta dell'utente è **auto-cattura + verifica** — il CLI genera l'overlay dal diff (modello section-marker di `overlay-merger.js`) e ri-merge contro il base upstream fresco; se il risultato riproduce l'edit a livello di sezioni+preamble procede muto, altrimenti **si ferma una sola volta** (`divergence-capture-incomplete`) elencando i blocker (section deletion, cambio frontmatter/preamble, file `src/`) — mai un'azione distruttiva silenziosa. La verifica è la rete di sicurezza che rende la cattura sicura a prescindere dalla fedeltà dell'algoritmo. Gli escape hatch espliciti (`--on-divergence pull|scaffold-overlays|abort`, `--reset --yes --i-know`) restano per CI/power-user, ma non sono più il path che l'utente deve scoprire: il default È quello seamless.

### Added

- **`src/utils/overlay-capture.js`** — nuovo modulo `captureAndVerify`: verifica che OGNI edit overlay-able divergente sia preservato da un overlay, esistente o catturato. Per agents/commands è **verificabile** — ri-merge l'overlay (esistente o generato dal diff via `[OVERRIDE]`/H2 plain) contro il base upstream fresco (`FETCH_HEAD`) e confronta sezioni ordinate + preamble; un overlay **esistente ma stale** (non riproduce l'edit `.framework` corrente) diventa un **blocker**, non una perdita silenziosa al reset. Per le skill (runtime-concat, non mergiabili dal CLI) un overlay esistente è fidato, l'assenza è un blocker. Path non mappabili → blocker. Il chiamante passa l'INTERO set di file toccati (mai un sottoinsieme pre-filtrato): un pre-filtro lascerebbe il reset cancellare un edit senza overlay che lo preservi. Esporta helper unit-testabili (12+ test).

### Changed

- **`src/commands/update.js`** — (a) **motore di stash unico** `autoStashNonFramework` (blanket `git stash push -u`, niente pathspec) riusato sia dalla normale flow sia da `runReset` (che prima *rifiutava* il tree sporco). (b) **`runReset` finalizer**: re-apply dello stash + `postUpdateAutoCommit` + emissione JSON (`action: 'reset'`) → tree pulito su tutti i path. (c) **divergenza overlay-able**: il default sotto `--yes` non rifiuta più — `resolveOverlayDivergenceSeamlessly` cattura+verifica gli uncovered (committandoli **prima** del reset così sopravvivono allo stash e alimentano il merge), poi reset-and-reapply; stop singolo solo su blocker. (d) **divergenza mixed/real-custom**: cattura il sottoinsieme overlay-able e si ferma una volta (i path `src/`/CHANGELOG non sono auto-risolvibili). (e) **preview reale**: `changelogSinceInstalled` mostra le entry del CHANGELOG upstream tra versione installata e remota (rimpiazza e **rimuove** il dead-code `git.diffWithRemote`, diff opt-in sempre-vuoto col subtree-squash). `--i-know` resta esplicito solo per il `--reset` standalone (implicito nel path seamless).
- **`src/utils/git.js`** — rimosso il metodo morto `diffWithRemote` (nessun chiamante dopo la nuova preview).
- **`src/commands/overlay.js`** — esporta `buildFrontmatter` (riusato da `overlay-capture.js` per scrivere overlay con frontmatter valido senza re-implementare la logica).
- **`framework/.claude/skills/baldart-update/SKILL.md`** — riscritta per il flusso seamless: rimossa la replica chat dei 2 decision-point (preview + stash) e lo step di preview morto; la skill ora lancia `npx baldart update --yes` e riporta, sorvegliando l'unica pausa legittima (`divergence-capture-incomplete`). Description frontmatter, Hard rule #2, tabella decision-point, workflow (Step 0/1/2) e DRIFT-CHECK (`update_js_prompts: 5`, `last_verified: v4.8.0`) allineati.



**`/new` smette di rifare per-card 4 review che la Final FULL gate già ripete sull'intero batch.** La Final Review F.3 gira sempre, incondizionata, su tutto il diff di batch: Codex + **doc-reviewer** + **api-perf-cost-auditor** + **qa-sentinel** (suite completa + build + audit). Quindi le passate *per-card* di quegli stessi reviewer sono, per la sicurezza del merge-gate, già coperte dalla rete finale — il loro unico valore residuo è il fail-fast/attribuzione. Quattro pruning, tutti deterministici e con la Final come rete: (1) **plan-auditor grounding** saltato (sequential) quando `/prd` ha già fatto l'holistic audit (`metadata.holistic_audit`) e non c'è drift sui path della card — mirror esatto del gate v4.3.0 Step 3d; (2) **doc-review per-card** deferita alla Final per le card `light` il cui diff non tocca file di documentazione; (3) **qa-sentinel per-card** che gira solo per `deep` (o `balanced` escalata da risk-drift) — `balanced` non-rischiose deferiscono la suite alla Final; (4) **api-perf-cost-auditor per-card** droppato via nuovo flag lean-contract `skip_api_perf_auditor`, deferito alla Final. **MINOR** (comportamento `/new` più lean; nessuna chiave `baldart.config.yml` — si appoggia su `review_profile` + `metadata.holistic_audit` già esistenti).

> **Why.** Il criterio è netto: si potano solo le review che la **Final FULL gate rifà** (doc, api-perf, qa, codex). Ciò che la Final NON copre — **AC-Closure** (scope discipline), **Simplify** (quality/reuse), **codebase-architect grounding**, **E2E** — NON si tocca: pruneare quelli sarebbe perdere copertura, non togliere ridondanza. Ogni skip è un gate-reason **enumerato e deterministico** (provenance `/prd`, oppure `review_profile`), mai un `time budget`. ⚠️ **Tradeoff #3 (esplicito)**: la maggior parte delle card `feature`/`enhancement` è `balanced` per Rule C, quindi un batch tutto-balanced esegue il primo run della suite di test solo alla Final gate — si perde il fail-fast/attribuzione per-card; il merge-gate resta salvo (Final = suite completa sull'intero batch). Reversibile in 1 riga (ripristina `balanced → SCOPED`). Il percorso sequenziale e quello team restano allineati (stessa logica di defer). `/codexreview` riceve un campo lean-contract simmetrico a `skip_doc_reviewer` (api-perf è un breadth-agent indipendente, NON un validatore del FP-gate → safe).

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — (1) Phase 1 step 4: skip provenance-based del plan-auditor (P1 holistic_audit presente + P2 no-drift `git log <audited_commit>..<trunk> -- <card paths>`; fail-safe → RUN). (2) Phase 3 + team D.1.5/D.2/D.4a: nuova sotto-classe `DOC_DEFER_CARDS` (light + diff senza doc) esclusa dal doc-reviewer, deferita alla Final. (3) Phase 3.5 step 21b + team D.4: `balanced` deferisce qa-sentinel alla Final (eccetto risk-escalation → FULL per-card); solo `deep`/escalated gira per-card. (4) lean-contract a Phase 3.7 Step C + D.4b: `skip_api_perf_auditor: true`. + Step D coverage assertion (accetta i nuovi defer come entry valide) + clausola in testa (enumera i nuovi gate-reason).
- **`framework/.claude/commands/codexreview.md`** — Step -0.5 + Step 2 lean-agent note: nuovo campo `skip_api_perf_auditor` che omette agent #2 (`api-perf-cost-auditor`) anche a `full`. Non tocca il FP-gate (dipende da `code-reviewer`, non da api-perf).

## [4.6.0] - 2026-06-03

**`/new` ora ha una corsia "trivial" per le card minime: niente review di codice su un diff che non contiene codice.** Dopo il de-dup v4.5.0 (profondità scalata su `review_profile`), restava un pavimento alto: una card `skip`-profile ma implementabile — un fix di copy in un `.md`, un update di una riga di `ssot-registry`, un bump di config non-API — passava comunque per codebase-architect + plan-auditor + code-reviewer + AC-closure + QA + doc + Final. Una nuova **fast-lane** scatta quando `IS_TRIVIAL(card)` = `review_profile == skip` **AND** 0 trigger del detector Step-A **AND** diff **non-source** (check deterministico per estensione, riusando il predicato "backend-only" di Phase 2.6 — solo `.md`/`.json`/`.yml`/`.txt`/`.csv`). Una card trivial esegue solo il tail minimo sicuro: `coder → gate meccanici inline (markdownlint/lint/build) → AC-closure → doc-review → commit`, saltando **codebase-architect + plan-auditor + code-reviewer + simplify + Codex + la suite di test QA**. Vale in **entrambe le modalità** (sequential: gate in Phase 1 step 2c + 2.55/3.5/3.7; team: sotto-classe `TRIVIAL_CARDS ⊆ LIGHT_CARDS` esclusa dallo scope del code-reviewer a D.2). **MINOR** (nuova capability per-card; nessuna chiave `baldart.config.yml` — si appoggia su `review_profile` già autorato in PRD).

> **Why.** `review_profile: skip` NON è riservato agli epic: `prd-card-writer.md § Rule C` lo assegna deterministicamente a card `docs`/`chore`/`config` o puramente cosmetiche (typo/rename/copy) senza logica. È quindi un segnale autorato in PRD, persistito nello YAML, su cui ancorare la corsia **senza inventare una soglia a righe** (coerente col principio "data-driven sopra threshold arbitrarie"). Il classificatore è un AND a 3 condizioni con **guard fail-safe**: un solo file source nel diff, un trigger Step-A, o un profilo non-`skip` → `IS_TRIVIAL=false` e la card torna al percorso normale. La triviality può solo **togliere review a una card che non ha codice da revisionare**, mai il contrario — una `skip` mis-autorata su codice, o una il cui testo cita `auth`/`payment`, viene sempre ricatturata. Gli invarianti restano: **AC-closure** (scope discipline) e **doc-review** non si saltano mai, e la **Final Review F.3** copre incondizionatamente ogni card trivial nel gate cross-batch prima del merge (stessa rete su cui si appoggia già il percorso `light`). Tutti gli skip sono gate-reason enumerati deterministici, mai `time budget`.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — (a) nuova sotto-sezione SSOT "Trivial-card fast-lane" (classificatore `IS_TRIVIAL` a 3 condizioni AND + guard + tabella di cosa salta/resta). (b) Clausola in testa: `IS_TRIVIAL` aggiunto all'elenco dei gate-reason enumerati. (c) Sequential Phase 1: nuovo step **2c** che salta codebase-architect + plan-auditor (step 3/3a/4) per card provvisoriamente trivial; il briefing del coder degrada a card-spec + target file. (d) Phase 2.55/3.5/3.7: gate `IS_TRIVIAL` ri-confermato sul diff reale → skip Simplify + QA-suite + Codex, gate meccanici inline; carve-out esplicito nell'"UNCONDITIONAL" di Phase 3.7. (e) Team D.1.5: sotto-classe `TRIVIAL_CARDS ⊆ LIGHT_CARDS`; D.2 esclude le trivial dallo scope del code-reviewer. (f) Step D coverage assertion: skip trivial come entry valide documentate. Non toccati: `codexreview.md`, `prd-card-writer.md`, AC-closure, doc-review, Final Review F.3, percorso light/full di v4.5.0.

## [4.5.0] - 2026-06-03

**`/new` Team Mode non revisiona più lo stesso diff fino a tre volte: la profondità della code-review ora scala col `review_profile` della card.** Lo Step D del team mode faceva passare ogni card "full" sotto code-review tre volte — `D.2` (code-reviewer di gruppo) → `D.4b` (`/codexreview` per-card, che rispawna code-reviewer + Codex + CoVe + FP-gate) → Final Review F.3 (Codex sull'intero batch) — e ogni card "light" due volte (D.2 + un D.4b che a `light` esegue solo code-reviewer + FP, puro doppione di D.2). Un nuovo step **D.1.5** calcola UNA volta il profilo effettivo per card (floor da `review_profile` + escalation-only dal detector Step-A sul diff reale) e partiziona il gruppo in `LIGHT_CARDS` / `FULL_CARDS`. **D.2** ora esegue il code-reviewer solo sull'union dei diff `LIGHT_CARDS` (skippato del tutto se non ce ne sono); **D.4b** gira `/codexreview` solo per i `FULL_CARDS`, mentre i `LIGHT_CARDS` lo saltano (la loro review adversarial è garantita dalla Final FULL gate, sempre on). Esito: ogni card riceve **una sola** code-review per-card/gruppo proporzionale al rischio + la Final batch gate — light: D.2 + Final; full: D.4b + Final (giù da 3). **D.3b Simplify** salta sulle card `review_profile == skip` (quality-only, niente diff sostanziale). Tutti gli skip sono **gate-reason enumerati e deterministici** (driven dal campo `review_profile` autorato in PRD), non skip "per time budget": una nuova clausola "REVIEW DEPTH SCALES WITH `review_profile`" in testa alla skill lo formalizza. Eliminato anche il polling `sleep N; echo "waiting..."` ad ogni barrier (i background agent re-invocano l'orchestratore automaticamente). **MINOR** (cambia il comportamento di review di `/new` team mode senza rimuovere agenti o rompere install; nessuna chiave `baldart.config.yml` nuova — si appoggia su `review_profile` già esistente nella card).

> **Why.** Il team mode parallelizza la parte economica (scrivere il codice) e serializzava la parte costosa (la coda di review): i coder partivano in parallelo, poi tutto convergeva in un imbuto sequenziale di 11 sotto-passi con la code-review ripetuta. La de-duplicazione si appoggia su due invarianti già presenti: (1) la **Final-review FULL gate** (dal v3.37.0) gira SEMPRE Codex (+ code-reviewer fallback) sull'intero diff di batch, senza riduzioni — quindi ogni riga di ogni card light riceve comunque Codex-adversarial prima del merge; (2) l'**escalation-only** del detector Step-A (`light`→`full`, mai l'inverso) garantisce che un path rischioso emerso a implementazione promuova comunque la card a `full`. Tradeoff accettato (scelta utente "aggressivo 3→2"): per i `FULL_CARDS` il rilevamento cross-card si sposta dal per-gruppo (D.2) al per-batch (Final F.3) — stessa sicurezza di merge-gate, feedback a fine batch invece che per-gruppo. `/codexreview` NON è toccato (il suo FP-gate dipende architetturalmente da `code-reviewer`, quindi la leva è l'invocazione condizionale di D.4b, non un flag interno). Il **percorso sequenziale è invariato** (non ha un code-reviewer di gruppo, quindi il suo `/codexreview` a Phase 3.7 resta l'unica code-review per-card). Allineato al principio "data-driven sopra threshold arbitrarie": nessuna soglia nuova, si onora un campo deterministico già esistente. AC-Closure (D.3a), D.4 QA e D.4a doc restano invariati.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — (a) nuova clausola in testa "REVIEW DEPTH SCALES WITH `review_profile`" che enumera il profilo come gate-reason deterministico (distinto dallo skip "per time"); la clausola "NO PHASE SKIP" cita il profilo tra le ragioni valide e ammorbidisce "never reduce" in "never reduce for any model-invented reason". (b) Nuovo **D.1.5** (Effective per-card review profile) che calcola la partizione `LIGHT_CARDS`/`FULL_CARDS` una volta sola e la logga. (c) **D.2** code-reviewer scoped a `LIGHT_CARDS` (skip se vuoto); doc-reviewer invariato su tutto il gruppo; tradeoff cross-card→Final documentato inline. (d) **D.3b** salta su `review_profile == skip`. (e) **D.4b** condizionale: `/codexreview` solo per `FULL_CARDS` (profilo sempre `full`, niente recompute), `LIGHT_CARDS` skip con reason enumerata. (f) **Step D coverage assertion** accetta gli skip profilo-driven come entry valide documentate. (g) **Step C** direttiva anti-polling (no `sleep N; echo waiting`). (h) **D.3b Simplify** ora fa **fan-out in parallelo** su tutte le card eleggibili del gruppo (analisi read-only su diff file-disgiunti → niente più coda sequenziale per-card; i fix si applicano dopo, lint/tsc una volta sola); **D.3c E2E** valuta il gate per tutte le card insieme e parallelizza le eleggibili quando colpiscono route disgiunte (fallback sequenziale su contesa del dev-server). Percorso sequenziale, `codexreview.md`, AC-Closure D.3a (resta sequenziale per via dell'`AskUserQuestion`), D.4 QA, D.4a doc, Final Review F.3: non toccati.

> **Nota sullo scope dei tagli.** La de-dup si ferma qui di proposito. La **Final Review F.3** ri-revisiona l'intero batch (Codex + doc + perf + qa-sentinel suite completa) anche dopo il lavoro per-gruppo: è ridondante su batch a gruppo singolo MA è la rete di sicurezza incondizionata su cui ora si appoggiano le card light, e il framework ha **già tentato di ridurne lo scope in v3.35.0 e l'ha rollbackata in v3.37.0** dopo un incidente — quindi resta piena. Stessa logica per il doppio run della suite di test (D.4 per-gruppo + Final). Questi NON sono toccati.

## [4.4.0] - 2026-06-03

**`/prd` mockup intake ora risolve le schermate di una feature in modo *strutturale* da un bundle Claude Design, invece di tirare a indovinare con lo scoring per keyword.** L'export di Claude Design non è controllabile dall'utente: si scarica solo il bundle intero del workspace multi-feature. Finora Step 1.6.4b gestiva il bundle `.zip` con un solo scoring di rilevanza su filename/`<title>` (pieno di `[CALIBRATION-NEEDED]`), e la sola modalità d'ingresso era l'archivio `.zip`. Ora un **passaggio strutturale** (Step b2) sfrutta i segnali nativi che il bundle già porta: **un canvas `*.html` per feature** (anchor sul titolo feature), **slot `id="s<N>"` = screen-ID** del § 3 del prompt di handoff (`s31` ↔ schermata `3.1`), e il manifest nativo **`.design-canvas.state.json`** (arricchimento label, mai autorità). Le schermate che vivono solo come artboard nel canvas (non esportate come PNG) vengono **renderizzate deterministicamente** via http-server + Playwright (nuova `canvas-intake-recipe.md`), risolvendo il CORS che blocca Babel su `file://`. Il trigger è allargato anche alle **cartelle già estratte** (non solo `.zip`). A valle, la routing Hybrid di Step 3 passa da fuzzy substring match a **exact join** su `slot_id`. Lo scoring per keyword resta come **fallback** (bundle generico o canvas ambiguo) e il **gate di conferma STOP umano è invariato**. **MINOR** (nuovo reference `canvas-intake-recipe.md` + nuovi campi di state-schema `mockups.archive.bundle_kind`/`canvas_ref`/`discarded[]`, `mockups.canvas_render`, `screens_in_scope[].slot_id`; trigger allargato; NON una chiave `baldart.config.yml` → nessuna cascade di schema-propagation).

> **Why.** Il mapping 1-a-1 tra schermate prodotte da Claude Design e schermate del PRD non va *imposto* nell'handoff (l'export è monolitico e fuori controllo): **esiste già** nella struttura del bundle, semplicemente il framework non la leggeva. Spostare il peso dall'handoff (dove non si può fare nulla) all'intake (dove i segnali sono tutti presenti) elimina l'archeologia manuale — CORS, http-server e Playwright re-improvvisati a ogni rientro — e la sostituisce con un resolver ripetibile. Contratto fail-safe in ogni direzione: firma assente / slot non-`s<N>` / manifest parziale / feature su 2 canvas / Playwright assente / porta occupata ricadono tutti sul comportamento attuale (scoring + STOP) — **il PRD non aborta mai sull'intake**. Allineato al principio "data-driven sopra threshold arbitrarie": lo scoring `[CALIBRATION-NEEDED]` diventa il ripiego, non la prima linea.

### Added

- **`framework/.claude/skills/prd/references/canvas-intake-recipe.md`** — ricetta archiviata per renderizzare gli artboard canvas-only di un bundle Claude Design in PNG: preflight (no auto-install), http-server con port-loop 8080→8090, render loop Playwright su `[data-dc-section]`/`[data-dc-slot]` (attributi runtime, non statici), dedup vs pre-render del designer, tabella di degradazione. Invocata da Step 1.6.4b/A5, mai standalone.

### Changed

- **`framework/.claude/skills/prd/references/discovery-phase.md`** — Step 1.6.4b: (a) Step a allargato per accettare una **directory** oltre al `.zip` (`mockups.archive.kind: zip | directory`, convergono su `root`); (b) nuovo **Step b2 — Claude-Design structural resolution** (A0 signature gate → A1 canvas anchor con `AskUserQuestion` su ambiguità → A2 estrazione slot `id="s<N>"` → screen-ID → A3 read manifest enrich-only → A4 sidecar jsx/PNG → A5 render via recipe → A6 discard-and-log) inserito tra enumerate e score; (c) Step c demoto a **fallback** esplicito; (d) Step d con variante tabella *screen-centric*; (e) Step e materializza anche jsx-deps + asset condivisi + PNG renderizzati e scrive `screens_in_scope[]` con `slot_id`/`screen_id`/`canvas_ref`. Appendice schema estesa coerentemente.
- **`framework/.claude/skills/prd/references/ui-design-phase.md`** — routing Hybrid (Step 2 della per-screen table): **exact-join-first** su `slot_id`/`screen_id` quando presenti (no STOP), fuzzy substring match conservato come ramo `else` con il suo STOP-and-ask invariato.
- **`framework/.claude/skills/prd/SKILL.md`** — HARD RULE 16: lead-in allargato a "`.zip` OPPURE cartella estratta di un bundle Claude Design (firma `.design-canvas.state.json` e/o `*.html` con slot `id="s<N>"`)" + frase su passaggio strutturale, scoring come fallback, STOP invariato, no-abort.
- **`framework/.claude/skills/prd/assets/state-template.md`** — `mockups.archive` estesa con `bundle_kind`/`canvas_ref`/`canvas_candidates[]`/`discarded[]`, nuovo blocco `mockups.canvas_render`, `kind` ora `zip | directory`; `screens_in_scope[]` con `slot_id`/`screen_id`/`canvas_ref`.
- **`framework/.claude/skills/prd/assets/claude-design-handoff-prompt.template.md`** — nudge minimo a fine § 3 (titola il canvas col nome feature; numera gli artboard coerenti con gli screen-ID): unica richiesta realistica dato che l'export non è controllabile.

## [4.3.0] - 2026-06-03

**`/new` skips its pre-flight cross-card Codex check when `/prd` already audited the same batch jointly and nothing drifted since.** The `/prd` Step 6.6 holistic audit (4-agent team + Codex adversarial pass, 6.6d) already reviews the full card set for `files_likely_touched` conflicts, dependency shadows, implicit ordering, and shared-state mutation — the *exact* lens of `/new` Step 3d's cross-card check. Running Step 3d unconditionally re-paid that Codex call on every multi-card batch, even when implementing the freshly-audited PRD set with zero intervening change. Now the audit phase **stamps provenance** on each card (`metadata.holistic_audit`: `audited_at` + `audited_commit` + `audited_set`), and `/new` Step 3d evaluates a four-condition skip: provenance present (S1), batch jointly audited (S2), batch ⊆ audited set (S3), and no commit touched the batch's claimed paths since the audit baseline (S4). All four hold → SKIP (logged, non-decision); any fails → RUN as before. **MINOR** (new card-template field + new conditional capability; NOT a `baldart.config.yml` key — `metadata.holistic_audit` is a backlog-card field written by the audit phase).

> **Why.** The two checks are not "different lenses" — Step 3d's `FILE_CONFLICT / IMPLICIT_DEP / ORDER_RISK / STATE_MUTATION` questions map 1:1 onto 6.6d's `attack_surface`, run over the same card set + adjacent cards. The only residual value of re-running at implementation time is the two cases the authoring-time audit *cannot* have covered: (a) the `/new` batch is not the audited set (a subset, or cards stitched from separate PRD runs), or (b) the codebase drifted on the claimed paths between audit and implementation. The gate isolates exactly those two cases via provenance + a `git log <audited_commit>..<trunk> -- <claimed_paths>` drift probe, and **degrades safely in every direction** — missing/partial provenance, mismatched sets, empty/unreachable commit, or any git error all fall through to RUN, never to a silent skip. Aligns with the standing "data-driven over arbitrary thresholds" and "no silent confirmation gates for non-decisions" principles: the skip is deterministic and auditable (the RUN/SKIP reason is logged in the tracker), not a heuristic.

### Added

- **`framework/templates/feature-card.template.yml`** — documents the optional `metadata.holistic_audit` block (`audited_at` / `audited_commit` / `audited_set`), written by the audit phase (not at authoring), consumed by `/new` Step 3d.

### Changed

- **`framework/.claude/skills/prd/references/audit-phase.md`** — Step 6.9 (Apply Findings to Cards) adds a "Holistic-audit provenance stamp" subsection: capture three run-level values once (`audited_at` = UTC now, `audited_commit` = `git merge-base HEAD <trunk>` fork point, `audited_set` = sorted jointly-audited card IDs) and write `metadata.holistic_audit` on each card. New per-card workflow step (was 5 steps, now 6). Fail-safe contract: the stamp is an optimization hint, never a gate — never blocks the PRD if it can't be written.
- **`framework/.claude/skills/new/SKILL.md`** — Step 3d replaces the flat "Skip if batch has only 1 card" with a provenance-aware four-condition skip decision (S1 provenance present, S2 jointly audited, S3 batch ⊆ audited set, S4 no drift on claimed paths via `git log <audited_commit>..<trunk> -- <paths>`). SKIP logs `SKIPPED (provenance) — …` and proceeds with the file-ownership map unchanged; otherwise RUN with the reason recorded. Tracker `## Cross-Card Conflicts (Codex)` placeholder updated to note the possible skip. Every ambiguity falls through to RUN.

## [4.2.2] - 2026-06-03

**`/new` Phase 0 stops blocking on framework-owned noise: telemetry-only dirty trees and `baldart update`/subtree commits are auto-ignored.** Every batch tripped its own Phase 0 gates: (1) the dirty-tree gate fired on `${paths.metrics}/archive/batch-tracker-*.md` + `skill-runs.jsonl` + `sessions/` — telemetry that Phase 8 writes *after* the merge and never commits, so the next batch always saw a "dirty" tree and forced a stash/commit decision; (2) the divergence gate fired on `.framework/` subtree commits + the `[CHORE] baldart update …` chore, treating expected framework-management divergence as the FEAT-0006 orphan-commit pattern and forcing a push/cherry-pick decision. Both gates now **partition the working tree / the ahead-commits before prompting** and surface only genuine in-progress user work. The FEAT-0006 protection is intact — real orphan commits and real uncommitted work still gate. **PATCH** (bugfix to an existing gate's classification; no new agent/skill/template, no `baldart.config.yml` key — `paths.metrics` already exists).

> **Why.** Phase 0 exists because of a real incident (FEAT-0006: an unpushed orphan commit + diverged trunk caused data loss), so the gates can't be removed. But they were classifying *framework-generated artifacts as if they were user state* — the skill's own Phase 8 telemetry and the `baldart update` subtree commits — and raising a non-decision to the user on every single `/new` run. The fix is surgical classification, not relaxation: the dirty-tree gate filters paths under `$METRICS` (resolved from `paths.metrics`) and gates only on the remainder; the divergence gate reclassifies each ahead commit (subtree merge into `.framework/`, `baldart update`/`add` chore, or `.framework/`-only diff) as framework-management and counts only genuine unpushed commits. Phase 6c's batch-close assertion gets the same reclassification so the noise doesn't re-appear at the end. `/prd` and `/nw` were already lean (read-only pre-flight, branch from `origin/$TRUNK`) and needed no change. Aligns with the standing "respect explicit user intent — no silent confirmation gates for non-decisions" principle.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — Phase 0 step 0 now resolves `$METRICS` (`paths.metrics`, default `docs/metrics`). Step 3 (dirty-tree gate) partitions `git status --porcelain`, dropping framework telemetry under `$METRICS`; gates only on `$USER_DIRTY`, otherwise logs `Dirty: framework-telemetry-only (auto-ignored)` and proceeds with no `AskUserQuestion`. Step 4 (divergence gate) and Phase 6c step 4 (batch-close divergence assertion) reclassify ahead commits, count only `$GENUINE_AHEAD`, auto-ignore `$FW_SKIPPED` framework-management commits, and branch on `$BEHIND`/`$GENUINE_AHEAD`. Step 5 log template extended with the two new auto-ignored states.

## [4.2.1] - 2026-06-03

**`/prd` pre-dev audit: `code-reviewer` swapped for `plan-auditor` (card grounding) in the `check-audit` team.** The Step 6 pre-development audit ran `code-reviewer` on backlog `.yml` cards — but `code-reviewer` is built around a *diff* (scope = `git diff --name-only`, diff-simulation pass, completion-report cross-check, line-range/`repro_steps` findings schema, `PASS/FAIL/NEEDS_REWORK` verdict), none of which applies before any code exists. The audit had to re-purpose it with an ad-hoc mandate and override its native output format. `plan-auditor` is the agent purpose-built for this surface (native `.yml` INPUT TYPE DETECTION + BACKLOG CARD ATTACK SURFACE + native `[Target: <field>]` tagging + project memory), so it now fills that slot. **PATCH** (refinement of an existing capability — swaps which already-shipped agent audits cards; no new agent/skill/template, no `baldart.config.yml` key).

> **Why.** Step 6.6d already replaced `plan-auditor`'s *plan-logic* role with a cross-model Codex adversarial audit, and kept `code-reviewer` as the Claude-side reviewer. But `code-reviewer`'s repurposed mandate (conflicts with existing patterns, reuse gaps, adjacent-card file conflicts) overlapped ~90% with the Codex `attack_surface` plus the Step 6.4 adjacent-card retrieval already fed to every agent. The only non-overlapping sliver — project memory of recurring code-level pitfalls — is covered *better* by `plan-auditor` (which has its own `.claude/agent-memory/plan-auditor/`). The swap aligns fit without losing coverage; the Maintenance Note's ablation rubric flagged `code-reviewer`-on-cards as the prime removal candidate. `code-reviewer` is unchanged and remains the agent for post-dev review (`/new` Phase 3, `/codexreview`).

### Changed

- **`framework/.claude/skills/prd/references/audit-phase.md`** — Step 6.6c agent-type mapping swaps the `code-reviewer` row for `plan-auditor` (card grounding, Always). New 6.6e agent-specific block runs `plan-auditor` with **full card-attack-surface coverage but team-aware spawn suppression**: no `codebase-architect` re-spawn (`/prd` already primed context at kickoff), no specialist auto-spawn (`security-reviewer` / `api-perf-cost-auditor` run as first-class teammates in the same team), and an output-format override to the teammate contract (`### [CARD-ID] — Plan Audit Findings` + `## FINDINGS` via `TaskUpdate`) — `plan-auditor`'s never-suppress list (`git_strategy: TBD`, missing-auth, `claimed_path_collision`, `adr_required_missing`, injection) still binds.
- **Codex fallback dedup (Step 6.6d)** — the Codex-unavailable fallback now spawns a **FULL-mode** `plan-auditor` and explicitly SKIPS the card-grounding `plan-auditor` teammate, preventing two `plan-auditor` instances on the same cards.
- **Consolidated report + Audit Engine Summary (Step 6.7)** — `### Code Review Findings` → `### Plan Audit Findings (card grounding — Claude-side)`; the engine summary now distinguishes the cross-model Codex pass from the Claude-side card-grounding pass.

## [4.2.0] - 2026-06-03

**Per-skill reasoning effort: every shipped skill now declares an `effort:` baseline + honors an inline `effort=<level>` override.** Different skills should reason at different depths — routine narration shouldn't pay for deep deliberation, and structured planning shouldn't be starved of it. Each of the 29 skills gets a frontmatter `effort:` baseline (`prd` → `high`, the other 28 → `medium`) and a `## Effort` body block letting the user dial reasoning depth per-run, e.g. `/prd effort=high <content>`. **MINOR** (new protocol module + new template = capability added; no `baldart.config.yml` key, so the schema-change propagation rule does not apply).

> **Why.** Claude Code's native `effort:` frontmatter is the right knob for the *baseline*, but its precedence puts frontmatter **above** the session `/effort` level — only the global `CLAUDE_CODE_EFFORT_LEVEL` env var beats it, and a running skill cannot re-set the knob itself ([anthropics/claude-code#47275](https://github.com/anthropics/claude-code/issues/47275)). So a naive "default + inline flag that overrides" is impossible with the native knob alone. The inline override is therefore implemented as a **body-level reasoning escalation**: the skill reads `effort=<level>` from its arguments and engages the matching thinking directive (`think hard` → `ultrathink`), which the harness interprets from the text. Fully portable — no env var, no consumer config. Escalation upward is reliable; de-escalation below the frontmatter baseline is best-effort (the body can skip optional deliberation but can't lower the native budget).

### Added

- **`framework/agents/effort-protocol.md`** (new protocol module, #21) — the textual SSOT: native-knob precedence chain, tiering rubric for choosing a baseline, the `effort=<level>` inline-override parsing contract (detect once at kickoff, strip the token before consuming user input — modeled on the existing `-stats` flag in `prd`), and the level→behavior mapping (`low` → no extended thinking; `high`/`xhigh`/`max` → `think hard`/`think harder`/`ultrathink`). Routed from `framework/agents/index.md`.
- **`framework/templates/skill-effort.snippet.md`** (new template) — copy-paste starter for the `effort:` frontmatter key + the `## Effort` body block, twin of `skill-project-context.snippet.md`.
- **`effort:` frontmatter + `## Effort` body block on all 29 skills** — `prd` is `high`; `new`, `bug`, `simplify`, `frontend-design`, `ui-design`, `design-system-init`, `e2e-review`, `issue-review`, and the remaining 19 are `medium`. Any `medium` skill can be pushed to `high`/`xhigh`/`max` per-run via the inline override.

### Changed

- **`skill-creator`** now instructs authors to set the `effort:` baseline (per the tiering rubric) and to paste the `## Effort` block for every new skill — the previous "do not include any other fields in YAML frontmatter" guidance is updated to permit `effort`.

### Out of scope (future)

- **Agents** (`framework/.claude/agents/*.md`) also support `effort:` in frontmatter, but they receive no inline arguments (they are spawned by orchestrators via the Task tool). Their override story — static baseline per agent + orchestrator passing the desired level in the Task prompt — is deferred to a separate release to avoid mixing two override models.

## [4.1.1] - 2026-06-02

**Fix: `/prd` (and `/new`) silently stopped creating their worktree after updating to ≥4.0.0.** Regression introduced in v4.0.0: `worktree-manager` was rewritten to resolve the worktree base from the **new `git.trunk_branch` config key** and to **hard-`exit 1`** when it is absent ("never assume develop"). Because `baldart update` never overwrites `baldart.config.yml`, every consumer that updated without re-running `configure` had no `git.trunk_branch` → the `nw-docs` / `nw` pre-flight aborted → no worktree. v4.0.0 also downgraded the `/prd` Step 1 worktree gate from BLOCKING to a non-blocking "MONITORING SIGNAL", so the abort no longer surfaced as a hard STOP — the session drifted into Discovery on the main checkout instead. **PATCH** (regression fix, no new capability, no schema change).

> **Why.** Before v4.0.0 the worktree base was hardcoded `origin/develop`, so creation always worked. Making it config-driven was correct, but the missing-key path should degrade to autodetection, not a hard failure — and the safety gate that would have caught the failure was removed in the same release. The fix resolves the trunk the same way `configure.js` already does (`origin/HEAD` → local `develop`/`main`/`master`), which is *not* "assume develop": it resolves the branch the repo actually uses.

### Fixed

- **`worktree-manager` `nw-docs` + `nw` pre-flight** — when `git.trunk_branch` is absent from `baldart.config.yml`, autodetect the real default branch (`git symbolic-ref refs/remotes/origin/HEAD` → first existing local `develop`/`main`/`master`) and warn, instead of `exit 1`. Hard-abort only when even autodetection finds nothing (bare repo / no branches). This unblocks `/prd` and `/new` worktree creation on consumers that updated without re-running `configure`.
- **`/prd` Step 1 worktree gate** (`prd/references/discovery-phase.md`) — restored from "MONITORING SIGNAL (non-blocking)" back to a **BLOCKING gate**: the kickoff "Ho creato la sessione PRD … in un worktree dedicato" message is emitted only after the worktree is verified on disk; if `nw-docs` aborted, HALT and surface the error verbatim instead of proceeding to Discovery on the main checkout (HARD RULE 17).
- **`/new` Phase 0 trunk resolution** (`new/SKILL.md`) — when `git.trunk_branch` is absent, `/new` now autodetects the trunk with the **same** `origin/HEAD` → `develop`/`main`/`master` logic as `worktree-manager`, instead of hard-assuming the `develop` literal. Keeps `/new`'s `$TRUNK` (used by the divergence gate and every `git diff "$TRUNK...HEAD"` gate) in lockstep with the branch `nw` actually bases the worktree on — a `main`-trunk repo no longer diverges between the two.
- **`update.js` schema-drift notice** — when the user declines `configure` and `git.trunk_branch` is among the missing keys, the CLI now states that `/prd` and `/new` will autodetect the trunk and warn until the key is set, replacing the previous (inaccurate for this key) "skills will prompt for the missing keys on first use" reassurance.

## [4.1.0] - 2026-06-02

**Portability: framework-wide path-literal de-contamination (`docs/…`, `backlog/`, `src/components/…` → `${paths.*}`) + new `paths.docs_dir` config key.** Closes the open follow-up forward-referenced in v4.0.3: a dedicated, per-occurrence judgment pass over the path-literals `contamination.js` flags across `framework/`. The scanner's `autofixable` count drops **219 → 49**; the 49 survivors are legitimate and intentionally kept. The new `paths.docs_dir` config key makes this **MINOR**.

> **Why.** A hardcoded path in a shipped agent/skill is a frozen spot — a consumer whose docs live at `documentation/` (not `docs/`) gets a wrong directive. But not every literal is contamination: `${paths.*}` is resolved *at runtime* by the skill/agent that reads it, so a token only belongs where the file is an operative instruction. A literal correctly stays literal when it is a `# e.g.` config-default comment (converting is circular), a clickable markdown link (a token breaks the link), a concrete dated worked-example (`ADR-20260115-…` — the literal *is* the teaching), a negative instruction (`do not assume src/components/shared/`), or an overlay example (overlays are verbatim, never runtime-resolved). The judgment was made per-occurrence by a workflow (one judge per file → an adversarial bidirectional verifier) and applied deterministically through the `contamination.js` rule SSOT scoped to approved lines. The adversarial verifier also surfaced contamination the per-file judges missed — including the identity leak below.

### Added

- **`paths.docs_dir`** — umbrella docs-search root, used by the `rg`-fallback in `code-reviewer`, `senior-researcher`, and `worktree-manager` (where bare `docs/` had no config key). Propagated end-to-end per the schema-change rule: `baldart.config.template.yml` + `configure.js` (autodetection + prompt) + `project-context.md` autodetection table + `PROJECT-CONFIGURATION.md` key table. `update.js` detects it automatically via the generic `missingPaths` diff, so existing consumers are offered `configure` on their next update.

### Changed — path-literals → `${paths.*}` across the framework payload (38 files changed, +162/−151)

- **Agents** — `REGISTRY.md`, `codebase-architect`, `coder`, `doc-reviewer`, `prd`, `prd-card-writer`, `api-perf-cost-auditor`, `code-reviewer`, `plan-auditor`, `qa-sentinel`, `security-reviewer`, `senior-researcher`, `skill-improver`, `visual-designer`, `motion-expert`, `wiki-curator`, `hybrid-ml-architect`, `legal-counsel-gdpr`: operative read/write/search paths now resolve `${paths.design_system | references_dir | prd_dir | adrs_dir | wiki_dir | backlog_dir | components_* | docs_dir}`.
- **Skills / commands** — `new`, `prd`, `bug`, `simplify`, `worktree-manager`, `check`, `codexreview`, `qa`, prd reference phases, and the prd asset templates (`card-template.yml`, `epic-template.yml`, `prd-template.md` — resolved by `prd-card-writer` at fill time).
- **Routines** — `ds-drift`, `wiki-review`.
- **`bug/SKILL.md`** — Project Context header now declares `paths.wiki_dir` (companion to the converted Phase-0 doc-rag fallback, keeping declared config deps in sync with the body).

### Fixed

- **Identity leak** — `prd-template.md` hardcoded `**Owner**: Antonio Baldassarre`, shipped into every consumer's PRD document; now the `{{Owner}}` placeholder.

### Kept literal (intentional)

- **`breaking-change-checklist.md`** — a human-read user-edit template (not agent-resolved), where a `${paths.*}` token would be inert text; left literal by design.
- **Opt-out marker** (`<!-- contamination-scan: skip -->`) added to 3 wall-to-wall pedagogical files so the scanner stops re-flagging intentional teaching literals: `doc-writing-for-rag/references/line-count-targets.md`, `schemas-and-errors.md`, `templates/overlays/commands/codexreview.example.md`.
- Config-default `# e.g.` comments (`baldart.config.template.yml`), worked examples, clickable links, and documented defaults across `design-system-init`, `prd/SKILL`, `design-review`, and others.

## [4.0.4] - 2026-06-02

**Portability fix (visual identity): genericize the last fidelity-app visual-identity flavor in doc-RAG worked examples.** Confirms the graphic dimension is clean: `Neo-Brutalism` (plus the `Press Start 2P` font and `amber accents`) appeared as bare styling labels in the `doc-writing-for-rag` before/after examples — now generalized to "the project's design system (per `identity.design_philosophy`)". No behaviour change → **PATCH**.

> **Why / scope confirmation.** After this, project-specific visual identity exists in the shipped payload ONLY where it should: as a config-driven example value (`identity.design_philosophy` e.g. "Neo-Brutalism" / "Minimalist" / "Glassmorphism"), in the config documentation (`PROJECT-CONFIGURATION.md`, template comments, the contamination scanner's own meta-docs), and in the intentional `templates/overlays/*.fidelity-example.md` starter overlays (named and shipped precisely so a consumer copies/adapts/deletes them). Zero project-specific design-philosophy, chart-wrapper, font, or brand-colour literal remains baked into portable skill/agent logic. (Hard-coded `${paths.*}` literals remain a separate, larger open cleanup — see v4.0.3 note.)

## [4.0.3] - 2026-06-02

**Portability fix (round 3 — final): remove fidelity-app DOMAIN example-flavor from shipped framework files.** Closes the T13 portability sweep started in v4.0.1/4.0.2 by genericizing the consumer-domain vocabulary (`merchant`/`booking`/`reservation`/`merchant-theming`, `/merchant/dashboard` routes, `customer-facing vs merchant-facing`, `FEAT-0127-merchant-points-guard`, `area:customer|merchant|admin`) that survived as illustrative example flavor. No behaviour change → **PATCH**.

> **Why.** Even in prose examples, a consumer-domain noun teaches the wrong vocabulary and trips the `contamination.js` `requires-decision` rules. The fix distinguishes genuine fidelity-domain flavor (generalized) from legitimate generic English (`customer` in marketing/security "multi-tenant … customers"), the scanner's own meta-references (`framework-edit-gate.js`, `contamination.js`), and `project-context.md`'s explanation that these opinionated tokens belong in the overlay — all of which are correct and untouched.

### Changed — genericized domain example flavor

- **Example queries/routes** → placeholders: `wiki-curator.md` / `doc-reviewer.md` search examples (`<feature-X>` / `<entity>`); `visual-fidelity-verifier.md`, `commands/design-review.md`, `e2e-review/SKILL.md` route envelopes (`/dashboard` instead of `/merchant/dashboard`); `impact-analysis.md` worked-CR endpoint.
- **`prd.md`** — example card filename and `area:` labels reference `identity.audience_segments` instead of literal `merchant`/`customer`/`admin`.
- **`prd-card-writer.md`** — "split by audience segment (per `identity.audience_segments`)" instead of "customer-facing vs merchant-facing"; generic subsystem example.
- **`doc-writing-for-rag` reference files** (`before-after-examples.md`, `compact-templates.md`, `schemas-and-errors.md`) — added an "illustrative sample domain" disclaimer (the worked before/after corpus uses a sample API; only the *technique* is normative), rather than rewriting 290+ lines of self-contained illustration.

> **Note — out of scope, surfaced for a future pass.** Running `contamination.js` over the whole `framework/` payload also flags ~200 hard-coded **path literals** (`backlog/`, `docs/design-system/`, `src/components/…`) that should be `${paths.*}` config keys. A meaningful fraction are legitimate (config-template defaults, `PROJECT-CONFIGURATION.md` docs, canonical protocol examples), but the rest are a genuine, larger portability cleanup deserving its own dedicated pass — not bundled here.

## [4.0.2] - 2026-06-02

**Portability fix (round 2): remove hard-coded auth/deployment project identifiers from shipped framework files.** Companion to v4.0.1 — extends the de-contamination from the database dimension to the auth and deployment dimensions, where fidelity-app-specific identifiers (`withAuth`, `withAuthNoParams`, `src/lib/auth/middleware.ts`, `src/app/api/v1/...`, `BookingTable`, `ADMIN + MERCHANT`, "Vercel Functions") leaked into self-verification examples, detection-signal lists, and template fills. No behaviour change → **PATCH**.

> **Why.** A project-specific symbol baked into a shipped detection list or CoVe example is the same frozen-spot defect as a hard-coded datastore: the consumer's auth wrapper, role names, and deploy platform are `stack.*` / overlay facts. Genuinely platform-named integrations were left untouched because they are correct: the `vercel:deploy` skill documentation (that skill *is* Vercel-specific), the multi-platform deploy tables in `/new` ("firebase: … vercel: … aws: …"), the per-platform timeout/limit tables in `api-perf-gate.md`, and the config-driven `stack_signature` example in `research-phase.md` (which already shows both a Firestore and a Supabase example).

### Changed — generalized auth/deployment identifiers to placeholders / config keys

- **`code-reviewer.md`** + **`plan-auditor.md`** — the chain-of-verification example findings now use `<auth-wrapper>` / `<route-file>` / `<auth-module>` placeholders (resolved from `${paths.high_risk_modules}`) instead of `withAuth` + literal `src/app/api/v1/...` / `src/lib/auth/middleware.ts` paths.
- **`audit-phase.md`** — the auth-change detection signal cites "the project's auth-guard wrapper / permission helper (e.g. a `withAuth*`-style wrapper)" rather than the literal `withAuth` / `checkPermission` symbols.
- **`doc-writing-for-rag`** (SKILL + `compact-templates.md`) — endpoint template fills use `<authWrapper>` / `<ROLE_A>` / `<ROLE_B>` instead of `withAuthNoParams` / `ADMIN + MERCHANT`.
- **`api-perf-gate.md`** — the upload red-flag cites "the platform's request-body limit (e.g. 4.5MB on Vercel Functions — per `stack.deployment`)".
- **`validation-phase.md`** — the logic-marker grep list covers datastore-access patterns per `stack.database` (not only `prisma` / `firestore.collection`).
- **`wiki-curator.md`**, **`doc-reviewer.md`** — illustrative example queries/paths use `<authWrapper>` / `<DomainType>` / `src/lib/<module>.ts#<symbol>` placeholders.

## [4.0.1] - 2026-06-02

**Portability fix: remove hard-coded `Firestore` assumptions from shipped framework files (T13 follow-up).** The v4.0.0 portability pass guarded most stack-specific checks behind `stack.database`, but a residual set of agent checklists and methodology blocks still *assumed* Firestore as the datastore — contamination that is dead-false (or misleading) on any non-Firestore consumer. No behaviour change for correctly-configured projects → **PATCH**.

> **Why.** A hard-coded datastore literal inside portable framework logic is a frozen-spot where the framework requires a hot-spot (Pree's taxonomy): the datastore is a per-consumer `stack.database` fact that belongs in config/overlay, not baked into a shipped checklist. The genuinely multi-DB surfaces were already correct and are untouched: the gated per-DB blocks in `plan-auditor.md` (§ Persistence-Specific) and `api-perf-gate.md` (§ Stack-specific addenda + per-DB pricing snapshots), the per-`stack.database` switch arms, and the `/new` Production-Readiness example (which already carries an explicit "adapt to your stack" note + multi-DB equivalents).

### Changed — generalized assumed-Firestore checklists to per-`stack.database`

- **`plan-auditor.md`** — 8 general-checklist items (dependency enumeration, transaction/batch concurrency, quota/rate-limiting, N+1, state-machine, async-propagation clock, the never-suppress exception list, common-missing-indexes) now say "datastore … per `stack.database`" instead of naming Firestore. The gated § Persistence-Specific per-DB block is unchanged.
- **`code-reviewer.md`**, **`api-design-principles`**, **`doc-reviewer.md`**, **`prd-card-writer.md`**, **`prd-template.md`** — review/cost/invariant checklist lines and template placeholders generalized to "datastore … per `stack.database`".
- **`api-perf-cost-auditor.md`** — the MANDATORY Load-Simulation methodology ("count exact reads/sec", quota/hot-partition ceilings, tail-latency, never-demote anti-patterns) is now datastore-neutral with Firestore kept only as an inline example; the Firestore-specific hot-doc 1-write/s limit is generalized to "hot-partition / hot-document write limits".
- **`bug` skill** + **`logging-patterns.md`** — the data-bug triage row, the data-debug tooling block, and the "Datastore-Specific Debugging" section (formerly "Firestore-Specific") are now gated on `stack.database` with Firestore + the Firebase MCP shown as the example, and the equivalent primitives for Postgres/Supabase/Mongo named.

## [4.0.0] - 2026-06-01

**Framework-wide architectural alignment of `/prd` and `/new` (the two most-used skills) and their child agents/skills, driven by a full coherence audit benchmarked against the scientific literature.** A multi-agent audit mapped every phase of `/prd` and `/new` one-by-one and found **439 findings** (23 critical, 145 high) concentrated in **6 systemic fault lines** — the canonical one being exactly the bug that triggered this work: the `qa-sentinel` agent was *referenced everywhere but architecturally orphaned*, invoked by callers for capabilities its own system prompt forbids. A research team then benchmarked the 14 architectural themes against the peer-reviewed literature (multi-agent orchestration, automated/LLM code review, CI quality gates, self-repair loops, requirements traceability, observability), producing 8 north-star principles. This release aligns the framework to them. **MAJOR** because it narrows agent capability contracts, demotes a legacy entry point (`commands/new.md`) to a redirect stub, and adds config keys.

> **Why.** The single meta-finding across 8 of 14 themes: BALDART repeatedly enforced critical invariants by **social contract** (an LLM or maintainer is trusted to remember, sync, or self-police) where the literature unanimously requires a **technical contract** (a build-time check, a tool-grant boundary, a tracker precondition, an append-only ledger, or a config-key indirection that fails loudly). The fix extends patterns BALDART already ships correctly (the `framework-edit-gate` hook, `contamination.js`, the worktree stash ban, the AC-Closure Gate, the schema-change-propagation rule — all *validated as correct* by the research) to the surfaces that bypassed them. The research also **refuted three assumptions** the framework held: (1) a cheap per-card "light" review is *not* recovered by an expensive full-batch backstop (Lost-in-the-Middle, TACL 2024 — large batches miss centrally-positioned findings); (2) single-model self-FP-checking and same-vendor "model diversity" are *not* independent validation (Kim et al., ICML 2025 — two LLMs agree 60% of the time *when both err*); (3) LLM diff-simulation is *not* a high-leverage runtime-invariant detector (Chen et al., ICSE 2025 — 44% on intermediate state). Fabricated citations that had become load-bearing rationale (`Arora 2023`, `Devin 2025`, an internal "Anthropic Harness Design" doc, a non-existent `GPT-5.4` model id) were stripped framework-wide.

### BREAKING

- **Agent capability contracts reconciled (the qa-sentinel class).** `qa-sentinel` is now consistently a mechanical gate-runner (PASS/FAIL, no source reads, no severity taxonomy, no test-writing, no e2e) across **every** caller and routing doc — `REGISTRY.md`, `/qa`, `/codexreview` (it is no longer the false-positive "Behavior Validator", which it structurally cannot perform — a code-aware reviewer is), and `/new` Phase 3.5 / Phase 7. The same reconciliation hits `plan-auditor` (a real `QUICK` mode now exists and callers pass `mode: QUICK`), `code-reviewer` (capability row vs `bypassPermissions` reconciled to one truth), `ui-expert` (`model: opus` + `Can Edit Code: Yes` now match the REGISTRY and its real implementer role), and `security-reviewer` / `api-perf-cost-auditor` (YAML-mergeable teammate output schema; never spawned via the prohibited `general-purpose` type). Consumers with overlays on these agents should re-check their overlays.
- **`commands/new.md` is now a thin redirect stub.** The legacy slash-command copy of the `/new` pipeline (which silently omitted the BLOCKING Phase 0 hygiene, AC-Closure, Simplify, E2E, Pre-Merge Codex, post-merge hygiene, and Metrics phases, ran `code-reviewer` where the skill forbids it, and pointed at a non-existent `/qa/README.md`) is removed. `/new` now has exactly one source of truth: the `new` skill. `prd-add-phase.md` is likewise a redirect stub to the worktree-aware `prd-add` skill (the `/prd` auto-trigger now targets the canonical skill, not the stub).
- **New `baldart.config.yml` keys** (schema-change-propagation: template + `configure` autodetect + `update` detector): `git.trunk_branch` (replaces every hard-coded `develop`/`main` integration-trunk literal in the worktree flow; autodetected from `origin/HEAD`), `paths.high_risk_modules` (config-driven risk detectors — no project paths baked into shipped gate logic), `paths.metrics`, `paths.wiki_log`. Existing consumers are prompted to populate them on next `update`.

### Changed — single source of truth (no more divergent twins)

- **`review_profile`** decision criteria now live ONLY in `prd-card-writer.md` § Rule C (the SSOT); `/new`, `backlog-phase.md`, `card-template.yml`, `validation-phase.md`, `testing.md`, `qa.md` carry a pointer, not a duplicate table. The drifted `>15-file → DEEP` trigger is reconciled into the SSOT.
- **`owner_agent` enum** and a new **card `status` enum** are machine-readable SSOT blocks in `REGISTRY.md`; the inline copies and the `prd.md:538`-line-number citations are replaced with pointers. `execution_strategy` is uniformly `groups[]/level`; `git_strategy` is uniformly an object `{branch, base, target}` (a string broke `/new`'s worktree-slug derivation); the duplicate `impact-analysis.md` was deleted (one canonical copy under `prd/references/`).

### Changed — gates made real, worktree hygiene fixed (P0)

- Abolished "MANDATORY NON-BLOCKING" labels (relabelled to MONITORING SIGNAL or given an executable consequence); fixed vacuous post-commit `git diff --name-only HEAD` (now `HEAD~1..HEAD`); added `tsc --noEmit` as its own gate (guarded by `stack.language`); re-run static checks after code-mutating phases; Final Review BLOCKER findings now block the merge.
- Removed every `git stash` that ran **inside** a worktree (`refs/stash` is globally shared — the skill's own Safety Rule, now honoured everywhere); externalized `$MAIN`/`$TRUNK`/slug to the tracker with presence guards (consistent variable name + resolution across both `/prd` and `/new`); task-scoped scratch filenames (no date-only collisions).

### Changed — P1 hardening (Wave 6)

- Retry loops are explicit three-branch state machines (SUCCESS / STUCK-ESCALATE / ABORT) with error-fingerprint stuck-detection, a single cap (3), `git status --porcelain` before commit, and TIMED_OUT branches for background tasks. `codebase-architect` is deduplicated (reuse per-card baselines instead of re-spawning for the batch); fan-out has explicit completion barriers; `run_in_background` is never combined with `--wait`.
- Telemetry is instrument-at-source: every Phase-8 metric field has a named producer (no post-hoc prose parsing); JSONL is a single atomic append (no append-then-rewrite); `cycle_time_mins` is anchored to a strategy-independent event. Thresholds carry provenance tags; `COMPREHENSION_DIMENSION_COUNT` is a single named constant so the research-launch trigger fires at a stable point.

### Added

- **`scripts/check-reference-integrity.js`** — a build-time reference-integrity gate over the shipped `framework/` payload: every `subagent_type` dispatch must resolve to a real agent, the `general-purpose` fallback is banned, and a denylist regression-guards the dead references this release removed (fabricated citations, `/qa/README.md`, `obsidian-sync`). Runs in CI on every push/PR (`.github/workflows/check-reference-integrity.yml`).
- A new `agents/coding-standards.md § Reference-Aliasing Mutation Patterns` section (the SSOT that `coder.md` / `code-reviewer.md` / `new` already cited but that did not exist), and graceful-degradation for optional/unshipped dependencies (`huashu-design`, `doc-reviewer-support`) instead of MANDATORY reads of missing files.

> **Process note.** The audit, literature benchmark, alignment, a correctness review, the P1 pass, and a holistic `/prd`+`/new` review were each run as multi-agent workflows; every change was adversarially re-reviewed before landing (the holistic review caught a stale card-contract twin in `agents/prd.md` that the per-file passes missed). Two genuine security findings surfaced and were fixed as P0: an `xargs … git show HEAD:{}` injection vector and a hard-coded vulnerability baseline that created false conformance.

## [3.41.0] - 2026-06-01

**`qa-sentinel` is now profile-driven: `balanced` cards run SCOPED tests (only the tests related to the touched files), and the FULL suite is reserved for `deep` cards (or a `balanced` change whose diff reveals undeclared risk).** This is a **behavioral change to QA depth** — the default for an ordinary feature/refactor changes from "run the whole suite" to "run the related tests" — not an additive feature. **No new `baldart.config.yml` keys** — it reuses the existing `review_profile` field; scoped/smoke test selection is auto-probed at runtime from the project's test runner.

> **Why.** The `review_profile` (skip/light/balanced/deep) is computed deterministically upstream by `prd-card-writer` (Rule C) and passed to `qa-sentinel` by `/new` — but the agent **ignored it**. Its prompt literally said `Run FULL VALIDATION MODE` while the line below passed `QA profile: balanced`, and the agent re-decided QUICK-vs-FULL from a crude `>5 files` heuristic. Net effect: a 6-file UI-only `balanced` card triggered the entire regression suite. The fix makes the profile **authoritative**: it selects the tier; file *count* no longer escalates (the framework already treats file count as advisory-only — `prd-card-writer.md`). The only upgrade path is **risk drift** — the diff itself touching auth/permission/payment/schema/migration/API-contract — which catches a card the author under-classified. An adversarial review of the first design rejected three over-reaches kept out of this release: putting the executable tier table in `agents/testing.md` (the agent never reads it), a ghost `SMOKE` tier mapped to `light` (which never invokes the agent), and a `>15 files` escalation (re-introducing the very full-too-often problem).

### Changed — `qa-sentinel` honors `review_profile` (scoped-by-default)

- **[framework/.claude/agents/qa-sentinel.md](framework/.claude/agents/qa-sentinel.md)** — Step 0 + `Operating Modes`: tier is now selected from the `QA profile` in the prompt, not file count. MODE 1 is **SCOPED VALIDATION** (profile `balanced`): a stack-aware test-impact cascade (`jest --findRelatedTests` → `vitest related` → `pytest <paths>` → `--testPathPattern`, monorepo-workspace-aware) that runs only related tests and **never** falls through to the full suite. MODE 2 is **FULL VALIDATION** (profile `deep`, or `balanced` escalated by risk drift). Confidence rule added: a SCOPED run that finds no related tests for a *logic* change records `Tests: SKIP — no related tests` and caps confidence at ≤70% instead of reporting PASS 95%. The H2 headings are unchanged (preserves consumer agent overlays); only the section bodies were rewritten.
- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** — Phase 3.5 qa-sentinel prompt no longer says "FULL VALIDATION MODE"; it passes the tier contract (`balanced → SCOPED`, `deep → FULL`, risk-drift-only escalation) and derives the E2E depth from the profile (`deep` full+blocking, `balanced` `--grep` advisory). Duplicate step "22" renumbered. Team-mode **D.4** documents that a mixed group with any `deep` card runs FULL (max-profile intentionally overrides the per-card SCOPED tier).
- **[framework/.claude/commands/qa.md](framework/.claude/commands/qa.md)** — `/qa` now passes `QA profile` to qa-sentinel (so the agent never defaults to absent), and the LIGHT path no longer falls back to the full suite when the test pattern is unclear.

### Changed — documentation aligned

- **[framework/agents/testing.md](framework/agents/testing.md)** — new stack-neutral **Scope-aware Test Selection** section documenting the principle (balanced→scoped / deep→full / risk-drift→escalate), explicitly pointing at `qa-sentinel` as the executable SSOT and fencing TIA as heuristic (no coverage maps / AST / ML). Runner-specific commands deliberately kept out (Codex / non-Node consumers read this module).
- **[framework/agents/index.md](framework/agents/index.md)** + **[README.md](README.md)** — routing pointer and the qa-sentinel one-liner updated (`Quick/Full/Deep profiles` → `profile-driven: scoped-by-default, full on deep`).

## [3.40.0] - 2026-06-01

**Doc fixes surfaced by `/new`'s doc-review are now applied by the `doc-reviewer` itself, not delegated to `coder`.** The `doc` fix-domain owner changes from `coder` to `doc-reviewer` (write mode) across every `/new` site that previously spawned a coder for documentation: sequential Phase 3, team-mode D.4a, and the post-batch Final review. **No new `baldart.config.yml` keys** — this is a fix-routing change internal to `/new`.

> **Why.** When `/new`'s post-implementation doc-review found stale/missing docs, the remediation was routed to `coder` — the wrong agent on two counts. (1) It's *overkill*: the `doc-reviewer` already ran the audit and has the full doc context in-context; a coder spawn re-derives it from scratch. (2) It's *wrong*: the doc invariants the orchestrator must not break (freshness markers, linking protocol, frontmatter standard, tabular formatting, SSOT/registry coverage, dependency-topological order, SCIP/code refs) are encoded in the **`doc-reviewer`** system prompt, NOT the coder's — so doc fixes were going to the agent *least* equipped for them. The `doc-reviewer.md` Constraints already mandate "WRITE missing docs directly … do not defer to other agents"; `/new` Phase 3 contradicted that by invoking it read-only and handing the fix to `coder`. The v3.28.3 Domain-Override decision was framed as "orchestrator-inline **vs** coder" and simply never put "doc-reviewer writes" on the menu. This release closes that gap. (`/prd` is unaffected — it *produces* docs directly via `prd-card-writer`/the writing phases and has no read-only-audit→remediation loop.)

### Changed — `doc` fix-domain owner is now `doc-reviewer` (not `coder`)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** — Domain-Override Domains table: the `doc` domain now carries an explicit **Owning agent = `doc-reviewer` (write mode)** column, with a rationale block ("Why `doc` is owned by `doc-reviewer`, not `coder`"). `security` and `migration` stay with `coder`. The mechanical `CHANGELOG.md`/`ssot-registry.md` append edge case now goes through `doc-reviewer`, never `coder`.
- **Sequential Phase 3** collapses from two spawns (doc-reviewer read-only audit + coder apply) to **one** — the doc-reviewer audits AND applies in a single invocation (it runs alone; code-review moved to Phase 3.7, so the old read-only/parallel-safety constraint no longer applies). The only output it does NOT fix itself is a doc-drift→bug finding rooted in CODE, which follows the code fix path.
- **Team-mode D.2/D.4a**: D.2 keeps the doc-reviewer **read-only** (it runs in parallel with `code-reviewer` — parallel-safety preserved); D.4a now re-invokes the **doc-reviewer in write mode** over the group to apply the per-card-attributed findings, instead of spawning a fix-coder.
- **Final review F.5**: verified `>= MEDIUM` findings are now partitioned by domain — `doc`-domain findings → `doc-reviewer`, all other findings → `coder`.

### Changed — Fix Application telemetry recognises the new owner

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** Fix Application Log Schema: `decision` and `applied_by` gain the `doc-reviewer` value; `phase` gains `D.4a`. (The analyzer regex in `framework/scripts/analyze-fix-application.js` already accepts arbitrary `[\w-]+` tokens — no code change needed.)
- **[framework/docs/FIX-APPLICATION-TELEMETRY.md](framework/docs/FIX-APPLICATION-TELEMETRY.md)**: violation/healthy pattern tables updated — `applied_by=coder` on a Phase 3 `doc` row is now itself a violation; the healthy doc pattern is `decision=doc-reviewer | applied_by=doc-reviewer`.

## [3.39.0] - 2026-06-01

**A `/prd` or `/new` run now CONCLUDES clean — it never ends by handing the user a list of "azioni tue, non bloccanti" (an uncommitted file blocking the local `develop` fast-forward, a merged remote branch left undeleted).** Every workspace-hygiene leftover is either auto-resolved by the finalizer or put behind ONE explicit `AskUserQuestion` gate. No passive manual TODO ever survives into the final summary. **No new `baldart.config.yml` keys.**

> **Why.** A `/prd` finalization closed with: *"il develop locale non si è fast-forwardato per una modifica non committata pre-esistente a `docs/metrics/skill-runs.jsonl` — gestisci quel file e poi `git pull --ff-only`"* and *"il branch remoto non è stato eliminato — puoi cancellarlo a mano"*. The orchestrator declared the run "completato" while leaving the workspace half-finished. Root cause: `worktree-manager mw-docs` printed `⚠️ … Leaving as-is` on a blocked fast-forward, and `/prd` (unlike `/new` Phase 6c) had no workspace-hygiene phase to consume that — it forwarded the warning to the user verbatim.

### Changed — the merge finalizer auto-resolves instead of giving up

- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)** (Common — sync local develop): the blocked-ff path no longer prints `Leaving as-is`. It partitions the dirty tree: the framework-owned append-only telemetry log (`docs/metrics/skill-runs.jsonl`) is reconciled autonomously and losslessly (commit + `pull --rebase` + best-effort push) so the ff completes; ANY foreign file emits a new `[SYNC-NEEDS-DECISION]` marker the orchestrator MUST convert into one explicit gate. Work this run does not own is never auto-committed.
- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)** (Cleanup step 6.3 + report): merged remote-branch deletion is now an explicit, reported step (was `2>/dev/null || true` silent) — a denial is retried, never degraded into a "puoi cancellarlo a mano" note.

### Added — `/prd` Step 7.5 Workspace Hygiene & finalization (BLOCKING)

- **[framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md)**: new Step 7.5 mirrors `/new` Phase 6c — consumes the `[SYNC-NEEDS-DECISION]` / `[SYNC-DEFERRED]` markers, raises ONE `AskUserQuestion` per unresolved residue, ensures the merged remote branch is gone. New **HARD RULE** in Final output: the summary MUST NOT contain a "azioni tue / note non bloccanti" section — a run that prints "gestisci tu il file e poi pulla" has FAILED Step 7.5.

### Changed — `/new` Phase 6c parses the new marker

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** (Phase 6c step 2): now also parses `[SYNC-NEEDS-DECISION]` (foreign-file ff block / rebase conflict) as a blocking gate, alongside the existing `[SYNC-DEFERRED]`.

### Changed — AGENTS.md scopes branch-deletion approval to unmerged branches

- **[framework/AGENTS.md](framework/AGENTS.md)**: owner approval is required before deleting an **unmerged** branch; deleting an **already-merged** feature branch is routine cleanup the merge skill performs automatically (no separate approval). The "prune remote branches only when instructed" rule is clarified to exclude a feature branch's own ref on successful merge.

### Revert

- Restore the `git -C "$MAIN" pull --ff-only origin develop || echo "⚠️ … Leaving as-is."` one-liner in `worktree-manager` and drop `/prd` Step 7.5.

## [3.38.0] - 2026-06-01

**The per-card review depth (QA + Codex) is now decided deterministically in the backlog card via a new `review_profile` field, computed once at PRD authoring time — instead of being re-derived heuristically by `/new` at runtime.** This removes the implementing agent's interpretation latitude: the card carries `skip|light|balanced|deep`, `/new` READS it, and only falls back to runtime computation for legacy cards that predate the field. **No new `baldart.config.yml` keys** — this is a backlog-card schema change, propagated end-to-end across templates + card-writer + validation + `/new` in the same release.

> **Why.** Previously the QA Profile Selector (Phase 3.5) and the Codex Review Profile Selector (Phase 3.7) both recomputed the profile inside `/new` from soft signals (title keywords, the self-admittedly-unreliable `files_likely_touched`), so the same card could be reviewed differently depending on the orchestrator's reading. The card already encodes every hard signal needed (`owner_agent`, type, `areas`, `data_fields`, `db_indexes`, `acceptance_criteria`, `estimated_complexity`) — so the decision is now made once, in the card, by the deterministic rules in `prd-card-writer.md § Rule C` (the SSOT).

### Added — `review_profile` card field (deterministic review depth)

- **[framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml)** + **[framework/templates/feature-card.template.yml](framework/templates/feature-card.template.yml)** + **[framework/.claude/skills/prd/assets/epic-template.yml](framework/.claude/skills/prd/assets/epic-template.yml)**: new `review_profile: skip|light|balanced|deep` field. Epics are always `skip` (trackers, no code work). Authors MAY override the computed value by hand.
- **[framework/.claude/agents/prd-card-writer.md](framework/.claude/agents/prd-card-writer.md)**: new **Rule C** — the canonical mapping table (metadata → profile) and the SSOT for the decision. Added `review_profile` to the required-fields list. `feature`/`enhancement` cards are floored at `balanced` (never `light`).
- **[framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md)**: Step 6.0 now BLOCKS on a missing/out-of-enum `review_profile` (mirrors the `owner_agent` gate), with a soft WARN+normalize for `feature`-card-`light` violations and epic non-`skip` values.
- **[framework/.claude/skills/prd/references/backlog-phase.md](framework/.claude/skills/prd/references/backlog-phase.md)**: documents that `prd-card-writer` populates `review_profile`.

### Changed — `/new` reads the field instead of recomputing it

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** (QA Profile Selector): primary path now READS `review_profile` from the card verbatim; the existing rule table is demoted to a **legacy-card fallback** with an explicit "keep in sync with `prd-card-writer.md § Rule C` (SSOT)" note.
- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** (Phase 3.7 Step C.0): the Codex profile derives from the card's `review_profile` (`skip`/`light` → Codex `light`, `balanced`/`deep` → Codex `full`). New **escalation-only invariant**: the Step A high-risk diff detector can only promote `light` → `full` (a risk surfaced at implementation time the authored profile didn't anticipate), never downgrade `full` → `light`. The card's `review_profile` is the deterministic floor; the diff detector is the safety net on top.
- **[framework/.claude/commands/new.md](framework/.claude/commands/new.md)**: thin mirror updated to the read-from-card primary path with compute-fallback.

### Changed — `light` profile no longer runs the per-card Codex adversarial pass

- **[framework/.claude/commands/codexreview.md](framework/.claude/commands/codexreview.md)** (Step -0.5) + **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** (Phase 3.7 Step C.0, Team D.4b): at `profile: light`, `/codexreview` now omits agent #5 (Codex adversarial) in addition to qa-sentinel + api-perf-cost-auditor + Step 3.5 CoVe. **`light` per-card = `code-reviewer` + the false-positive gate only.** The Codex adversarial review becomes the prerogative of `full` and of the final cross-check.
- **Why it's safe (BUG-0530 invariant relocated, not dropped)**: the "never miss a blocker on a low-risk path" guarantee is now carried by (1) the **escalation-only** Step A detector — any high-risk path in the actual diff promotes `light` → `full`, so risky code still gets per-card Codex adversarial immediately; and (2) the unconditional **Final-review FULL gate** (v3.37.0), which runs Codex adversarial over the entire batch diff before merge. So every `light` card's code is still Codex-reviewed at least once pre-merge — the pass is paid once at the batch gate instead of N times per-card. Trade-off: a Codex-class finding on a 0-trigger `light` card surfaces at the final gate rather than per-card (later fix-cycle, still pre-merge).
- **Revert**: re-add agent #5 to the `light` branch of `/codexreview` Step -0.5.

### Backward compatibility

- **Non-breaking**: backlogs authored before v3.38.0 have no `review_profile`. `/new` silently falls back to the runtime rule table (logging `computed — review_profile absent`), so existing cards run exactly as before. New PRD-generated cards get the field automatically.

## [3.37.1] - 2026-06-01

**`/prd` Step 7 merge is fully seamless — auto merge + worktree cleanup with zero end-of-session questions; the user is asked ONLY for conflicts `mw-docs` genuinely cannot auto-resolve.** Fixes a drift between the `/prd` `validation-phase.md` and the real `worktree-manager mw-docs` behaviour that made the skill over-ask. No new `baldart.config.yml` keys.

### Fixed — `/prd` no longer over-asks at merge time

- **[framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md)**: the cross-PRD-collision note and the Step 7 conflict-resolution table claimed structured registries (`project-status.md`, `ssot-registry.md`, …) "abort the rebase and ask for manual resolution". That contradicted `worktree-manager` `mw-docs`, which actually strips markers + runs a structural validation and only aborts when the validation fails (the rare same-row collision). Aligned both spots to the real behaviour: append-only registries auto-resolve and merge with zero questions; escalation happens **only** on a genuinely unresolvable conflict (code/test/fixture, or a structured registry whose strip+validate failed).
- **[framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md)** (Step 7 step 9): reinforced "no confirmation prompt" — explicitly forbids the "vuoi che mergi e pulisca il worktree?" pre-emptive gate. Merge **and** cleanup run automatically; the model lets `mw-docs` run and only surfaces what it actually reports as unresolvable, instead of asking "just in case".

## [3.37.0] - 2026-05-30

**Two-tier review for `/new`: fast `light` per-card during the loop, one guaranteed FULL pass over the whole batch before merge.** This activates the dormant v3.35.0 Review Profile Selector AND re-introduces an unconditional final full review — together they cut tokens/wall-clock on low-risk cards *without* weakening the merge gate. **No new `baldart.config.yml` keys** — both are changes inside the skill, so the schema-propagation rule does not apply.

> **The design.** `light` (per-card Phase 3.7) is now an **early-feedback optimization**, not the safety gate. The safety gate is the **Final review**, which since this release ALWAYS runs a single FULL `/codexreview` (full agent set) over the **entire batch diff** before merge — no N=1 skip, no cross-card scope reduction. So every line of every card, including any reviewed at `light`, gets a full-depth Codex review at least once before merge. This is what makes activating `light` safe: it shifts the breadth passes (qa-sentinel, api-perf-cost-auditor, doc-reviewer) + CoVe from the per-card loop to the single final pass, rather than dropping them from the merge gate. (The v3.35.0 data gate — which had kept the selector hard-coded to `full` pending consumer telemetry — is lifted by explicit maintainer decision; an adversarial review showed that gate's intended post-hoc telemetry could never have validated it anyway, since the Fix Application Log is blind by construction to blockers `light` *misses*. The final full gate makes that validation question moot.)

### Changed — Review Profile Selector live (per-card early feedback)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: Phase 3.7 Step C (sequential) and Team Mode D.4b no longer hard-code `profile = full`. The selector runs as designed — `light` ⟺ Step A matched **0** high-risk triggers **AND** the card's QA profile ∈ {`skip`, `light`}; `full` otherwise. **BUG-0530 invariant preserved**: `light` ≠ skip — Codex adversarial + `code-reviewer` + false-positive gate ALWAYS run per-card; `light` drops only the breadth passes + CoVe.

### Changed — Final review is now an unconditional FULL batch-diff gate (supersedes v3.35.0 scope-reduction)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** (Final review / Step F.1): the v3.35.0 final-review scope gate (N=1 → skip F.1–F.4; N>1 → reduce to the cross-card surface) is **removed**. The final review now runs Steps F.1–F.5 for **every batch including N=1**, with `review_scope_files` = the **full union** of all touched files, never reduced. F.3 reviews the entire batch diff at full depth (Codex + doc-reviewer + api-perf-cost-auditor + qa-sentinel); F.5 runs the final build. Team mode reaches the same gate via "Post-batch — same as sequential mode" (its prior "cross-card-only scope" note is updated accordingly).
- **Rationale**: v3.35.0 de-duplicated the final full pass on the assumption that Phase 3.7 had already full-reviewed every card. Selecting `light` breaks that assumption, so the post-batch full pass is re-introduced as the guaranteed merge gate.

### Cost & trade-offs (documented)

- **Net cost**: one full `/codexreview` per batch is re-introduced (the deliberate price of the per-card `light` speed-up). The saving comes from cheaper per-card reviews during the loop; the final pass is paid once per batch, not per card.
- **Later feedback on `light` cards**: a defect on a `light`-reviewed card may now surface at the final pass rather than at the per-card step — possibly a later fix-cycle. Acceptable: it is still caught before merge.
- **Revert**: to make Phase 3.7 always-full again, restore `profile = full` (no config kill-switch by design). The final full gate is independent and stays regardless.

## [3.36.0] - 2026-05-30

Adds **opt-in session telemetry** to `/new` and `/prd`: append `-stats` (or `--stats`) to a run and, at the end, the skill writes a per-**agent-role** breakdown of REAL token consumption and wall-clock time. This finally answers "what took the most time and tokens in a `/new` run" — and feeds the v3.35.0 Review Profile Selector **data gate** (still pending precisely because `docs/metrics/` carried no per-role cost data). **No new `baldart.config.yml` keys** — the opt-in is a flag, not a config feature, so the schema-propagation rule does not apply.

> **Zero model-token overhead.** The measurement is a **post-hoc Node script**, not a hook and not LLM self-reporting (the model has no reliable token counter). It reads the Claude Code session transcripts — the orchestrator's `<session>.jsonl` plus every `subagents/agent-*.jsonl` — which already carry `message.usage` (the four token counters) and per-message `timestamp`. On a real `/new` run the bulk of cost lives in the subagents (measured ~53M cumulative subagent tokens vs ~47k orchestrator output), and it is all captured. Cache-read tokens are kept **separate** from fresh input (they bill ~10×cheaper) so the breakdown doesn't mislead.

### Added — `-stats` / `--stats` opt-in session telemetry

- **[framework/scripts/analyze-session-tokens.js](framework/scripts/analyze-session-tokens.js)** (new): post-hoc aggregator. Locates the live transcripts from `$CLAUDE_CODE_SESSION_ID` + the cwd slug, sums the four `usage` counters (`fresh_input` / `cache_read` / `cache_creation` / `output`) **kept separate**, computes per-file wall-clock spans, and maps each subagent file → its role by matching the subagent's first prompt to the `Agent` tool_use `subagent_type` in the main transcript (verified 20/20 on a real `/new` batch; unmatched files bucket as `unknown`). Emits a human-readable `docs/metrics/sessions/<run_id>.md` (per-role table sorted by a weighted `cost_units` figure + totals + caveats) and a machine-readable `TELEMETRY_JSON=` line. **Fail-safe**: any error prints `stats: SKIPPED (<reason>)` and exits 0 — it can never abort a run.
- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: arg parser strips `-stats` / `--stats` like `-full` and sets an internal `STATS` flag (composes with `-full`). **Phase 8 — Metrics Log** gains a `STATS`-gated step that invokes the script and enriches the same `skill-runs.jsonl` row with a `"cost"` object (`by_role` + `totals` + `run_wall_ms`). Without `-stats`, Phase 8 is byte-for-byte unchanged.
- **[framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md)**: new **HARD RULE 18** — since `/prd` is conversational (no flag parser), the opt-in is detected once at **Step 1 (Kickoff)** from `$ARGUMENTS` and persisted as `stats_enabled` in the session state file (survives context compression / Step 0 resume). The **Step 7 Metrics Log** reads it and runs the same script (`--skill prd`), enriching its JSONL row identically.
- **[framework/.claude/skills/prd/assets/state-template.md](framework/.claude/skills/prd/assets/state-template.md)**: new `stats_enabled` header field.

### Notes & caveats

- **Accurate in sequential mode.** Team-mode L0/L1 orchestrators can spawn background / nested sub-orchestrators in *separate* session dirs; their subagents are not under the top-level `<session>/subagents/`. The script detects the gap (Agent spawns in main with no local subagent file) and prints an explicit caveat with the excluded count rather than silently undercounting.
- The per-role `wall(sum)` reflects active compute time; the headline run wall-clock is the orchestrator span and **includes human idle gaps** — both are labelled in the report.

## [3.35.2] - 2026-05-30

Resolves the **3.35.1 known follow-up**: a flag introduced in a newer release (e.g. `--on-divergence`, v3.29.0) was rejected by an older global CLI with a cryptic `error: unknown option` **before** the `npx baldart@latest` auto-relaunch could self-upgrade — commander validates options during `program.parse()`, which runs *before* the action handler where the relaunch lives. The `update` command now parses **permissively**, so an unknown flag is captured instead of crashing the parser; the action then transparently relaunches under `@latest` (forwarding the raw argv so the unrecognized flag survives) or, when that isn't possible, emits an **actionable** error instead of commander's cryptic one. **No new `baldart.config.yml` keys** — the schema-propagation rule does not apply.

> **Forward-looking by nature.** This cannot retroactively fix CLIs *already* installed at a pre-3.35.2 version — their commander still throws before any code we control runs. The fix rescues flags introduced *after* this release for anyone on 3.35.2+. For older installs the only path remains `npm i -g baldart@latest` (now surfaced by the helpful error + the update notifier).

### Fixed — self-upgrade survives flags newer than the installed CLI

- **[bin/baldart.js](bin/baldart.js)**: the `update` command gains `.allowUnknownOption()` + `.allowExcessArguments()` (scoped to `update` only — every other command keeps strict validation). Unknown tokens now land in `command.args` instead of throwing `commander.unknownOption` during parse; they are forwarded to `update(options, command.args)`.
- **[src/commands/update.js](src/commands/update.js)** (`update`): an early unknown-flag block runs **before** any work. When unknown tokens are present it relaunches under `npx baldart@latest`, forwarding `process.argv.slice(3)` (raw flags, with the leading `update` subcommand token stripped so the child isn't spawned as `update update …`). The unrecognized flag never enters `options`, so raw-argv forwarding is the only faithful path.
- **[src/commands/update.js](src/commands/update.js)** (`maybeRelaunchUnderLatest`): accepts `extra.rawArgs` (forward verbatim instead of reconstructing from parsed `options`) and `extra.unknown` (skip the interactive "proceed with current CLI" prompt — pointless when the CLI literally can't parse the flag; relaunch whenever a newer version exists). The normal stale-CLI relaunch path is unchanged (still reconstructs known flags, still forwards `--on-divergence` / `--json`).
- **[src/commands/update.js](src/commands/update.js)** (`failOnUnknownArgs`): new helpful terminal error replacing commander's cryptic message — hedges between "recently-added flag → `npm i -g baldart@latest`" and "typo", and **suppresses the upgrade suggestion** when `BALDART_RELAUNCHED=1` (we are already the `@latest` child, so a typo is the only explanation). `--json`-aware: emits a `usage-error` JSON object on stdout.
- **[src/commands/update.js](src/commands/update.js)**: `--json` mode is now **armed before** the unknown-flag block so the relaunch notice / helpful error route to STDERR — preserving the single-object STDOUT contract for `update --json --yes --<new-flag>` (the agent/CI scenario this machinery exists for).

## [3.35.1] - 2026-05-30

Fixes a **HIGH-severity** bug where `baldart update --reset` could leave a consumer **committed-without-`.framework/`** (broken install, manual recovery only). The reset path wiped `.framework/`, committed the deletion, then handed off to `add()` **programmatically with no repository** — the bin-layer default (`antbald/BALDART`) only applies on the CLI path, so `repo` arrived `undefined`, crashing at `repo.startsWith(...)` ("Cannot read properties of undefined") *after* the wipe was already committed. Three layers now make this unreachable: (1) repo resolved **before** any destructive step, (2) reinstall failures roll back to the pre-reset backup tag, (3) defensive guards + legible errors. **No new `baldart.config.yml` keys** — the schema-propagation rule does not apply.

### Fixed — `update --reset` never leaves the consumer without `.framework/`

- **[src/commands/update.js](src/commands/update.js)** (`runReset`): the upstream repo is now resolved via `State.resolveRepo()` **before** the backup tag / `rm -rf` / reset commit, then threaded **explicitly** into `add()` (was `addCmd(undefined, { yes: true })` — the root cause). The stale v3.28.1 comment claiming "undefined → use default from package.json" (which never existed in `add.js`) is corrected.
- **[src/commands/update.js](src/commands/update.js)**: the reinstall is wrapped in a `try/catch` that, on **any** failure (undefined repo, network, ignored path, missing hooks), auto-runs `git reset --hard <backupTag>` to restore the pre-reset `.framework/` and prints an explicit recovery command. Safe because reset is gated on a clean working tree — the backup tag holds everything but the reset commit.
- **[src/commands/add.js](src/commands/add.js)**: new `options.throwOnError` makes `add()` **re-throw** instead of `process.exit(1)` at every reachable failure point (download catch, gitignore-still-ignored, hooks-missing post-flight), so a programmatic caller can recover. CLI behavior is unchanged (flag absent → historical exit-1). Also resolves a falsy `repo` from the ledger as a last resort and fails fast **before** download with a recovery hint.
- **[src/utils/git.js](src/utils/git.js)** (`normalizeRepoUrl`): guards a falsy/non-string `repo` with a legible `Error` ("repository non specificato …") instead of the cryptic `Cannot read properties of undefined (reading 'startsWith')`.

### Changed — shared repo resolver (de-dup)

- **[src/utils/state.js](src/utils/state.js)**: new exported `resolveRepo(cwd)` (cascade `state.framework_repo → DEFAULT_REPO`, never `undefined`, preserves a consumer's fork) + exported `DEFAULT_REPO` constant. `recordInstall` now uses the constant.
- **[src/commands/version.js](src/commands/version.js)**: drops its local `REPO_DEFAULT` duplicate in favor of the shared `State.DEFAULT_REPO`.

### Known follow-up (separate fix)

- `--on-divergence` (v3.29.0) is rejected by an **older global CLI** with "unknown option" *before* the `npx baldart@latest` auto-relaunch can self-upgrade (commander validates options first). Tracked separately — workaround: `npm i -g baldart@latest` before using new flags.

## [3.35.0] - 2026-05-30

De-duplicates the review chain in `/new` and adds a **Review Profile Selector** so the per-card pre-merge gate scales its *depth* by risk instead of running the full multi-agent `/codexreview` on every card — including trivial ones. The per-card gate stays **unconditional** (it always runs); only its breadth varies, and the safety-critical Codex adversarial pass + false-positive gate run on every card regardless of profile. Copertura invariata: every line is still code-reviewed once, every doc doc-reviewed once. **No new `baldart.config.yml` keys** — the profile is computed at runtime from signals already present (the Phase 3.7 high-risk-trigger detector + the QA profile), and lean mode is an internal caller→skill file contract, so the schema-propagation rule does not apply.

### Changed — `/new` review chain (de-dup, coverage-invariant)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** + **[framework/.claude/commands/codexreview.md](framework/.claude/commands/codexreview.md)**: per-card Phase 3.7 (and Team Mode D.4b) now hand `/codexreview` a lean contract file (`/tmp/codexreview-lean-<CARD-ID>.json`) so it **reuses the Phase-1 architecture baseline** (`/tmp/arch-baseline-<CARD-ID>.md`) instead of re-spawning `codebase-architect`, and **skips the duplicate `doc-reviewer`** (Phase 3 already audits the same diff). `/codexreview` gains a backward-compatible **Step -0.5 (lean-mode detection)** with consume-once semantics; standalone `/codexreview <CARD-ID>` (no contract file) behaves exactly as before.
- **doc-reviewer lens-merge**: Phase 3's `doc-reviewer` mandate now also carries the *spec/docs-drift → bug* lens that was previously exclusive to `/codexreview` agent #4 — so dropping the duplicate loses no coverage.
- **Final review de-dup**: single-card batches short-circuit F.1–F.4 (the lone card already passed Phase 3.7 unchanged) while still running the F.5 final build; multi-card batches scope the cross-card review to a **computed** surface (files touched by ≥2 cards ∪ `depends_on`-linked pairs), never a model-decided skip.
- **Team Mode doc-reviewer collapse**: `doc-reviewer` went from ~3× per card (D.2 + D.4a + D.4b) to ~1× per group — D.2 attributes findings per-card, D.4a consumes them (no re-spawn), D.4b skips its copy via the lean contract.

### Added — Review Profile Selector (data-gated)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: Phase 3.7 Step C selects `light` ⟺ (0 high-risk triggers AND QA profile ∈ {skip, light}), else `full`. `light` drops the breadth agents (qa-sentinel, api-perf-cost-auditor, doc-reviewer) + Step 3.5 CoVe but ALWAYS keeps `code-reviewer` + Codex adversarial + the false-positive gate (BUG-0530 invariant preserved — `light` ≠ skip). Mirrors the existing QA Profile Selector pattern. Ships behind a **data gate**: the mapping stays hard-coded to `full` until consumer telemetry (`docs/metrics/` Fix Application Log) confirms `light`-eligible cards have not historically produced verified BLOCKER/HIGH at full depth.

## [3.34.0] - 2026-05-30

Adds an **assert-active git-hooks health check** to `baldart doctor` and the `worktree-manager` verify step. A consumer had `core.hooksPath` pointing at `.git/hooks` (empty) while the repo's versioned hooks lived in `scripts/git-hooks/` — the pre-commit schema guard and pre-push were **silently inactive for days**. The existing worktree-manager verify (`git config core.hooksPath || true`) only *read* the value without asserting anything, so it never caught the drift. Same class as the LSP bug: the framework "verifies" by reading a proxy instead of asserting the real state. This moves it from read-only to assert-active. **WARN-only** — the check never mutates `core.hooksPath` (a wrong heuristic + auto-fix would be worse than the original bug); it surfaces the exact remediation command and lets the user apply it. **No new `baldart.config.yml` keys** — it's a universal, always-on, filesystem-autodetected check, so the schema-propagation rule does not apply.

### Added — git-hooks health check (assert-active)

- **[src/utils/githooks.js](src/utils/githooks.js)**: new, self-contained, fail-safe utility (distinct from `src/utils/hooks.js`, which manages `.claude/settings.json` hooks). `getStatus(cwd)` detects the repo's versioned-hooks directory (`.husky`, `.githooks`, `scripts/git-hooks`, `githooks` — first that exists and contains a standard-named hook), resolves the **active** hooks dir explicitly (`core.hooksPath` → relative resolved against `git rev-parse --show-toplevel`, absolute verbatim; unset → `<git-dir>/hooks`, handling the worktree `.git`-is-a-file case), and asserts the active dir serves the versioned hooks. **`git rev-parse --git-path hooks` is deliberately NOT used** — empirically it returns the configured value relative/unresolved and historically some git versions ignored `core.hooksPath` for it. Returns `checked:false` (silent, `ok:true`) when no managed hooks exist or on any error, so the check never produces a false WARN.
- **[src/utils/githooks.js](src/utils/githooks.js)**: pure exported `isActiveServing(activeAbs, expectedAbs)` — identical-or-nested decision (covers husky v9, which sets `core.hooksPath=.husky/_` with user hooks in `.husky/`). Paths are canonicalized (symlink-resolved, handling non-existent leaves) before comparison so a symlinked ancestor (macOS `/var`→`/private/var`, symlinked repo/home) never yields a false INACTIVE. Executability (`-x`) is asserted only for raw dirs git execs directly; husky user hooks (not required `+x` in v9) skip the assert.
- **[src/commands/doctor.js](src/commands/doctor.js)**: probe wired into `detectState` (`state.gitHooks`), a `Git hooks` line in `renderDiagnostic` (silent when `checked:false`), and an **advisory print-only** action in `planActions` (`githooks-health`) whose `run()` prints the WARN + fix command and **never mutates** `core.hooksPath` (safe under `--auto`).
- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)**: step 4 (isolation setup) replaces the silent `git config core.hooksPath || true` read with a lean **assert** — detects the versioned-hooks dir, resolves the active dir, and echoes a loud `⚠️ WARNING` + the `git config core.hooksPath <dir>` fix on mismatch (matches husky v9 `.husky/_` and `<dir>/*` as active). The worktree shares `core.hooksPath` via `.git/commondir`, so the same resolution applies.

### Tests

- **[src/utils/__tests__/githooks-active.test.js](src/utils/__tests__/githooks-active.test.js)**: new — 8 pure fixtures for `isActiveServing` (the mayo bug, husky v9 nesting, identical dirs, prefix-not-a-boundary, trailing-sep normalization).
- **[scripts/test-githooks.sh](scripts/test-githooks.sh)**: new — drives **real** git repos: misconfig (`hooksPath=.git/hooks`, hooks in `scripts/git-hooks`) → inactive + fix; correctly wired → ok; non-executable raw hook → ok=false + chmod fix; husky v9 → active, `-x` assert skipped; no managed hooks → `checked:false` (silent).

## [3.33.1] - 2026-05-30

Fixes a **false-positive in the LSP layer's verify step** that let `/lsp-bootstrap` (and `baldart configure` / `doctor`) report a language server as `verified` while every real LSP call failed with `Executable not found in $PATH` — a dead layer masked by `codebase-architect`'s silent Grep fallback. Root cause was a **resolution mismatch**: verify used `npx --no-install typescript-language-server --version` (resolves from `node_modules/.bin`), but the consumer — Claude Code's LSP tool — spawns the server **by name from `$PATH`**. The npm-dev install (`npm install --save-dev`) put the binary in `node_modules/.bin`, which is *not* on `$PATH`, so verify passed and spawn failed. The npm-managed adapters now install **globally** so the binary lands on `$PATH`, and verify checks `$PATH` reachability the *same way the consumer spawns it*. **No new `baldart.config.yml` keys** — reachability is computed, not configured, so the schema-propagation rule does not apply.

### Fixed — LSP verify no longer reports unreachable servers as `verified`

- **[src/utils/lsp-installer.js](src/utils/lsp-installer.js)**: `verifyServers()` rewritten around a 3-state taxonomy. It first probes `$PATH` reachability via `command -v <binary>` (new `_isOnPath` helper — mirrors the consumer's by-name spawn exactly), then runs the functional `verifyCommand()`. Results carry `status` ∈ `verified` (on `$PATH`, probe passed) / `installed-but-unreachable` (present in `node_modules` but not on `$PATH` — new `_isInstalledLocally` helper fingerprints the legacy devDep install; `fix` = global-install command) / `unreachable` / `not-installed` / `unknown-adapter`. Never throws.
- **[src/utils/lsp-adapters/typescript.js](src/utils/lsp-adapters/typescript.js)** + **[src/utils/lsp-adapters/python.js](src/utils/lsp-adapters/python.js)**: `installMode` `npm-dev` → `npm-global`; `installCommand()` → `npm install -g …`; `verifyCommand()` → `<binary> --version` (resolved from `$PATH`, mirroring the consumer) instead of `npx --no-install …`. System adapters (go/rust/ruby) were already PATH-correct and are unchanged.

### Changed — `configure` records unreachable servers; `doctor` fix is now real

- **[src/commands/configure.js](src/commands/configure.js)**: **every** known-adapter result that isn't cleanly `verified` is now recorded in `lsp.installed_servers` (like system-mode pending installs) instead of being dropped by the old `filter(v => v.ok)` — otherwise re-runs reinstall blindly and `doctor` never sees the mismatch. This covers both `installed-but-unreachable` and the npm-global happy-path trap where `npm install -g` exits 0 but npm's global bin dir isn't on the current `$PATH` (→ status `not-installed`, since `npx` can't see the global prefix either) — without this, the primary configure path would have silently reintroduced the dead layer. The fix command (+ `npm prefix -g` PATH hint) is surfaced inline for each.
- **[src/commands/doctor.js](src/commands/doctor.js)**: the `lsp-fix` remediation now reinstalls globally (a real fix that puts the binary on `$PATH`), where before it re-ran `npm install --save-dev` — a no-op for the reachability problem. The `why` text explains the `$PATH` contract, and a post-reinstall re-verify surfaces the `npm prefix -g` PATH escape when a server stays unreachable (so a misconfigured global bin dir doesn't trap the user in a reinstall loop).

### Docs

- **[framework/.claude/skills/lsp-bootstrap/SKILL.md](framework/.claude/skills/lsp-bootstrap/SKILL.md)**: Output Contract redefines `verified` as `$PATH`-reachable, adds the `installed-but-unreachable` status, and documents the **cold-start caveat** (first query may under-report references until indexing completes).
- **[framework/agents/code-search-protocol.md](framework/agents/code-search-protocol.md)**: new fallback rule for cold-start under-reporting — re-issue the first query before trusting a suspiciously-low reference count; never declare a helper unused off one cold find-references.
- **[framework/docs/LSP-LAYER.md](framework/docs/LSP-LAYER.md)** + **[framework/docs/PROJECT-CONFIGURATION.md](framework/docs/PROJECT-CONFIGURATION.md)**: adapter-contract table, install-mode descriptions, `verifyServers` contract, and troubleshooting §6.3/§6.4 updated for global install + `$PATH` reachability + the PATH-mismatch cause.

## [3.33.0] - 2026-05-30

Closes a **structural silent-partial-install** failure mode (observed on the `mayo` consumer, 3.28→3.31, still live in 3.32.0; same class as the historic v3.27.1 bug): `add` was **non-atomic** — the hook registration ran *after* optional blocking prompts (git aliases, configure). In a non-TTY context (Claude's Bash tool never has a TTY on stdin) or on any interruption, those `inquirer` prompts throw/hang, so `.framework/` + `state.json` landed on the new version while `Hooks.registerAll` never ran — a half-install that `baldart status` still reported as "Valid". The fix makes install atomic (correctness-critical mutations run **before** every optional prompt), adds a **post-flight hook assert** that refuses to declare success with any active hook missing, and finally exposes `add --yes` so agents/CI can install unattended. Two additive, non-breaking follow-ups ship alongside. **No new `baldart.config.yml` keys** — `--yes` is a CLI flag, so the schema-propagation rule does not apply.

### Fixed — `add` is now atomic + never declares a half-install successful

- **[src/commands/add.js](src/commands/add.js)**: `Hooks.registerAll` is **reordered ahead of** the optional git-aliases and configure prompts. Hook registration determines install correctness and must never sit behind a blocking prompt — if `add` is interrupted at an optional prompt, the hooks are already in place. The drift callback now passes `{ autoYes: nonInteractive }` (parity with `update.js`), so a drifted hook in non-TTY no longer throws.
- **[src/utils/hooks.js](src/utils/hooks.js)**: new exported `verifyAll(cwd)` — a pure read over `getStatus` returning `{ ok, malformed, missing, path }`. It iterates only the **active** `HOOK_REGISTRY` (commented-out entries like `capture-detect` excluded for free). A malformed `settings.json` counts as a failure (registration was skipped).
- **[src/commands/add.js](src/commands/add.js)**: **post-flight assert** before the success summary — re-reads `settings.json`, attempts a single best-effort re-registration if anything is missing, then exits non-zero ("Install is NOT complete") if a hook is still absent. Covers `update --reset` too, since that path delegates to `add(undefined, { yes: true })`.

### Added — `add --yes` (non-interactive install for CI / agents)

- **[bin/baldart.js](bin/baldart.js)**: registers `-y, --yes` on the `add` subcommand (the underlying `add(repo, options)` already honored `options.yes`, it was just unreachable from the CLI). With `--yes`, `add` auto-confirms optional prompts and passes `{ nonInteractive: true }` to `configure` so its wizard writes autodetected values without blocking.
- **[src/commands/add.js](src/commands/add.js)**: `add --yes` on a repo that **already has** `.framework/` now fails fast with a clear pointer to `npx baldart update --reset --yes --i-know` instead of hitting the deliberately-interactive destructive "Remove and reinstall?" prompt (which would throw under non-TTY and surface as a generic failure). Clean installs only; reinstall goes through `--reset`.

### Added — divergence classifier flags overlay-covered edits (`--reset` is safe)

- **[src/utils/git.js](src/utils/git.js)**: `classifyDivergence` now tags each `custom-overlay-able` commit with `overlayCovered` — true when **every** framework file it touches already has a matching overlay in `.baldart/overlays/`. New pure static `overlayRelForFrameworkFile(path)` mirrors the `overlay.js` path conventions (unit-tested).
- **[src/commands/update.js](src/commands/update.js)**: in the overlay-able divergence branch, when all commits are `overlayCovered` the UI/JSON now recommend `--reset --yes --i-know` as the **safe** resolution (intent already preserved in the overlays) and note that `scaffold-overlays` would create redundant skeletons. New `overlay_already_covered` field in the `--json` output.

### Added — deterministic `overlay drift` via base-sha backfill

- **[src/utils/overlay-merger.js](src/utils/overlay-merger.js)**: new `ensureBaseFileShaInOverlay(overlayPath, baseSha)` — **additive-only, idempotent, fail-safe**: stamps `base_file_sha` into an overlay's frontmatter when absent (never overwrites, never fabricates frontmatter). Called from the merger in **[src/utils/symlinks.js](src/utils/symlinks.js)** `_generateOverlayedFile` at `add`/`update` time, so `baldart overlay drift` becomes deterministic for **every** overlay — not just freshly-scaffolded ones (hand-written / pre-v3.19.0 overlays were "unknown" to drift).

### Tests

- **[scripts/test-add-postflight.sh](scripts/test-add-postflight.sh)**: new — encodes the post-flight **invariant** (a `settings.json` missing an active hook → `verifyAll().ok === false`; the dropped hook in the test is literally `baldart-overlay-telemetry`, the mayo failure). Fails on pre-v3.33.0 code (no `verifyAll`).
- **[scripts/test-overlay-backfill.sh](scripts/test-overlay-backfill.sh)**: new — backfill is additive, idempotent, never overwrites an existing sha, never fabricates frontmatter, fail-safe.
- **[src/utils/__tests__/classify-divergence.test.js](src/utils/__tests__/classify-divergence.test.js)**: extended with overlay-path-mapping fixtures (38 assertions, all green).

### Changed — `/baldart-update` skill

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: post-flight report (step 4) documents the new CLI hook assert in `add`/`--reset`; `last_verified` v3.32.0 → v3.33.0 (`hook_registry_entries`/`update_js_prompts` unchanged).

## [3.32.0] - 2026-05-30

Makes the CLI **agent-native at the output layer**: `npx baldart version --json` and `npx baldart update --json --yes` now emit a single **machine-readable JSON object** on stdout (schemas `baldart.version/1` / `baldart.update/1`), with every human line routed to stderr. This closes the last fragile seam in the autonomous-update path — the `/baldart-update` skill previously had to **regex-scrape human-readable prose** (the v3.26.0 wording change had already forced a dual-format tolerant parser in Step 0). An agent now reads a field instead of matching a sentence. Borrowed from the [CLI-Anything](https://github.com/HKUDS/CLI-Anything) philosophy (`--json` on every command) — but BALDART is already a headless CLI, so this is a thin, hand-rolled output layer, not a generated wrapper. **No new `baldart.config.yml` keys** — `--json` is a CLI flag, not a config key, so the schema-propagation rule does not apply.

### Added — `--json` structured output

- **[src/commands/version.js](src/commands/version.js)**: `--json` emits `{schema:"baldart.version/1", installed, installed_version, remote_version, aligned, offline, fetched, uncommitted_files, commits_ahead/behind, cli_version, cli_latest, cli_update_available, repository, last_update_date, last_pushed_version, last_push_date}` — exactly the facts the human box shows. Framework-not-installed and error paths emit `{installed:false|null, error, …}` so the agent always gets parseable stdout.
- **[src/commands/update.js](src/commands/update.js)**: `--json` emits a single `{schema:"baldart.update/1", ok, action, …}` result at **every** terminal point through one `emitUpdateJson` chokepoint (all-or-nothing: a missed exit would leave stdout empty). `ok:true` ⟺ the framework is now at the remote version (`action` `updated` / `already-current`); every refusal/failure is `ok:false` with `action` (`divergence-refused`, `divergence-scaffolded`, `aborted`, `failed`, `framework-missing`, `usage-error`), `reason`, and a `next_command` the human can re-run. `--json` **requires `--yes`** (so no prompt can block) and **rejects `--reset`** (the nuclear path stays an interactive human escape hatch).
- **[src/utils/ui.js](src/utils/ui.js)**: new `UI.setJsonMode(on)` redirects every `console.log` the UI helpers emit to **stderr** (reversible), keeping stdout clean for the JSON. `ora` spinners already write to stderr, so they need no change. The JSON itself is written with `process.stdout.write`, never `console.log`, so it survives the redirect.

### Changed — autonomous `/baldart-update` reads JSON, not prose

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: Step 0 now **prefers** `version --json` (field reads) and **falls back** to the dual-format text parser when the global CLI predates v3.32.0 (`--json` unknown / non-JSON stdout) — the same "new skill + old global CLI" transition window the text parser already guards. Autonomous mode runs `update --json --yes`, parses the single result object (`ok` / `action` / `installed_after` / `backup_tag` / `next_command`), and keeps the **exit code as the safety backstop**. The post-flight assert cross-checks `installed_after` against on-disk `.framework/VERSION` and the Step 0 `remote_version`. `last_verified` v3.31.0 → v3.32.0; added a **JSON-OUTPUT NOTE** recording that `--json` is a flag, not a config key.

### Fixed — self-relaunch flag forwarding

- **[src/commands/update.js](src/commands/update.js)** `maybeRelaunchUnderLatest`: the stale-CLI self-relaunch now forwards `--json` **and** `--on-divergence <strategy>` to the `npx baldart@latest` child. The latter was a pre-existing latent drop (an agent's non-interactive divergence strategy would silently vanish on relaunch); both are fixed together.
- **[bin/baldart.js](bin/baldart.js)**: the CLI self-update notifier's stdout `announce` is suppressed under `--json` so it cannot pollute the single JSON object.

## [3.31.0] - 2026-05-30

Gives `/baldart-update` an **autonomous mode** so agents and scheduled routines can bring the framework up-to-date with **zero human interaction** — closing the gap that the skill, as shipped, was *purely* human-in-the-loop (it replicated every CLI decision point as a chat confirmation, which an unattended agent cannot answer). The new mode is **auto-detected** from a positive automation signal and keeps every one of the CLI's safety gates: it drives `npx baldart update --yes` and **stops hard** (never forces a merge or reset) the moment the CLI refuses on custom divergence or a conflict — exactly the calls a human should own. **No new `baldart.config.yml` keys** — `BALDART_AUTONOMOUS` is a runtime env var, not a config flag, so the schema-propagation rule does not apply.

### Added — autonomous (unattended) mode for `/baldart-update`

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: new **Mode detection** first step (`printenv | grep -qE '^(BALDART_AUTONOMOUS|CI|GITHUB_ACTIONS)='`) branches the skill into INTERACTIVE (a human is present → existing chat-confirmation flow, unchanged) or AUTONOMOUS (no human → new flow). Detection is **positive-signal-only**: absence of every signal defaults to INTERACTIVE so a human never silently loses control. Documents *why* the TTY test (`[ -t 0 ]`) cannot be used — Claude's Bash tool never has a TTY on stdin, even with a human present, so it can't tell the two situations apart.
- **`## Autonomous mode`** section: pre-flight (shared Step 0 + CLI version drift gate) → `npx baldart update --yes` (bare, `timeout: 600000`) → exit-code interpretation → shared post-flight assert → structured log-friendly report. **Hard guardrail**: on a non-zero exit (custom-overlay-able / real-custom divergence, subtree-pull conflict, stale-CLI halt) the skill STOPS and reports the exact flag a human would re-run with — it **never** retries with `--on-divergence pull|scaffold-overlays` or `--reset` to force past a refusal. A refusal is a *correct* outcome that needs a human.

### Added — automation marker injected by the routine-adapters

- **[src/utils/routine-adapters/github-actions.js](src/utils/routine-adapters/github-actions.js)**: the generated workflow's run step now sets `BALDART_AUTONOMOUS: '1'` in its `env:` block (alongside `ANTHROPIC_API_KEY`), inherited by the `claude` process and every child Bash it spawns.
- **[src/utils/routine-adapters/cron.js](src/utils/routine-adapters/cron.js)**: the generated shell wrapper `export`s `BALDART_AUTONOMOUS=1` before invoking `claude`.
- **[src/utils/routine-adapters/claude-code-cloud.js](src/utils/routine-adapters/claude-code-cloud.js)**: the generated routine config now carries `env: { BALDART_AUTONOMOUS: '1' }` (best-effort — relies on the RemoteTrigger surfacing `env` into the run environment; github-actions / cron inject it directly).

### Changed — drift-check counter + honest framing

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: `last_verified` v3.26.0 → v3.31.0; added an **AUTONOMY NOTE** recording that `BALDART_AUTONOMOUS` is a runtime env var (not a config key) so the schema-propagation rule does not apply. Hard rules gain a rule #0 (run mode detection first; default-on-absence is INTERACTIVE); rule #2 now qualifies the chat-replication mandate as INTERACTIVE-only. The closing "both keep the user in the loop" claim is corrected to reflect that autonomous mode drops confirmations while keeping the safety gates.

## [3.30.0] - 2026-05-29

Hardens the **overlay** system — the layer that turns generic framework skills into project-specific knowledge. An audit surfaced a structural asymmetry: agent/command overlays are *compile-time merged* (the overlay is materially fused into the generated file → guaranteed delivery), but skill overlays use *runtime concatenation* (a line of prose tells the agent to read `.baldart/overlays/<skill>.md` → best-effort, nothing verifies it loaded). The most important overlays for knowledge transfer had the weakest guarantee. This release ships the high-certainty authoring/DX fixes now and adds **observation-only telemetry** to gather real data on whether skill overlays actually load — before deciding whether a stronger enforcement mechanism is worth building. **No new `baldart.config.yml` keys** — the telemetry hook is always-on (like `framework-edit-gate` / `agent-discovery-gate`), so the schema-propagation rule does not apply.

### Added — overlay-load telemetry (observation only)

- **[framework/.claude/hooks/overlay-telemetry.sh](framework/.claude/hooks/overlay-telemetry.sh)**: a fail-safe `PostToolUse` hook (matcher `Read`) that appends one JSON line to `.baldart/telemetry/overlay-loads.jsonl` every time a `.baldart/overlays/*.md` file is read. Objective, not self-reported — the failure mode being measured is "the agent forgot the overlay", which a self-report would corrupt in exactly that case. Written in **bash, not node**, on purpose: it fires on every `Read` (the hottest tool) in the critical path, so a per-read node cold-start (~50ms) would tax every file read — a bash fork (~5ms) keeps the hot path cheap. **Numerator-only by design**: skill invocation is prompt-expansion, not a tool call, so the denominator (skill runs that *should* have loaded an overlay) is not hook-observable (claude-code [#43630](https://github.com/anthropics/claude-code/issues/43630), [#22655](https://github.com/anthropics/claude-code/issues/22655)). Registered in `HOOK_REGISTRY` (`baldart-overlay-telemetry`) → auto-registered by `add`, backfilled by `update`, reported by `doctor` — no call-site changes (registry-driven).
- **[framework/scripts/analyze-overlay-loads.js](framework/scripts/analyze-overlay-loads.js)**: zero-dependency aggregator that cross-references the read log against the overlays that exist on disk. Headline signal: a **skill** overlay that exists but is read ~never (`NEVER-LOADED`). Agent/command overlays (compile-time merged) are deliberately not flagged on zero reads.
- **[framework/docs/OVERLAY-TELEMETRY.md](framework/docs/OVERLAY-TELEMETRY.md)**: full reading guide, with an explicit "conservative proxy — UNDER-counts" section (in-context reads and prompt-expansion are invisible to the hook; a low count is a reason to investigate, not proof of a miss).

### Changed — overlay authoring/DX parity for skills

- **[src/commands/overlay.js](src/commands/overlay.js)** `validate` (skill path): now renders a **merge dry-run** for skill overlays — section counts, concatenated size, and a per-marker check flagging `[OVERRIDE]/[APPEND]/[PREPEND]` headings that don't anchor to any base H2. Previously it printed only "valid, concat model", so a skill-overlay author never saw what the merge would look like (agent/command overlays already got this). Powered by a new pure, unit-tested `analyzeSkillConcat()` helper.
- **[src/commands/overlay.js](src/commands/overlay.js)** `drift` (skill path): when the base file changed, now lists *which* overlay markers were orphaned by the change (heading renamed/removed upstream) — the same actionable feedback agent/command overlays get from the merger, instead of a bare sha-mismatch warning.
- **[src/commands/overlay.js](src/commands/overlay.js)** `scaffoldBody` (skill path): the `overlay scaffold skill/<x>` skeleton is now a guided template — explains the runtime-concat model honestly, enumerates the typical project-fact sections (canonical paths, domain vocabulary, BLOCKING rules, mandated stack), and shows a ready-to-edit `[OVERRIDE]` block. Improves authoring for all skills, not just the two that ship an example overlay.
- **[framework/agents/project-context.md](framework/agents/project-context.md)** §5 precedence rules: corrected the overclaim that a skill overlay's `## [OVERRIDE] <topic>` "replaces" the base section. For skills nothing parses the markers — they are textual conventions the executing agent honours (best-effort), unlike agent/command overlays which are materially fused (§5.bis). The doc now matches `scaffoldBody`'s honest framing instead of the middle.

### Changed — drift-check counter

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: `hook_registry_entries` 3 → 4 (the new telemetry hook). The `add`/`update`/`doctor` integration is registry-driven, so no narration changes were needed.

### Added — tests

- **[src/utils/__tests__/analyze-skill-concat.test.js](src/utils/__tests__/analyze-skill-concat.test.js)**: 5 cases covering marker matching, orphan detection, plain (new) sections, line counts, and the empty-overlay edge case.

## [3.29.0] - 2026-05-29

Closes a two-part process gap in `npx baldart update` reported from a non-TTY agent context on a shared consumer repo: (1) plumbing/collaboration merge commits were mislabeled `custom-other`, downgrading an otherwise overlay-able divergence set into `mixed`; (2) an automated agent had **no** non-destructive path to complete an update when the consumer carried local `.framework/` commits — `--yes` refused, the interactive prompt needed a TTY, and `--reset` is gated on a clean working tree. **No new `baldart.config.yml` keys** — `--on-divergence` is a CLI flag, not a config flag, so the schema-propagation rule does not apply.

### Fixed — divergence classifier no longer poisoned by merge commits

- **[src/utils/git.js](src/utils/git.js)** `classifyDivergence()`: merge commits are now detected **structurally by parent count** (`%P` in the `git log` format → `parents.length >= 2`) instead of relying solely on the brittle subject regex `^Merge … into <branch>$`. That regex missed real `git subtree pull --squash` merges that present as a bare `Merge commit '<sha>'` (git omits the `into <branch>` suffix when the merge lands on `master`) and shared-repo collaboration merges shaped like `Merge branch 'x' of <url> into <branch>` / `Merge remote-tracking branch '…'`. Such merges fell through to the path classifier, which returns `custom-other` for the empty diff-tree of a merge — collapsing an `overlay-able` set into `mixed` and hiding the overlay-migration option. Merges are now labeled `subtree-merge` (noise) and excluded before the aggregate decision.
- **[src/utils/git.js](src/utils/git.js)**: aggregation extracted to a pure, unit-testable static `GitUtils.aggregateDivergenceClass(commits)`. Each divergence commit record now also carries `parents` and `touched` (the latter reused by the new overlay-scaffold flow below).
- **[src/utils/__tests__/classify-divergence.test.js](src/utils/__tests__/classify-divergence.test.js)**: +9 fixtures — bare-merge / `of <url>` / remote-tracking subjects (documenting the regex blind spot now covered by parent count) and an `aggregateDivergenceClass` suite asserting `overlay-able + plumbing-merge → overlay-able` (not `mixed`).

### Added — `--on-divergence` non-interactive escape hatch for agents / CI

- **[src/commands/update.js](src/commands/update.js)** + **[bin/baldart.js](bin/baldart.js)**: `npx baldart update --on-divergence <strategy>` resolves a divergent update without a TTY:
  - `scaffold-overlays` → auto-creates overlay skeleton(s) for the overlay-able framework files touched by the divergent commits (reusing `overlay.scaffoldFile`), then stops so the edits can be filled in and the update re-run. For `mixed`, scaffolds the overlay-able subset and refuses to pull while `custom-other` commits remain (would be destructive).
  - `pull` → keeps the commits and lets the subtree pull create a merge (non-destructive; conflicts are surfaced, never auto-resolved). Works for `overlay-able`, `mixed`, and `real-custom`.
  - `abort` → explicit no-op.
- Bare `--yes` on a divergent class still refuses (nothing destructive by accident) but now prints the exact flag to re-run with instead of dead-ending.

### Changed — interactive `overlay` choice now actually scaffolds

- **[src/commands/update.js](src/commands/update.js)**: the interactive "migrate to overlay" option previously printed *"run `/overlay` then re-run"* and exited (a no-op). It now scaffolds the overlay skeleton(s) in-place via `overlay.scaffoldFile` and points the user to `/overlay` for guided filling.
- **[src/commands/overlay.js](src/commands/overlay.js)**: extracted a non-exiting `scaffoldFile(cwd, target, options)` (returns a structured result) from the `process.exit`-driven `scaffold()` CLI wrapper, so it can be called in a loop from `update` without a single missing/existing overlay terminating the run. Exported alongside `parseTarget`.

### Changed — `/baldart-update` skill narration synced

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: documents the parent-count merge detection and the new `--on-divergence` strategies; corrects the stale "real user-custom commits ABORT in --yes mode" claim.

## [3.28.3] - 2026-05-28

Release B of the 3-release plan on `/new` fix-application authority. Closes two inline-apply violation patterns observed in real `/new` runs (doc fixes applied by the orchestrator instead of delegated to coder; Codex HIGH security findings written inline) and shuts the structural loophole in `SKILL.md`'s sub-agent failure protocol that was being used to bypass delegation even when no agent had crashed. Telemetry from v3.28.2 will measure the effectiveness of these changes; Release C (conditional threshold rule) waits for that data.

### Added — Domain-Override Domains sub-section

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: new sub-section under "Fix Application Log Schema" enumerates tassativamente the domains where inline orchestrator apply is **never** safe, regardless of patch size: `doc` (any `*.md` under references/, prd/, root CHANGELOG, ssot-registry), `security` (Phase 3.7 detector Triggers #2/#3 + SQL RLS policy mutations), `migration` (`supabase/migrations/*.sql` or `${paths.migrations_dir}`). The coder agent's system prompt + project overlay enforces invariants (freshness, tabular formatting, RLS structure) that orchestrator inline edits routinely break. Edge case explicit: mechanical CHANGELOG / ssot-registry append-a-row stays `doc` (uniformity > spawn cost).

### Changed — Phase 3 doc delegation reinforced

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** Phase 3 step 15: explicit chiosa "Doc fixes are NEVER applied inline by the orchestrator, regardless of size or perceived triviality" added below the existing coder-spawn instruction. The rule already existed (line 1116, pre-v3.28.2: "invoke the coder agent once") but was being bypassed in practice — the chiosa makes the rule non-discretionary.

### Changed — Phase 3.7 coder applies Codex patches verbatim

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** Phase 3.7 step 4 sub-bullet "spawn coder": coder now receives the report path + list of VERIFIED bugs **plus the patches Codex suggested inline in the report**, and applies them verbatim. The coder does NOT re-do the analysis, does NOT re-grep, does NOT re-read files extensively. Codex has already produced both the diagnosis and the patch; the coder's value-add is the right system prompt (project conventions, naming, testing patterns), not redoing security/correctness review. Saves an estimated 5-10k tokens per Codex-driven coder spawn.

### Changed — Sub-agent failure protocol (loophole closed)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** "Sub-agent failure protocol" section: replaced single bullet "attempt the work yourself directly" with a 4-step protocol. (1) log failure with full trace, (2) retry once (transient errors are common), (3) **if domain-override (doc / security / migration) and retry still fails: STOP and AskUserQuestion — never inline-fallback for these domains**, (4) non-override domains may fall back inline but MUST log `applied_by=orchestrator-fallback` for telemetry visibility. Razionale: the pre-v3.28.3 wording was being used as a shortcut to bypass delegation rules even when no agent had crashed; closing that defeats the violation pattern at the structural level.

### Rationale

Release B addresses 4 of the 17 critiques raised in the adversarial review of the original threshold-rule plan (#4-5 domain override definition, #7 Codex re-analysis ridondante, #10 escape valve aperta). The remaining critiques are either addressed by Release A telemetry, deferred to conditional Release C, or accepted as design trade-offs documented in the plan file. **No new `baldart.config.yml` keys** — domain-override is enforced by content/path matching, not by config flag (KISS).

## [3.28.2] - 2026-05-28

Telemetry-only release. Adds structured logging for fix-application decisions in `/new` Phases 2.55 / 3 / 3.5 / 3.7. **Zero behavior change** — the orchestrator's decisions remain identical to v3.28.1. The goal is to measure, with real data, whether the per-phase delegation rules need to change in a future release.

### Added — `## Fix Application Log` section in tracker

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: new section **"Fix Application Log Schema (telemetry)"** inserted before Phase 2.55. Defines a pipe-separated row format (`<phase> | <domain> | est_lines=<n> | decision=<...> | applied_by=<...> | <key=val>...`) appended to the tracker's `## Fix Application Log` section. Phases 2.55 / 3 / 3.5 / 3.7 each instructed to log one row per finding processed (including skipped). Cost: ~1-2 KB per card, negligible context impact.
- Per-phase logging instructions added to Phase 2.55 step 4 (simplify reuse/quality/efficiency), Phase 3 step 15 (doc), Phase 3.5 step 23 (qa-sentinel blockers), Phase 3.7 step 4 (codex BLOCKER/HIGH + false-positive-filtered). Each phase classifies finding `domain` consistently with the schema enum.

### Added — Aggregator script

- **[framework/scripts/analyze-fix-application.js](framework/scripts/analyze-fix-application.js)**: zero-dependency Node script. Reads N tracker files (shell-expanded glob), parses rows in the `## Fix Application Log` section, aggregates by `phase × domain × decision × applied_by`, outputs a tabular summary plus an automatic violation footer flagging inline-apply on doc / security / migration domains. Usage: `node framework/scripts/analyze-fix-application.js docs/references/trackers/*.md`.

### Added — Telemetry reading guide

- **[framework/docs/FIX-APPLICATION-TELEMETRY.md](framework/docs/FIX-APPLICATION-TELEMETRY.md)**: new doc. Explains why telemetry exists, the row schema (cross-references SKILL.md as SSOT), how to read the analyzer output, violation patterns vs delegation candidates vs healthy patterns, and the trigger conditions for the future Release C decision (threshold rule vs Phase 2.55→3 merge).

### Rationale

This release is the first of a 3-release sequence (A: telemetry → B: loophole + 2 known violations → C: conditional threshold rule). The decision to instrument *before* changing rules avoids inventing arbitrary thresholds without data — see plan file at `/Users/antoniobaldassarre/.claude/plans/sto-notando-una-cosa-misty-lightning.md` for full reasoning. Schema-change propagation rule: **no new `baldart.config.yml` keys** in this release.

## [3.28.1] - 2026-05-28

Patch su tre drift segnalati dal curator dopo v3.28.0 install in `mayo`:
(1) `--yes` non si propaga da `update --reset` al child `baldart add`,
(2) errore "Configuring alias is not permitted without enabling allowUnsafeAlias" stampato come install-failed mentre l'install era OK,
(3) due template `ui-guidelines.template.md` + `brand-guidelines.md` lasciati untracked senza istruzioni chiare all'utente. Tutti e tre cosmetici/UX, nessuno affecta la correctness del framework installato — ma confondono l'utente durante il post-install.

### Fixed — `baldart add` honors `--yes` from parent caller

- **[src/commands/add.js](src/commands/add.js)** + **[src/commands/update.js:237](src/commands/update.js)**: `add(repo, options)` ora legge `options.yes` e auto-conferma i prompt non-distruttivi (`Install framework?`, `Configure git aliases?`, `Configure project context now?`). Il prompt distruttivo `Remove and reinstall?` resta interattivo anche con `--yes` (safety). `update --reset` ora invoca correttamente `addCmd(undefined, { yes: true })` — prima passava `{ yes: true }` come **primo argomento** (cioè `repo`), facendo `options` undefined e ogni prompt scattava.

### Fixed — Git alias config failure non più mascherato da install-failure

- **[src/commands/add.js](src/commands/add.js)**: la configurazione degli alias `fw-version`/`fw-update`/`fw-push` passa da `git.git.addConfig()` (simple-git, che rifiuta i bang-alias senza `allowUnsafeAlias`) a `spawnSync('git', ['config', key, value])` diretto. Il bypass del check simple-git è esplicito qui: l'utente sta opt-in agli alias durante l'install. Failure isolato per alias (non bubble-up al catch globale), reported come WARNING se non-zero count, install continua. Il messaggio "✗ Installation failed: Configuring alias is not permitted…" non appare più.

### Fixed — Template editable post-install: istruzioni chiare

- **[src/commands/add.js](src/commands/add.js)** NEXT STEPS box: i due template `docs/references/ui-guidelines.template.md` e `docs/references/brand-guidelines.md`, copiati incondizionatamente durante `copyCustomizableFiles()`, ora sono elencati esplicitamente con cosa farne (rename / edit / delete) e con il punto chiave: "Commit them to track your project's design language, or add them to .gitignore if you don't need them." Evita la confusione del curator durante v3.28.0 install ("non chiari se debbano essere scelti/scartati dall'utente o committati. Non c'è prompt che lo chieda").

## [3.28.0] - 2026-05-28

Tre additions indipendenti accumulate nel working tree e pronte per il rilascio: convenzione testuale per le dipendenze MCP, supporto `.zip` per il mockup intake del PRD, e nuovo hook `capture-detect` come parte del proactive mode dello skill `/capture` (registrazione opt-in, non auto-attivata). Tre temi distinti ma raggruppati in un'unica release di accumulo perché tutti additive e zero-schema-change.

### Added — Convenzione MCP integration

- **[framework/docs/MCP-INTEGRATION.md](framework/docs/MCP-INTEGRATION.md)**: nuovo documento centralizzato. Spiega quando un agent/skill deve appoggiarsi a un MCP server, fornisce un catalogo di MCP raccomandati per i casi d'uso comuni (doc-rag, playwright, firestore/db, LSP, …), e definisce la convenzione testuale "MCP dependencies" per dichiarare le dipendenze nel body dell'agent/skill. Zero nuove chiavi config in `baldart.config.yml` — la convenzione resta testuale finché una skill concreta non dovrà gate-arne il comportamento (regola di schema-change propagation: niente layer drift).
- **[framework/.claude/agents/codebase-architect.md](framework/.claude/agents/codebase-architect.md)**: aggiunta sezione `## MCP dependencies` con classificazione Required (`doc-rag` come primo step dell'Investigation Protocol) vs Optional (LSP find-references, DB MCP per inspect runtime). Documentato il degrade path quando l'MCP non è disponibile.
- **[framework/.claude/skills/bug/SKILL.md](framework/.claude/skills/bug/SKILL.md)**: aggiunta sezione `## MCP dependencies` con Required (`playwright` per Phase 1 reproduction di UI/Client bug — senza è impossibile catturare stato deterministico) e Optional (`doc-rag` per known patterns, DB MCP per Data bugs, `codebase-architect` agent).

### Added — PRD mockup intake da archivio `.zip`

- **[framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md)** Step 1.6: nuovo formato `local-archive` accettato accanto a `chat-images` / `local-paths` / `mixed`. Caso d'uso: export "Download as .zip" da Claude Design (o qualsiasi altro tool che produce bundle di mockup) già scaricato sul disco dell'utente.
- **[framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md)** § Mockup Intake: Step 1.6.4b descrive l'intero flow per `.zip` — `unzip` in scratch dir `mockups/_archive/`, enumerazione candidati visivi, scoring di rilevanza contro lo scope del Kickoff, selezione auto-proposta, STOP per conferma, materializzazione del sottoinsieme confermato in `${paths.prd_dir}/<slug>/mockups/`. Evita di copiare PNG irrilevanti nell'asset directory definitiva.
- **[framework/.claude/skills/prd/assets/state-template.md](framework/.claude/skills/prd/assets/state-template.md)**: aggiornato `mockups.format` enum per includere `local-archive`.

### Added — Hook `capture-detect` (proactive mode, opt-in)

- **[framework/.claude/hooks/capture-detect.sh](framework/.claude/hooks/capture-detect.sh)**: nuovo Stop hook che drop una candidate JSON in `.claude/capture-queue/` quando rileva ragionamento cross-document nella sessione appena chiusa. La distillazione vera in synthesis page wiki resta sempre una human-approved manual step via lo skill `/capture`. Failure-isolated (exit 0 sempre, non blocca lo Stop event).
- **Registrazione opt-in (NOT auto-attivata)**: la entry corrispondente in [src/utils/hooks.js](src/utils/hooks.js) resta commentata (linee 90-97). Lo skill `/capture` funziona già in modalità manual senza il hook; il proactive mode è una feature opzionale. Per attivarlo nel proprio consumer: scommentare la entry e re-installare via `npx baldart update`. Una futura release potrà rendere il proactive il default.

### Added — Smoke test del hooks registry

- **[scripts/test-hooks.sh](scripts/test-hooks.sh)**: nuovo smoke test per `src/utils/hooks.js` (multi-hook registry introdotto v3.18.0). Verifica register/unregister/isRegistered/merge idempotente con fixture in tmpdir. Dev tooling, non user-visible.

## [3.27.1] - 2026-05-28

Patch follow-up alla v3.27.0 dopo l'incidente reale in `mayo`: l'utente aveva CLI globale v3.24.0 e ha eseguito `/baldart-update` per aggiornarsi a v3.27.0. La skill ha rilanciato `npx baldart update`, che ha aggiornato il framework payload da v3.24 a v3.27, ma `update.js` in esecuzione era il binario v3.24.0 — quello che precede sia il meccanismo di auto-relaunch (introdotto v3.26.0) sia la logica di symlink-indirection (v3.27.0). Risultato: i 6 agent overlay-merged sono stati rigenerati con la **vecchia logica** (file regolari direttamente in `.claude/agents/`), e quindi il fix v3.27.0 NON si è attivato — `npx baldart doctor` ha riportato "tutto healthy" ma in realtà gli agent erano ancora invisibili a Claude Code (bug #20931). Chicken-and-egg: il CLI vecchio non sa rilanciare il nuovo, il nuovo non è ancora in esecuzione.

### Fixed — /baldart-update skill: CLI version drift gate

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)** Step 0: la skill ora rileva esplicitamente il CLI version drift e HALTA pre-update quando il CLI globale è < v3.26.0 (no auto-relaunch) e il payload sta per attraversare il boundary v3.27.0 (migration breaking). Decision logic articolata: se CLI è già >= v3.26.0 lascia che l'auto-relaunch faccia il suo lavoro; altrimenti istruisce `npm i -g baldart@latest` PRIMA di ri-eseguire la skill. Evita il silent-fail dove il CLI vecchio applica una pseudo-migration con codice obsoleto.

### Note al curator

Questo è un edge case strutturale che si manifesta SOLO quando un consumer attraversa il boundary v3.26.0 / v3.27.0 partendo da un CLI globale precedente alla v3.26.0. I consumer già su CLI v3.26.0+ hanno l'auto-relaunch (`npx baldart@latest`) e non sono affected. Per i consumer pre-v3.26.0 il fix in skill aiuta solo le sessioni future (skill caricata da framework payload >= v3.27.1); chi aveva già la skill vecchia in sessione al momento dell'incidente — come `mayo` ieri — può solo riparare manualmente: `npm i -g baldart@latest` + restart sessione CC + `/baldart-update`.

## [3.27.0] - 2026-05-28

Fix critico al meccanismo di overlay-merge degli agent. Sintomo riproducibile osservato in `mayo` durante `/new FEAT-0009 -full` (team mode, claude-opus-4-7): gli agent overlay-merged (`coder`, `ui-expert`, `code-reviewer`, `doc-reviewer`, `qa-sentinel`, `security-reviewer` quando hanno un overlay) **non comparivano nella lista "Available agent types"** del Claude Code session reminder. Sequenza concreta: `codebase-architect` (symlink) parte e completa; i 4 background agent (`coder×2` + `ui-expert×2`) lanciati subito dopo falliscono con `InputValidationError: subagent_type not recognized`; l'orchestratore halta con un AskUserQuestion fuori-protocollo. Repro: `ls -la .claude/agents/` mostra 21 symlink + 6 file regolari (i 6 overlay-merged); CC discoverizza i 21 symlink e omette i 6 file regolari. Confermato su sessioni indipendenti del 2026-05-27 e 2026-05-28 → bug strutturale, non transiente. Workflow `/new` in team mode, `/check`, `/qa`, `/codexreview` e skill `prd` erano **bloccati end-to-end** su qualsiasi consumer con overlay agent.

### Root cause — Claude Code upstream bug #20931

Il session-start discovery scanner di Claude Code che costruisce la registry di `subagent_type` validi **skippa silenziosamente i file regolari** in `.claude/agents/` (solo i symlink vengono caricati). Il bug è già documentato in `framework/.claude/hooks/agent-discovery-gate.js:10-15`:

> #20931 — file-based discovery of `.claude/agents/*.md` is documented as BROKEN in some scenarios (files ignored at session start)

Dalla v3.8.0 in poi BALDART ha generato i file overlay-merged come **file regolari** direttamente in `.claude/agents/<name>.md`, intersecando esattamente la zona di rottura di CC. Tutti gli overlay agent attivi nei consumer sono stati silenziosamente invisibili per ~3 mesi, e il sintomo emergeva solo quando una skill provava ad invocare lo `subagent_type` (es. team-mode `/new`).

Diagnosi precedente erroneamente attribuita a "HTML comment prima del frontmatter rompe il parser YAML" è stata rivista: il riordino del marker post-frontmatter resta una pulizia di formato corretta ma **non era la causa primaria** del sintomo.

### Fixed — Symlink indirection per agent/command overlay-merged (root cause fix)

- **[src/utils/symlinks.js](src/utils/symlinks.js)** `_generateOverlayedFile()`: il content merged ora vive in `.baldart/generated/<kind>s/<name>.md` (file regolare), e `.claude/agents/<name>.md` (o `.claude/commands/<name>.md`) viene creato come **symlink relativo** verso quel file. Claude Code vede il path consumer come un symlink come tutti gli altri e lo discoverizza correttamente. Stesso pattern per Codex (`.agents/skills/`) tramite il tool adapter.
- **[src/utils/symlinks.js](src/utils/symlinks.js)** `_installPerItemSymlink()`: quando `hasOverlay` viene rimosso e il path consumer è un symlink che punta in `.baldart/generated/`, il CLI relinka al framework base e cancella lo stale generated file. Coerenza forte tra "overlay esiste" e "esiste un generated file".
- **[src/utils/symlinks.js](src/utils/symlinks.js)**: migrazione automatica del legacy layout. Se il consumer ha un file regolare baldart-generated direttamente in `.claude/agents/<name>.md` (pre-v3.27.0), la prossima `npx baldart update` lo riconosce dal marker, sposta il content in `.baldart/generated/agents/<name>.md` e crea il symlink. Nessun intervento manuale richiesto. User-fork (file regolare senza marker baldart) restano protetti.

### Added — Safety net SessionStart hook

- **[framework/.claude/hooks/agent-discovery-info.sh](framework/.claude/hooks/agent-discovery-info.sh)**: oltre a rilevare gli agent **mancanti** (file assente), il hook ora rileva anche gli agent in **stato degradato** — ovvero file regolari con marker `<!-- baldart-generated: -->` presenti in `.claude/agents/` (pre-v3.27.0 layout). Quando rilevati, emette `additionalContext` esplicito al modello con il nome degli agent affected, la causa (upstream bug #20931) e l'azione richiesta (`npx baldart update`). Difesa-in-profondità: il `mergeAgents` migra automaticamente e questo hook copre l'edge case in cui un consumer non lanci `update` subito dopo l'upgrade.

### Fixed — Formato file generato (cleanup secondario, non root cause)

- **[src/utils/overlay-merger.js](src/utils/overlay-merger.js)**: i file generati hanno il frontmatter YAML su riga 1 e il marker `<!-- baldart-generated: -->` subito dopo il `---` di chiusura (era prima invertito). Allinea il formato alla convenzione standard CommonMark + YAML frontmatter, indipendentemente dal fatto che CC sia o meno strict sul parser.
- **[src/utils/overlay-merger.js](src/utils/overlay-merger.js)**: nuovo export `isGeneratedFile(content)` posizione-agnostico. `readMarker()` cerca il marker in tutto il file (cap 4KB) per leggere sia file v3.27.0+ sia legacy pre-v3.27.0 (retrocompat fondamentale per la migration).
- **[framework/.claude/hooks/framework-edit-gate.js](framework/.claude/hooks/framework-edit-gate.js)**: il PreToolUse hook controlla il marker via `head.includes` su slice 5KB, così intercetta edit ai file generati sia che il marker sia su riga 1 (legacy) sia dopo il frontmatter (v3.27.0+).

### Migration path per i consumer

Su un consumer con overlay agent in layout legacy (es. `mayo`):

1. `npx baldart update` rileva la nuova VERSION e lancia `mergeAgents`.
2. Per ogni file regolare baldart-generated in `.claude/agents/`: marker letto → content spostato in `.baldart/generated/agents/<name>.md` → simlink creato.
3. `git status` mostra: `.claude/agents/<name>.md` modificato (file → symlink), `.baldart/generated/agents/<name>.md` nuovo. La `_autoCommitFn` del CLI committa con `chore(baldart): post-update reconcile to v3.27.0`.
4. **Restart della Claude Code session** richiesto perché la registry di `subagent_type` viene letta al session-start.
5. Verifica: `Available agent types` ora include tutti i nomi (sia symlink che generated-symlink). `/new` in team mode funziona.

Nessuna perdita di overlay: i file in `.baldart/overlays/agents/*.md` sono consumer-owned e non vengono toccati. Lo skill `/baldart-update` cita esplicitamente nel report post-flight gli agent migrati e ricorda di restartare la sessione CC.

## [3.26.0] - 2026-05-27

Fix definitivo del processo `baldart update` dopo che un'esecuzione reale su un consumer (`mayo`) ha esposto un workflow strutturalmente rotto. Sintomi: `update --yes` stampava "Already up to date!" mentre `version`/`doctor` riportavano correttamente "Remote: 51 commit ahead"; CLI globale (v3.18.1) silenzioso vs npm latest (v3.24.0); divergenza subtree (19 commit locali + 51 upstream) mai menzionata; output sporcato da `ExperimentalWarning` di npm a ogni invocazione; skill `/baldart-update` che narrava il workflow rotto senza accorgersene. Root cause: tre superfici (`update`, `version`, `doctor`) implementavano TRE check diversi per la stessa domanda, e nessuna era giusta — `update.js:193` usava `hasChangesToPush()` (commit LOCALI non pushati al CONSUMER's origin, niente a che vedere con BALDART upstream); `version`/`doctor` usavano `HEAD...FETCH_HEAD` full-repo che include sempre subtree-merge noise → falso positivo perenne "52 commit ahead".

### Fixed — Bug "Already up to date" su consumer con update disponibili

- **[src/utils/git.js](src/utils/git.js)** + **[src/commands/update.js](src/commands/update.js)**: nuova `getUpdateStatus(repo, branch)` SSOT basata su **VERSION compare** (`git show FETCH_HEAD:VERSION` vs `.framework/VERSION`), non sul commit count (che è strutturalmente rumoroso sul subtree). `update.js` Step 2 ora usa `status.isAligned` come segnale autoritativo. Eliminata la chiamata a `git.hasChangesToPush()` (rimane per compatibilità di `push.js`).

### Fixed — Falsi positivi "52 commit ahead / 20 unpushed" perenni

- **[src/commands/version.js](src/commands/version.js)** + **[src/commands/doctor.js](src/commands/doctor.js)**: messaging riscritto su `isAligned`. `Remote: v3.26.0 (aligned)` quando OK, `Remote: v3.26.0 — update available (installed v3.22.1)` altrimenti. Il commit count migra dietro `version --verbose` (etichettato "subtree commit history — includes auto-generated merges"). Doctor `Local changes` mostra `unpushed commit(s)` SOLO se non-aligned (altrimenti tutto subtree-noise → riga `clean`).

### Added — Classifier divergenza subtree (Fix 2)

- **[src/utils/git.js](src/utils/git.js)**: nuova `classifyDivergence()` che ispeziona i commit consumer non in upstream e li categorizza branch-agnostic:
  - `subtree-merge` (`Merge commit '*' into <any-branch>`) → auto-resolve silente.
  - `subtree-squash` (`Squashed '.framework/' changes from *`) → auto-resolve silente.
  - `chore-wrapper` (`[CHORE]` / `chore(baldart):` con `baldart` keyword nel subject o scope) → auto-resolve silente.
  - `custom-overlay-able` (modifiche SOLO a file in `.framework/.claude/{agents,skills,commands}/`) → prompt: "Migra a overlay (raccomandato)".
  - `custom-other` → prompt 3-way (push first / pull anyway / abort).
- Aggregate class `all-noise | overlay-able | real-custom | mixed | unknown`. Default conservativo `custom-other` per commit non riconosciuti — mai data loss silente.
- Test fixture-driven: **[src/utils/__tests__/classify-divergence.test.js](src/utils/__tests__/classify-divergence.test.js)** con 24 fixture inclusi i pattern commit reali osservati nel debug di `mayo` (`[CHORE] reconcile baldart state ledger v3.22.1 -> v3.25.0`, `[CHORE] register baldart agent-discovery hooks`, ecc.).

### Added — CLI self-upgrade auto-relaunch (Fix 4)

- **[src/commands/update.js](src/commands/update.js)** `maybeRelaunchUnderLatest()`: se il CLI globale è dietro npm latest, `update` rilancia trasparentemente `npx baldart@<latest> update <flags>` con `stdio: inherit` (preserva prompt). Loop guard via env `BALDART_RELAUNCHED=1` (impedisce ricorsione se npm cache serve ancora un vecchio). Fallback al CLI corrente se npx fallisce. Opt-out: `BALDART_NO_RELAUNCH=1`. Razionale rispetto a `update-notifier.js`: l'avvertenza "no auto-update globale" si riferisce a `npm i -g` silente; `npx@latest` è per-invocation, esplicito, opt-out → policy intatta.

### Added — `--reset` nuclear option (Fix 5)

- **[bin/baldart.js](bin/baldart.js)** + **[src/commands/update.js](src/commands/update.js)** `runReset()`: `npx baldart update --reset` rimuove `.framework/` e reinstalla pulito, preservando `baldart.config.yml` + `.baldart/` (overlays, state.json) + `.claude/settings.json` + custom agents/skills/commands non-symlink. Safety gate obbligatorio: working tree clean (refusal se dirty), prompt esplicito per file untracked/ignored in `.framework/` (in `--yes` richiede ALSO `--i-know`), backup tag pre-rm. Post-restore sanity check: se un file user-owned è scomparso, refuse + suggest `git reset --hard <backup-tag>`.

### Added — Soppressione `ExperimentalWarning` (Fix 6)

- **[bin/baldart.js](bin/baldart.js)**: `process.env.NODE_NO_WARNINGS = '1'` impostato ALL'INIZIO (prima di ogni `require`). Propagato a tutti i child process spawnati dal CLI (git, npx, ecc.). Verificato su Node 23.3.0: silenzia il warning. Escape hatch: `BALDART_VERBOSE_WARNINGS=1` lo riabilita. **Limite noto**: non silenzia il warning emesso da `npx` PRIMA che il nostro codice carichi — l'unico workaround è `npm i -g baldart` (modalità raccomandata) + invocazione diretta `baldart`.

### Changed — Auto-backfill hook missing senza prompt (Fix 7)

- **[src/commands/update.js](src/commands/update.js)** logging: hook missing aggregati in una singola linea (`Auto-registered N missing hook(s): a, b, c`) invece di una per hook. La logica `Hooks.registerAll` era già corretta (auto-create silenzioso per missing, prompt solo per drift) — il cambio è cosmetico ma riduce noise post-update.

### Changed — Skill `/baldart-update` parser tollerante + post-flight assert (Fix 9)

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: Step 0 parser tollerante con 5 pattern in ordine (nuovo format `Remote: vX.Y.Z (aligned)`, format intermedio `update available`, legacy `N commit(s) ahead` / `up to date`, offline) — gestisce la finestra di transizione skill v3.26.0 + CLI globale v3.18.1. Se nessun pattern matcha → istruisce `npm i -g baldart@latest`. Eliminato Step 3 (hooks drift — ora auto-backfill in CLI) e Step 4 (auto-commit prompt — sempre yes in `--yes`). Step 6 ora include **post-flight assert obbligatorio**: confronta `.framework/VERSION` pre/post, FAIL LOUD se exit 0 ma VERSION invariata (safety net contro regression del Fix 1).

### Changed — Doctor planner usa version-compare authority

- **[src/commands/doctor.js](src/commands/doctor.js)** `planActions()`: `remoteAhead = state.remote.fetched && state.remote.isAligned === false`. Label dell'action `update` ora `Update framework (vX.Y.Z → vA.B.C)` invece di `Pull N commit(s) from upstream`. Local-work detection sopprime commit count quando aligned.

### Internal

- Nessuna nuova chiave in `baldart.config.yml`. Schema-change propagation rule non applicabile.
- `getRemoteVersion()` esistente non più consumato — VERSION compare ora dentro `getUpdateStatus()` via `git show FETCH_HEAD:VERSION`. La funzione resta per back-compat (status.js la chiama).
- `hasChangesToPush()` resta in `git.js` per `push.js` che la usa correttamente (consumer's origin commits to push upstream).

## [3.25.0] - 2026-05-27

Strict-phase-gating + workspace-hygiene gates per `/new`. Risolve l'incident **FEAT-0006**: epic team-mode (4-layer L0→L1×5→L2×2→L3×2, 10 sub-card) che ha saltato per ogni card le fasi MANDATORY 2.5b (AC-Closure), 2.55 (Simplify), 2.6 (E2E-Review), 3 (Doc-Review), 3.5 (QA-Sentinel), 3.7 (Codexreview) — giustificazione documentata come "time budget", parola **inventata dal modello a runtime**, non esistente nella skill. In aggiunta, AC-IMG-3 "deferred dal coder 09" senza passare dal gate, e main repo lasciato con orphan commit non pushato + local `develop` diverged. Tre root cause radicate: (1) Phase 2.5b non era propagata in team-mode Step D, (2) coder `completion-report` accettava `status: partial`/`blocked` come terminale, (3) zero gate di workspace hygiene su `$MAIN`.

### Added — Phase 0: Workspace Hygiene Pre-flight (BLOCKING)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: nuova `## Phase 0` prima di `## Pre-flight (once)`. Su `$MAIN`: `fetch origin`, dirty-tree gate (`git status --porcelain` → `AskUserQuestion` con opzioni `[stash auto] / [commit ora] / [procedi senza stash] / [abort]`), divergence gate (`rev-list --left-right --count` → `AskUserQuestion` per ahead/behind/diverged). Stash ref salvato nel tracker sotto `## Workspace Snapshot` per il restore di Phase 6c. Anti-bypass guard finale che rifiuta di proseguire senza `## Phase 0: PASS` nel tracker. Auto Mode NON override.

### Added — Phase 6c: Workspace Hygiene Post-merge (BLOCKING)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: nuova `### Phase 6c` dopo Phase 6b. Parse del marker `[SYNC-DEFERRED]` emesso da `/mw`, clean-tree assertion, divergence assertion (cattura il pattern FEAT-0006 orphan commit con `AskUserQuestion` opzioni `[push adesso] / [cherry-pick] / [reset --hard con conferma esplicita] / [halt]`), restore dello stash di Phase 0 (con `AskUserQuestion` su conflict, mai silent drop). Anti-bypass guard come in Phase 0.

### Added — Team-mode Step D rewrite (FEAT-0006 anti-regression)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** righe ~1988: nuovo Step D con clausola di apertura esplicita ("La lean Step D è ora FORBIDDEN") + sub-step MANDATORY per-card propagati identici al sequential mode:
  - **D.3a** — Phase 2.5b AC-Closure Gate per-card (sequenziale), bloccante.
  - **D.3b** — Phase 2.55 Simplify per-card (diff per-card per attribuibilità).
  - **D.3c** — Phase 2.6 E2E-Review per-card (honora il Gate table esistente — skip backend-only/docs/chore con motivazione documentata, mai con "time budget").
  - **D.4a** — Phase 3 Doc Review per-card (oltre al D.2 combined).
- **Coverage assertion** end-of-group MANDATORY: prima di Step E, il tracker deve contenere per OGNI card le entry `AC Closure Ledger`, `simplify`, `e2e-review`, `doc-review`, `qa`, `codex-review`. Entry mancante → ritorno alla sub-step mancante. Documentare "skipped per time budget" è violazione esplicita.

### Added — Top-of-file rigidity clause + tracker header

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: nuovo blockquote `> **NO PHASE SKIP FOR PERCEIVED TIME, EVER (BLOCKING)**` subito sotto SCOPE CLOSURE DISCIPLINE. Enumera esplicitamente le phrase forbidden (`time budget`, `context budget`, `context pressure as override`, `for time`, `to save tokens`, `to save iterations`). Self-correction trigger: se il modello produce una di queste phrase nel proprio output, deve fermarsi e routare via `AskUserQuestion`.
- **Tracker template** (~riga 50): HTML comment header che sopravvive alle compaction e ricorda la regola dopo ogni context recovery.

### Changed — `coder` agent completion report (semantica binaria)

- **[framework/.claude/agents/coder.md](framework/.claude/agents/coder.md)**: card-level `status:` enum ridotto a `done | failed` (era `done | partial | failed`). Per-requirement e per-AC `status:` ridotto a `done | not_implemented` (era `done | partial | blocked`). Le legacy values `partial`/`blocked` erano usate come canali di silent deferral. Aggiunta a `Forbidden Actions`: "Don't defer acceptance criteria" + "Don't invent skip reasons". Retro-compatibile: coder che emettono ancora `partial` faranno triggerare il gate Phase 2.5b invece di passare in silenzio.

### Changed — `framework/agents/workflows.md` Scope Closure Discipline

- Nuovo `### Coder authority (no defer)` in coda alla sezione: chiarisce che coder e orchestratori L0/L1 relay-ano i signal nel gate, non li aggregano/filtrano. Citable da skill e agent.

### Changed — `/mw` `[SYNC-DEFERRED]` structured marker

- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)** righe ~835–855: quando `$MAIN` HEAD ≠ develop, oltre al `fetch` esistente emette ora `[SYNC-DEFERRED] main repo HEAD=<branch>, ...` su stdout. `/mw` standalone preserva l'isolamento terminale (no auto-checkout); orchestratori upstream (Phase 6c di `/new`) parsano il marker e routano la riconciliazione via user gate.

### Why MINOR (not MAJOR)

Tutte le fasi MANDATORY esistevano già — vengono solo propagate (team-mode Step D) e rese non-bypassabili. Le nuove Phase 0/6c sono additive bloccanti, non rimuovono né rinominano capability. La semantica del `completion-report` cambia ma è retro-compatibile: coder pre-3.25 che emettono `partial`/`blocked` triggerano il gate invece di passare silently, e il gate ora esiste in team-mode. Per la decision tree in MAINTAINING.md "Tightened existing MANDATORY gates without removing capabilities → MINOR". L'incident FEAT-0006 era un fallimento di propagazione, non di design — la fix riallinea l'implementazione all'intent originale di v3.24.0.

### Why no schema-change propagation

Modifiche testuali a 4 file framework. Zero chiavi nuove in `baldart.config.yml`, zero modifiche CLI (`src/commands/*`), zero hook nuovi, zero adapter, zero template. Considerato e scartato `policy.strict_phase_gating: true`: sarebbe 5 punti di propagazione (template + configure + update + skill + CHANGELOG) per zero beneficio tunable dall'utente — gli overlay `.baldart/overlays/new.md` restano la via per personalizzazioni progettuali. Verifica: `grep -r "policy.strict_phase_gating\|features.strict_phase_gating" framework/ src/` deve restituire vuoto.

## [3.24.0] - 2026-05-27

Nuovo gate **AC-Closure** (Phase 2.5b) in `/new` + regola **Scope Closure Discipline** in `framework/agents/workflows.md`. Risolve il pattern ricorrente in cui l'orchestratore — sotto pressione di context-window o complessità inattesa — defer-iva unilateralmente acceptance criteria numerate marcandole `deferred` negli `implementation_notes` YAML con razionalizzazione post-hoc ("ApproveSheet già copre il confirm"). Caso reale che ha innescato la fix: sessione `/new` su FEAT-0005 (2026-05-27) dove 5 AC su 22 sono state silenziosamente deferite con razionalizzazione poi rivelata falsa alla verifica funzionale. Le decisioni di deferral sono ora **sempre routate all'utente via `AskUserQuestion`** una-per-AC.

### Added — Phase 2.5b AC-Closure Gate (BLOCKING)

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: nuova Phase 2.5b inserita tra Phase 2.5 (Completeness Check) e Phase 2.55 (Simplify). 6 step:
  1. **AC Closure Ledger** — tabella esplicita una-riga-per-AC con status enum binario `implemented` | `deferred` (nessuna terza categoria — l'AC è nel diff o sta venendo deferita), evidence `file:line`, campo `User-approved?`.
  2. **Rationalization scan** — per ogni AC `deferred`, se la giustificazione matcha pattern `covered by X` / `already handles` / `redundant with Y` / `out of scope`, l'orchestratore DEVE leggere `X/Y` e produrre una scope-mapping line-per-line nel tracker sotto `## Rationalization Verification`. Il pattern senza verifica NON auto-approva il deferral.
  3. **User gate** — `AskUserQuestion` BLOCKING uno-per-AC (mai batched), con testo AC verbatim + 4 opzioni: implementa adesso (spawna fix coder scoped al ownership map) / approva deferral (registra `[USER-APPROVED DEFERRAL] <date>: <reason>` in `implementation_notes`) / sposta su follow-up card (crea `<CARD-ID>-followup-AC<N>.yml` con `status: TODO`) / ferma il batch.
  4. **`implementation_notes` audit** — grep finale per phrase `deferred` / `out of scope` / `covered by` / `skipped` etc.; ogni match deve carry il prefisso `[USER-APPROVED DEFERRAL]` o tornare allo Step 3.
  5. **Context-pressure exception** — directive esplicita: sotto context pressure si ferma e si chiede, NON si compensa skippando AC. Auto Mode's "bias toward proceeding" non override questa regola.
- **Top-level directive** in `/new` SKILL.md (riga 13, subito dopo YOLO MODE banner): riassume la Scope Closure Discipline e linka workflows.md.

### Added — Scope Closure Discipline rule (citable protocol)

- **[framework/agents/workflows.md](framework/agents/workflows.md)**: nuova sezione `## Scope Closure Discipline (MUST)` tra "Status Tracking" e "Deploys and Publications". 3 violazioni esplicite (silent deferral, rationalization-without-verification, scope reduction under context pressure) + protocollo di enforcement via `AskUserQuestion`. Citable da skill/agent come autorità — un finding "silent AC deferral" è BLOCKING in code review.

### Why MINOR (not PATCH)

Aggiunta di capability: nuovo gate obbligatorio in `/new` che cambia comportamento osservabile per i consumer (Phase 2.55 non parte più finché Phase 2.5b non passa). Nessun breaking change — card senza AC-deferral passano il gate transparently in <1s con `ac-closure: implemented=N | deferrals=0 | follow-ups=0`. Per la decision tree in MAINTAINING.md "Added an agent / command / skill / routine / template (or capability)? → MINOR".

### Why no schema-change propagation

Modifiche testuali a due file framework. Zero chiavi nuove in `baldart.config.yml`, zero modifiche CLI (`src/commands/*`), zero hook, zero adapter, zero template. Il gate sfrutta strutture YAML già esistenti (`acceptance_criteria`, `implementation_notes`) e `${paths.backlog_dir}` già nella config schema da v3.0.0.

## [3.23.0] - 2026-05-27

Aggiunto il flag `-full` / `--full` a `/new` per espandere un epic ai suoi figli in un colpo solo — `/new FEAT-005 -full` lancia automaticamente `FEAT-005-1`, `FEAT-005-2`, … senza domande di grouping né di branch. Risolve il pattern ricorrente "ho un epic con 5 child card, le voglio tutte ora": prima richiedeva di listarle a mano (`/new FEAT-005-1 FEAT-005-2 FEAT-005-3 …`) o di usare l'hyphen-range solo se gli ID erano numericamente contigui.

### Added — `/new` epic expansion flag

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)**: nuova regola di parsing `Epic expansion (-full / --full)` nella sezione args. Discovery dei figli via union di due rules: (1) filename in `${paths.backlog_dir}` con prefisso `<PARENT-ID>-` (separator letterale, quindi `FEAT-005-1.yml` matcha ma `FEAT-005.yml` e `FEAT-0050.yml` no); (2) cards con `group.parent: <PARENT-ID>`. Il parent stesso è escluso (override esplicito: `/new FEAT-005 FEAT-005 -full`). Ordinamento per `group.sequence` ASC, fallback lessicografico. Batch-scoped: `/new FEAT-005 FEAT-008 -full` espande entrambi gli epic. Zero figli trovati → HALT esplicito (no silent no-op). Mixed batches supportati: `/new FEAT-005 -full FEAT-007` accoda i figli di 005 + FEAT-007 standalone.
- **[framework/.claude/commands/new.md](framework/.claude/commands/new.md)**: versione condensata della stessa regola, allineata al SKILL canonico.

### Why MINOR (not PATCH)

Capability addition pura: nuovo flag, nessuna modifica al comportamento esistente. Tutti gli invocation pattern pre-3.23 continuano a funzionare identici. Per la decision tree in MAINTAINING.md "Did you add new functionality? → YES → MINOR".

### Why no schema-change propagation

Modifica testuale a due file di skill/command. Zero chiavi nuove in `baldart.config.yml`, zero modifiche CLI (`src/commands/*`), zero hook, zero adapter. La discovery dei figli sfrutta `${paths.backlog_dir}` già presente nella config schema da v3.0.0.

## [3.22.1] - 2026-05-27

Rimosso il gate `AskUserQuestion` di conferma merge introdotto in v3.22.0 Step 7 — era friction inutile contro il "seamless" che è il punto della release. Quando l'utente lancia `/prd` ha già implicitamente accettato la pipeline completa (discovery → design → cards → commit → merge in develop). Chiedere "Procedo con merge?" a fine sessione equivale a chiedere "sei sicuro di voler completare quello che mi hai chiesto di fare?".

### Changed — `/prd` Step 7 merge è ora full-seamless

- **[framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md)**: Step 7 point 9 non è più un `AskUserQuestion` ma una nota che chiarisce il razionale del no-confirmation flow + escape hatch user-initiated ("non mergiare" / "lascia il worktree" honora l'override e ferma dopo il commit).
- **[framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md)** HARD RULE 17 lifecycle: la descrizione di Step 7 cita "Seamless by default, no confirmation gate" + l'escape hatch.
- **[framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md)**: il messaggio di kickoff visibile all'utente promette esplicitamente "commit + merge automatico in develop + cleanup worktree, seamless" — niente più sorprese a fine sessione.

### Why PATCH and not MINOR

Rimozione di un comportamento (gate di conferma) introdotto solo 0 minuti fa nella stessa giornata. Nessuna nuova capability, nessun breaking change per consumer: il gate non era una feature documentata né attesa dagli utenti pre-v3.22.0, era una sovra-cautela introdotta da me contro la richiesta esplicita ("seamless"). La decision tree di MAINTAINING.md classifica "bugfix di behaviour non desiderato senza breaking change" come PATCH.

### Why no schema-change propagation

Modifica puramente comportamentale dentro un singolo step di una singola skill. Zero chiavi nuove, zero modifiche CLI, zero hook.

## [3.22.0] - 2026-05-27

Le sessioni `/prd` ora vivono in un **docs worktree** dedicato, esattamente come fa `/new` per le card di codice. Il problema risolto: con 5-10 sessioni Claude in parallelo, ogni PRD inquinava il main checkout con `docs/prd/<slug>/PRD.md`, lo state file, le card YAML, l'eventuale `design.html`, e gli update a `ssot-registry.md` / `project-status.md`. `git status` diventava un campo minato, le sessioni `/new` rischiavano di stashare/committare file PRD non loro, e la review manuale era pressoché impossibile. Soluzione: estendere `worktree-manager` con una **modalità docs lean** (no `npm install`, no porta, no `.env` copy, no build verify, no lint/tsc gates) e integrarla nel ciclo di vita del PRD — creazione automatica al kickoff, merge automatico (previa conferma utente) al termine, cleanup seamless.

### Added — `worktree-manager` programmatic docs mode

- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)**: nuova sezione "Programmatic API — docs mode" con due funzioni:
  - **`nw-docs`** — input `{ slug, branchPrefix? }`, output `{ path, branch, kind: "docs" }`. Esegue solo `git fetch origin develop` + `git worktree add .worktrees/prd-<slug> -b prd/<slug> origin/develop` + entry nel registry con `kind: "docs"`, `port: null`, `buildVerified: null`. Target di costo: < 5 secondi (vs minuti per `/nw` standard a causa di `npm install`).
  - **`mw-docs`** — input `{ worktreePath, commitMessage? }`, output `{ merged, mergeCommit, prNumber, strategy, worktreeRemoved }`. Esegue safety commit + rebase su `origin/develop` (con il protocollo conflict-resolution doc-friendly esistente in `/mw` step 4b) + merge via `git.merge_strategy` configurato (`pr` o `local-push`) + cleanup worktree/branch. **Salta** i gate inapplicabili a un worktree di docs: `npm run build`, `npx eslint`, `npx tsc`, `npm run test`, post-merge build verify.
- Forme **FORBIDDEN** documentate esplicitamente: niente `npm install`/port scan/`.env*` copy/build in `nw-docs`; niente lint/tsc/build gates in `mw-docs`. Il senso della modalità lean è proprio non eseguirli.
- Registry condiviso con i code worktrees (`.worktrees/registry.json`) — `/lw` e `/cw` mostrano docs worktrees accanto ai code worktrees, distinti dal campo `kind`.

### Changed — `worktree-manager`: rebase conflict protocol hardened

- **[framework/.claude/skills/worktree-manager/SKILL.md](framework/.claude/skills/worktree-manager/SKILL.md)** — la tabella di conflict-resolution dello step 4b di `/mw` (riusata verbatim da `mw-docs`) è stata estesa con due categorie nuove **prima** della catch-all "Docs/config":
  - **Structured registries** (`project-status.md`, `ssot-registry.md`, `field-registry.*`, `traceability-matrix.md`, `REGISTRY.md`) → **STOP**: abort rebase, report all'utente. Senza questa riga, due PRD paralleli che toccano `project-status.md` farebbero auto-strip dei marker di conflitto producendo file con sezioni duplicate e tabelle malformate. Il fix vale anche per `/mw` su code worktrees — qualsiasi card che modifica una di queste registry ora chiede risoluzione manuale invece di corromperle silenziosamente.
  - **Append-only logs** (`**/*.jsonl` sotto `docs/metrics/`) → auto-resolve "keep both sides", che per JSONL è semanticamente corretto (file line-oriented additivo). Questo permette al Metrics Log del PRD (vedi sotto) di sopravvivere a merge concorrenti senza intervento.

### Changed — `/prd` skill: HARD RULE 17 (worktree isolation)

- **[framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md)**: nuova HARD RULE 17 obbligatoria — ogni sessione PRD apre un docs worktree allo Step 1 prima di scrivere qualunque file, e committa/mergia tutto allo Step 7 prima di terminare. La regola include:
  - **Path resolution explicit rule**: tutti i `Write` / `Edit` / `Read` di Claude Code DEVONO prefissare i path con `$WORKTREE_PATH` (Claude Code richiede path assoluti — senza prefisso esplicito i file finirebbero nel main repo, vanificando l'isolamento). Esempio `right vs wrong` nel testo della regola.
  - **Forbidden operations**: nessun `git checkout`/`switch`/`branch` sul main, nessun `git stash` (refs/stash è condiviso fra worktree), nessuna scrittura diretta sul main repo durante la sessione.
  - **Failure handling**: se `nw-docs` fallisce (e.g. `.worktrees/` non in `.gitignore`), la skill DEVE fermarsi e segnalare — niente fallback silenzioso al main repo.
- **HARD RULE 9 aggiornata**: il session-start update di `project-status.md` viene **skippato** quando la sessione gira in un docs worktree (sarebbe invisibile alle altre sessioni). L'update at-commit resta attivo, eseguito nel worktree come ultima scrittura prima del rebase + merge.
- **Step 1 (Kickoff)** in [discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md): aggiunto step 3 che chiama `worktree-manager nw-docs` subito dopo aver derivato lo `<slug>`, prima di creare lo state file. Output kickoff aggiornato per mostrare path/branch del worktree all'utente.
- **Step 0 (Context Recovery)** in [discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md): scan iniziale di `.worktrees/registry.json` per entries con `kind: "docs"` che matchano lo slug — se trovate, `cd` nel worktree e riprendi. Gestito anche il caso "state file legacy nel main da pre-v3.22.0" (chiede all'utente se migrare o continuare in legacy mode).
- **Step 7 (Resolution, Commit & Merge)** in [validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md) — rinominato da "Resolution and Commit" e riscritto end-to-end:
  - Update di `ssot-registry.md` / `project-status.md` avviene nel worktree (skip se file inesistente).
  - Commit con `git -C "$WORKTREE_PATH" commit` (file espliciti, mai `git add .`).
  - **Gate di conferma esplicito** (`AskUserQuestion`) prima del merge: "Procedo con merge automatico in `develop` e cleanup del worktree?". Opzione "lascia il worktree" per review manuale o test.
  - Sulla conferma: invocazione `worktree-manager mw-docs` che rebase + merge + cleanup. Failure handling: il worktree resta su disco, niente retry automatico.
- **[framework/.claude/skills/prd/assets/state-template.md](framework/.claude/skills/prd/assets/state-template.md)**: nuova sezione `## Worktree` (path, branch, kind, created_at, merge_commit).
- **Metrics Log riposizionato dentro il worktree.** La sezione "Metrics Log" di SKILL.md scriveva precedentemente in `docs/metrics/skill-runs.jsonl` con risoluzione del path rispetto al main repo root — incompatibile con HARD RULE 17 e con isolamento parallelo. Ora scrive in `$WORKTREE_PATH/docs/metrics/skill-runs.jsonl` come ultimo append PRIMA del commit di Step 7; il file è incluso nella lista di stage esplicita e il conflict protocol di `mw-docs` lo riconosce come append-only (vedi categoria "Append-only logs" sopra). Risultato: due PRD paralleli che terminano insieme producono un append di entrambe le righe nel file su `develop`, senza intervento manuale.

### Why MINOR and not PATCH

Nuova capability portata da una skill esistente (`worktree-manager` guadagna l'API docs mode) + cambio di workflow per `/prd` (gira sempre in worktree). Nessun breaking change per i consumer: le sessioni PRD esistenti pre-3.22.0 sono detettate in Step 0 e gestite con prompt di migrazione (legacy mode resta supportato). La decision tree di `MAINTAINING.md` classifica "added capability" come MINOR.

### Why no schema-change propagation

Zero nuove chiavi in `baldart.config.yml`. La modalità docs riusa:
- `git.merge_strategy` (già esistente da v3.9.0) per il merge — `pr` o `local-push`, rispetta la scelta del progetto.
- `paths.prd_dir`, `paths.backlog_dir`, `paths.references_dir` (esistenti) — risolti relativamente al worktree (HARD RULE 17).
- `features.has_prd_workflow` (esistente) — gating della skill.

Nota CHANGELOG-only (non normativa): per i PRD molti consumer preferiscono `merge_strategy: local-push` (no PR per documentazione). Chi vuole differenziare la strategia per skill può farlo via overlay `.baldart/overlays/prd.md` — la skill non impone scelte.

### Pre-release code review fixes (within v3.22.0)

Un `/code-review` high-effort eseguito prima del tag ha individuato 10 bug — tutti corretti in questa stessa release prima del rilascio. I più gravi:

- **`prd-card-writer` agent + `backlog-phase.md` invocation**: il subagent scriveva le card a `backlog/` consumer-root-relative senza alcuna `$WORKTREE_PATH` awareness — Step 5 del PRD spruzzava le card nel main repo, vanificando l'isolamento. Fix: nuova sezione "Working Directory (MANDATORY)" nel file dell'agent + il prompt da `backlog-phase.md` ora passa `WORKING_DIRECTORY` esplicito (regola di substitution documentata).
- **`ui-design/SKILL.md` Step G**: copia `design.html` a `${paths.prd_dir}/<slug>/design.html` consumer-root-relative. Fix identico: sezione "Working Directory" + `ui-design-phase.md` ora passa `WORKING_DIRECTORY=$WORKTREE_PATH` come arg.
- **`/prd-add` Step 0**: listava `${paths.prd_dir}/sessions/` sul main repo, zero scan del registry. Fix: Step 0 riscritto — scansiona `.worktrees/registry.json` per `kind: "docs"` prima di guardare il main; gestisce sia worktree-mode che legacy; documenta gli effetti collaterali downstream (delta commit nel worktree, sub-agents ri-invocati con `WORKING_DIRECTORY`, no auto-merge per CR).
- **Regressione `/new` introdotta dal conflict protocol hardening**: la prima versione della tabella step 4b in `worktree-manager` abortiva blanket il rebase su `ssot-registry.md` / `project-status.md`, ma `/new` AGGIORNA `ssot-registry.md` per ogni card (line 1131 di `framework/.claude/skills/new/SKILL.md`, hook pre-commit che blocca commit a `backlog/` senza ssot). Risultato: ogni batch multi-card di `/new` sarebbe fallito al rebase. Fix: sostituito il blanket abort con **structural validation** (`validate_structured_md` helper) — strip markers + check struttura (no marker residui, no header table duplicati, no frontmatter doppio). Se valido → accept (caso routine: append additivo); se invalido → abort (caso raro: due card hanno modificato la stessa riga). Restituisce a `/new` il comportamento additivo originale, mantiene la protezione contro corruzione reale.
- **`git stash` in `mw-docs`**: la prima versione riusava verbatim `/mw` step 4b che fa `git stash push --include-untracked` prima del rebase. Violazione diretta della Safety Rule "NEVER use git stash in worktrees" (refs/stash condiviso fra worktree — incidente FEAT-0522). Fix: la sezione `mw-docs` ora esplicita "NO stash" — il safety commit dello step 1 garantisce working tree clean, qualunque dirty post-safety è un errore che richiede STOP, mai uno stash di recovery.
- **`nw-docs` mancante `.gitignore` check**: la Safety Rule generale dice "Always verify `.worktrees/` is in `.gitignore` before creating", ma `nw-docs` la omette. Fix: step 0 di `nw-docs` è ora un pre-flight `.gitignore` check che fallisce con errore esplicito se mancante.
- **`nw-docs` pseudo-bash relativo invece che assoluto**: il path ritornato come `path` doveva essere assoluto (per HARD RULE 17), ma lo pseudo-code lo settava relativo. Fix: ancorato a `$MAIN_ROOT` esplicito via `git rev-parse --show-toplevel` + `WORKTREE_PATH="$MAIN_ROOT/$WORKTREE_REL"`.
- **Esempio HARD RULE 17 fuorviante**: l'esempio "RIGHT" mostrava `Write file_path = "$WORKTREE_PATH/..."` come stringa literal. Claude Code NON espande shell variables → o errore "path must be absolute" o creazione di una dir letterale `$WORKTREE_PATH/`. Fix: esempio riscritto con 3 forme "wrong" (main repo, literal var, relative) + 1 "right" (substitution textuale dell'absolute path).
- **Off-by-one cross-ref Metrics Log**: "Add the file to Step 7 point 5" ma il staging è point 6 (point 5 è il Metrics Log step stesso, self-referential). Fix: cross-ref corretto a point 6 con nota di disambiguazione.
- **Case pattern `*.jsonl` non raggiungibile per il secondo arm + auto-strip non desiderato per fixture**: `*.jsonl|docs/metrics/*.jsonl` matchava ogni `.jsonl` ovunque (anche fixture/seed). Fix: tabella sostituita con pattern espliciti — `docs/metrics/*.jsonl` auto-resolve, altri `*.jsonl` abort + manuale.
- **Numerazione stale "Step 1, point 7"** in `discovery-phase.md` § 1.6: dopo la renumerazione, point 7 è "Context Loading", non l'output kickoff. Fix: aggiornato a "point 9 (Business Rationale Extraction is the final point)".

### Verification

Modifiche applicate ma non testabili automaticamente (nessuna test suite in questo repo, come da `CLAUDE.md`). Validazione richiesta in un consumer reale:
1. Lanciare `/prd` su una feature qualunque → verificare la creazione `.worktrees/prd-<slug>/` + entry `kind: "docs"` nel registry.
2. Aprire una seconda sessione Claude in parallelo e fare `git status` sul main → deve essere pulito.
3. Procedere fino allo Step 7 → verificare il prompt di conferma merge, completare → worktree rimosso, branch eliminato, PRD su `develop`.
4. Resume test: chiudere/riaprire la sessione a metà PRD → Step 0 deve trovare il worktree dal registry e `cd` lì.
5. `/prd-add` su PRD attivo → deve riusare il worktree esistente, non crearne uno nuovo.

## [3.21.2] - 2026-05-26

Aggiunto un livello di automazione che intercetta il drift fra la skill `/baldart-update` (v3.21.0+) e le sorgenti CLI che essa replica in chat. La skill replica **5 decision point fissi** del comando `npx baldart update`; quando una release tocca `src/commands/update.js` (nuovi prompt), `src/utils/hooks.js` (nuove entry in `HOOK_REGISTRY`), o `framework/templates/baldart.config.template.yml` (nuove top-level key), la skill può silenziosamente disallinearsi. Il bug v3.21.0 → v3.21.1 (`git -C .framework fetch` errato) è già stato un esempio del problema: la skill prometteva un pattern che il test live ha smentito. Senza un meccanismo di alert, futuri drift sarebbero individuabili solo per accidente.

### Added — Drift detection script

- **[scripts/check-update-skill-drift.js](scripts/check-update-skill-drift.js)**: Node script standalone (no dipendenze) che misura:
  - Numero di prompt interattivi in `src/commands/update.js` (regex `await confirm(` + `inquirer.prompt(`).
  - Numero di entry attive in `HOOK_REGISTRY` di `src/utils/hooks.js` (linee `id: 'baldart-...'` non commentate).
  - Elenco delle top-level key in `framework/templates/baldart.config.template.yml`.
  Confronta con i valori attesi salvati nel marker `<!-- DRIFT-CHECK ... -->` dentro lo SKILL.md della skill, e stampa un report human-readable. Exit code sempre 0 — drift è warning, non errore: la decisione finale resta umana.

### Added — Marker DRIFT-CHECK nello SKILL.md

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: HTML comment in cima (subito dopo il frontmatter) con i conteggi/key attesi e il campo `last_verified` (versione in cui l'allineamento è stato confermato). Il marker è invisibile al consumer (HTML comment) ma viene letto dallo script.

### Added — GitHub Actions workflow

- **[.github/workflows/check-update-skill-drift.yml](.github/workflows/check-update-skill-drift.yml)**: workflow CI triggered su tre eventi:
  - `push tags: v*.*.*` (ogni release).
  - `pull_request` su `main` con paths filter sui file rilevanti (update.js, hooks.js, config template, SKILL.md, drift script).
  - `workflow_dispatch` (lancio manuale).
  Esegue lo script e annota eventuali drift come GitHub warnings (`::warning::`). Job non blocca mai il merge / la release.

### Schema change propagation

Nessuna nuova chiave in `baldart.config.yml`, nessuna modifica al CLI. La modifica è puramente CI/dev tooling.

### Versioning rationale

PATCH (3.21.1 → 3.21.2) per la decision tree di [MAINTAINING.md](MAINTAINING.md): tooling interno, nessun cambio di behaviour CLI, nessun consumer-facing impact. Il marker HTML nello SKILL.md è invisibile al consumer.

### Verification

Eseguito localmente: `node scripts/check-update-skill-drift.js` riporta baseline attuale `update_js_prompts=6, hook_registry_entries=3, config_template_top_level_keys=[features, git, identity, lsp, paths, stack, tools, version]` e "No drift detected" — coerente col marker `last_verified: v3.21.2`.

## [3.21.1] - 2026-05-26

Fix del Step 0 e Step 1 della skill `/baldart-update` introdotta in v3.21.0. La prima release usava `git -C .framework fetch origin main` per il pre-flight check del numero di commit ahead — pattern **sbagliato** perché `.framework/` è un git subtree all'interno del repo consumer, non un git repo separato: i comandi `git` eseguiti con `-C .framework` ricadono sul `.git` del consumer e fetcherebbero il remote del CONSUMER (es. `github.com/antbald/mayo`), non quello di BALDART. Bug individuato durante test live su `mayo` (l'output mostrava `Da https://github.com/antbald/mayo  * branch  main -> FETCH_HEAD` invece di BALDART).

### Fixed — Pattern git corretto nella skill `/baldart-update`

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: Step 0 ora delega completamente a `npx baldart version` per ottenere installed version, commit ahead count, e BALDART repo URL — nessuna git operation a mano. Step 1 usa il pattern corretto del CLI (`src/utils/git.js`): `git fetch <BALDART-repo-url> main && git log HEAD..FETCH_HEAD --oneline -- .framework/`, dove l'URL viene estratto dal campo `Repository:` di Step 0. Aggiunta callout box dopo la tabella dei 5 decision point con la regola esplicita: *"do NOT use `git -C .framework fetch`"*, motivazione (`.framework/` è subtree, non repo separato), e pattern corretto.

### Schema change propagation

Nessuna nuova chiave in `baldart.config.yml`. Nessun cambiamento al CLI.

### Versioning rationale

PATCH (3.21.0 → 3.21.1) per la decision tree di [MAINTAINING.md](MAINTAINING.md): bug fix interno alla skill, nessun cambio di behaviour del CLI, nessuna nuova capability, nessun cambio di directory layout o install command.

### Verification (manuale)

Eseguita in `~/mayo`:

1. Test del comando originale ERRATO: `git -C .framework fetch origin main` → l'output mostrava `Da https://github.com/antbald/mayo` (remote di mayo, NOT BALDART) — bug confermato.
2. Test del pattern CORRETTO: `npx baldart version` espone `Repository: https://github.com/antbald/BALDART`. `git fetch https://github.com/antbald/BALDART.git main && git log HEAD..FETCH_HEAD --oneline -- .framework/` rileva correttamente i commit BALDART ahead del subtree merge nel consumer.

## [3.21.0] - 2026-05-26

Fino a v3.20.x l'aggiornamento di un consumer all'ultima versione del framework passava da due strade: **CLI diretta** (`npx baldart update`, interattivo, con 5 prompt nativi gestiti dall'utente in TTY) oppure **chat ad hoc** (l'utente chiede a Claude di aggiornare, e Claude — senza protocollo standardizzato — improvvisa ogni volta, gestendo gli edge case in modo non riproducibile). La prima opzione è solida ma richiede di lasciare la chat e tornare al terminale; la seconda è comoda ma fragile: silenzia decisioni che il CLI farebbe esplicitamente.

Questa release introduce la skill `/baldart-update` come **wrapper conversazionale agent-driven** di `npx baldart update`, complementare a `/baldart-push`. La skill **replica in chat ognuno dei 5 decision point nativi del CLI** prima di lanciare il comando con `--yes` — non per silenziarli, ma perché ogni prompt è già stato risposto in chat. Il CLI resta source of truth: la skill non re-implementa la logica di `update.js`, la guida e la riporta.

### Added — Skill `/baldart-update`

- **[framework/.claude/skills/baldart-update/SKILL.md](framework/.claude/skills/baldart-update/SKILL.md)**: nuova skill (~220 righe) modellata strutturalmente su [/baldart-push](framework/.claude/skills/baldart-push/SKILL.md) ma con divergenza intenzionale sul modello esecutivo: `/baldart-push` è pure-guidance (l'utente lancia il CLI in TTY), `/baldart-update` è agent-driven con conferme replicate in chat. La skill:
  - **Pre-flight read-only** (Step 0): legge `.framework/VERSION` e conta i commit ahead con `git -C .framework rev-list HEAD..origin/main --count`. Se 0 → "Already up-to-date", stop senza lanciare il CLI.
  - **Replica 5 decision point** (Step 1–4) — surface in chat di: preview diff (delta versione + commit message), working tree dirty + stash automatico in ref `baldart-pre-update-<timestamp>`, hooks drift confrontato con `HOOK_REGISTRY` di `.framework/src/utils/hooks.js`, schema config drift confrontato con `.framework/framework/templates/baldart.config.template.yml`, auto-commit classification (BALDART-managed vs user-owned per path-prefix matching su `.framework/`, `.claude/agents/`, `.claude/skills/`, `.claude/commands/`, `AGENTS.md`, `agents/`).
  - **Run CLI** (Step 5): `Bash` con `timeout: 600000` (10 min — il default 120s è insufficiente per subtree pull + symlink reconcile + post-update wizard + hooks register + auto-commit su repo reali), comando `npx baldart update --yes`. NON passa `--auto-stash` (ridondante con `--yes`, segnalerebbe imprecisione).
  - **Report** (Step 6): estrae backup tag con `git tag -l 'baldart-pre-update-*' --sort=-creatordate | head -1` (più affidabile del parsing stdout), nuova versione da `cat .framework/VERSION`, e detection stash-pop conflict via grep su stdout/stderr per pattern `CONFLICT` / `stash pop failed` / `unmerged paths`. Fornisce sempre il comando di rollback `git reset --hard <backup-tag>`.

### Bootstrap paradox — esplicitato in cima al SKILL.md

I consumer pre-v3.21.0 incontreranno il paradosso al primo aggiornamento dopo questa release: la skill non è ancora installata, quindi `/baldart-update` non esiste. La skill mette in **box CAPS in prima posizione** l'istruzione di lanciare `npx baldart update` direttamente nel terminale una volta — dopo, `/baldart-update` diventa disponibile per ogni aggiornamento successivo. Non è un edge case nascosto: è il caso primario del rollout, e va trattato come tale.

### Hard rules nella skill

Sette hard rule chiudono i vettori di errore:

1. NEVER re-implementare la logica di `update.js`.
2. MUST replicare in chat tutti e 5 i decision point prima di passare `--yes`.
3. MUST passare `timeout: 600000` al Bash.
4. NEVER passare `--auto-stash`.
5. MUST riportare il backup tag.
6. NEVER lanciare update se `.framework/` manca → defer a `npx baldart add`.
7. NEVER assumere default quando un pre-check fallisce — chiedi all'utente o aborta.

### Schema change propagation

**Nessuna nuova chiave in `baldart.config.yml`** — la skill legge solo `.framework/VERSION`, `.baldart/state.json`, `.claude/settings.json` (file BALDART-owned, non gated da config). La regola di propagazione schema (memoria utente + CLAUDE.md) non si applica a questa release.

### Versioning rationale

MINOR (3.20.0 → 3.21.0) secondo la decision tree di [MAINTAINING.md](MAINTAINING.md):
- Nuova skill (capability additiva) → MINOR.
- Nessuna rimozione / cambio directory layout / cambio install command → non MAJOR.
- Non è un fix doc-only / perf / security senza behaviour change → non PATCH.

### Verification (manuale, no test suite)

Da verificare in un consumer downstream (dopo merge + push + `npx baldart update`):

1. **Discoverability**: dopo l'update, `.claude/skills/baldart-update` esiste come symlink a `.framework/framework/.claude/skills/baldart-update/SKILL.md`. La skill appare nella lista delle skill caricate alla prossima sessione Claude Code.
2. **Collision test (bootstrap edge case)**: creando `.claude/skills/baldart-update/` user-owned PRIMA dell'update, `SymlinkUtils.mergeSkills` logga in `.baldart/skill-conflicts.json` e NON sovrascrive. La skill ufficiale resta non disponibile — comportamento corretto e documentato.
3. **Happy path**: in repo allineata, `/baldart-update` risponde "Already up-to-date" via Step 0 (count = 0), senza lanciare il CLI.
4. **Behind path**: forzato `cd .framework && git reset --hard HEAD~1`, `/baldart-update` rileva 1 commit ahead, replica Step 1–4 in chat, lancia CLI con timeout 600000, riporta backup tag + nuova versione.
5. **Dirty path**: working tree modificato → Step 2 lista i file e chiede conferma esplicita per lo stash. Se rifiuto → STOP pulito.
6. **Conflict path**: file modificato sia upstream che locale → CLI esce non-zero, skill riporta verbatim + comando rollback con backup tag.

## [3.20.0] - 2026-05-26

Su `mayo` è stato osservato il seguente comportamento patologico: il modello ha chiamato `Agent({ subagent_type: "code-reviewer" })` e `Agent({ subagent_type: "doc-reviewer" })` (entrambi agent BALDART legittimi e fisicamente presenti su disco), il tool è ritornato `0 tool uses · Done` senza errore visibile, e il modello ha **silenziosamente fatto fallback** su `feature-dev:code-reviewer` + `plan-auditor` — proseguendo come se nulla fosse. Né l'utente né l'agent orchestratore hanno avuto modo di accorgersi che la review codice + doc review obbligatorie non erano state eseguite dagli agent attesi.

Ricerca su issue tracker Anthropic ha confermato due bug Claude Code che si combinano in questo scenario: [#56869](https://github.com/anthropics/claude-code/issues/56869) (sub-agent error non propagato al parent — `0 tool uses · Done` è il sintomo) e [#20931](https://github.com/anthropics/claude-code/issues/20931) (file-based discovery di `.claude/agents/*.md` documentato come BROKEN in certi scenari, anche con file validi). **Nessuno dei due è patchabile lato BALDART** — il framework può solo mitigare con difese stratificate.

Questa release introduce **quattro layer di difesa stratificati** contro la silent-substitution di agent BALDART, senza nuove chiavi in `baldart.config.yml`:

### Added — Hook `agent-discovery-gate` (PreToolUse, bloccante)

- **[framework/.claude/hooks/agent-discovery-gate.js](framework/.claude/hooks/agent-discovery-gate.js)**: nuovo PreToolUse hook con `matcher: 'Agent'` che intercetta ogni invocazione del tool `Agent` (alias `Task`). Estrae `subagent_type` (o `agent_type`) dall'input, lo confronta con la lista degli agent BALDART derivata via `fs.readdirSync` su `.framework/framework/.claude/agents/*.md` (esclusa `REGISTRY.md`). Se il nome è BALDART e il symlink `.claude/agents/<name>.md` del consumer è mancante o rotto (`fs.statSync` segue il link → ENOENT/ELOOP), emette `permissionDecision: "deny"` con reason esplicito che istruisce: *"riavvia la sessione, esegui `npx baldart doctor`, NON sostituire con un altro agent"*. Nomi non-BALDART (built-in `Explore`/`Plan`/`general-purpose`, marketplace `feature-dev:*`/`vercel:*`/`sentry:*`, plugin) passano sempre. Fail-safe contract identico a `framework-edit-gate.js`: ogni try-catch interno esce 0 (allow), mai bloccare per errore proprio.

### Added — Hook `agent-discovery-info` (SessionStart, informativo)

- **[framework/.claude/hooks/agent-discovery-info.sh](framework/.claude/hooks/agent-discovery-info.sh)**: nuovo SessionStart hook bash che enumera gli agent BALDART attesi (stesso `find` su `.framework/framework/.claude/agents/`) e li confronta con `.claude/agents/<name>.md` nel consumer. Quando uno o più sono mancanti, inietta `additionalContext` come **fatto dichiarativo** (best practice anti-prompt-injection da [Claude Code hooks docs](https://code.claude.com/docs/en/hooks)) elencando i nomi non disponibili. Silente quando tutto è OK. **Limite onesto**: il bug Claude Code [#10373](https://github.com/anthropics/claude-code/issues/10373) può saltare l'iniezione su new conversations — in quel caso la difesa scende sul gate bloccante + regola AGENTS.md.

### Added — Regola MUST `framework/AGENTS.md` § 4.a

- **[framework/AGENTS.md](framework/AGENTS.md)**: nuova regola in sezione *Universal MUST*, inserita immediatamente dopo la MUST su `codebase-architect`. Istruisce ogni agent che lavora in un consumer BALDART a **NON fare fallback silenzioso** quando una chiamata `Agent` ritorna `0 tool uses · Done`, va in timeout, o riceve `permissionDecision: deny`. Comportamento richiesto: STOP, replicare verbatim al user il messaggio *"Agent `<name>` non discoverato o fallito in questa sessione — riavvia Claude Code e verifica con `npx baldart doctor`"*, attendere direzione. Nessun auto-routing a `feature-dev:*`/`general-purpose`/altri marketplace. La regola chiude esplicitamente: *"This rule is the ONLY defense against bug #56869, which BALDART cannot patch from outside Claude Code; deviating from it reintroduces the exact failure mode the rule exists to prevent."* È inseparabile dall'hook bloccante — senza la MUST, il `deny` del gate produrrebbe esattamente la silent-substitution che si vuole prevenire.

### Added — `baldart doctor` check integrità symlink agent

- **[src/commands/doctor.js](src/commands/doctor.js)**: aggiunta detection `state.agentSymlinksBroken` nel blocco `detectState` (stesso pattern del check `hooksStatus`, non-blocking probe in `try`/`catch` silenzioso). Enumera `.framework/framework/.claude/agents/*.md`, per ogni nome verifica `fs.statSync('.claude/agents/<name>.md')`, raccoglie i broken. Nuova action `repair-agent-symlinks` con `autoOk: true` che chiama direttamente `new SymlinkUtils().mergeAgents({ tools })` — NON dispatcha `require('./update')` (l'update completo avrebbe side-effect indesiderati: pull, overlay merge, hook register, config drift verify). Tools letti via `loadConfig(cwd).tools.enabled` con fallback `['claude']`.

### Added — Registry idempotente per i due nuovi hook

- **[src/utils/hooks.js](src/utils/hooks.js)**: appese due entry a `HOOK_REGISTRY`: `baldart-agent-discovery-gate` (PreToolUse, matcher `Agent`, order 200 — corre dopo `framework-edit-gate`) e `baldart-agent-discovery-info` (SessionStart, matcher `*`, order 100). Nessuna modifica strutturale al modulo — l'infrastruttura multi-hook esistente (`registerAll`/`getStatus`/`createDriftPrompt`, introdotta in v3.18.0) gestisce automaticamente i due nuovi hook: `baldart add`/`update`/`doctor` li registrano idempotentemente in `.claude/settings.json` del consumer, e `unregisterAll` li rimuove pulitamente. Backfill garantito su consumer pre-3.20.0 via `baldart update`.

### Changed — Nota machine-readability in REGISTRY.md

- **[framework/.claude/agents/REGISTRY.md](framework/.claude/agents/REGISTRY.md)**: nuova nota in cima che dichiara la directory filesystem `framework/.claude/agents/` come fonte di verità machine-readable per i tre nuovi consumatori (gate hook, info hook, doctor). Aggiungere un agent richiede solo `<name>.md` nella directory — nessun manifest JSON da mantenere, nessuna drift da sincronizzare. La tabella Agent Map resta human-readable.

### Limiti dichiarati

Onestà tecnica > marketing: questa release **non risolve** i bug Claude Code che hanno causato il problema su `mayo`, li mitiga.

- **#56869 (silent sub-agent fallback)** non è patchabile dall'esterno di Claude Code. L'unica difesa è la regola MUST in `AGENTS.md` — soft, non meccanicamente bloccante. Il successo dipende dall'aderenza del modello alla regola.
- **#20931 (file-based discovery broken con file validi)** resta **scoperto dal gate** quando il file esiste e Claude Code lo ignora comunque: il gate vede il file → ritorna allow → la call torna `0 tool uses` (bug originale). In quello scenario worst-case, l'unica difesa attiva è la regola MUST.
- **#10373 (SessionStart non iniettato)** può saltare il warning informativo su new conversations — accettato come best-effort.

In pratica, il caso comune (symlink rotto / consumer mid-migration) è coperto al 100% dal gate. Il caso patologico (Claude Code ignora silenziosamente un file `.md` valido) è coperto solo dalla regola comportamentale.

### Schema change propagation

- **Nessuna nuova chiave in `baldart.config.yml`** — la lista degli agent vive nel filesystem del payload, gli hook sono auto-registrati via `HOOK_REGISTRY`, il check doctor è always-on quando `state.frameworkPresent`. Schema-change propagation rule non si applica a questa release.

### Versioning rationale

MINOR (3.19.0 → 3.20.0) secondo la decision tree di [MAINTAINING.md](MAINTAINING.md):
- Due nuovi hook (capability) → MINOR.
- Nuovo doctor check + nuova action (capability) → MINOR.
- Nuova MUST in AGENTS.md (regola comportamentale aggiuntiva ma additiva, non rompe install esistenti) → MINOR.
- Nessuna rimozione / cambio directory layout / cambio install command → non MAJOR.

### Verification (manuale, no test suite)

Eseguita in-place su questa working copy via smoke test del hook standalone (`node framework/.claude/hooks/agent-discovery-gate.js < input`):

1. **Happy path** — input `{ tool_name: 'Agent', tool_input: { subagent_type: 'definitely-not-an-agent' } }`: exit 0, output vuoto. ✓ (nome non in `baldartAgents` → allow).
2. **Deny path** — input `{ tool_name: 'Agent', tool_input: { subagent_type: 'code-reviewer' }, cwd: '/tmp' }`: exit 0 con stdout JSON `permissionDecision: "deny"` e reason verbose istruttivo. ✓.
3. **Allow path** — tmpdir con `.claude/agents/code-reviewer.md` simulato + invocation: exit 0, output vuoto. ✓ (symlink healthy → allow).
4. **Module load** — `node -e "require('./src/utils/hooks.js')"` riporta 3 entries in HOOK_REGISTRY. ✓.
5. **Doctor module load** — `node -e "require('./src/commands/doctor.js')"` carica senza syntax error. ✓.

Da verificare manualmente in un consumer downstream (dopo merge):
- `baldart doctor` con un symlink agent volontariamente rotto → propone action `repair-agent-symlinks` con `autoOk: true`, l'esecuzione ripristina il symlink.
- `baldart update` da consumer pre-3.20.0 → backfilla i due nuovi hook in `.claude/settings.json` senza toccare le entry esistenti.
- Sessione Claude Code in consumer con symlink rotto → `agent-discovery-info.sh` inietta `additionalContext` al boot; chiamata `Agent` verso quel nome → gate emette deny + il modello (se aderisce alla MUST) replica il messaggio all'utente invece di sostituire.

## [3.19.0] - 2026-05-26

`.baldart/overlays/` smette di essere un percorso "andare al buio". Fino a v3.18.x, un consumer che voleva personalizzare uno skill / agent / command doveva ricordare a memoria tre canali distinti, due modelli di merging (concat per le skill, marker-based per agent/command), il path overlay corretto per ogni canale, lo schema frontmatter (`base_<kind>`, `base_<kind>_version`, `mode`) e la versione framework da targetare — il tutto senza nessun feedback su drift quando il base file cambiava upstream. L'unico touchpoint guida era il messaggio del hook `framework-edit-gate` quando bloccava un edit dentro `.framework/`, che diceva "metti la roba in `.baldart/overlays/<skill>.md`" senza spiegare *quale path esatto*, *quale schema*, *quale versione*. Risultato: overlay scritti raramente o male.

Questa release introduce uno skill `/overlay` + tre nuovi sotto-comandi CLI (`baldart overlay scaffold|validate|drift`) che colmano il gap end-to-end, seguendo il pattern stabilito da `/baldart-push` (skill = guidance, CLI = source of truth). Nessuna nuova chiave in `baldart.config.yml`: gli overlay vivono su path universali che non dipendono dalla configurazione del consumer.

### Added — Skill `/overlay`

- **Nuova skill [framework/.claude/skills/overlay/SKILL.md](framework/.claude/skills/overlay/SKILL.md)**. Thin conversational wrapper attorno al nuovo sotto-comando CLI. Sette fasi: pre-check (`.framework/` esiste?), discovery (skill/agent/command + nome), inspect base + overlay esistente, scaffold via CLI, author body con Edit, validate via CLI, drift maintenance. La skill non re-implementa mai la logica del merger — invoca sempre `npx baldart overlay <sub>`. Quattro hard rules: mai editare `.framework/` direttamente, mai editare file con marker `<!-- baldart-generated: -->`, mai inventare `base_file_sha` o `base_<kind>_version` (computati dalla CLI), mai re-implementare i marker sezionali (vivono in `overlay-merger.js`). Frontmatter porta `contamination_scan: skip` perché contiene esempi illustrativi con token di overlay reali.

### Added — Sotto-comando CLI `baldart overlay {scaffold,validate,drift}`

- **Nuovo file [src/commands/overlay.js](src/commands/overlay.js)** + dispatcher aggiunto in [bin/baldart.js](bin/baldart.js) come gruppo `overlay` (stesso pattern di `routines`). Tre sotto-comandi:
  - `baldart overlay scaffold <type>/<name> [--mode extend|override]` — idempotente. Valida che il base esista in `.framework/`, computa `base_file_sha` (sha256 12-char) dal contenuto base, legge framework `VERSION`, scrive `.baldart/overlays/<path>` con frontmatter completo e body placeholder appropriato al kind (concat per skill, marker-based per agent/command). Refuse pulito se l'overlay esiste già — punta l'utente a `drift` o all'Edit diretto.
  - `baldart overlay validate <path>` — dry-run del merger. Per agent/command invoca `mergeOverlay()` e riporta numero linee generate + warning (es. heading non matchato nel base, che diventa nuova sezione in coda). Per skill valida frontmatter YAML + esistenza base. Confronta `base_file_sha` overlay vs SHA attuale e segnala se il base è cambiato.
  - `baldart overlay drift [<path>] [--all]` — drift report cross-overlay. SHA-based quando `base_file_sha` è presente nel frontmatter, version-based fallback altrimenti. Stampa summary `{clean, drifted, unknown}` su tutti gli overlay sotto `.baldart/overlays/` (root + `agents/` + `commands/`). Skip `*.example.md` e `README.md`.
- **Pre-check fail-safe**: ogni sotto-comando rifiuta con messaggio chiaro quando `.framework/` non esiste (consumer non installato). Non assume mai stato.

### Changed — `overlay-merger.js` arricchito (additivo, retro-compat)

- **[src/utils/overlay-merger.js](src/utils/overlay-merger.js)**: nuova funzione esportata `computeBaseFileSha(content): string` (12-char sha256 prefix) usata dal CLI scaffolder per scrivere `base_file_sha` nel frontmatter overlay. Il marker `<!-- baldart-generated: -->` ora include opzionalmente `base_sha=<sha>` come campo aggiuntivo; `readMarker()` lo legge se presente, altrimenti ritorna `null` per quel campo (retro-compatibile con file generati pre-v3.19.0). Il merger continua a produrre output bit-identico nei casi senza `base_file_sha` nel overlay frontmatter — l'arricchimento è opt-in via lo scaffolder.
- **YAML SHA quote**: il scaffolder quota il sha (`base_file_sha: "fd940fcc1786"`) perché un 12-char hex può essere tutto-digit, che YAML interpreterebbe come numero (`0` per `000000000000`). `parseFrontmatter()` lato CLI coerce sempre a string per difensività.

### Changed — Hook `framework-edit-gate` con handoff `/overlay`

- **[framework/.claude/hooks/framework-edit-gate.js](framework/.claude/hooks/framework-edit-gate.js)**: due piccoli edit testuali (logica del hook intatta). Nel messaggio "(B) Project-specific opinion" e nel messaggio "Cannot edit BALDART-generated file" la guida ora suggerisce di invocare `/overlay` per scaffolding guidato + drift check. La consistenza è garantita perché skill, CLI e hook vengono distribuiti insieme dal `baldart update`.

### Schema change propagation

- **Nessuna nuova chiave in `baldart.config.yml`** — gli overlay vivono su path universali (`.framework/`, `.baldart/overlays/`) non parametrizzati. La regola di propagazione schema (template + configure prompt + update detector + doctor) NON si applica a questa release.
- **Schema change additivo nel frontmatter overlay**: nuovo campo opzionale `base_file_sha`. Retro-compatibile — overlay scritti pre-v3.19.0 senza questo campo restano validi; `drift` li classifica come "unknown" (drift detection by version only) invece di errori. Documentato in `framework/templates/overlays/README.md` come campo opzionale via il body placeholder dello scaffolder.

### Versioning rationale

MINOR (3.18.1 → 3.19.0): added capability (nuova skill + nuovi sotto-comandi CLI). No breaking change: nessuna API esistente modificata semanticamente; `readMarker()` resta backward-compatible; il merger produce lo stesso output per ogni input pre-esistente. Decision tree CLAUDE.md: "Added an agent / command / skill / routine / template → MINOR" qualifica.

### Verification (manuale, no test suite)

Tested in sandbox `/tmp/overlay-test` con symlink `.framework/ → /Users/antoniobaldassarre/BALDART` per simulare un consumer:

1. `baldart overlay scaffold agents/coder` → overlay creato a `.baldart/overlays/agents/coder.md` con frontmatter completo (`base_agent: coder`, `base_agent_version: 3.19.0`, `base_file_sha: "<sha>"`, `mode: extend`).
2. `baldart overlay validate .baldart/overlays/agents/coder.md` → dry-run del merger, "Overlay valid (agent, marker model), 350 lines generated".
3. `baldart overlay drift --all` → cross-overlay drift report, summary corretto.
4. Drift positivo: alterazione manuale di `base_file_sha` nel frontmatter → `drift` segnala "drifted (sha mismatch)".
5. Scaffold idempotente: re-run con overlay esistente → "Overlay already exists, run drift to check it".
6. Skill scaffolding: `baldart overlay scaffold bug` → overlay skill creato a `.baldart/overlays/bug.md` con `base_skill: bug`, body concat-friendly.
7. Pre-check: `baldart overlay scaffold ...` in dir senza `.framework/` → exit 1 con messaggio chiaro.

## [3.18.1] - 2026-05-25

Chiusura del debito tecnico introdotto da v3.18.0. Il release v3.18.0 ha annunciato che `/design-review` supporta un programmatic JSON output mode invocabile da `/e2e-review` per gating uniforme — ma l'implementazione era una sola nota nel frontmatter ("se invocato con `mode: programmatic` ritorna JSON") senza body command che la concretizzasse, senza schema vincolato, senza mode-detection deterministica, senza taxonomy mapping tra le severity Markdown (Blockers/High/Medium/Nitpicks) e le severity programmatic (critical/major/minor) del `visual-fidelity-verifier`. Il risultato: una feature half-implemented, parsable only in teoria. Questo PATCH la rende reale.

L'adversarial review del piano v3.19.0 (multi-viewport + interactive states + Phase 4b design judgment pass) ha esplicitamente identificato il problema: "`/design-review` programmatic JSON è vaporware in v3.18.0 — va prima implementato bene come release a sé". Questa release fa esattamente quello, separando la chiusura del debito dall'aggiunta di nuove feature, in modo che v3.19.0 (quando arriverà, post-dogfooding) possa appoggiarsi a un contratto programmatic stabile invece di costruirsi sopra una stub.

### Fixed — `/design-review` dual-mode reale (Markdown ↔ JSON)

- **[framework/.claude/commands/design-review.md](framework/.claude/commands/design-review.md) riscritto** (frontmatter invariato, body completamente sostituito). Quattro sezioni nuove + workflow rivisto:
  - **§ Invocation Modes**: documenta esplicitamente Mode A (interactive, Markdown — il comportamento storico) e Mode B (programmatic, JSON-only — il nuovo contratto). L'envelope JSON di input per Mode B è specificato verbatim: `{mode, card_id, route, viewport, dev_server_port, registry_paths: {index, tokens, components_dir}, ui_guidelines_path, features: {has_design_system, multi_tenant_theming}, tolerance}`. Il comando lo riceve passato dall'orchestratore `/e2e-review` quando Phase 4b sarà introdotta.
  - **§ Programmatic output schema**: schema JSON di output vincolato, allineato verbatim a `framework/.claude/agents/visual-fidelity-verifier.md` § "Output Schema" per uniformità di parsing nell'aggregatore orchestratore. Campo `source: "design-review"` distingue le finding di questo command da quelle del verifier in modalità debug; i campi `findings[].severity`/`category`/`description`/`expected`/`actual`/`fix_hint`/`evidence`/`ds_drift_code`/`confidence` sono shape-identici. La deduplicazione cross-source avviene per `(route, category, ds_drift_code)` con precedenza al verifier quando duplicate (più strutturato).
  - **§ Severity normalization (Markdown ↔ programmatic)**: tabella di mapping deterministica Blocker → `critical`, High → `major`, Medium → `minor`, Nitpick → DROP. Il drop dei Nitpick in modalità programmatic è deliberato (troppo noise per gating; restano nel report Markdown per uso interattivo).
  - **§ Programmatic-mode category taxonomy**: enumera le 10 categorie ammesse in Mode B (`primitive-reinvented`, `primitive-missing-spec`, `token-bypass`, `component-stale`, `index-drift`, `authority-violation`, `brand-voice-drift`, `a11y-contrast`, `a11y-focus-visible`, `interactive-state-missing`) — sottoinsieme della tassonomia canonica del verifier adattato allo scope di design judgment di questo command. Categorie non in lista vanno emesse come stringhe libere in `gaps[]`, MAI come nuove categorie inventate (l'orchestratore non sa pesarle).
  - **§ Workflow (both modes)**: passi 0–5 unificati per Mode A e Mode B, con la divergenza solo nel formato di output finale (passi 4 e 5). Mode B usa `dev_server_port` dall'envelope per costruire URL, salva screenshot in `/tmp/design-review/<route-slug>.png` per cross-reference dall'orchestratore.
  - **§ Mode-detection contract**: regola deterministica anti-ambiguità — Mode B è triggered SOLO se l'input contiene la substring esatta `"mode": "programmatic"` (con quote, come appare nell'envelope JSON). Vietato auto-detection da contesto invocazione ("sto girando dentro a una skill") — produce silent Markdown leakage nel parser dell'orchestratore che rompe il gate. Su envelope malformato emette `{status: "error", source: "design-review", reason: "malformed_input_envelope", findings: []}` invece di fallback a Mode A (il fallback silenzioso bloccherebbe la card invece di segnalare).

### Notes

- **Nessun cambio framework payload oltre al command.** Nessun nuovo agente, nessuna nuova skill, nessuna nuova chiave config, nessuna modifica a `/new` o `/e2e-review`. Solo il command `/design-review` viene reso effettivamente dual-mode come annunciato in v3.18.0. `/e2e-review` non lo invoca ancora — quella connessione (Phase 4b) viene rimandata a una future release post-dogfooding, sulla base di dati reali su quanto valore aggiunge la judgment pass sopra il `visual-fidelity-verifier`.
- **Backward compatibility totale per uso interattivo**: utenti che invocano `/design-review /merchant/dashboard` continuano a ricevere il Markdown report identico a v3.17.x e prima. Mode A è il default; Mode B è opt-in via envelope JSON esplicito.
- **Severity taxonomy in lockstep**: il body del command include una nota esplicita "quando aggiorni la taxonomy qui, aggiorna in lockstep `framework/.claude/agents/visual-fidelity-verifier.md` § Severity Taxonomy". Le due fonti restano allineate per evitare drift cross-source nell'aggregatore.
- **Versioning rationale.** PATCH (3.18.0 → 3.18.1): no new capability. Il contratto Mode B era già annunciato in v3.18.0; questa release rende l'annuncio reale. Decision tree CLAUDE.md: "Doc fix, script bugfix, perf, security patch with no behaviour change → PATCH" — qualifica perché la behavior change è solo per chi sta ancora costruendo l'integrazione (nessun consumer esistente la invoca programmatically oggi).
- **Verification** (manuale, no test suite): (1) `/design-review /merchant/dashboard` interattivo → Markdown identico al pre-3.18.1 (regression check Mode A); (2) input contenente `"mode": "programmatic"` → JSON-only output, no preamble, parsable da `JSON.parse()`; (3) envelope malformato → emette `{status: "error", reason: "malformed_input_envelope"}` senza fallback a Markdown; (4) campi `route`/`dev_server_port` mancanti dall'envelope → stesso errore; (5) finding di severity Nitpick → presente in Mode A, assente da Mode B `findings[]`.
- **Path post-3.18.1**: dogfooding di v3.18.0 + 3.18.1 per 1–2 settimane su feature UI reali. Sulla base di falsi positivi/negativi raccolti, decidere data-driven quali estensioni includere in v3.19.0 tra: multi-viewport (mobile/desktop/tablet), interactive states (hover/focus/active/disabled + opt-in loading/empty/error), Phase 4b design judgment pass che invoca `/design-review` in Mode B all'interno di `/e2e-review`. L'adversarial review ha esplicitamente sconsigliato il bundle v3.19.0 senza dati di dogfooding — questa release rispetta quella raccomandazione.

## [3.18.0] - 2026-05-25

The end-to-end review fase post-implementazione diventa un **gate deterministico e BLOCKING** invece di un assortimento di check advisory. Until v3.17.x, la chiusura di una card via `/new` attraversava tre fasi eterogenee tutte non-bloccanti sul fronte UI: Phase 2.5 (Implementation Completeness Check) validava solo card AC / PRD alignment / code quality / schema / artifacts senza alcun gate visivo o funzionale runtime; Phase 2.6 (E2E Testing) era *condizionale* (solo con `test_plan.e2e_required: true` o QA profile BALANCED/DEEP) e invocava `qa-sentinel` che per charter (`qa-sentinel.md:151-161`) esegue solo gate meccanici (lint/type/test/build/audit) — non test browser-based né design; Phase 2.7 (Visual Design Review) invocava `ui-expert` ma era esplicitamente marcata "advisory, non-blocking" (`new/SKILL.md:897` — "findings logged but do NOT block commit"). Le skill `webapp-testing` e `playwright-skill` esistevano ma non erano mai auto-invocate da `/new`. La routine `ds-drift` (`framework/routines/ds-drift.routine.yml`) girava settimanalmente, catturando il drift visivo 1–7 giorni dopo il merge. **Conseguenza concreta** lamentata dall'utente: doveva rifare manualmente la review schermata-per-schermata di ogni card UI, e i dettagli grafici dei mockup venivano regolarmente ignorati.

Questo release risolve il problema con un **HYBRID architecture** validata da advisor: nuova skill orchestratore `/e2e-review` + nuovo agente specializzato `visual-fidelity-verifier`, integrazione BLOCKING in `/new` Phase 2.6 (unificando 2.6 + 2.7 attuali), gating tunabile via `features.e2e_review.*`, con riuso massimo dell'esistente (`coder` per scrivere `.spec.ts`, `playwright-skill` per esecuzione, `webapp-testing` per ispezione runtime in compliance-only, `/design-review` esteso a output JSON, `ui-expert` cascade riusata in compliance-only mode). L'architettura segue il **three-layer harness pattern** documentato da Anthropic in [Effective Harnesses for Long-Running Agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents): implementer (`coder`) e verifier (`visual-fidelity-verifier`) sono distinti, con aggregazione rule-based in mezzo — nessun agente valuta il proprio lavoro. Il pattern **Definition of Done machine-readable** ([Policy Cards, arXiv 2510.24383](https://arxiv.org/abs/2510.24383)) guida la sintesi Gherkin del piano di verifica. La combo **Playwright MCP + Vision** ([Playwright MCP](https://playwright.dev/mcp/introduction), [Giving Claude Code Eyes — round-trip screenshot testing](https://medium.com/@rotbart/giving-claude-code-eyes-round-trip-screenshot-testing-ce52f7dcc563)) è il pattern canonico 2025–2026 per browser+vision in agentic dev.

### Added — Nuova skill `/e2e-review` (orchestratore BLOCKING)

- **Nuova skill [framework/.claude/skills/e2e-review/SKILL.md](framework/.claude/skills/e2e-review/SKILL.md)** (~530 righe). Sei fasi sequenziali:
  - **Phase 1 — Verification Plan Extraction**: legge `card.acceptance_criteria[]`, `card.test_plan` (quando presente, schema compatibile con quello già scritto da `/prd` Step 6, senza modifiche al PRD), `card.requirements[]`, `card.links.design[]`, `card.scope.routes[]` (con fallback derivato dal diff via `codebase-architect`). Sintetizza scenari Gherkin (Given/When/Then) e li persiste in `.baldart/e2e-review/<CARD-ID>/plan.json` come contratto machine-readable consumato dalle fasi successive.
  - **Phase 2 — Mockup Source Cascade (4 livelli)**: (a) URL Figma + Figma MCP runtime probe → `get_design_context` + `get_screenshot`; (b) path locale referenziato in `card.links.design[]` (PNG/JPG/PDF — PDF estratto con `pdftoppm`, degrada gracefully); (c) no mockup ma `features.has_design_system: true` → modalità **compliance-only** che riusa la cascade BLOCKING del [design-system-protocol.md](framework/agents/design-system-protocol.md) (INDEX.md + tokens-reference.md + components/<Name>.md); (d) nulla → skip visual + warning in `gaps[]`, functional E2E continua comunque. Nessun hard fail per assenza di Figma — la dipendenza è genuinamente opzionale.
  - **Phase 3 — Functional E2E**: spawn `coder` con contratto preciso (file in `${paths.e2e_tests_dir}/<CARD-ID-slug>.spec.ts`, screenshot path fissato a `.baldart/e2e-review/<CARD-ID>/screenshots/<route-slug>.png` per consumo da Phase 4, no chromium.launch, no edit fuori `${paths.e2e_tests_dir}`). Esegue `npx playwright test --reporter=json` headless in modalità programmatica (headed riservato a flussi OTP in invocazione manuale). Mappa scenari falliti in finding funzionali con severity rule-based (auth/payments/data-mutation = Critical; display-only = Major; flaky-passed-on-retry = Minor + `confidence: low`).
  - **Phase 4 — Visual Fidelity**: pre-filter con pixel-diff (`pixel_diff_threshold` configurabile, default `0.02`) prima di invocare Vision — quando la diff è sotto soglia, il route passa senza spendere token Vision (primary latency/cost saver). Sopra soglia o in compliance-only, spawn `visual-fidelity-verifier` con payload strutturato; deduplicazione findings per `(route, category, region)`.
  - **Phase 5 — Aggregation + Strict Severity Gate**: combina functional + visual findings, applica filtro tolerance (`strict` default → Critical+Major+Minor block; Minor con `confidence: low` auto-demote ad advisory — escape hatch documentato contro font-rendering false positive). Self-heal loop bounded (`max_self_heal_iterations`, default `2`): re-spawn `coder` con findings come fix instructions, re-esegui Phase 3+4, re-decidi. Override path: dopo esaurimento iterazioni in modalità manuale chiede reason obbligatoria (`require_override_reason: true`); in modalità programmatica restituisce `verdict: "blocked"` lasciando la decisione a `/new`.
  - **Phase 6 — Return Protocol**: scrive `.baldart/e2e-review/<CARD-ID>/report.json` con schema strutturato `{status, iterations, findings[], ds_drift_codes[], gaps[], override_reason, transcript_paths[]}`. Modalità manuale renderizza anche Markdown summary; programmatica emette solo JSON come messaggio finale (parsato direttamente da `/new`).
- **Gate di skip pre-flight** per evitare overhead inutile: skill REFUSES quando `features.has_e2e_review: false` (preserva backward compat); auto-skip quando card non ha `links.design` E `git diff` non mostra file UI (`.tsx`/`.css`/`.scss`/`.svelte`/`.vue` sotto `${paths.app_dir}` o `${paths.components_primitives}`); auto-skip su card type `docs`/`chore`/`config`/`backend`/`api`/`db`/`infra`. Lo skip non è failure — `/new` procede a Phase 3 normalmente.
- **Re-run trigger**: se `/codexreview` (Phase 3.7) modifica file sotto `${paths.app_dir}`/`${paths.components_primitives}` o qualsiasi `*.css`/`*.scss` dopo il primo run, l'orchestratore re-invoca `/e2e-review` PRIMA del commit di Phase 4. Protegge contro code-review changes che reintroducono silenziosamente drift visivo.

### Added — Nuovo agente `visual-fidelity-verifier`

- **Nuovo agente [framework/.claude/agents/visual-fidelity-verifier.md](framework/.claude/agents/visual-fidelity-verifier.md)** (`model: sonnet`, ~280 righe). Worker stateless multimodale invocato esclusivamente da `/e2e-review` Phase 4. Unico nuovo agente del release — l'advisor ha bocciato la creazione di un secondo `e2e-functional-verifier` perché duplicherebbe il pattern Phase 2.6 attuale (`coder` scrive `.spec.ts`, Playwright esegue).
  - **Hard prohibitions non-negotiable**: NEVER reads application source code (anti assertion-fitting bias — antipattern documentato Meta/Capgemini 2025: LLM-generated assertions che lockano il bug come "expected behavior"), NEVER edits files, NEVER proposes architectural refactors, NEVER declares "done" o "passed" (la decisione di gate è dell'orchestratore).
  - **Input contract strict JSON**: `{card_id, route, viewport, implementation_screenshot_path, mockup_source: {level, mockup_path | figma_node_id | design_system_index_path | tokens_reference_path | components_in_scope}, tolerance, fix_hints_enabled}`.
  - **Severity Taxonomy canonica** (SSOT consumata anche da `/e2e-review` e `/design-review`): **Critical** (`layout-break`, `element-order`, `responsiveness-break`, `component-missing`, `component-duplicated`, `unreachable-action`) blocca in ogni tolerance; **Major** (`spacing-off-scale`, `typography-{family,weight,size,line-height}`, `color-mismatch`, `color-opacity`, `gradient-direction`, `token-bypass`, `interactive-state-missing`, `a11y-contrast`, `a11y-focus-visible`) blocca in strict+balanced, advisory in lenient; **Minor** (`border-radius-off`, `shadow-off`, `micro-misalignment`, `font-rendering-variance`) blocca solo in strict (con auto-demote `confidence: low` ad advisory anche in strict).
  - **Output schema strict JSON**: `{status, route, mockup_source_level, viewport, compliance_score, findings: [{severity, category, description, expected, actual, fix_hint, evidence: {region, screenshot_crop_path}, ds_drift_code, confidence}], ds_drift_codes[], gaps[], transcript_path}`. Citazioni token rigorose (mai paraphrase — `--color-action-primary` non "the primary blue"). Anti-hallucination discipline (no finding emessi senza evidenza visibile; `status: error` invece di guessing quando screenshot insufficiente).
  - **Mockup source cascade behaviour**: in modalità `figma`/`local` il mockup è ground truth visuale; in `compliance-only` riusa la cascade BLOCKING del [design-system-protocol.md](framework/agents/design-system-protocol.md) (INDEX + tokens-reference + per-component specs) e mappa findings ai drift code canonici `DS_TOKENS_DRIFT`/`DS_COMPONENT_STALE`/`DS_INDEX_DRIFT` (definiti nel protocollo, non qui — un solo SSOT).

### Added — Schema config propagation (regola interna del repo)

- **Template [framework/templates/baldart.config.template.yml](framework/templates/baldart.config.template.yml)**: aggiunto `features.has_e2e_review: false` (default opt-in esplicito — backward compatibility totale) e nuova sezione top-level `e2e_review:` con 4 chiavi commentate. La chiave esistente `paths.e2e_tests_dir` è ora consumata anche da `e2e-review` oltre che da `playwright-skill`.
- **`features.e2e_review` tuning sub-section**: `fidelity_tolerance` (`strict` default — Critical+Major+Minor block con escape hatch low-confidence; `balanced` blocca Critical+Major; `lenient` solo Critical), `max_self_heal_iterations` (default `2`, allineato a Phase 2.5 retry policy in `new/SKILL.md:699`), `pixel_diff_threshold` (default `0.02`, range 0.0–1.0 — pre-filter pixel-diff per skippare la Vision call quando implementazione e mockup sono pixel-identici), `require_override_reason` (default `true` — la reason loggata in tracker `## Issues & Flags` e in `report.json`).
- **[src/commands/configure.js](src/commands/configure.js)** prompt + status: aggiunto `has_e2e_review` al loop dei feature prompt (linea ~563 nuovo); nuova sezione "E2E review tuning" prima della sezione LSP (linea ~624 nuovo) che gating su `has_e2e_review: true` chiede via `UI.select` la tolerance e poi le 3 chiavi nested; aggiunta riga "E2E review:" nel box AUTODETECTED. Il detector schema-drift in `update.js` (linee 462–503) cattura automaticamente la nuova chiave `features.has_e2e_review` via template scan generico — nessun cambio dedicato necessario. `doctor.js` `configSchemaDrift()` (linee 70–82) usa lo stesso meccanismo.
- **[framework/AGENTS.md](framework/AGENTS.md)** tabella MUST aggiornata (linea 18): aggiunta riga `E2E review BLOCKING gate (Phase 2.6 of /new invokes /e2e-review) | features.has_e2e_review: true`.

### Changed — `/new` Phase 2.6 unificata e BLOCKING

- **[framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md) linee 814–901**: sostituite integralmente la legacy Phase 2.6 (E2E Testing condizionale invocante `qa-sentinel`) e Phase 2.7 (Visual Design Review advisory invocante `ui-expert`) con una singola Phase 2.6 unificata "End-to-End Review (BLOCKING, since v3.18.0)" che invoca `/e2e-review` in modalità programmatica. Il gate documentato sostituisce il branching condizionale precedente (`test_plan.e2e_required` + QA profile + legacy fallback heuristic): 3 condizioni di skip (feature disabled / backend-only card / docs|chore|config card), payload JSON strutturato all'invocazione, contract di ritorno JSON parsato programmaticamente. Mapping verdict→azione: `passed`/`skipped`/`overridden` → procede; `blocked` STOPpa la card senza procedere a Phase 3 e chiede all'utente override/escalation/abandon; `error` chiede retry/skip/abandon. Re-run trigger dopo Phase 3.7 `/codexreview` se file UI modificati. Tracker output strutturato per audit cross-card.

### Changed — `/design-review` command estensione output JSON

- **[framework/.claude/commands/design-review.md](framework/.claude/commands/design-review.md)** prologo: aggiunta nota di invocazione che documenta il dual-mode. Interactive (default) produce Markdown report (Blockers/High/Medium/Nitpicks) come oggi; programmatic mode (input contiene `"mode": "programmatic"`) ritorna un singolo JSON object matching lo schema di `visual-fidelity-verifier` per gating dell'orchestratore. Nessuna logica nuova nel command body — solo flag di output format. Permette al `/e2e-review` skill di riusare `/design-review` come tooling alternativo quando appropriato senza duplicare la cascade design-system.

### Changed — Documentazione integrazione skill tooling

- **[framework/.claude/skills/playwright-skill/SKILL.md](framework/.claude/skills/playwright-skill/SKILL.md)**: nuova sezione "E2E Review Integration (since v3.18.0)" prima di Quick Reference. Documenta il contratto consumer: spec generata da `coder` partendo da `plan.json`, screenshot path fissato per consumo da Phase 4, headless in modalità programmatica, mapping severity di scenari falliti. Nessuna logica nuova nella skill — solo documentazione del nuovo consumer.
- **[framework/.claude/skills/webapp-testing/SKILL.md](framework/.claude/skills/webapp-testing/SKILL.md)**: nuova sezione "E2E Review Integration (since v3.18.0)" in testa. Documenta l'uso in compliance-only mode (cascade level c) per fetch dei computed style runtime quando manca mockup-as-image; ruolo del server lifecycle helper. Nessuna logica nuova.
- **[framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml)** linea 134 ca.: aggiunto commento sopra `test_plan:` che documenta il nuovo dual consumer (`/new` legacy fallback + `/e2e-review` Phase 1 — il mapping verso Gherkin scenarios). Nessuna modifica allo schema PRD; le card scritte prima della v3.18.0 sono pienamente compatibili (Phase 1 di `/e2e-review` ha fallback da `acceptance_criteria` quando `test_plan` manca).

### Changed — Registry e routing

- **[framework/.claude/agents/REGISTRY.md](framework/.claude/agents/REGISTRY.md)**: aggiunta riga per `visual-fidelity-verifier` immediatamente sotto `ui-expert` (sezione Design). Specializzazione: visual diff, severity taxonomy, design-system compliance via registry cascade. Auto-invoke yes via `/e2e-review` — non spawnabile ad-hoc (l'orchestratore è il solo entry point).

### Changed — Documentation utente

- **[README.md](README.md)**: agent count aggiornato da 24 a 27 (visual-fidelity-verifier + numerazione resa coerente Design&UX→Specialized); skill count da 24 a 25 (e2e-review aggiunto sotto "Code quality"); nuova sezione "End-to-End Review BLOCKING Gate (new in v3.18.0)" prima della sezione "UI Excellence + Post-Intervention Coherence Gate (v3.12.0)" che riassume l'intera architettura, citando le fonti del 2025–2026 e documentando il tuning `e2e_review.*`.
- **[framework/docs/PROJECT-CONFIGURATION.md](framework/docs/PROJECT-CONFIGURATION.md)**: tabella paths aggiornata (e2e_tests_dir ora cita anche e2e-review); nuova sezione "§ 4.2.1 `e2e_review` — tuning for the BLOCKING end-to-end review gate (v3.18.0+)" che documenta ogni chiave con tipo/default/effetto e citazioni cross-reference al protocollo `design-system-protocol.md` per i drift code.

### Notes

- **Backward compatibility totale.** Il default `features.has_e2e_review: false` significa che progetti pre-3.18.0 non vedono nessun cambio di comportamento finché non eseguono `npx baldart configure` (che proporrà la nuova feature) o editano manualmente la config. `/new` Phase 2.6 silenziosamente skippa la chiamata a `/e2e-review` quando il flag è false, preservando il workflow legacy con un breve log informativo nel tracker. Le card scritte prima della v3.18.0 (senza `test_plan` block, o con `test_plan.e2e_required: false`) sono pienamente supportate dalla nuova fase: Phase 1 di `/e2e-review` deriva Gherkin scenarios direttamente da `acceptance_criteria` quando `test_plan` è assente, e la mockup cascade degrada gracefully a compliance-only / skip quando `links.design` manca.
- **Riuso massimo, footprint minimo.** Un solo nuovo agente (`visual-fidelity-verifier`) e una sola nuova skill (`/e2e-review`). Riusati: `coder` (scrittura `.spec.ts` + self-heal loop), `playwright-skill` (esecuzione spec), `webapp-testing` (compliance-only computed styles fetch), `codebase-architect` (mapping route→file fallback), `/design-review` command (esteso a output JSON per gating), `ui-expert` cascade (referenziata da `visual-fidelity-verifier` in compliance-only mode), `design-system-protocol.md` (SSOT della cascade BLOCKING e dei drift code — referenziato, mai duplicato), `ds-drift` routine (resta safety net settimanale per edit umani che bypassano `/new`). Bocciata da advisor la creazione di un secondo agente `e2e-functional-verifier` (avrebbe duplicato il pattern Phase 2.6 attuale) e di un terzo "evaluator" three-agent-harness completo (overkill quando l'aggregazione è rule-based su severity strutturata, da introdurre solo se dogfooding rivela troppi falsi positivi).
- **Dependency contract.** Il consumer deve avere `@playwright/test` installato e `paths.e2e_tests_dir` valorizzato; lo skill rileva il gap in Phase 3 step 1 e fornisce setup hint actionable (`npm i -D @playwright/test && npx playwright install chromium`). Figma MCP è opzionale — la cascade Phase 2 degrada a level (b)/(c)/(d) senza errori. Vision usage è governata dal `pixel_diff_threshold` pre-filter — la maggior parte dei route passa senza spendere token Vision.
- **Override path con audit trail.** Quando self-heal esaurisce iterazioni in modalità manuale e l'utente vuole comunque procedere, una reason è obbligatoria (`require_override_reason: true` default). La reason è loggata sia nel `/tmp/batch-tracker-<FIRST-CARD-ID>.md` come `[E2E-OVERRIDE] <reason>` sia in `.baldart/e2e-review/<CARD-ID>/report.json` come campo `override_reason`. Il rationale: il gate deve essere strict per default ma deve avere una via di fuga documentata — gate troppo rigidi senza escape hatch generano user override fatigue e perdita di fiducia (failure mode documentato).
- **Re-run automatico dopo `/codexreview`.** Se Phase 3.7 modifica file UI, `/e2e-review` re-esegue PRIMA del commit di Phase 4. Lo state directory `.baldart/e2e-review/<CARD-ID>/` viene riusato (no re-fetch mockup, no re-build plan) — solo Phase 3 (functional) + Phase 4 (visual) vengono ri-eseguite. Protezione contro la classe di bug dove il code-review introduce silenziosamente drift visivo per "consistency cleanup".
- **Verification è manuale (no test suite — `npm test` resta no-op stub).** Dogfood scenarios da girare dopo l'install:
  1. `npx baldart configure` su progetto pre-3.18.0 → la propagazione schema drift detector cattura `features.has_e2e_review` mancante e offre `configure` automaticamente. Il prompt "Enable BLOCKING end-to-end review?" appare; rispondendo `Yes` la sezione "E2E review tuning" chiede tolerance/max-iter/threshold/require-reason.
  2. `/e2e-review CARD-ID` manuale su una card UI esistente con `features.has_e2e_review: true` + `links.design` valorizzato → il piano viene scritto in `.baldart/e2e-review/<CARD-ID>/plan.json`, il `.spec.ts` generato in `tests/e2e/`, screenshot scattate, `visual-fidelity-verifier` invocato per ogni route, report JSON + Markdown summary mostrato all'utente.
  3. `/new <CARD-ID>` su card UI dopo opt-in → Phase 2.6 invoca `/e2e-review` programmatically; verdict `passed` procede a Phase 3 normalmente; verdict `blocked` STOPpa la card e chiede override/escalation/abandon all'utente.
  4. `/new <CARD-ID>` su card backend-only (no `links.design`, no file UI nel diff) → Phase 2.6 auto-skip con log `"e2e-review: SKIP (backend-only card)"`, procede a Phase 3 senza errori.
  5. Test sintetico iniettivo: card con 3 deviazioni note pre-iniettate (colore button hardcoded, padding fuori scala token, border-radius literal) → report contiene esattamente 3 finding con severity (Major, Major, Minor in strict mode) e gate blocca il commit.
  6. Negative path Figma MCP non disponibile → cascade falls through a level (b) (immagine locale) o (c) (compliance-only) senza hard fail; warning loggato in `gaps[]`.
  7. Self-heal loop: card con token bypass fixabile → il loop converge in ≤ 2 iterazioni e il commit procede senza override richiesto.
  8. Override path: card con falso positivo evidente → l'utente fornisce reason, commit procede con `[E2E-OVERRIDE] <reason>` loggato nel tracker e in `report.json`.
- **Versioning rationale.** Impact MINOR (3.17.1 → 3.18.0): aggiunge una skill (`/e2e-review`), un agente (`visual-fidelity-verifier`), una fase BLOCKING nella skill esistente `/new`, e nuove chiavi di configurazione (`features.has_e2e_review` + sub-section `e2e_review.*`). Non rompe installazioni esistenti — il default `false` mantiene il comportamento legacy fino a opt-in esplicito via `configure`. Non rimuove agenti / cambia layout di directory / cambia comandi CLI (decision tree MAJOR non si applica).

## [3.17.1] - 2026-05-23

CLI quality-of-life patch shaken loose by a hands-on consumer update simulation (mayo, 3.13.0 → 3.16.0). Four real bugs surfaced when actually driving `baldart update` end-to-end from a scripted context. Two were show-stoppers for any non-interactive use (CI, AI-driven, headless update agent), one was cosmetic but ugly, one was a stale stub that lied to the user. All four are fixed; nothing in this release changes framework payload behaviour for end-users running BALDART interactively in a TTY.

### Fixed

- **`baldart update` gains `--yes`/`-y` and `--auto-stash` flags** in [bin/baldart.js](bin/baldart.js) and [src/commands/update.js](src/commands/update.js). Until now, every `UI.confirm` and the `UI.select` arrow-key stash prompt blocked indefinitely on a non-TTY stdin (inquirer's pattern detection mis-handles re-rendered prompts when fed via pipe or `expect`). Both flags route through a single internal `confirm()` wrapper that short-circuits to the prompt's default when `--yes` is set; `--auto-stash` is the narrower variant for users who want manual control over confirms but auto-stash on dirty trees. `--yes` implies `--auto-stash` and additionally passes `--non-interactive` down to the `baldart configure` sub-invocation triggered by the schema-drift detector, plus switches symlink reconcile from `mode: 'prompt'` to `mode: 'force'` (the conflict-resolution prompts inside `SymlinkUtils` would otherwise reintroduce blocking I/O even with `--yes`).
- **`baldart status` actually checks the framework against the upstream tip** in [src/commands/status.js](src/commands/status.js). Previously the "Update Status" section was a hard-coded `UI.success('Up to date')` stub regardless of the installed version — a 4-versions-stale consumer was told everything was fine. The fix calls `git.fetch('antbald/BALDART')` + `git.getRemoteVersion(...)` (already exported by `src/utils/git.js`), compares semver with a tiny inline `cmpSemver` helper, and renders one of four outcomes: update available (`X → Y`), local ahead of remote (unreleased commit, common during BALDART self-development), up to date, or "Cannot check (offline?)" on fetch failure. The SUMMARY box now carries the update line too, so it's visible without scrolling past the per-section output.
- **`baldart --version` and `baldart --help` exit cleanly** in [bin/baldart.js](bin/baldart.js). Commander 9 throws a `CommanderError` with `exitCode: 0` from these flags as its internal flow-control mechanism; the previous `bin/baldart.js` had no try/catch around `program.parse(process.argv)`, so the throw escaped to Node and printed a 12-line `CommanderError: 3.17.0` stack trace right after the version itself. Now wrapped in try/catch that recognises the three known sentinel codes (`commander.version`, `commander.help`, `commander.helpDisplayed`) and exits 0 silently; real errors keep the previous behaviour (log message + non-zero exit).
- **CLI version is read from the `VERSION` file, not `package.json`** in [bin/baldart.js](bin/baldart.js). `package.json.version` is only synced at publish time via `npm run sync-version` (see `scripts/sync-version`), so in dev / local-source / pre-publish contexts it lagged the actual `VERSION` file — `node bin/baldart.js --version` was printing `3.14.1` while VERSION was `3.17.0`. A new `resolveVersion()` reads `VERSION` at module load with `package.json.version` as fallback. Published `baldart@<x.y.z>` on npm is unaffected (the workflow still runs `sync-version` first, so both files agree in published artefacts).

### Notes

- **No framework payload changes** — all four fixes are inside `bin/` and `src/`. Skills, agents, templates, hooks, routines, docs are untouched. The next `baldart update` for consumers picks up the new flags and the working status check without any change to the framework files they install.
- **Not a v3.18.0 because no new capability shipped** — this is purely "what should already have worked". The new flags are additive (zero-config end-users in a TTY see no change) and the status fix is replacing a stub with the real check it was always supposed to be.
- **Bug #4 from the mayo simulation (schema-drift detector skipped with `--no-commit`) was a false positive** — the drift detector at lines 462–503 runs before `postUpdateAutoCommit` (line 506), so `--no-commit` doesn't reach it. Verified by reading the call sequence, no code change needed.
- **Verification** (manual, no test suite): (1) `node bin/baldart.js --version` prints `3.17.1` and exits 0 silently; (2) `node bin/baldart.js update --help` shows the three options (`--no-commit`, `--yes/-y`, `--auto-stash`); (3) in a stale consumer repo, `baldart status` reports "Update available" with the version diff; (4) `baldart update --yes` on a dirty consumer auto-stashes, pulls, reconciles symlinks in force mode, runs `configure --non-interactive`, and exits 0 with no blocking prompts.

## [3.17.0] - 2026-05-23

The `/prd` skill gains a third option in the design phase. Until v3.16.0, when a user reached Step 3 without having provided mockups (`mockups.status: none`), the only path forward was internal generation via the `ui-design` skill (3 HTML options → iteration → save). With Claude Design now offering a design-system-sync'd visual generator, users who prefer iterating in that surface had to abandon the PRD flow halfway, work externally, and then re-enter as if they'd had mockups from the start. This release adds **Step 3.0 — Mockup Source Decision**: when `mockups.status: none` AND UI impact ≠ N/A, the skill explicitly asks "internal generation or handoff to Claude Design?". The handoff branch generates a context-rich prompt (discovery + brand voice + design-system tokens + screens in scope + stack target) ready to paste, then STOPs. On user re-entry with local paths, it re-uses the existing Step 1.6 Mockup Intake verbatim — copy → analyze → populate state — and Step 3 falls naturally into Hybrid mode. Zero code paths downstream of the bivio are new; the entire feature is a decision point + a prompt generator on top of pre-existing infrastructure.

### Added — PRD Step 3.0 Mockup Source Decision

- **§ Step 3.0 in [framework/.claude/skills/prd/references/ui-design-phase.md](framework/.claude/skills/prd/references/ui-design-phase.md)** before the precondition decision tree. Activates only when `mockups.status: none` AND UI impact ≠ N/A. Poses the binary question (Internal `ui-design` vs Handoff Claude Design) with explicit STOP gate — silence or ambiguity re-asks rather than defaulting silently. Branch A sets `mockups.source: internal` and falls through unchanged. Branch B collects context from state file + `baldart.config.yml` + `.baldart/overlays/prd.md` + `${paths.design_system}/INDEX.md` + `tokens-reference.md` (when `features.has_design_system: true`), renders the handoff prompt template, presents it inside a fenced code-block with copy-paste instructions, persists `mockups.status: pending_external` + `mockups.source: claude-design` + `mockups.handoff_prompt_path` in the state file (prompt saved to `${paths.prd_dir}/<slug>/handoff/claude-design-prompt.md` for session-restart survival), and STOPs.
- **Re-entry handler in same file**. When the user returns with local paths after handoff, the skill invokes the existing Step 1.6 Mockup Intake procedure verbatim (no code duplication) — copy into `mockups/` with collision-safe rename, run Mockup analysis schema, populate `## UI Design`. Then flips `mockups.status` from `pending_external` to `provided` or `partial` based on coverage of `screens_in_scope[]`, keeps `mockups.source: claude-design` as audit trail, and falls through to the precondition decision tree which routes to Hybrid mode unchanged.
- **Precondition decision tree extended** with two new rows: `mockups.status: pending_external` ⇒ STOP (handoff in flight); `mockups.status: none` AND user picked Branch A ⇒ Full (preserves backwards-compatible default). The existing `provided` / `partial` ⇒ Hybrid and UI = N/A ⇒ Skip rows are unchanged.
- **New template asset [framework/.claude/skills/prd/assets/claude-design-handoff-prompt.template.md](framework/.claude/skills/prd/assets/claude-design-handoff-prompt.template.md)**. Seven-section structure populated at runtime: (1) feature objective, (2) user flow + personas, (3) screens in scope with required states + primary actions + data shown, (4) brand & voice (conditional on overlay), (5) design system constraints with tokens excerpt + registered primitives + numeric tables from `framework/agents/design-system-protocol.md` (conditional on `features.has_design_system: true`), (6) stack target + output format preference derived from `stack.framework`, (7) deliverable spec (PNG exports per state + code/Figma + inline notes on new components / token deviations / layout decisions). Generic and project-agnostic — all project-specific content is injected as placeholders.
- **Hard Rule 16 note in [framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md)** added to the existing Mockup Intake rule. Codifies that a NO answer at intake routes through Step 3.0 (not directly to Full mode), with cross-reference to the ui-design-phase section. The rule itself remains unchanged for backwards compatibility — Step 3.0 is purely additive.

### Notes

- **No new `baldart.config.yml` keys** — the schema-change propagation rule does not apply to this release. The bivio is runtime-only per the user's preference (zero persistent flag); no template/configure/update/doctor updates needed. All context for the handoff prompt is read from existing config keys (`paths.*`, `stack.*`, `features.has_design_system`).
- **Backwards compatibility.** PRDs in flight on v3.16.x finish on v3.16.x semantics. New PRD sessions started on v3.17.0+ see Step 3.0 only when they actually reach Step 3 with `mockups.status: none` — the default (Branch A) preserves the v3.16.x behavior exactly. Existing PRDs that already provided mockups never enter Step 3.0.
- **`ui-design` skill is unchanged** — Step 3.0 sits in front of it, doesn't modify it. The Full pipeline (3a → 3d) is reached identically when the user picks Branch A or when no mockup-related state exists at all. The `ui-expert` agent and `huashu-design`/`frontend-design` skills are equally untouched.
- **`mockup_analysis` schema is unchanged** — the new `pending_external` value of `mockups.status` is purely transitional and is overwritten by `provided` / `partial` on user re-entry, so existing PRD templates and state files keep validating without migration.
- **Files modified outside the Step 3.0 scope** (`card-template.yml`, `backlog-phase.md`, `validation-phase.md`, `prd-card-writer.md`, `framework/.claude/skills/new/SKILL.md`) are the v3.16.0 in-flight release (agent routing + UI specialist briefing — see entry below) and are NOT part of v3.17.0. v3.16.0 should be tagged and pushed before v3.17.0 in the release sequence.
- **Verification is manual** (no test suite). Dogfood scenarios: (1) `/prd` with mockups provided at intake → never enters Step 3.0, Hybrid mode as today (regression check); (2) `/prd` without mockups, answer "A" at Step 3.0 → Full mode runs identically to v3.16.x; (3) `/prd` without mockups, answer "B" → prompt rendered with discovery + tokens + screens + stack, saved to `handoff/`, state shows `pending_external`, skill STOPs; (4) simulate re-entry with a PNG path → file copied to `mockups/`, status flips to `provided`, Step 3 enters Hybrid mode invoking `ui-design` only for uncovered screens; (5) `features.has_design_system: false` → § 5 of the rendered prompt is omitted entirely; (6) no overlay present → § 4 brand voice omitted.

## [3.16.0] - 2026-05-23

The `/new` orchestrator becomes specialist-aware: it now reads `owner_agent` from each card YAML and dispatches the right implementation agent — `coder` for backend/logic, `ui-expert` for UI cards, with `plan`/`visual-designer`/`motion-expert` reserved as stubs that fall back to `coder` with a WARN. UI cards (`owner_agent: ui-expert`) receive a dedicated mission briefing that enforces the registry-first BLOCKING cascade (read `${paths.design_system}/INDEX.md` + `tokens-reference.md` + per-component specs BEFORE coding), the implementation-only constraint (no mockup redesign — STOP and ask on registry conflicts), and the Post-Intervention Coherence Check (reconcile INDEX + components + tokens in the same commit). The PRD side enforces the contract end-to-end: `owner_agent` is now a mandatory enum on every child card, UI cards have UI-only scope (logic moves to a sibling `coder` card), and the validation phase blocks commit on enum violations + warns on mixed-scope UI cards.

Until now, every card from `/new` was spawned to the generic `coder` agent regardless of `owner_agent`, and mockups arrived as a textual "match this reference" hint without triggering the registry/tokens/specs cascade that `ui-expert` mandates. The symptom reported by users: graphical code generated from mockups was "not quite right" — wrong shadows, off-scale font sizes, off-palette colors. Root cause: the right specialist existed and PRD cards declared it (`owner_agent: ui-expert`), but the dispatcher never read the field.

### Added — Agent router + UI specialist briefing

- **§ Agent Routing section in [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md)** between Pre-flight and Per-card pipeline. Documents the `owner_agent` → spawn-target mapping in a single table (coder | ui-expert | plan | visual-designer | motion-expert → resolved agent + briefing variant) with a cross-reference list of the 5 layers that must stay in lockstep (template, prd.md enum, prd-card-writer rule, validation-phase check, REGISTRY.md decision tree).
- **Step 6b "Agent dispatch (router)" in Phase 2** of [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md). Reads `card.owner_agent`, resolves the spawn target via a switch with explicit fallbacks: `coder` → coder briefing; `ui-expert` → ui-expert briefing; `plan`/`visual-designer`/`motion-expert` → coder briefing + WARN ("specialized briefing not yet implemented"); empty/`claude`/missing → coder briefing + WARN; unknown value → HALT and ask the user. The resolved spawn + briefing pair is logged in the tracker as the audit trail.
- **`[UI-ONLY]` mission briefing sections** in Phase 2 step 7 of [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md). Three new sections inserted between "Design Reference" and "File Permissions", included ONLY when the router resolves `briefing = "ui-expert"`:
  - **Registry-first BLOCKING pre-work** — reproduces verbatim the cascade defined in [framework/.claude/agents/ui-expert.md](framework/.claude/agents/ui-expert.md) § "Registry-first pre-work (BLOCKING)": read `INDEX.md` + `tokens-reference.md` + `components/<Name>.md` for every primitive in scope, declare which Authority Matrix rows govern the card in the first response. Orchestrator rejects the implementation if the declaration is missing.
  - **Implementation-only constraint** — explicit DO NOTs: don't redesign the mockup, don't silently adapt registry-conflicting details (STOP + `[DS-DRIFT] mockup-vs-registry` log + ask the user), don't extend scope into business logic / API / state / validation (split-into-sibling-card is upstream PRD concern).
  - **Post-Intervention Coherence Check** — citation of [framework/agents/design-system-protocol.md](framework/agents/design-system-protocol.md) § "Post-Intervention Coherence Check": reconcile `INDEX.md` (drift code `DS_INDEX_DRIFT`) + `components/<Name>.md` (`DS_COMPONENT_STALE`) + `tokens-reference.md` (`DS_TOKENS_DRIFT`) in the same commit. The completion report MUST include a `design_system_coherence:` block enumerating the artifacts updated.
- **`[UI-ONLY]` design-system conformance step in Phase 2.5** of [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md). New row in the Completeness Check verification depth, runs only when `briefing = "ui-expert"` AND `features.has_design_system: true`. Greps changed files for hardcoded design values that bypass the registry (hex colors, `rgb(`, inline `box-shadow:` offsets, raw px/rem `border-radius:`/`font-size:`/spacing); verifies the `design_system_coherence:` block in the completion report; verifies new primitives ship their `components/<Name>.md` spec in the same commit; verifies new tokens land in `tokens-reference.md`. Each gap classified Partial → fix-coder sub-loop.

### Added — PRD-side enforcement (split-scope + mandatory enum)

- **§ MANDATORY Card Rules (Rule A + Rule B)** in [framework/.claude/skills/prd/references/backlog-phase.md](framework/.claude/skills/prd/references/backlog-phase.md). New section between MANDATORY Card Structure and Specialist Audits. **Rule A**: every child card MUST declare `owner_agent` with one of the five enum values; missing/empty/placeholder/non-enum BLOCKS commit. **Rule B**: cards with `owner_agent: ui-expert` MUST be UI-only — components, layout, styling, motion, accessibility — and MUST NOT bundle business logic, API integration, state management, validation logic, or backend changes; mixed steps split into a `ui-expert` card (visual, consumes mocked contract) + a `coder` card (logic, API, state) wired via `depends_on:`/`blocks:`. Ships a canonical Login feature split example (bad: single mixed card; good: `FEAT-XXXX-02` UI + `FEAT-XXXX-01` logic).
- **§ Agent Specialization Rules in [framework/.claude/agents/prd-card-writer.md](framework/.claude/agents/prd-card-writer.md)** before Card Atomicity Rules. The same Rule A + Rule B formulated from the agent writer's POV (which value to pick per work type, what NOT to bundle, when not to assign `ui-expert` to a tight-knit mixed card). Cross-references the backlog-phase split example for canonical guidance.
- **§ Step 6 Agent Specialization Audit (new step 0)** in [framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md). Runs FIRST, before the Field Name Audit. Step 0a: `owner_agent` enum check (BLOCKING) — halts the audit on missing/empty/placeholder/non-enum with an actionable message. Step 0b: UI-only scope heuristic (WARN, soft gate) — for each `ui-expert` card, greps `scope.summary` + `requirements` + `acceptance_criteria` for logic markers (`api call`, `endpoint`, `route handler`, `state machine`, `validation logic`, `business rule`, `database`, `migration`, `prisma`, `auth flow`, `token storage`, `webhook`, `cron`) and emits `[UI-SCOPE-WARN]` lines for human review. Intentionally soft: legitimate UI references to mocked contracts ("consumes `useAuth()` mock") don't trip the WARN thanks to the consume/mock/contract-reference exception.
- **§ Card-template fields updated** in [framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml). The placeholder `owner_agent: claude` is replaced with `owner_agent: <REQUIRED — coder | ui-expert | plan | visual-designer | motion-expert>` so a card left at the placeholder fails validation. A leading comment block documents each enum value (one line each) and a comment under `scope:` reminds the writer of the UI-only constraint for `ui-expert` cards.
- **§ Epic-template aligned to the tracker-only contract** in [framework/.claude/skills/prd/assets/epic-template.yml](framework/.claude/skills/prd/assets/epic-template.yml). `owner_agent` changes from the legacy `claude` placeholder to `""` (empty string) with an inline comment clarifying that epics are tracker-only — the dispatcher skips them and the validation-phase enum-check exempts them — so the empty value is the correct signal that no specialist is assigned to an epic. Prevents the false-positive of writing a child enum value (e.g. `coder`) on an epic and implying it carries implementation work.

### Notes

- **No new `baldart.config.yml` keys** — the schema-change propagation rule does not apply to this release. The router consumes an existing field (`owner_agent`) already present in the PRD card schema.
- **Backwards compatibility for cards written before v3.16.0.** The dispatcher accepts `owner_agent: ""` / missing / legacy `claude` value via the WARN-and-fallback-to-coder branch, so existing backlogs keep working — the WARN surfaces the gap on first use without blocking. New PRD sessions on v3.16.0+ produce enum-compliant cards from the start; the validation phase's enum BLOCKER applies only to NEW cards generated under v3.16.0+.
- **Specialist briefings for `plan` / `visual-designer` / `motion-expert` are stubs** in this release — they fall back to the coder briefing with a WARN. Specialized briefings will follow in future iterations without further schema changes (the router structure is already in place).
- **`ui-expert` in `/new` is implementation-only by design** — refinement of the mockup remains a PRD-phase responsibility (the `ui-design-phase` skill handles design decisions). If the implementation step finds a mockup-vs-registry conflict, `ui-expert` STOPs and asks the user; it does not auto-adapt the mockup. This boundary is documented in the `[UI-ONLY] Implementation-only constraint` briefing section.
- **REGISTRY.md decision tree** in [framework/.claude/agents/REGISTRY.md](framework/.claude/agents/REGISTRY.md) gains a header note clarifying that `/new` dispatches by `owner_agent` (not by re-running the decision tree) — the tree remains authoritative for ad-hoc spawns and for `prd-card-writer`'s writing decisions, with a lockstep-with-skill warning to keep the 5 layers in sync.
- **Verification is manual** (no test suite — `npm test` remains a no-op stub). Dogfood scenarios to run after install: (1) `/new <CARD-ID>` on a card with `owner_agent: coder` → identical behavior to v3.15.x (no regression); (2) `/new <CARD-ID>` on a card with `owner_agent: ui-expert` + valorized `links.design` → `ui-expert` is spawned (not coder), the briefing contains the BLOCKING cascade verbatim, the implementation reads `INDEX.md` + `tokens-reference.md` + `components/<Name>.md` BEFORE coding, the completion report ships the `design_system_coherence:` block, the commit reconciles the three design-system sources; (3) `/new` on a legacy card with `owner_agent: claude` or missing field → fallback to coder + WARN line in tracker; (4) `/prd` on a feature involving login (UI + logic) → produces TWO cards (`UI-LOGIN-*` ui-expert + `LOGIC-LOGIN-*` coder) wired via `depends_on:` rather than one mixed coder card; (5) PRD validation phase on a card with `owner_agent: ""` → enum BLOCKER, commit refused.

## [3.15.0] - 2026-05-23

Two changes ship together. **(1) PRD becomes mockup-aware** — formal intake of existing mockups between Kickoff and Discovery, with a per-screen decision tree in Step 3 that skips generation for screens already covered. **(2) Framework becomes stack-aware** — four new `stack.*` cardinal scalars (`database`, `auth_provider`, `framework`, `deployment`) drive vocabulary, gate sections, anti-pattern checklists, and deploy commands across `/prd`, `coder`, `api-perf-cost-auditor`, `/new`, `code-reviewer`, `security-reviewer`, `bug`. Project-identity assumptions previously hard-coded (mayo personas, real test credentials, Italian merchant business names) are removed from the framework and now resolved from `identity.audience_segments[]` + `.baldart/overlays/prd.md`. Sanitization sweep removed all PII (phone numbers, real usernames, real passwords, real store names) from the published payload.

### Added — Mockup-aware PRD

- **Step 1.6 — Mockup Intake** in [framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md). New mandatory step between Kickoff (Step 1) and Discovery (Step 2). Asks the user "Hai dei mockup a disposizione?" — if yes, asks for format (chat images / local paths / mixed), STOPs, and on receipt: copies local files into `${paths.prd_dir}/<slug>/mockups/` (with collision-safe rename), analyzes every mockup against the **Mockup analysis schema** (screens, user_flow, components, states_visible, copy_excerpts, gaps, design_system_alignment with violations when `features.has_design_system: true`), and populates `## UI Design` in the state file before entering Discovery. If "no", marks `mockups.status: none` and proceeds with the standard flow — fully backwards-compatible for projects without mockups.
- **Step 3 decision tree (Full / Hybrid / Skip)** in [framework/.claude/skills/prd/references/ui-design-phase.md](framework/.claude/skills/prd/references/ui-design-phase.md). Replaces the legacy "skip if UI N/A" precondition. The tree reads `mockups.status` + `screens_in_scope[]` from the state file and routes per-screen: covered-by-mockups screens skip 3a (options) and 3b (generation) — only Component Registry Lookup, approval gate (3c), and inventory (3d) run; uncovered screens invoke the `ui-design` subskill scoped to that single screen, passing the existing mockups as "design language anchor" so the new screen stays consistent. Includes a canonical 5-screen example (3 covered + 2 new) showing exactly which subskill invocations happen and how the final inventory merges.
- **"Design Reference & Mockup Inventory" section in the PRD template** ([framework/.claude/skills/prd/assets/prd-template.md](framework/.claude/skills/prd/assets/prd-template.md)). New section inserted between Documentation Impact and Section 1 (Problem Statement). Tracks `mockups.status`, `step_3_mode`, the screen inventory table (mapping each screen to its user story + mockup path + origin), the component-mapping table when a design system is in play (reused + new components + token violations), and the source-files traceability list (original user paths → canonical `mockups/` paths, or `chat://image-N` for inline images).
- **Mockup-aware fields in the state template** ([framework/.claude/skills/prd/assets/state-template.md](framework/.claude/skills/prd/assets/state-template.md)). Extends `## UI Design` with `mockups.{status,format,original_paths,canonical_paths}`, `mockup_analysis.{screens,user_flow,design_system_alignment}`, `step_3_mode`, `design.html_path`, `screens_in_scope`, `ui_inventory`. Single source of truth for everything Step 1.6 produces and Step 3 + PRD Writing consume.
- **PRD writing acknowledges mockup canonicality** in [framework/.claude/skills/prd/references/prd-writing-phase.md](framework/.claude/skills/prd/references/prd-writing-phase.md). The canonical sources resolution step now adds the `mockups/` directory as a canonical visual reference when `mockups.status` ∈ {`provided`, `partial`}, and the PRD sections checklist includes a new "Design Reference & Mockup Inventory" entry plus revised UI Specifications guidance to avoid duplication.
- **Hard Rule 16 (Mockup Intake) + Step 1.6 in the flow overview + progress-bar row** in [framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md). Codifies the new step as MANDATORY in the hard-rules section, adds the row `1.6 Mockup intake` to both the in-progress and completed progress-bar templates (with the `Mockup` and `Step 3 mode` summary lines), and adds a "Mockup-driven pre-population" note to the Quick Reference Comprehension Dimensions explaining how dimension 5 (UI impact) and partially dimension 2 (User journey) are pre-resolved from `mockup_analysis` when mockups are provided.
- **New-Screen Check during Discovery** in [framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md). When a user answer reveals a UI screen NOT in `mockup_analysis.screens[]`, the skill no longer treats it as a generic scope expansion (which would offer `/prd-add`). Instead it appends the screen to `screens_in_scope[]` with `covered_by_mockups: false` and logs inline that Step 3 will generate that single screen in Hybrid mode. The Scope Expansion Check still applies for non-UI entities (new endpoints, roles, collections).

### Added — Stack-aware framework

- **Four new `stack.*` scalar keys** in [framework/templates/baldart.config.template.yml](framework/templates/baldart.config.template.yml): `stack.database` (enum: firestore/supabase/postgres/mysql/mongodb/dynamodb/sqlite/none), `stack.auth_provider` (firebase-auth/supabase-auth/clerk/auth0/cognito/nextauth/lucia/custom/none), `stack.framework` (nextjs/remix/sveltekit/astro/nuxt/rails/django/fastapi/express/none), `stack.deployment` (vercel/firebase/aws/gcp/cloudflare/render/fly/self-hosted/none). Empty string preserves the always-ask contract — skills prompt the user when a needed scalar is unset and suggest `npx baldart configure`.
- **`baldart configure` autodetection + prompts for the 4 stack scalars** in [src/commands/configure.js](src/commands/configure.js). Autodetection probes `package.json` deps (`firebase`/`@supabase/*`/`pg`/`mongoose`/`@clerk/*`/`@auth0/*`/`next-auth`/`lucia`/`@aws-sdk/client-cognito-identity-provider`/`next`/`@remix-run/*`/`@sveltejs/kit`/`astro`/`nuxt`/`express`), file presence (`vercel.json`/`firebase.json`/`wrangler.toml`/`fly.toml`/`render.yaml`/`Gemfile`/`manage.py`/`requirements.txt`), and content patterns (Django/FastAPI in requirements, rails gem in Gemfile). Detected values pre-fill the interactive prompt. AUTODETECTED box surfaces all 4 values to the user.
- **`baldart update` schema-drift detector covers scalar `stack.*` keys** in [src/commands/update.js](src/commands/update.js). Extends the existing detector (which already flagged missing `features.*`, `paths.*`, `git.*` keys after an update) to enumerate `tpl.stack[*]` and report only the string-typed top-level scalars as `stack.<key>`. Sub-object keys (`charting`/`animation`/`testing`/`monorepo`/`design_system_signals`) are intentionally ignored — their drift is handled by their dedicated configure prompts. On missing scalars, the detector auto-offers `baldart configure`.
- **Database vocabulary branching in PRD writing** in [framework/.claude/skills/prd/references/prd-writing-phase.md](framework/.claude/skills/prd/references/prd-writing-phase.md). The Schema Verification Gate header now reads "MANDATORY if feature touches the persistence layer" instead of "if feature touches Firestore"; vocabulary adapts ({entity} = collection|table, {field} = field|column); the schema-registry path resolves via overlay → `stack.database` convention (Firestore `field-registry.json`, SQL `prisma/schema.prisma` or migrations dir, Mongo collection markdown). The PRD-template's Section 5 "Data Model" now ships **5 variants** (Firestore Composite, Relational B-tree+RLS, MongoDB compound, DynamoDB GSI/LSI, None) — the writer keeps only the variant matching `stack.database`.
- **api-perf-gate refactored as universal-core + stack addenda** in [framework/.claude/skills/prd/references/api-perf-gate.md](framework/.claude/skills/prd/references/api-perf-gate.md). Gate 5 (keyword scan): universal keywords + "Stack-specific keyword addenda" table activated only by the matching `stack.database`. Gates 1–3: universal checklist + "Stack-specific addenda" per `stack.database` (firestore listener/index/ID-hotspot rules, postgres/supabase RLS+EXPLAIN+JSONB rules, mongo compound-field-order+embed rules, dynamodb GSI+hot-partition rules) and per `stack.framework` (nextjs Route Handler `revalidate`/`use cache`, remix `headers`+`shouldRevalidate`, sveltekit runtime declaration, astro `prerender`). Reference Data section split into "Database pricing & cost-model snapshots" (4 stacks), "Runtime limits by `stack.deployment`" (Vercel/Cloudflare/AWS/Firebase table), "Framework caching primitives by `stack.framework`" (4 frameworks). User cites only the variant matching their config.
- **coder.md Database Index Invariant** in [framework/.claude/agents/coder.md](framework/.claude/agents/coder.md). Renamed from "Firestore Index Invariant"; ships per-stack instructions (Firestore composite + `firestore.indexes.json`, Postgres/Supabase `EXPLAIN` + migration with covering/GIN/RLS, MongoDB `createIndex` with equality→range→sort field order, DynamoDB GSI/LSI in CDK/Terraform, None/unset → skip with warning).
- **prd-card-writer Field Grounding Rule per-stack registry resolution** in [framework/.claude/agents/prd-card-writer.md](framework/.claude/agents/prd-card-writer.md). The Grep target for field verification is no longer hard-coded to `docs/references/field-registry.json` — it resolves via `.baldart/overlays/prd.md § Schema Registry` first, then by convention matching `stack.database` (Firestore field-registry, Prisma schema, SQL migrations, Mongo collection docs, DynamoDB table definition). Universal identifiers (`id`/`uuid`/PK, `createdAt`/`created_at`, `updatedAt`/`updated_at`) are exempt regardless of stack.
- **api-perf-cost-auditor reads stack scalars at session start** in [framework/.claude/agents/api-perf-cost-auditor.md](framework/.claude/agents/api-perf-cost-auditor.md). The agent's "Project Context" section is no longer hard-coded to "Next.js 16 + Firestore + Vercel Fluid Compute" — it now lists universal hard rules + addenda tables for `stack.framework`, `stack.database`, `stack.deployment`, and cites the resolved values in every audit header so the user knows which variant applied. Perf budgets move out of the agent into `${paths.references_dir}/perf-budgets.md` (consumer-owned); absence is flagged as a `BUDGETS_GAP` warning instead of using framework-baked numbers.
- **security-reviewer multi-stack access-rule coverage** in [framework/.claude/agents/security-reviewer.md](framework/.claude/agents/security-reviewer.md). Rule 6 (cloud/infra risks) lists the access-rule variant per `stack.database`: Firebase security rules, Supabase RLS policies, MongoDB validators + collection access, DynamoDB IAM policies, Postgres GRANT/REVOKE + RLS.
- **`/new` Production Readiness Checklist becomes stack-aware** in [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md). Detection table and auto-executable commands branch on `stack.deployment` + `stack.database`. Firebase commands are no longer the default — they're one branch among many (vercel deploy, supabase db push, CDK/Terraform apply, firebase deploy). When `stack.deployment` is empty, the skill infers from config-file presence and falls back to asking the user — never auto-executes a guess.
- **`/new` end-to-end flow generalized for non-Firestore stacks** in [framework/.claude/skills/new/SKILL.md](framework/.claude/skills/new/SKILL.md). Five additional spots got the same treatment so a Supabase / Postgres / Mongo / DynamoDB project doesn't see Firestore-only instructions: pre-flight `data_fields` warning triggers on `db_indexes` (any dialect) and reads "touches the persistence layer" instead of "touches Firestore"; Codex parallel-conflict analyzer asks about same-{entity} writes in the project's vocabulary; DEEP-review trigger keys off `db_indexes` (legacy `firestore_indexes` still recognized); the Coder Briefing's index section ships per-`stack.database` instructions (Firestore JSON snippet, SQL migration / Prisma `@@index`, Mongo `createIndex`, DynamoDB GSI/LSI in IaC) and explicitly tells the agent which artifact to stage; post-implementation verification of compound queries branches into 4 dialects with their own missing-index severity; Docs Routing Reminder uses `{entity}` vocabulary and points at the stack-matched index artifact; and the post-deploy "DB Index Verification" section (previously Firestore-only with curl to the Firestore REST API) now opens with a Skip rule and 4 dialect variants (Firestore REST API, Postgres `pg_indexes` + `pg_stat_progress_create_index`, Mongo `getIndexes` + `currentOp`, DynamoDB `describe-table` GSI status). The Firestore variant retains its full procedure verbatim — no regression for Firestore consumers.
- **bug/logging-patterns.md framework-agnostic** in [framework/.claude/skills/bug/references/logging-patterns.md](framework/.claude/skills/bug/references/logging-patterns.md). The Next.js 16 logging section is now one of four (Next.js, Remix, SvelteKit, Astro/Nuxt) plus a generic `pino`/`winston` fallback. DB-tracing env var examples document the per-stack equivalent (`DEBUG_FIRESTORE=true`, `SUPABASE_DEBUG=true`, `DEBUG=knex:query`, `DEBUG=mongoose:*`).
- **Persona references generalized to `identity.audience_segments[]`** in [framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md), [framework/.claude/skills/prd/references/prd-writing-phase.md](framework/.claude/skills/prd/references/prd-writing-phase.md), [framework/.claude/skills/prd/references/audit-phase.md](framework/.claude/skills/prd/references/audit-phase.md), [framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml), [framework/.claude/skills/prd/assets/epic-template.yml](framework/.claude/skills/prd/assets/epic-template.yml), [framework/.claude/skills/prd/assets/prd-template.md](framework/.claude/skills/prd/assets/prd-template.md). Hard-coded persona literals (`CUSTOMER`/`MERCHANT`/`MERCHANT_STAFF`/`SUPER_ADMIN`) are gone from the framework — replaced by `{{from identity.audience_segments[]}}` placeholders and explicit instructions to loop over the project-declared segments. AC grouping in PRD writing now reads "one group per persona drawn from `identity.audience_segments[]`, plus cross-cutting groups".
- **`db_indexes` field replaces `firestore_indexes` in card YAML** in [framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml). The card template renames the section; entries gain a `dialect` field matching `stack.database` and a `kind` field (composite/covering/partial/GIN/GSI/LSI/unique/fulltext). The legacy `firestore_indexes` name remains accepted as an alias so v3.14.x cards keep working — the validator in [framework/.claude/skills/prd/references/validation-phase.md](framework/.claude/skills/prd/references/validation-phase.md) recognizes both. Metric `has_firestore_indexes_pct` in [framework/.claude/skills/prd/SKILL.md](framework/.claude/skills/prd/SKILL.md) renamed to `has_db_indexes_pct`.
- **PROJECT-CONFIGURATION.md § 4.4 documents the 4 new scalars** in [framework/docs/PROJECT-CONFIGURATION.md](framework/docs/PROJECT-CONFIGURATION.md). New "Stack-cardinal scalars" subsection enumerates the 8 skills/agents that branch on them, the always-ask contract on empty values, and the schema-drift detector behavior.

### Removed — PII sanitization

- **Test credentials (phone, username, password, store name) removed from the published payload.** [framework/.claude/skills/prd/references/discovery-phase.md](framework/.claude/skills/prd/references/discovery-phase.md), [framework/.claude/skills/prd/assets/prd-template.md](framework/.claude/skills/prd/assets/prd-template.md), [framework/.claude/skills/prd/assets/card-template.yml](framework/.claude/skills/prd/assets/card-template.yml) no longer ship any real credential as "example". The discovery Test Strategy step now (a) consults `.baldart/overlays/prd.md § Test Credentials` first, (b) asks the user one structured question per persona if the overlay is silent, (c) stores a credential REFERENCE (secret-manager path, `.env.test:VAR_NAME`) in the state file instead of the literal secret, and (d) explicitly forbids the framework from carrying real PII.

### Notes

- **One CLI change — `baldart configure`.** v3.15.0 adds autodetection + prompts for the 4 `stack.*` scalars and extends the `baldart update` schema-drift detector to flag missing scalar `stack.*` keys. Consumers upgrading from v3.14.x get the new prompts on the next `baldart configure` invocation; the detector triggers automatically on the next `baldart update`.
- **The schema-change propagation rule applied end-to-end** for the 4 new scalars: template ✓, configure prompt ✓, update detector ✓, doctor not extended (the always-ask contract in skills already covers the empty case), CHANGELOG ✓.
- **Backwards-compatibility for v3.14.x cards.** Cards with the legacy `firestore_indexes` field still validate (the writer-side renamed it to `db_indexes`, but the validator and downstream agents accept both names). Cards with `type: firestore` in `data_sources` are equivalent to `type: db, db_dialect: firestore`.
- **Active PRD sessions started under v3.14.x** do not have `## UI Design` populated with the mockup fields; the skill treats `pending`/missing as `mockups.status: none` and falls into the Full Step-3 branch (legacy behavior). New sessions on v3.15.0+ pick up Step 1.6 automatically.
- **Step 3 Hybrid mode preserves the design-system cascade.** When `features.has_design_system: true`, the BLOCKING reads of `${paths.design_system}/INDEX.md` + `tokens-reference.md` run BOTH in Step 1.6 (during mockup analysis, to flag violations early) AND in Step 3 (during per-screen Component Registry Lookup, to validate covered + new screens uniformly). Violations bubble up to the unified 3c approval gate.
- **Sanitization is final, not a heuristic.** Grep across `framework/` for `3486417303`, `antonio.baldassarre2336`, `bRkS2bzqyesK`, `Bar Roma` returns zero hits after this release. Consumers who already had `.baldart/overlays/prd.md` with their real credentials are unaffected — the overlay never leaves the consumer repo.
- **Verification is manual** (no test suite — `npm test` is no-op stub). Scenarios to dogfood after install: (1) `npx baldart update` on a v3.14.x consumer → schema-drift detector flags 4 new `stack.*` keys → `baldart configure` autodetection pre-fills them; (2) `/prd` on a Supabase project (`stack.database: supabase`) → Schema Verification Gate uses Variant B (Relational), api-perf-gate adds postgres+supabase addenda, audit-phase enforces RLS; (3) `/prd` on a Firestore project (`stack.database: firestore`) → behavior identical to v3.14.x (no regression); (4) `/prd` with mockups → mockup intake + Hybrid Step 3; (5) `/new` on a Cloudflare deployment (`stack.deployment: cloudflare`) → no `firebase deploy` commands proposed.

## [3.14.1] - 2026-05-23

Hotfix for the `.framework/` gitignore self-heal shipped in v3.14.0. The detector misread `git check-ignore -v` output: when `check-ignore` matched a negation pattern (`!.framework`), it still reported the match, and `checkFrameworkIgnored` flagged the path as ignored. The result: every consumer who hit the self-heal on first run got `.framework/ is STILL ignored after auto-heal` and `process.exit(1)` — the heal *had* actually worked (the negation was in place), but the verifier didn't recognize its own fix. Reported during the v3.13.0 → v3.14.0 update in a downstream consumer.

### Fixed

- **`GitUtils.checkFrameworkIgnored()` recognizes negation patterns** in `src/utils/git.js`. Per `gitignore(5)`, a leading `!` flips the match semantics: `check-ignore` still prints the pattern (matching ≠ ignored), but the path is RE-INCLUDED, not ignored. The detector now inspects the pattern, returns `{ ignored: false, negatedBy }` on a negation match, and `ensureFrameworkNotIgnored`'s post-heal recheck reports `stillIgnored: false` as expected. Smoke-tested end-to-end: `.gitignore` with `.framework` → heal → recheck reports not-ignored → idempotent second call is a no-op → `git add .framework/foo` succeeds.
- **CHANGELOG consolidation for v3.14.0.** The original 3.14.0 entry had a duplicate `### Added` / `### Notes` block (LSP-only Notes + gitignore-only Notes side by side). Merged into a single Notes section so the release reads as one feature shipped, not two.

### Notes

- **v3.14.0 is functionally broken on first install of the heal.** Consumers who upgrade `npm i -g baldart@latest` get 3.14.1 directly and never observe the regression. Consumers who already updated to 3.14.0 (CLI binary or `npx baldart`) should reinstall: `npm i -g baldart@latest`, then re-run `baldart` — the heal converges on the second pass because the negation block already in `.gitignore` short-circuits the detector cleanly.
- **No framework payload changes.** Single-file fix in `src/utils/git.js` plus CHANGELOG.

## [3.14.0] - 2026-05-23

BALDART closes the **LSP install completeness gap** AND adds a long-overdue **`.framework/` gitignore self-heal** so install / update / push never trip over a stale or pre-existing ignore rule.

### LSP install completeness

Since v3.10.0, opting into `features.has_lsp_layer: true` installed the language-server binary (typescript-language-server, pyright, gopls, rust-analyzer, ruby-lsp) into the consumer repo — but the matching native Claude Code plugin (`<lang>-lsp@claude-plugins-official`) was left for the user to install manually via `/plugin install`. Almost nobody did, so Claude itself never invoked the LSP and code-exploration silently fell back to Grep. The "opt-in to LSP" UX was effectively broken: users thought they had it, agents behaved as if they didn't. v3.14.0 wires the two layers together — when the user opts into LSP, BALDART now installs the binary AND enables the native plugin in the same `configure` step, idempotently and gated on `claude` being in `tools.enabled`.

### `.framework/` `.gitignore` self-heal

The other historical sharp edge in BALDART: consumers who had `.framework` (or a parent pattern like `framework/` or `*framework*`) in `.gitignore` saw every `baldart push` and most `baldart update` autofix flows blow up with `paths are ignored by one of your .gitignore files`, leaving half-finished `chore:` commits on HEAD and requiring a manual `git reset --hard` to recover. The path is a git subtree managed by BALDART — it MUST be tracked — and the failure mode was both confusing (the user never added `.framework` to ignore on purpose) and unfixable from inside the broken flow. v3.14.0 detects the collision and rewrites `.gitignore` idempotently at the top of every install / update / push, so a single `baldart` invocation always converges to a working state.

### Added

- **`GitUtils.checkFrameworkIgnored()` + `GitUtils.ensureFrameworkNotIgnored()`** in `src/utils/git.js` — central, idempotent self-heal for the `.framework/` ↔ `.gitignore` collision. `checkFrameworkIgnored` runs `git check-ignore -v --no-index -- .framework` and returns `{ ignored, source, line, pattern, raw }` (degrade-open: simple-git throws on non-zero exit, treated as "not ignored"). `ensureFrameworkNotIgnored` strips standalone `.framework` / `.framework/` / `/.framework` rules from the project `.gitignore`, appends a single negation block (`!.framework` + `!.framework/**`) with a marker comment so re-runs are no-ops, then re-runs the check and surfaces `stillIgnored` if a rule still wins from outside the repo (parent-dir `.gitignore`, exotic `core.excludesFile`). Returns `{ wasFixed, stillIgnored, removedLines, appendedNegation, source, pattern, recheck }`. Never throws — callers can safely wrap it in best-effort try/catch.
- **`baldart add` / `baldart update` / `baldart push` auto-heal `.gitignore`** — each command calls `ensureFrameworkNotIgnored()` at the top of its flow, before `git subtree add` / `git subtree pull` / any `git add` runs. Healing prints a yellow warning so the user sees what changed, then continues without prompting (purely local file rewrite, fully idempotent). If a rule outside the project still wins, the flow exits with a clear, actionable message instead of failing deep in the autofix step with the cryptic `paths are ignored` git error.
- **`baldart doctor` `.gitignore` diagnostic + `unignore-framework` action** — `detectState` populates `state.frameworkIgnored` with the same payload; the status table renders a yellow `.gitignore IGNORES .framework/ (<source>:<line>) — will be auto-healed` row when triggered, and the planner inserts an `autoOk: true` action that calls the same helper. So `baldart` (no-arg doctor mode) detects, reports, and self-heals in a single invocation. Backfills the heal for consumers predating v3.14.0 that hit the broken-push state.
- **`claudePluginId()` returns real plugin ids** in every LSP adapter (`src/utils/lsp-adapters/{typescript,python,go,rust,ruby}.js`). Previously a `null` placeholder; now resolves to `typescript-lsp@claude-plugins-official`, `pyright-lsp@claude-plugins-official`, `gopls-lsp@claude-plugins-official`, `rust-analyzer-lsp@claude-plugins-official`, `ruby-lsp@claude-plugins-official` respectively. All plugins live in the Anthropic-published [`anthropics/claude-plugins-official`](https://github.com/anthropics/claude-plugins-official) marketplace.
- **`LspInstaller.installClaudePlugins(serverNames, opts)`** in `src/utils/lsp-installer.js` — high-level installer that (1) verifies the `claude` CLI is on PATH, (2) idempotently registers the `claude-plugins-official` marketplace via `claude plugin marketplace add anthropics/claude-plugins-official --scope user` (no-op when already registered — list is parsed first), (3) reads `claude plugin list --json` to skip plugins already installed at any scope, (4) runs `claude plugin install <id> --scope user` per missing plugin. Defaults to `--scope user` so a single install is reused across every project on the machine (LSP plugins are machine-wide useful — project scope would create N copies). Always returns `{ installed, skipped, manual, marketplace }`; never throws.
- **`LspInstaller.ensureClaudeMarketplace()`, `isClaudeCliAvailable()`, `listInstalledClaudePlugins()`, `verifyClaudePlugins(serverNames)`** — supporting primitives, all idempotent and degrade-open (parsing failure returns empty set → next install is attempted rather than incorrectly skipped).
- **`baldart configure` companion-plugin step** — after installing the language servers, the configure flow now calls `installClaudePlugins(merged.lsp.installed_servers)`. Surfaces clear messaging for each branch: success (✓ Installed plugin ids), idempotent skip ("Already installed"), Claude CLI missing (manual install commands printed verbatim per plugin), or `claude` opted out of `tools.enabled` (silent skip — Codex-only users don't want Claude marketplace mutation).
- **`baldart doctor` companion-plugin diagnostic** — `detectState` now calls `verifyClaudePlugins` when LSP is enabled and the `claude` CLI is available, populating `state.lspMissingClaudePlugins`. The `LSP layer` row in the diagnostic table reports `N server(s) verified, M Claude plugin(s) missing (…)` when the layer is half-installed. Severity `warn`.
- **`lsp-install-claude-plugins` action** in the doctor planner — `autoOk: true` (user-scope plugin install is idempotent and fully reversible via `claude plugin uninstall`). Backfills missing plugins for installs predating v3.14.0 and recovers from manual plugin uninstalls. Pairs with other actions; never short-circuits.

### Changed

- **LSP opt-in is now end-to-end.** Pre-v3.14.0: `features.has_lsp_layer: true` installed the binary but the integration was dormant until the user manually enabled the plugin. Post-v3.14.0: opting in installs both layers in the same step. The framework's "always-ask, never-assume" contract still applies — when `claude` is not in `tools.enabled` OR the `claude` CLI is missing, the install gracefully degrades (no error, clear hint, deferred until the prerequisites are in place). `baldart doctor` then surfaces the half-install state on every subsequent invocation.

### Notes

- **No new `baldart.config.yml` keys.** The plugin install reuses `features.has_lsp_layer`, `lsp.installed_servers`, and `tools.enabled` already in place — so the [schema-change propagation rule](../.claude/projects/-Users-antoniobaldassarre-BALDART/memory/feedback_schema_change_propagation.md) does not require a new template / configure prompt / update detector for this release. Scope is hard-coded to `user` (machine-wide LSP plugins) rather than exposed as `lsp.claude_plugin_scope` — the decision is opinionated and consistent with how the official marketplace tracks its own installs. If a future release needs per-project plugin scope, *that* would trigger the propagation rule (template + prompt + update detector + this changelog).
- **Marketplace and plugins live at `--scope user`, not project.** Three reasons: (1) LSP plugins are machine-wide useful — typescript-lsp installed once at user scope is reused by every TS project on the machine, not duplicated N times; (2) marketplace install at user scope means subsequent `baldart configure` runs in sibling projects skip the marketplace add entirely (idempotent); (3) project-scope plugins live under `.claude/` in the consumer repo, which would either need committing (noisy diffs) or `.gitignore` adjustments (extra protocol burden). User scope avoids all three.
- **Idempotency is checked twice.** Marketplace add reads `claude plugin marketplace list --json` first; plugin install reads `claude plugin list --json` first. Re-running `baldart configure` (or `baldart doctor` with the action accepted) is a no-op when everything is in place — no spurious install commands, no error noise.
- **Gate on `claude` in `tools.enabled`.** When the consumer has `tools.enabled: [codex]` (Codex-only), the plugin install is a silent no-op — Codex has no equivalent plugin model and Claude marketplace mutation would be off-target. Doctor diagnostics for missing plugins are similarly gated.
- **No framework payload changes.** All work is inside `src/` — adapters, installer, configure command, doctor command, and `GitUtils`. No framework agent/skill/command/protocol module touched, so existing consumers see the new behavior the next time they run `baldart configure` or `baldart doctor` after upgrading the CLI (`npm i -g baldart@latest`).
- **`.gitignore` heal is idempotent across re-runs.** The negation marker is checked before appending; the strip step is a `filter` over current lines. Running `add` / `update` / `push` repeatedly on a healthy repo is a no-op (the negation block is added once, future calls match the marker and skip).
- **`.gitignore` heal has no false positives.** The strip filter matches only the exact forms (`.framework`, `.framework/`, `/.framework`, `/.framework/`, `.framework/*`, `/.framework/*`) — `.frameworks/` or `framework/something.bak` patterns are preserved untouched.

## [3.13.0] - 2026-05-23

BALDART closes **two drift visibility gaps in one release**. (1) The CLI binary (installed globally via `npm i -g baldart`) and the framework payload (`.framework/` inside each consumer repo, updated by `baldart update`) drifted silently — users on an old CLI never learned that new flags, new doctor checks, or new adapters had shipped. (2) The framework payload on disk and the `.baldart/state.json` ledger could drift silently when the ledger write failed (the catch in `update.js` was silent until now), when an update flow exited mid-way, or when a manual commit excluded `.baldart/state.json` — `baldart doctor` had no way to detect or fix this. Both gaps are now closed: a **non-invasive CLI self-update notifier** for the first, and **ledger drift detection + self-heal** in the doctor for the second. The unifying theme is "no silent drift" — every layer that can fall out of sync now reports it and offers a one-step fix.

### Added

- **`src/utils/update-notifier.js`** — best-effort CLI self-update check. HTTPS GET to `https://registry.npmjs.org/baldart/latest` with a 1.5s timeout, cached at `~/.baldart/cli-update-check.json` (TTL 24h). Banner is printed from the cache on subsequent invocations (the very first run shows nothing — that is intentional; the check never blocks the command). The background refresh fire-and-forget: its timer is `unref()`'d and the underlying socket dies with the process, so the check can never hold the event loop open. Suppression: `BALDART_NO_UPDATE_CHECK=1`, `CI=true`, `NODE_ENV=test`, non-TTY stdout, or `--offline` on `baldart` / `baldart doctor` / `baldart version`.
- **Top-of-run banner in `bin/baldart.js`** — when a newer baldart npm package has been detected on a previous run, every command prints a 3-line hint at startup (yellow arrow + bold version, then two gray lines): `↑ baldart X.Y.Z available (you have A.B.C)` + `Update with: npm i -g baldart@latest` + `Suppress with: BALDART_NO_UPDATE_CHECK=1`. Zero-latency (reads cache only); no banner if cache is missing or up-to-date.
- **`CLI` row in `baldart doctor` diagnostic table** — always shown (independent of git/framework state), surfaces drift inline alongside framework version / config / overlays / hooks so the doctor remains the single pane of glass. Severity `warn` when an update is available, `ok` otherwise.
- **CLI drift in `baldart version` box** — the existing `CLI: vX.Y.Z` line now appends `→ vA.B.C available (npm i -g baldart@latest)` when a newer version was detected.
- **`State.reconcileInstalledVersion({ to, reason })` in `src/utils/state.js`** — new ledger reconciliation helper. Idempotent (no-op when `from === to`), audit-trailed (writes a `ledger-reconcile` history entry distinct from real `update` events), and reversible (pure local file write). Used by the doctor's new `reconcile-state-ledger` action and available for direct call from future recovery flows.
- **Ledger drift detection in `baldart doctor`** — `detectState` now compares `.framework/VERSION` against `.baldart/state.json:installed_version` and surfaces a `stateLedgerDrift` state field when they diverge. The action planner inserts a `reconcile-state-ledger` action (`autoOk: true`, idempotent, no remote effects) that one-shot fixes the drift via `State.reconcileInstalledVersion()`. Drift is non-blocking — it pairs with other proposed actions instead of short-circuiting them.

### Changed

- **`src/commands/update.js`** — replaced the silent `try { State.recordUpdate(...) } catch (_) {}` on the ledger write with a visible warning + actionable next step. When the write fails or returns an unexpected `installed_version`, the user now sees a yellow warning explaining the drift (framework on disk is at vNEW but ledger still records vOLD) and is told to run `baldart doctor` to reconcile. This is the root cause of most "my doctor says I'm on the wrong version" reports — silent catches were hiding write failures (permissions, EBUSY) for weeks. New behavior makes the failure immediately visible and self-healable.

### Notes

- **CLI vs framework — two independent layers.** `npm i -g baldart@latest` updates the binary on disk (this package); `baldart update` updates the *framework payload* inside the consumer repo (the `.framework/` subtree). They are *independent* by design — a consumer may want to update the framework without touching the global CLI (or vice versa). The notifier surfaces drift on the *CLI* layer; the existing `Remote` row in doctor (and the `Remote:` line in `version`) surfaces drift on the *framework* layer. Both layers are now visible at a glance from the same diagnostic surface.
- **Why not auto-install.** Three reasons we deliberately reject `npx baldart` auto-running `npm i -g`: (1) blast radius — global install impacts *every* project on the user's machine, not just the one they invoked from; (2) permissions — `npm i -g` often needs sudo (without nvm/volta), and a no-args `baldart` cannot trigger sudo prompts silently; (3) reproducibility — `--auto` runs in CI must be deterministic, so auto-updating the CLI mid-run would change behavior unpredictably. The notifier is the right level: visible, dismissible, never invasive.
- **No new `baldart.config.yml` keys.** The notifier cache lives in `~/.baldart/cli-update-check.json` (per-user, not per-project), so the schema-change propagation rule (per `feedback_schema_change_propagation`) does not apply to this release. Suppression is via environment variable (`BALDART_NO_UPDATE_CHECK=1`) which is intentionally global — disabling the notifier in one project but enabling it in another is not a use case worth modeling.
- **Backwards-compatible.** Pre-3.13 consumers see the new banner on their next `baldart` run after the npm package upgrades. Consumers that never want the check set `BALDART_NO_UPDATE_CHECK=1` in their shell rc; CI runs are auto-suppressed.
- **No framework payload changes.** All work is inside `src/` and `bin/` — the framework agents/skills/commands/protocols are unchanged. This is purely a CLI-layer improvement; no consumer files are touched by `baldart update`.

## [3.12.0] - 2026-05-23

BALDART completes the **UI excellence loop**: the `ui-expert` agent is upgraded from a "good but generic" baseline to a world-class UI/UX reviewer/designer, and — crucially — the registry-first discipline gains a **post-intervention coherence gate** so visual coherence is enforced not only when work begins (BLOCKING reads) but also when work ends (BLOCKING completion check). The asymmetry of v3.11.0 (registry read before, no verification after) is closed: every UI change introduced by `ui-expert`, the `ui-design` skill, the `frontend-design` skill, or caught by `code-reviewer` must now reconcile the registry (INDEX + per-component specs + tokens-reference) in the **same change**, never deferring drift to the weekly `ds-drift` routine. This release contains zero schema changes — the cascade reuses `features.has_design_system` + `paths.*` already in place.

### Added

- **`framework/agents/design-system-protocol.md` § "Reference Tables (embed)"** — canonical numeric ranges that every agent/skill must cite verbatim: type scale (6 modular ratios + line-height pairing rule), color contrast targets (WCAG 2.x + APCA Lc as secondary perceived-legibility gate), spacing scale (base-4 vs base-8 with use-case mapping), density tiers (compact/cozy/comfortable/spacious with row-height + padding numbers). Replaces the previous "agents pick their own numbers" failure mode.
- **`framework/agents/design-system-protocol.md` § "Token Cascade: primitive → semantic → component"** — three-layer token discipline (raw → purpose-named → optional component alias) that makes components theme-agnostic and multi-brand-ready by construction. Documents OKLCH > HSL for palette generation and forbids `filter: invert()`-style dark mode shortcuts.
- **`framework/agents/design-system-protocol.md` § "Post-Intervention Coherence Check"** — new BLOCKING completion gate (4 verifications + decision matrix + reporting format + relationship with the weekly `ds-drift` routine as safety net). This is the SSOT consumed by `ui-expert`, `ui-design`, and `code-reviewer`.
- **`framework/.claude/skills/motion-design/reference/accessibility-and-modern-apis.md`** — new reference file covering `prefers-reduced-motion` collapse semantics (0.01ms, not 0 — Safari bug), the View Transitions API (same-doc + shared-element + cross-doc), spring vs cubic-bezier decision matrix, choreography budget (≤ 600ms total), and animations that communicate state and must NOT be collapsed under reduced-motion.
- **`framework/.claude/skills/frontend-design/SKILL.md` § "Performance Gates (Core Web Vitals 2026)"** — CWV thresholds table (LCP / INP / CLS / TTFB / FCP) with p75 good/poor cuts, canonical LCP image pattern (preload + fetchpriority + AVIF/WebP + explicit dimensions), INP fixes (`scheduler.yield`, debounce-to-frame, off-main-thread parsing), CLS budget (dimensions on every replaced element, font-display strategies, reserved space for late-injected UI).
- **`framework/.claude/skills/frontend-design/SKILL.md` § "Modern CSS (2025–2026)"** — container queries as default for component-level responsiveness, `:has()` for styling-only state reactions (never business logic), View Transitions API, subgrid, logical properties for free i18n, `color-mix()` for derived states, `@scope`.
- **`framework/.claude/skills/frontend-design/SKILL.md` § "Form quality (always-on requirements)"** — non-optional form contract (label association, `autocomplete=`, validation on blur, actionable error messages, `aria-invalid` + `aria-describedby`, `inputMode=`, paste normalization, submit-disabled-only-after-failed-attempt).
- **`framework/.claude/skills/ui-design/SKILL.md` Step H — Registry Coherence Reconciliation** — new BLOCKING completion gate inside the `ui-design` workflow that walks the UI Element Inventory from Step G and reconciles each item with INDEX + per-component spec + tokens-reference before sign-off. Catches drift at design time when it's cheapest to prevent.
- **`framework/docs/UPGRADE-3.12-UI-COHERENCE.md`** — onboarding doc consumable by Claude in any consumer repo. Step-by-step alignment of an existing project's UI organization with the new discipline: registry presence check, overlay drift audit, token/primitive hardcoding scan, one-shot `ds-drift` baseline, prioritized reconciliation backlog.

### Changed

- **`framework/.claude/agents/ui-expert.md` (massive upgrade — from ~250 to ~500 lines)** — § "Your Expertise" expanded to cite WCAG 2.2 SCs (2.4.11/2.4.12 focus, 2.5.5 target size, 1.3.5 input purpose) + APCA, OKLCH-based palette generation, modern CSS surface (container queries, `:has()`, View Transitions, subgrid, logical properties), motion intent-driven curve selection, Core Web Vitals literacy, AI-era patterns, inclusive design beyond AA, i18n with text-expansion buffers. **New § "UI States Taxonomy"** (8 states explicitly: idle / loading / empty / partial / success / error / offline / optimistic) with loading rules (sub-1s → no indicator, 1–10s → skeleton, > 10s → progress with cancel) and skeleton-must-match-real-dimensions discipline. **New § "Token Cascade Discipline"** (semantic-only consumption, primitive → semantic → component, OKLCH preferred). **New § "Composition over Configuration"** (Radix Slot / `asChild`, controlled+uncontrolled contract). **New § "AI-Era UI Patterns"** (streaming-without-CLS, always-visible stop/regenerate, ARIA live regions, confidence affordance, hallucination guardrails). **New § "Performance Gates"** with UI-level red flags. **Red flag list** expanded from 14 generic items to **60+ categorized findings** (registry & tokens / minimalism / architecture / mobile / a11y / state design / forms / motion / data viz / modern CSS / i18n). **Workflow steps** gain step 10 (Review) and step 12 (Design) — both BLOCKING completion gates that run the Post-Intervention Coherence Check. **Linked Skills** section enabled (was commented out) with explicit "MUST use" routing to `frontend-design`, `motion-design`, `ui-design`, `ui-ux-pro-max`, `webapp-testing` / `playwright-skill`.
- **`framework/.claude/agents/code-reviewer.md` § "Design System Compliance"** — gains rule 8 (Post-Intervention Coherence Check as BLOCKING merge gate). If the check trips at code-review time, it means the upstream `ui-design`/`ui-expert` checks were skipped — the violation is flagged alongside the drift. Code-reviewer is the third and final gate before the weekly safety net.
- **`framework/.claude/skills/ui-design/SKILL.md` § "Design Principles (quick reference)"** — gains a BLOCKING bullet point pointing at Step H so the registry-coherence rule is visible even to readers who skim only the quick reference.

### Notes

- **No new `baldart.config.yml` keys.** The post-intervention gate, the upgraded `ui-expert`, the new motion / frontend-design sections, and the `code-reviewer` rule 8 all gate on the existing `features.has_design_system` and read paths already declared in `paths.*`. The schema-change propagation rule is satisfied — this release only adds behaviors, expands an agent, and adds reference content.
- **Backwards-compatible.** Projects with `has_design_system: false` see no behavioral change for the registry-coherence gate (it short-circuits when the flag is false). Projects with `has_design_system: true` get the gate active immediately on next agent/skill invocation — if their codebase already has drift (silent primitives, missing specs, hardcoded tokens), the agent will surface it as findings on the first UI task post-update rather than fixing it silently. Use the `UPGRADE-3.12-UI-COHERENCE.md` doc to baseline and prioritize the reconciliation.
- **Three-layer gating design.** Drift detection is now redundant by design: (1) `ui-design`/`ui-expert`/`frontend-design` post-intervention check (per-task) → (2) `code-reviewer` rule 8 (per-merge) → (3) `ds-drift` routine (weekly safety net). The failure of any one layer is caught by the next. The per-task check is the cheapest reconciliation point — the weekly routine should now find very little to report when the upstream gates run correctly.
- **SSOT discipline preserved.** All numeric reference tables (type scale, contrast, spacing, density, CWV thresholds, motion durations / easings, spring parameters) live in exactly one place. The agent is explicitly instructed to cite from there verbatim and not invent numbers. The four post-intervention conditions are defined once in `design-system-protocol.md` and referenced from `ui-expert`, `ui-design`, and `code-reviewer` rather than restated.
- **Migration for consumers.** No file the user owns is touched. `baldart update` symlinks the upgraded framework payload; `.baldart/overlays/ui-expert.md` (if present) keeps overriding sections by name as before — but consumers with overlays should run the `UPGRADE-3.12-UI-COHERENCE.md` walkthrough to spot overlay sections that may now conflict with the expanded base (e.g. overlays that previously codified red flags or workflow steps the base now defines authoritatively).

## [3.11.0] - 2026-05-23

BALDART makes the **component registry the SSOT for visual coherence**. Until now, the registry-first discipline (`${paths.design_system}/INDEX.md` + `tokens-reference.md` + `components/<Name>.md`) was BLOCKING only inside the `ui-design` and `frontend-design` skills, while the `ui-expert` agent — the most-used entry point for ad-hoc UI edits — and the `/design-review` command had only a generic "follow the style guide" prompt. This release closes the gap end-to-end (agent + skill + command + onboarding + bootstrap) so every UI change passes through the same inventory and token contract, regardless of which entry point routes the work. A new `/design-system-init` skill scaffolds a registry from the codebase's current state, so projects without one can adopt the discipline with a single command.

### Added

- **`framework/agents/design-system-protocol.md`** — new protocol module that is the textual SSOT of the registry-first cascade (Component Discovery, token contract, Authority Matrix, "primitive missing" decision tree). Referenced from `framework/agents/index.md`. All UI-touching agents/skills/commands cite it instead of duplicating the BLOCKING-reads block. Same pattern as `code-search-protocol.md`.
- **`framework/.claude/skills/design-system-init/`** — new invokable skill (`/design-system-init`). Inventories primitives via `codebase-architect`, extracts tokens from `${paths.global_styles}` + Tailwind config, scaffolds `INDEX.md` + `tokens-reference.md` + a `components/<Name>.md` per primitive (template under `scripts/component-spec.template.md`), and flips `features.has_design_system: true` + `paths.design_system` in `baldart.config.yml`. Idempotent — refuses to overwrite an existing registry, supports incremental re-runs.
- **`scripts/component-spec.template.md`** — dense per-component spec template (props, variants, tokens consumed, accessibility, anti-patterns, MUST rules, changelog) used by `design-system-init` and by hand when extending the registry.
- **Design-system bootstrap hint in `configure`** — when the user opts in to `has_design_system: true` but autodetect found no path on disk, configure now surfaces `/design-system-init` as the next step so the flag doesn't end up pointing at nothing. When the user opts out, a soft pointer to `/design-system-init` is shown for future adoption.

### Changed

- **`framework/.claude/agents/ui-expert.md`** — now follows the registry-first protocol. Gains a Project Context header citing `baldart.config.yml` keys, replaces the lone `ui-guidelines.md` BLOCKING-read with the full cascade (`INDEX.md` + `tokens-reference.md` + per-component specs) gated by `has_design_system`, prepends Component Discovery (registry → source-tree → catalog → create) to both `When Reviewing UI Code` and `When Designing New Interfaces`, and adds red flags for re-implementing existing primitives and hardcoding values that bypass the token contract.
- **`framework/.claude/commands/design-review.md` + `framework/agents/design-review.md`** — both gain a step 0 BLOCKING when `has_design_system: true` (read `INDEX.md` + `tokens-reference.md` + `components/<Name>.md` for primitives on the route under review). The checklist gains a new top section **Design-System Conformance** (primitives reused, primitives reinvented HIGH, tokens bypassed HIGH, spec drift MEDIUM, new primitives proposed). The agent module ships a Report Template for that section; the command file points the subagent at it.
- **`framework/.claude/agents/code-reviewer.md`** — the existing § "Design System Compliance (MANDATORY for UI work)" now cites `design-system-protocol.md` as the SSOT, makes the BLOCKING reads conditional on `features.has_design_system: true` (instead of "if the project has a design system"), and adds two explicit HIGH findings: re-implementation of an existing registry primitive, and new primitives introduced without their per-component spec.
- **`framework/.claude/skills/ui-design/SKILL.md` + `framework/.claude/skills/frontend-design/SKILL.md`** — both gain a one-paragraph header on the Prerequisites section declaring `framework/agents/design-system-protocol.md` as the SSOT when local rules drift. Existing BLOCKING-read content kept verbatim (skills remain readable stand-alone).
- **`framework/agents/index.md`** — new routing entry: "if you touch ANY visual surface and `features.has_design_system: true` → read `agents/design-system-protocol.md`". `design-system-protocol.md` added to the Modules list.
- **`src/commands/configure.js`** — after the `has_design_system` prompt, surfaces `/design-system-init` either as the recommended next step (when the user opts in but no path is detected) or as a soft pointer (when the user opts out).

### Notes

- **No new `baldart.config.yml` keys.** The cascade gates on the existing `features.has_design_system` and reads paths already declared in `paths.*`. The schema-change propagation rule (per `feedback_schema_change_propagation`) is therefore satisfied without template/configure/update detector churn — this release only adds behaviors, a skill, and a protocol module.
- **Backwards-compatible.** Projects with `has_design_system: false` see no behavioral change (the BLOCKING reads in the agent/skill/command are no-ops when the flag is false). Projects with `has_design_system: true` already had the discipline inside `ui-design`/`frontend-design`/`code-reviewer` — this release just propagates it to the remaining surfaces (`ui-expert`, `/design-review`).
- **Bootstrap path.** Projects that want the discipline but lack a registry now have a single command (`/design-system-init`) that inventories existing primitives, scaffolds the registry from them, and flips the flag — eliminating the chicken-and-egg gap where the BLOCKING reads pointed at files that didn't exist.
- **SSOT trade-off.** `design-system-protocol.md` is the textual SSOT, but the BLOCKING-reads block is still inlined verbatim in `ui-design/SKILL.md`, `frontend-design/SKILL.md`, `code-reviewer.md`, and `ui-expert.md` (each file gains an SSOT-pointer header instead of being structurally deduplicated). Skills/agents stay readable stand-alone and don't depend on runtime module loading; the cost is that future protocol changes need to touch the consumers as well. Acceptable for v3.11.0 — revisit if the cascade churns frequently.

## [3.10.0] - 2026-05-23

BALDART grows an **LSP symbol-search layer**. When enabled, `codebase-architect` and the code-exploration skills (`context-primer`, `bug`, `prd`, `new`, `simplify`) prefer LSP `find-references` / `go-to-definition` over Grep for identifier-shaped queries. The filtering happens **before** Claude reads files — a function-name search that previously returned thousands of textual matches now resolves to the actual callsites of the same symbol. Grep remains the fallback for free-text queries and when the LSP layer is unavailable, so existing flows degrade gracefully.

The layer is opt-in: at install / update / configure, BALDART asks "Enable LSP symbol-search layer?" and (on yes) detects the project's languages and installs the matching language servers — npm devDeps for TypeScript and Python, system commands printed for Go, Rust, Ruby.

### Added

- **`baldart.config.yml` → `features.has_lsp_layer`** — new boolean (default `false`). Gates the LSP retrieval tier in agents and skills. When `true`, also populates the new `lsp` section with the installed language servers.
- **`baldart.config.yml` → `lsp.installed_servers` + `lsp.auto_verify`** — new top-level section tracking which language servers are recorded for the project and whether `baldart doctor` should re-verify them on every run.
- **`framework/agents/code-search-protocol.md`** — new protocol module defining the retrieval hierarchy (RAG hybrid → LSP for symbols → Grep → Git), query-type discrimination, LSP budget (max 3 calls/task), and fallback rules. Referenced from `framework/agents/index.md`.
- **`framework/.claude/skills/lsp-bootstrap/SKILL.md`** — new invokable skill (`/lsp-bootstrap`) that runs the install + verify flow as a standalone workflow. Refuses to run when `features.has_lsp_layer: false`. Idempotent.
- **`src/utils/lsp-adapters/`** — new pluggable adapter registry (TypeScript, Python via Pyright, Go via gopls, Rust via rust-analyzer, Ruby via ruby-lsp). Each adapter exposes `detect(cwd)`, `installCommand()`, `verifyCommand()`, and `claudePluginId()`. Same pattern as `src/utils/tool-adapters/`. Add a new language by dropping in a `<lang>.js` file and registering it in `index.js`.
- **`src/utils/lsp-installer.js`** — high-level installer used by `configure.js` and the `doctor` LSP step. `detectLanguages()` / `recommend()` / `installServers()` / `verifyServers()`. Pure detection has no side effects; install ops respect `--non-interactive`.
- **`src/commands/configure.js` LSP block** — when the user opts in, autodetects languages, prompts per server, installs npm-mode adapters in-process, prints system-mode install commands, verifies binaries, and persists `lsp.installed_servers`. Re-running configure on an existing project re-detects and updates the list.
- **`src/commands/doctor.js` LSP step** — new "LSP layer" line in the diagnostic table + a `lsp-fix` action that reinstalls broken servers. Surfaces missing binaries without breaking other doctor flows.
- **`framework/templates/overlays/agents/codebase-architect.lsp-example.md`** — starter overlay demonstrating how to add project-specific LSP tuning (monorepo quirks, generated-code exclusions, callsite thresholds) on top of the shipped protocol.

### Changed

- **`framework/.claude/agents/codebase-architect.md`** — Investigation Protocol step 3 now points at `agents/code-search-protocol.md` for LSP escalation. Behavior under `features.has_lsp_layer: false` is unchanged.
- **`framework/.claude/agents/REGISTRY.md`** — `codebase-architect` row gains "LSP symbol resolution (when `features.has_lsp_layer: true`)" as a specialization.
- **`framework/.claude/skills/context-primer/SKILL.md`** — step 3 of the retrieval loop redirects to the code-search protocol for symbol-shaped queries when the flag is on.
- **`framework/.claude/skills/bug/SKILL.md`** — Phase 0 step 3 references the protocol; `codebase-architect` now automatically uses LSP for symbol lookups during code-path mapping.
- **`framework/.claude/skills/simplify/SKILL.md`** — deduplication step 3 uses LSP `find-references` to confirm callsites before recommending consolidation; eliminates false positives from textual name collisions and dead-code candidates.
- **`framework/.claude/skills/prd/SKILL.md` / `new/SKILL.md`** — one-line note that LSP usage is transitive via `context-primer` / `codebase-architect` when the flag is on.
- **`framework/templates/baldart.config.template.yml`** — adds `features.has_lsp_layer` and the new `lsp:` block with inline documentation.
- **`framework/docs/PROJECT-CONFIGURATION.md`** — new § 4.5 row for `has_lsp_layer` and a new § 4.6 documenting the `lsp.*` keys, per-adapter install modes, and the fallback contract.

### Notes

- **Schema-change propagation** (matches the `v3.9.0` pattern): the new key is wired into the template, the configure prompt, the update-detector, the doctor table, the relevant skills, and the changelog.
- **Backwards-compatible.** Pre-v3.10 consumers running `npx baldart update` see the standard "new config keys" warning (`has_lsp_layer`) and are offered `configure`. Agents and skills behave exactly as before when the flag is absent or `false`.
- **Fallback contract.** When the layer is enabled but a server is missing, broken, or returns no references for a known symbol, agents silently fall back to Grep. Code search never fails because of LSP issues; `baldart doctor` surfaces the mismatch.
- **Languages out of scope** (Java, C/C++, PHP, Elixir, …) can be added by dropping a new file under `src/utils/lsp-adapters/`. No core change required.

## [3.9.0] - 2026-05-22

`worktree-manager` (`/mw`) gains a configurable merge strategy. The previous flow forced every consumer through `gh pr create` + `gh pr merge`; in repos where `develop` is not protected on GitHub, that roundtrip is unnecessary and was observed to stall (PR opened, never merged). The new `git.merge_strategy` key lets each project pick its lane.

### Added

- **`baldart.config.yml` → `git.merge_strategy`** — new top-level config section with one key (default `pr`):
  - `pr` (default, backwards-compatible) — push feature branch, open PR via `gh pr create`, merge via `gh pr merge`. Required if `develop` is protected on origin.
  - `local-push` — after the rebase in `/mw` step 4b, fast-forward push the feature branch directly to `origin/develop` (`git push origin <feat>:develop`). No PR, no `gh` dependency. Fails hard with a clear hint if `develop` is protected on origin.
- **`framework/.claude/skills/worktree-manager/SKILL.md` step 4c** — fully rewritten into two clearly labelled strategies (A: `pr`, B: `local-push`) plus a shared local-develop-sync block. Both strategies remain strictly worktree-isolated — neither runs `git checkout` / `git switch` / `git branch` on the main repo.
- **Programmatic API output** — `/mw` now returns `strategy: "pr"|"local-push"`. `prNumber` is `null` in `local-push` mode.

### Changed

- **`framework/.claude/skills/new/SKILL.md`** — Phase 6 description and fail-safe rules reference the configured strategy instead of hard-coding `gh pr merge`.
- **`worktree-manager` Project Context header** — added at the top of the skill, citing `git.merge_strategy` + `paths.backlog_dir` per the always-ask-never-assume contract in `framework/agents/project-context.md`.
- **`src/commands/configure.js`** — gained autodetection + interactive prompt for `git.merge_strategy`. The probe runs `gh api repos/:o/:r/branches/develop/protection` to recommend `pr` (when develop is protected) or `local-push` (when not). Falls back to `pr` safely when `gh` is missing or the remote is unreachable. The detection reason is surfaced inline so the user knows *why* a strategy is recommended.
- **`src/commands/update.js`** — schema-drift detection now also covers new top-level sections (currently `git.*`). When new keys are detected, `update` proactively offers to run `configure` instead of just warning. This means an existing consumer running `npx baldart update` will be **asked** about the new `git.merge_strategy` key, not silently defaulted.

### Notes

- No migration is required for existing consumers. `git.merge_strategy` defaults to `pr` when the key is absent, preserving the v3.4.0+ behaviour exactly.
- On `npx baldart update`, existing consumers are **prompted** to run `configure`, where the merge-strategy autodetection probes their repo and recommends `pr` or `local-push` with a one-line reason. The user picks; nothing is auto-applied.
- `develop` remains the only supported integration branch; configurable `base_branch` is out of scope for this release.

## [3.8.2] - 2026-05-22

`baldart` (smart doctor) no longer double-confirms safe actions. Previously, running `baldart` with a remote-ahead state would ask *"Run: Pull N commit(s)?"* and then the `update` command would ask *"Proceed with update?"* immediately after — two prompts for the same intent.

### Fixed

- **`src/commands/doctor.js` action runner** — every action now carries an `autoOk` flag. Safe & idempotent actions (`update`, `configure*`, `migrate`, `register-edit-gate`) run without the outer doctor confirmation (their own internal prompts take over). Actions that change external state or need explicit version classification (`add`, `push`) keep the outer confirmation so the user signals intent explicitly.

Result: `baldart` in a repo with "remote ahead" now goes straight into the update flow (with its own pre-flight stash + diff preview + proceed prompt), instead of double-asking.

## [3.8.1] - 2026-05-22

Documentation: promote the "one command for everything" install path in the README and surface the global-install recommendation in CLAUDE.md.

### Changed

- **`README.md` Quick Start** — restructured. Primary path is now `npm i -g baldart` once, then `baldart` (no args) in any project. Added a table that maps every repo state to the action `baldart` proposes (install / migrate / configure / update / push). The `npx baldart …` form remains documented as the no-global-install alternative.
- **`CLAUDE.md`** — added the recommended end-user install (`npm i -g baldart`) to the distribution note.

No behavioural changes; the smart-doctor entry point has worked this way since v3.2.0, this release just makes it discoverable.

## [3.8.0] - 2026-05-22

`.claude/agents/` and `.claude/commands/` move from bulk symlink to **per-item merge** with an **overlay system** (compile-time merge of base + overlay). Closes the long-standing gap where specializing a framework agent forced you to either fork the whole file (and lose upstream updates) or contaminate `.framework/`.

### Added

- **`src/utils/overlay-merger.js`** — new pure-function module. `mergeOverlay({ kind, name, baseContent, overlayContent, frameworkVersion })` returns the generated file content + warnings. Section markers: `## [OVERRIDE]`, `## [APPEND]`, `## [PREPEND]`; un-marked H2 sections in the overlay are appended at the end. Base frontmatter (`name`, `description`, `tools`, `model`) is preserved verbatim. Output carries a leading `<!-- baldart-generated: kind=… name=… base_version=… overlay_sha=… -->` marker so future updates know it's regeneration-safe.
- **`SymlinkUtils.mergeAgents(opts)` / `mergeCommands(opts)`** — per-item merge with overlay awareness:
  - No overlay → install as symlink (same model as skills).
  - Overlay present → write a real file with the merged content + marker.
  - Existing generated file → detected by marker → safely regenerated.
  - User-authored fork (no marker) → conflict logged, file left untouched, user told to migrate to overlay.
  - Broken symlink → silently re-linked (same fix-pattern as v3.5.1).
- **`tool-adapters/claude.js`** declares `agentsDir()` and `commandsDir()`. Codex returns `null` (no equivalent), so the dispatch skips it cleanly.
- **`framework/.claude/hooks/framework-edit-gate.js`** extended: blocks `Edit`/`Write` of any file whose first line is `<!-- baldart-generated:` unless the new content also starts with the same marker. Forces the user to edit the overlay instead.
- **Example overlays**: `framework/templates/overlays/agents/coder.example.md` and `framework/templates/overlays/commands/codexreview.example.md` showing all three section markers in use.
- **`framework/agents/project-context.md` § 5.bis** — new section documenting the agent/command overlay protocol, layout, frontmatter, section markers, generation rules, drift handling, and hard rules.
- **`baldart doctor` Overlays line** now counts skill overlays + agent overlays + command overlays.

### Changed

- **`SymlinkUtils.createBulkSymlinks()`** — no longer creates bulk symlinks for `.claude/agents/` and `.claude/commands/`. If a legacy bulk symlink is detected (v3.7.x and earlier layout), it's unlinked and replaced with a real directory, populated per-item by `mergeAgents`/`mergeCommands` immediately after.
- **`SymlinkUtils.createAllSymlinks({ tools })`** — calls `mergeSkills` + `mergeAgents` + `mergeCommands` in sequence.
- **`SymlinkUtils.verifySymlinks()`** — `AGENTS.md` and `agents/` stay bulk symlinks; `.claude/agents/` and `.claude/commands/` are now validated as real dirs containing at least one framework-managed file (symlink or generated). Per-tool, only for tools whose adapter declares a non-null `agentsDir()` / `commandsDir()`.
- **`src/commands/update.js`** — post-update reconcile now runs the new agent + command merges in addition to skills.
- **`src/commands/migrate.js`** — additionally converts legacy `.claude/agents` and `.claude/commands` bulk symlinks to per-item.

### Migration (pre-3.8.0 → 3.8.0)

On the next `npx baldart update`, an old consumer sees:

```
⚠ Detected legacy bulk symlink at .claude/agents. Converting to per-item layout (v3.8.0+).
⚠ Detected legacy bulk symlink at .claude/commands. Converting to per-item layout (v3.8.0+).
✓ [Claude Code] agent linked: .claude/agents/coder.md
…
```

Zero data loss (the bulk symlink was just one pointer; the real content stays in `.framework/`). After the convert, the user can create `.baldart/overlays/agents/<name>.md` to specialize any framework agent without forking.

### Verification (proven in sandbox)

1. **Legacy → per-item migration**: bulk `.claude/agents` symlink → real dir + per-item symlinks ✓
2. **Overlay create → regenerate**: real generated file with marker + overlay content visible ✓
3. **Overlay remove → revert to symlink**: generated file deleted, symlink re-created ✓
4. **User fork without marker → conflict logged**: framework version NOT overwritten, user file preserved, guidance printed ✓
5. **Heading mismatch in overlay**: warning emitted, content appended as new section (no silent loss) ✓
6. **Marker round-trip**: SHA-256-truncated overlay hash + framework version recoverable from generated file header ✓

### Out of scope (explicitly deferred)

- **Codex agent/command equivalents**: Codex doesn't have subagents or slash commands (its custom prompts are deprecated). The overlay system is Claude-only by mechanism; the cross-tool flat skill model (v3.7.0) continues to handle Codex.
- **Skill overlay refactor**: skills today use runtime concatenation (SKILL.md body reads overlay), which works. Migrating skills to compile-time merge is opt-in for a future release.
- **Drift-section diff during update**: structural support is in place via `base_agent_version` in overlay frontmatter and `base_version` in generated marker, but the actual section-level diff against the previous-version base content is a future enhancement (the framework just emits a one-line "overlay vs vX → vY" warning today).

## [3.7.1] - 2026-05-22

`baldart update` now actively migrates pre-v3.7.0 configs to the new multi-tool layout instead of silently leaving them on `[claude]` only.

### Fixed

- **`src/commands/update.js` legacy-config migration** — after the subtree pull, if `baldart.config.yml` exists but `tools.enabled` is missing/empty, the CLI now:
  1. Detects the legacy state.
  2. Surfaces the new capability and shows the autodetected tool list (Claude always, plus Codex if `~/.codex/` exists).
  3. Asks: *"Backfill `tools.enabled: [claude, codex]` into baldart.config.yml?"*.
  4. On accept: writes the array into the config AND re-runs `mergeSkills` for the new tools so `.agents/skills/` is populated immediately — not deferred to the next update.

Without this, a pre-v3.7.0 install upgrading to v3.7.x would never get the Codex skill dir populated, because the per-tool dispatch falls back to `['claude']` when `tools.enabled` is absent.

## [3.7.0] - 2026-05-22

BALDART is now AI-tool agnostic. Every framework skill is installed into both Claude Code (`.claude/skills/`) AND OpenAI Codex CLI (`.agents/skills/`) when both are enabled — same source, two symlinks, zero duplication. Codex skill format is structurally identical to Claude's (`SKILL.md` with `name` + `description` frontmatter + optional `scripts/`/`references/`/`assets/`), so the same file works for both runtimes.

### Added

- **`src/utils/tool-adapters/`** — new adapter registry (mirrors the existing `routine-adapters/` pattern). Each adapter declares the directory its tool reads skills from, plus capability flags (subagents, slash commands, hooks). Currently: `claude.js`, `codex.js`. Adding Cursor / Aider / Cline later is a one-file addition.
- **`tools.enabled` in `baldart.config.yml`** — array of enabled tool names. Default: `[claude]`, plus `codex` if `~/.codex/` is detected on the user's machine. Honored by `add`, `update`, `migrate`, and `verifySymlinks`.
- **`baldart configure` AI-tools section** — prompts per available tool with the autodetected default. Refuses to disable Claude (the framework's primary target).
- **`AUTODETECTED` summary** now shows `AI tools: claude, codex`.

### Changed

- **`SymlinkUtils.mergeSkills({ tools })`** — runs the per-item merge once per enabled tool, resolving each tool's target directory via its adapter. The same source under `.framework/framework/.claude/skills/<skill>/` is linked into every enabled tool's expected location.
- **`SymlinkUtils.createAllSymlinks({ tools })`** — accepts the tool list and passes it through to `mergeSkills`.
- **`SymlinkUtils.verifySymlinks()`** — reads `tools.enabled` from `baldart.config.yml` and validates each tool's skills dir separately. Output lines are now tagged `[Claude Code]` / `[OpenAI Codex CLI]`.
- **Per-item skill merge resilience** — `_mergeSkillsForTool` now detects broken symlinks via `fs.lstatSync` and re-links them silently instead of crashing on EEXIST (same fix-pattern as v3.5.1, now applied per-tool).

### Why it matters

Until now BALDART was Claude-Code-only by construction. A consumer using Codex saw nothing — Codex's discovery walks `.agents/skills/`, and the framework had no awareness of that path. v3.7.0 makes the install universal: declare `tools.enabled: [claude, codex]` (or both, autodetected) and Codex picks up the same 24 skills, indexed by its own discovery, with `AGENTS.md` at the root continuing to work for both as it did before. Adding more tools later (Cursor `.cursor/rules/`, Aider `CONVENTIONS.md`, Cline, etc.) only requires writing a new adapter — the rest of the install pipeline is already tool-pluggable.

### Out of scope (Codex-specific limitations)

- **Subagents** (Claude's `.claude/agents/<name>.md`) have no Codex equivalent. They remain Claude-only. AGENTS.md sections cover the cross-tool coordination protocol.
- **Slash commands** (Claude's `.claude/commands/`) — Codex deprecated custom prompts in favor of skills, so `/commandX` patterns stay Claude-only. If you want a workflow callable from Codex, author it as a skill.
- **Hooks** (Claude's `.claude/hooks/`) — Codex has no equivalent hook system. `framework-edit-gate` remains a Claude-only safety net.

## [3.6.4] - 2026-05-22

Two configure-flow fixes surfaced by a real-world `mayo` install audit.

### Fixed

- **API docs false negative** — `has_api_docs` only checked for `docs/.../api/schemas.md` (and OpenAPI/GraphQL specs). Projects that ship API docs as one markdown per resource (e.g. `docs/references/api/{index,admin,push,…}.md`) were marked `has_api_docs: false` despite having extensive API documentation. Added a populated-directory probe: if `docs/references/api/` or `docs/api/` exists and contains ≥1 `.md` file, the feature flag flips to `true` and `api_index` falls back to `<dir>/index.md`.

### Changed

- **Prompt label clarity** — `"Charting libraries (canonical, comma-separated)"` → `"Approved charting libraries — comma-separated package names (empty if none)"`. Same edit for forbidden and animation. The previous label read like the user was being asked to type the word `canonical`; one tester actually did.

## [3.6.3] - 2026-05-22

Smarter project autodetection in `baldart configure`. The previous probe only looked at a single hardcoded path (`docs/design-system/INDEX.md`), which silently failed on the vast majority of real projects (monorepos, inline `src/ui/`, `packages/ui/`, shadcn/ui setups, Tailwind-theme-only projects, etc.). The redundant "Design philosophy" prompt is now skipped automatically when a design system is detected.

### Changed

- **`src/commands/configure.js` `detect()`** — multi-signal design system detection:
  - **Path probes**: `docs/design-system`, `packages/ui`, `packages/design-system`, `packages/components`, `packages/ds`, `src/design-system`, `src/ui` (plus per-package expansion in monorepos).
  - **Storybook**: `.storybook/` at root or in any monorepo package.
  - **shadcn**: `components.json` marker.
  - **Tailwind theme**: `tailwind.config.{ts,js,mjs,cjs}` with `theme.extend` (regex inspection).
  - **DS libraries**: `@radix-ui/*`, `@chakra-ui/*`, `@mui/material`, `antd`, `@mantine/core`, `@nextui-org/*`, `daisyui`, `flowbite-react`, `@headlessui/react`, etc.
- **Monorepo awareness** — detects `pnpm-workspace.yaml`, `lerna.json`, `nx.json`, `turbo.json`, `rush.json`, or `package.json` `workspaces`. Expands every path probe across `packages/*`, `apps/*`, `services/*`, `libs/*`.
- **`components_primitives`** — when a design system is detected but no `components/ui` directory exists, probes the DS's own subdirs (`<ds>/src/components`, `<ds>/components`, `<ds>/src/ui`) and only returns paths that actually exist on disk.
- **UI guidelines** — widened to also detect `STYLEGUIDE.md`, `BRANDING.md`, `BRAND.md`, `docs/style-guide.md`, `docs/brand/README.md` plus monorepo variants.
- **API docs** — widened to detect OpenAPI specs (`openapi.{yaml,yml,json}`) and GraphQL schemas (`schema.graphql`) in addition to the existing markdown locations.
- **Brand name** — fallback chain: `package.json` `name` (scope stripped) → first H1 in `README.md` → `path.basename(cwd)`. No more empty `brand_name` for projects without a `package.json` name.
- **E2E framework** — also detects via dependency (`@playwright/test`, `cypress`) and `playwright.config.mjs`.
- **Charting/animation libraries** — added `chart.js`, `d3`, `visx`, `@tremor/react`, `@react-spring/web`, `auto-animate`, `motion`.
- **`interactivePrompts()` design philosophy** — gated on `!detected.features.has_design_system`. If a DS is detected, the prompt is skipped and a heuristic hint (e.g. `"Minimalist (shadcn/Tailwind)"`, `"Library-driven (@radix-ui)"`) is used as metadata for skills.
- **AUTODETECTED summary box** — now shows brand name, DS yes/no with its concrete signals (path/storybook/shadcn/tailwind-theme/library), monorepo detection + package count, API docs yes/no.
- **New config fields** under `stack.*`:
  - `stack.monorepo: { detected, roots }` — exposes detected monorepo packages to skills.
  - `stack.design_system_signals: { path, storybook, shadcn, tailwind_theme, library }` — granular signals for skills that want to adapt to the specific DS flavor.

### Why it matters

Before this change, a project with a real design system at `src/ui/` (or any non-canonical location) was told `Design system: — (none found)` and then asked the wrong question (`Design philosophy?`). Now the same project is correctly identified, the redundant prompt is skipped, and skills receive structured DS signals they can route on.

## [3.6.2] - 2026-05-22

BALDART is now published to npm as [`baldart`](https://www.npmjs.com/package/baldart) on every `v*.*.*` tag. End-users can run `npx baldart <cmd>` (without the `-y github:antbald/BALDART` prefix), which resolves through the npm registry and avoids the stale-tarball cache that plagued the GitHub-source installation.

### Added

- **`.github/workflows/publish-npm.yml`** — triggers on `v*.*.*` tag push (and `workflow_dispatch`). Steps: checkout the tagged ref, setup Node 20 with npm registry auth, sync `package.json` version to the `VERSION` file, verify tag matches VERSION (hard fail otherwise), `npm ci`, `npm publish --access public --provenance`. Requires `NPM_TOKEN` repo secret.
- **`package.json` scripts** — `sync-version` reads `VERSION` and writes it into `package.json`. `prepublishOnly` runs it so manual `npm publish` is also safe.
- **`package.json` files** — added `VERSION` and `CHANGELOG.md` to the published tarball; removed dangling `LICENSE` reference; removed unused `main` field (this is a CLI package, `bin` is what matters).

### Changed

- **README.md** — primary install/usage examples now use `npx baldart <cmd>` (npm form). The legacy `npx -y github:antbald/BALDART <cmd>` is mentioned as the fallback for unreleased commits on `main`.
- **CLAUDE.md** — documented the dual distribution model (npm registry + Git subtree for the framework payload) and the workflow's tag-vs-VERSION gate.
- **MAINTAINING.md** — release checklist now mentions the npm publish trigger and the `npm view baldart version` verification step.

### Why it matters

Before this, `npx -y github:antbald/BALDART` was the only install path, and npx's GitHub-tarball cache regularly served stale CLI code even after `update` to a newer framework version — producing confusing errors like *"Template not found at .framework/templates/baldart.config.template.yml"* that referenced bugs already fixed on `main`. Publishing to the npm registry sidesteps that entire caching layer.

### Operational note

The first tag push to v3.6.2 will fail if `NPM_TOKEN` is not yet configured as a GitHub Actions secret. Generate an **Automation token** on npmjs.com (Settings → Access Tokens → Generate New Token → Automation) and add it as `NPM_TOKEN` under repo Settings → Secrets and variables → Actions.

## [3.6.1] - 2026-05-22

Fix a long-standing path bug: the bulk symlinks (`AGENTS.md`, `agents/`, `.claude/agents/`, `.claude/commands/`), the customizable-template copies (hooks, ui-guidelines, brand-guidelines, backlog templates), and the `baldart.config.template.yml` lookup all referenced `.framework/<sub>` (single `framework/` segment). The Git subtree pull nests the entire BALDART repo under `.framework/`, and the repo itself nests its shippable content under a top-level `framework/` directory — so the real path is `.framework/framework/<sub>`. The bulk symlinks were therefore being created broken (pointed nowhere) and `configure` couldn't find its template, surfacing the misleading *"Run `npx baldart add` first"* error during `update`.

### Fixed

- **`src/utils/symlinks.js`** — introduced `FRAMEWORK_PAYLOAD = .framework/framework` constant. All bulk symlinks (`AGENTS.md`, `agents`, `.claude/agents`, `.claude/commands`), the `copyCustomizableFiles()` source paths (hooks, docs/references, templates copy loop), and the relative-mode symlink callsites now resolve through it. The relative symlink callers no longer mis-prepend `..` (which previously combined with `path.join(cwd, target)` to navigate ABOVE the consumer repo).
- **`src/commands/configure.js`** — `TEMPLATE_PATH` corrected to `.framework/framework/templates/baldart.config.template.yml`.
- **`src/commands/update.js`** — schema-drift template lookup corrected to the same path.
- **`src/commands/doctor.js`** — `loadConfigTemplate()` corrected to the same path.

### Why it matters

Until now, every consumer had **broken** `AGENTS.md`/`agents/`/`.claude/agents/`/`.claude/commands/` symlinks (pointing at paths that don't exist) — Claude Code silently saw "no shared agents/commands" and used only what the user had authored locally. Same story for `configure`: it could never find its template, so it bailed and the user's `baldart.config.yml` was never generated automatically during `update`. After v3.6.1, on the next `npx baldart update` the symlinks get re-created with the correct `.framework/framework/` target and `configure` runs end-to-end.

## [3.6.0] - 2026-05-22

`npx baldart update` now handles a dirty consumer worktree gracefully instead of crashing with `fatal: working tree has modifications. Cannot add.` from `git subtree pull`. Before the subtree pull runs, the CLI detects uncommitted changes, lists them, and offers to auto-stash + re-apply automatically after the update.

### Added

- **`src/commands/update.js` pre-flight check** — after the user confirms "Proceed with update?", if the working tree is dirty: the CLI lists the dirty paths and asks "Auto-stash now and re-apply after the update (recommended) / Abort". Stash is labelled `baldart-pre-update-<ISO timestamp>` for traceability.
- **Auto-pop after success** — on a successful subtree pull the stash is popped. If the pop conflicts with the framework update, the stash is left intact and the user gets a `STASH CONFLICT` box with recovery commands (`git stash list`, `git stash pop`).
- **Auto-restore on failure** — if the subtree pull itself fails after stashing, the stash is restored automatically so the user's pre-update changes never appear lost.

### Why it matters

Combined with the v3.5.0 post-update auto-commit, the full update flow now keeps the worktree clean on entry AND on exit, without ever silently dropping the user's work-in-progress.

## [3.5.1] - 2026-05-22

Fix `npx baldart update` crashing with `EEXIST: file already exists, symlink …` when the bulk symlinks (`AGENTS.md`, `agents`, `.claude/agents`, `.claude/commands`) were broken (target sparito dopo un cleanup o move). `fs.existsSync` segue il symlink e tornava `false` per i broken link → il ramo di cleanup veniva saltato → `fs.symlinkSync` falliva creando il nuovo link sopra quello rotto.

### Fixed

- **`src/utils/symlinks.js` `createSymlink()`** — switched from `fs.existsSync` to `fs.lstatSync` (with try/catch) to detect broken symlinks and treat them as "existing symlink to wrong target" → unlink + recreate.
- **`src/utils/symlinks.js` `verifySymlinks()`** — same fix on the diagnostic path, so broken bulk symlinks are now reported as `Broken symlink: <link> → <target>` instead of being misclassified as `Missing`.

## [3.5.0] - 2026-05-22

`npx baldart update` now offers to commit the post-update reconcile (symlinks, hook registration, state.json, routine adapters, etc.) in a single dedicated commit, so the consumer's worktree stays clean and there is a clear revert point per upgrade.

### Added

- **`src/commands/update.js` `postUpdateAutoCommit()`** — runs after all post-update reconcile steps. Scans `git status`, filters to a BALDART-managed allowlist (`.framework/`, `AGENTS.md`, `agents`, `.claude/{agents,commands,skills,hooks,routines,settings.json}`, `.baldart/`, `baldart.config.yml`, `docs/references/{ui-guidelines.template,brand-guidelines}.md`, `templates/`, `.github/workflows/baldart-*`, `scripts/routines/`), shows a preview and asks before committing as `chore(baldart): post-update reconcile to vX.Y.Z`. User-owned dirty paths are never staged.
- **`bin/baldart.js` `--no-commit` flag on `update`** — skips the prompt entirely for CI/scripted runs.

### Why it matters

Before v3.5.0 every `update` left a dirty worktree (the subtree-pull commit only covers `.framework/`; symlink reconcile + hook + state writes happen outside it). Consumers either had to `git status` and craft a commit by hand or live with a noisy `git status`. The new flow keeps the post-update side effects bundled in one self-describing commit you can `git revert` if anything goes wrong.

## [3.4.1] - 2026-05-22

Fix `npx baldart update` crashing with `ReferenceError: newVersion is not defined` after a successful framework update. The update completed correctly but the final "WHAT CHANGED" summary referenced an out-of-scope `const`, surfacing the error through the outer catch and exiting non-zero — confusing the user and breaking CI signal.

### Fixed

- **`src/commands/update.js`** — hoisted `newVersion` declaration above the inner try/catch (`let newVersion;`) so the "WHAT CHANGED" box can read it. The `const` lived inside the inner try block (lines 105–138) and was unreachable from the success summary below it.

## [3.4.0] - 2026-05-22

Worktree manager now respects terminal isolation. The `/mw` skill no longer runs `git checkout develop` on the main repo (which is a shared resource across parallel terminals in worktree-driven projects). Same for `/nw` pre-flight. Reported from a production failure in the `mayo` consumer where CLAUDE.md declares terminal-isolation as an absolute rule and Claude Code's classifier preemptively blocked the checkout.

### Changed

- **`framework/.claude/skills/worktree-manager/SKILL.md` `/mw` step 4c** — replaced local-checkout merge (`git checkout develop && git pull --ff-only && git merge && git push`) with remote merge via GitHub PR (`gh pr create` + `gh pr merge --merge --delete-branch`). The main repo's `HEAD` is never touched. Local `develop` is synced via `git -C <main> pull --ff-only` ONLY when the main repo HEAD is already `develop`; otherwise we just fetch and leave the user's working state alone.
- **`/mw` step 4c fallback** — if `gh pr merge` CLI fails (typically on the local-branch cleanup when `develop` is checked out in another worktree), fall back to the REST API (`gh api -X PUT repos/:owner/:repo/pulls/:n/merge`). Hard-fails (no local-checkout fallback) — the user installs `gh` and re-runs.
- **`/nw` pre-flight (step 1) + step 3** — no longer runs `git checkout develop` on the main repo. Pre-flight is now `git fetch origin develop` (read-only); worktrees branch from `origin/develop` directly.
- **`/mw` Programmatic API output** — now returns `prNumber` alongside `mergeCommit`.
- **Safety Rules** in worktree-manager SKILL.md — added explicit `NEVER git checkout/switch/branch on main repo` rule with rationale (terminal isolation, classifier pre-blocking).
- **`framework/.claude/skills/new/SKILL.md` Phase 6 description + Fail-safe rules** — aligned to the new flow.

### Why it matters

Many consumer projects declare a "shared main repo" / terminal-isolation rule in their CLAUDE.md. The classifier in Claude Code pattern-matches the rule text (not the runtime state) and preemptively denies `git checkout develop` on the main repo. The legacy flow therefore failed silently in those projects, and the `/new` orchestrator could not close the merge loop — the user had to merge by hand. With v3.4.0 the merge happens entirely server-side via the GitHub API and the main repo is never touched.

### Note for consumers

`gh` CLI must be installed and authenticated (`gh auth login`). If it isn't, `/mw` STOPs and asks the user to install it — it does NOT fall back to the legacy local-checkout flow.

## [3.3.0] - 2026-05-21

Real-time contamination gate. A Claude Code `PreToolUse` hook now intercepts every `Edit`/`Write`/`MultiEdit` whose target resolves (via symlink) to a path inside `.framework/`, runs the contamination scanner on the new content, and blocks the call with an informative reason if project-specific tokens are detected. Auto-registered in every consumer at install time.

### Added

- **`framework/.claude/hooks/framework-edit-gate.js`** — Claude Code `PreToolUse` hook (Node, cross-platform, fail-safe). Reads tool input from stdin, resolves the real path with `fs.realpathSync` (follows symlinks), and checks if it lands inside `.framework/`. If so, scans the new content via the existing contamination utility and returns `permissionDecision: "deny"` with a structured reason walking the user through three resolution paths: (A) reformulate generically using `${paths.X}`/`identity.X`/`stack.X`, (B) move to `.baldart/overlays/<skill>.md`, (C) declare the file opt-out with the contamination-scan marker.
- **`src/utils/hooks.js`** — registration utility. `register()` / `unregister()` / `isRegistered()` operate on `.claude/settings.json` with idempotent JSON merge (preserves user-authored hooks). The registered entry carries a stable `id: "baldart-framework-edit-gate"` for future detection.
- **Auto-registration**:
  - `baldart add` registers the hook after symlinks are created (creates `.claude/settings.json` if absent, merges otherwise).
  - `baldart update` re-registers (idempotent) — backfills consumers installed before v3.3.0.
  - `baldart doctor` detects missing registration and proposes "Register framework-edit-gate hook" as a remediation action.
- **`baldart doctor` status line**: new "Edit gate" row in the diagnostic table showing registered vs not-registered.

### How it works

A typical flow now looks like this. Claude is editing a skill in fidelity-app:

```
Claude → Edit .claude/skills/ui-design/SKILL.md (new_string contains "Neo-Brutalism")
       ↓
       framework-edit-gate.js fires (PreToolUse)
       ↓
       Resolves symlink → /…/.framework/framework/.claude/skills/ui-design/SKILL.md
       ↓
       Detects ".framework/" segment → triggers
       ↓
       Runs contamination.scan() on new_string → finds "Neo-Brutalism"
       ↓
       Returns permissionDecision: "deny" with structured reason
       ↓
Claude reads the reason, moves the rule to .baldart/overlays/ui-design.md instead.
```

The hook is fail-safe: any internal error (missing scanner, parse error, IO error) exits 0 and lets the tool call through. The scanner is a safety net, not a tribunal.

### Disabling

If the gate gets in the way for a specific session, remove the entry from `.claude/settings.json` under `hooks.PreToolUse[]` (look for `id: "baldart-framework-edit-gate"`). `baldart doctor` will detect the absence and offer to re-register on next run.

For per-file legitimate exceptions (docs that legitimately quote canonical paths as illustrative examples), add the existing opt-out marker:

```markdown
<!-- contamination-scan: skip -->
```

or in YAML frontmatter:

```yaml
contamination_scan: skip
```

## [3.2.0] - 2026-05-21

Single smart entry point. `npx baldart` with no arguments now runs a doctor that diagnoses the install and proposes the next sensible action.

### Added

- **`baldart doctor` command** + **no-argument shortcut**. Detects (in priority order): not-a-git-repo, framework absent, legacy v2.0.x bulk-skills layout, malformed config, missing config, config schema drift, remote ahead, local framework changes, overlay version drift. Renders a status table, then proposes the matching action(s) and runs each with confirmation. Acts as the new informational default — `npx baldart` no longer prints commander help, it prints a diagnostic.
- **`--auto` flag**: skips yes/no confirmations for CI; errors out with exit 2 if multiple actions are proposed (avoids guessing in ambiguous states). Pairs cleanly with `--offline` for fully sandboxed runs.
- **`--offline` flag**: skips the upstream fetch (no remote drift detection).
- Both flags work without typing `doctor`: `npx baldart --auto --offline` is equivalent to `npx baldart doctor --auto --offline`.

### Changed

- `npx baldart` with no arguments previously dumped the commander help. It now runs the doctor flow instead. The help is still accessible via `npx baldart --help` or `npx baldart <subcommand> --help`.

## [3.1.0] - 2026-05-21

Contribution flow upgrade. Centralized framework versioning via `.baldart/state.json`. New `/baldart-push` skill orchestrates safe upstream contributions with automatic contamination autofix, version bump, CHANGELOG generation, and tag push.

### Added

- **`/baldart-push` skill** (`framework/.claude/skills/baldart-push/`) + slash command (`framework/.claude/commands/baldart-push.md`). Conversational orchestration of the contribution flow with hard rules: no secrets, no overlay/config push, always show current version first.
- **`src/utils/contamination.js`** — scanner + autofix library. Catalogues every fidelity-app-specific token pattern that v3.0.0 cleaned up. Three severities: `autofixable` (safe substitutions like `docs/design-system/` → `${paths.design_system}/`), `requires-decision` (identity/stack tokens — user reviews per file), `block` (secrets — hard fail).
- **`src/utils/state.js`** — `.baldart/state.json` read/write. Centralized versioning state: installed_version, install_date, last_update_date, last_pushed_version, last_push_date, rolling history of the last 20 events. Updated automatically by `add`/`update`/`push`.
- **`baldart version`** command rewritten — now shows installed version + install date, remote drift (ahead/behind), local uncommitted-files count, last-push info, CLI version. New `--offline` flag to skip the upstream fetch.
- **`baldart push`** command rewritten — auto-pull when behind, file triage, new-skill discovery, contamination scan with per-file decisions, auto VERSION bump, auto-generated CHANGELOG entry (Added/Changed/Removed grouped from diff), annotated tag push, state.json update. Replaces the previous "next steps required" manual workflow.
- **Triage rules**: hooks customizations (`.claude/hooks/*.sh` non-template) are blocked from push automatically; overlay starter examples (`framework/templates/overlays/*`) are exempted from contamination scan.

### Changed

- `src/commands/add.js`, `src/commands/update.js` now record events into `.baldart/state.json` after successful operations.
- The push flow now exits with `Description is required.` when the user provides empty description (was previously a silent failure mode).
- `AGENTS.md` `§ 4 Non-negotiables (MUST)` split into **4.a Universal MUST** (apply everywhere) and **4.b Workflow MUST** (gated by `features.has_backlog` / `features.has_adrs` / existence of `project-status.md` / overlay-declared conventions). Pre-v3 absolute MUST rules around backlog / ADRs / GitHub issue labels / `git_strategy` are now correctly feature-gated. Duplicate ADR completeness MUST removed.
- `framework/agents/project-context.md` § 6 autodetection table rewritten as the source of truth, listing all 20 probes that `configure.js detect()` actually runs.
- README "Migrating an Existing Install" rewritten as `v1.x / v2.x → v3.x` (was v2.1.1 framing), with explicit `npx baldart configure` step and links to `PROJECT-CONFIGURATION.md` § 9.
- README duplicate `push` command section removed; Quick Start now lists `configure`.
- `MAINTAINING.md` prefaced with v3.1.0+ automated flow TL;DR (the CLI now does steps that used to be manual: VERSION bump, CHANGELOG entry, tag).

### Fixed

- `src/commands/push.js` "skip file" branch now distinguishes status `A` (newly added — `git rm -f`) from `M`/`R`/`D` (existing upstream — `git checkout origin/main`). Previously the `checkout` failed silently on newly-added files and they were pushed despite the user's skip decision.
- `src/commands/push.js` autofix commit now stages only the files autofix actually modified (was sweeping all pushable files, capturing unrelated edits).
- `src/commands/push.js` partial failure surfaces a rollback SHA snapshot (`git reset --hard <preFlightSHA>`) so the user can undo chore commits cleanly.
- `src/commands/push.js` path comparisons normalized to forward slashes — fixes Windows where `simple-git` returns `/` but `path.join` produces `\`, so `CHANGELOG.md` and `VERSION` previously fell into the `pushable` bucket and were double-committed.
- `src/commands/push.js` `buildChangelogEntry` strips the `.framework/` mount-point prefix from path entries (was `.framework/framework/.claude/skills/foo/SKILL.md`; now `framework/.claude/skills/foo/SKILL.md`).
- `src/commands/configure.js` `mergePreserving` preserves deliberate empty-string user values (was overwriting them with detected defaults).
- `src/utils/contamination.js` `stack-recharts` rule case-insensitive (was missing lowercase `import 'recharts'`).
- `src/utils/contamination.js` `secret-api-key` rule catches unquoted assignments (`API_KEY=abc123...` in `.env` files).
- `src/utils/contamination.js` added secret patterns: GitHub PATs (`ghp_`/`ghs_`/`ghu_`/`gho_`), Slack tokens (`xox[baprs]-`), JWTs.

## [3.0.0] - 2026-05-21

This is a MAJOR release. Skills no longer hard-code paths, brand identity, or technology stack choices — they read them from `baldart.config.yml` and project-specific overlays. The same skills now work across projects with different design systems, audiences, and stacks.

### Added

- **Project context system** (3 layers):
  - `baldart.config.yml` at repo root — schema versioned, populated by the new `npx baldart configure` command. Defines `paths.*`, `identity.*`, `stack.*`, `features.*` for the project.
  - `.baldart/overlays/<skill>.md` — consumer-authored per-skill extensions with `base_skill_version` frontmatter for drift detection. Sections marked `## [OVERRIDE] <topic>` replace base skill sections.
  - `framework/agents/project-context.md` — protocol module loaded as MANDATORY pre-read for any skill that touches project-specific facts.
- **`npx baldart configure` command** — interactive prompts + filesystem autodetection (probes for design-system index, components dirs, backlog, ADRs, wiki overlay, package.json dependencies for charting/animation/E2E framework). Idempotent: re-running merges without clobbering user values. Available as `--non-interactive` for CI.
- **Project context header convention** — every refactored skill now declares its config dependencies in a dense 4-line header citing `framework/agents/project-context.md` for the protocol.
- **Overlay starter templates** under `framework/templates/overlays/`:
  - `ui-design.fidelity-example.md` — Neo-Brutalism, Recharts/nivo stack, merchant theming, Safari open command, evaluation thresholds.
  - `copywriting.fidelity-example.md` — 4-pillar brand voice, Italian-only rule, audience register matrix, forbidden vocabulary.
- **`framework/docs/PROJECT-CONFIGURATION.md`** — comprehensive guide for consumers: schema reference, overlay system, per-skill behaviour when keys are missing, ground-up examples, v2→v3 migration cheat sheet, FAQ.
- **Status command extensions** — reports `baldart.config.yml` presence + schema version, lists authored overlays with version-drift warnings.
- **Update command extensions** — never overwrites `baldart.config.yml`; warns when new framework version adds config keys not yet in the consumer's file (suggests re-running configure).

### Changed (BREAKING)

- **Skills no longer assume the fidelity-app shape**. Sixteen skills with significant hard-coded paths or identity tokens were refactored to resolve them from `baldart.config.yml`; the remaining skills (`api-design-principles`, `remotion-best-practices`, `seo-audit`, `skill-creator`, `worktree-manager`, `find-skills`) had no significant project-leaks and were left untouched. The refactored skills:
  - `ui-design`, `ui-design/references/*` — design-system reads, charting stack, identity-driven aesthetic
  - `copywriting` — brand voice pillars now overlay-driven
  - `prd`, `prd-add`, `prd/references/*`, `prd/assets/*` — backlog/PRD/references paths, design-system gating
  - `new` — backlog and references paths, validation tooling generalized
  - `bug`, `bug/references/logging-patterns.md` — project-specific debug entry points moved to overlay
  - `kie-ai` — design-system reads gated by `features.has_design_system`
  - `motion-design` — illustration-motion path resolved from config
  - `capture` — wiki overlay path from `paths.wiki_dir`, domain vocabulary from overlay
  - `context-primer` — backlog/references/wiki paths from config; feature-gated probing
  - `playwright-skill`, `webapp-testing` — design-system reads gated, theming pairing from overlay
  - `gamification-design` — design system + audience segment from config, customer surface guidance overlay-driven
  - `doc-writing-for-rag` — schemas/errors paths from config, examples marked as illustrative
  - `frontend-design`, `simplify` — design-system reads gated by `features.has_design_system`
- **Skills refuse to run when their gating feature is `false`**. E.g. `/new` requires `features.has_backlog: true`, `/prd` requires `features.has_prd_workflow: true`, `/capture` requires `features.has_wiki_overlay: true`. This replaces silent degradation with explicit refusal + suggestion to enable the feature in config.
- **Always-ask on missing keys**. When a skill encounters a required path that is `""`, or a `features.*` flag absent from the YAML, it prompts the user and suggests persisting the answer via `npx baldart configure`. Never defaults to `false` from absence.
- **`templates/` install excludes framework-internal templates**. `baldart.config.template.yml`, `skill-project-context.snippet.md`, and the `overlays/` example directory live in `.framework/templates/` only — they are consumed by CLI commands and references, not copied to the user's `templates/`.

### Migration from 2.x

`npx baldart update` warns if `baldart.config.yml` is missing and offers to run configure. Existing v2.x installs need:

1. `npx baldart update` (pulls v3.0.0).
2. `npx baldart configure` (interactive — autodetects ~80% of values from typical fidelity-app layouts).
3. For projects that relied on the pre-v3 hard-coded opinions (Neo-Brutalism, merchant/customer, Recharts-only): copy the matching `*.fidelity-example.md` from `.framework/templates/overlays/` into `.baldart/overlays/` and adapt.
4. `npx baldart status` — verify config + overlays.
5. Commit `baldart.config.yml` and `.baldart/overlays/`.

Full migration guide: `.framework/docs/PROJECT-CONFIGURATION.md` § 9.

## [1.0.0] - 2026-02-13

### Added
- Complete npm package implementation with CLI
- 5 commands: add, update, push, version, status
- Interactive prompts with colored terminal output
- Git subtree bidirectional sync (pull updates + push improvements)
- Professional utility classes (git, ui, symlinks)
- Comprehensive README with npm/npx usage
- 9 generic AI agents (codebase-architect, coder, code-reviewer, doc-reviewer, prd, plan-auditor, senior-researcher, api-perf-cost-auditor)
- 17 domain modules (architecture, workflows, testing, security, etc.)
- 3 commands (/new, /design-review, /issue-review)
- Templates (feature-card, spec, breaking-change-checklist, ui-guidelines, brand-guidelines)
- Symlink-based auto-update mechanism
- Customizable files (hooks, UI guidelines, templates)

### Changed
- **BREAKING**: Replaced bash scripts with npm package
- **BREAKING**: Installation now via `npx baldart add` instead of bash script
- **BREAKING**: All commands now use `npx baldart <command>` syntax
- **BREAKING**: Framework files moved to `framework/` subdirectory

### Removed
- install-framework.sh (replaced by src/commands/add.js)
- update-framework.sh (replaced by src/commands/update.js)
- push-improvements.sh (replaced by src/commands/push.js)

## [1.1.0] - 2026-02-27

### Added
- 12 new specialized agents: hybrid-ml-architect, hyper-gamification-designer, legal-counsel-gdpr, marketing-conversion-strategist, motion-expert, onboarding-architect-lead, qa-sentinel, remotion-animator-orchestrator, seo-analytics-strategist, ui-expert, visual-designer, website-orchestrator
- `/qa` command: Standardized QA workflow with 3 profiles (Light, Balanced, Deep)
- `/check` command: Pre-development parallel quality audits on backlog cards using agent teams
- QA Protocol in REGISTRY.md with profile selection guide and self-healing loop
- Batch tracking protocol in `/new` command for parallel session safety
- Context recovery protocol in `/new` for resilience after context compaction
- Phase 3.5 (QA Validation) in `/new` pipeline with automatic profile selection
- Phase 7 (Production Readiness Checklist) in `/new` pipeline
- Branching Strategy section in AGENTS.md (Simplified Git Flow with develop branch)
- Execution Modes section in AGENTS.md (local/cloud/hotfix)
- Pre-commit vs Pre-PR build check separation in AGENTS.md
- Macro Feature Identification Protocol in doc-reviewer
- Single Source of Truth (SSOT) Protocol in doc-reviewer
- Post-Development Doc Debt tracking in doc-reviewer
- Multi-phase review with risk matrix in plan-auditor
- PRD Review & Confirmation phase (Phase 2.1) with [TO CONFIRM] markers in prd
- API Performance & Cost Audit phase (Phase 2.5) in prd
- Autonomous technical decisions (Phase 1A) in prd — agent decides tech, asks user only product questions
- Parent/child card structure with epic support in prd
- Bug card template with mandatory clarity analysis in prd

### Changed
- **REGISTRY.md**: Expanded from 9 to 21 agents with compact table format, decision tree, and QA Protocol section
- **doc-reviewer**: Major rewrite — now fully responsible for documentation (audit + write), with SSOT sync and doc debt tracking
- **plan-auditor**: Enhanced with 4 expert personas (Staff Engineer, Tech Lead, Security Engineer, SRE), anti-patterns checklist, and pre-mortem scenario
- **prd**: Major expansion — one-question-at-a-time discovery, gamification validation for B2C, performance audit integration, detailed card templates
- **senior-researcher**: Improved with AI-readable output format (retrieval index, evidence map, structured reading notes)
- **code-reviewer**: Added structured output format (Critical/Major/Minor/Recommendations), review checklist, and Linked Skills pattern
- **coder**: Added explicit build/test/lint gates, task scoping guidelines, and Linked Skills pattern
- **api-perf-cost-auditor**: Added project context integration section and Linked Skills pattern
- **new.md**: Added worktree grouping, QA validation phase, production readiness checklist, and context recovery
- **AGENTS.md**: Added branching strategy, execution modes, testing gates separation, commit hygiene rules

## [2.1.1] - 2026-05-18

Critical migration-safety fix. v2.0.0 / v2.1.0 introduced a bulk symlink for
`.claude/skills/` that — on existing installations — silently renamed the
user's personal skills directory to `.claude/skills.backup` and replaced it
with the framework symlink. This release fixes the merge strategy and ships
a migration command for repos already affected.

### Added

- **`npx baldart migrate` command** — idempotent recovery for repos that ran
  v2.0.0 / v2.1.0 update with the old bulk-symlink strategy:
  - Detects legacy `.claude/skills` bulk symlink and converts it to a real
    directory.
  - Re-merges framework skills as per-item symlinks alongside user skills.
  - Restores `.claude/skills.backup/*` user skills into `.claude/skills/`
    (skipping name collisions, which stay in `.backup/` for manual review).
  - Removes empty `.backup/` directory at the end.
  - Records remaining collisions in `.baldart/skill-conflicts.json`.
- **Per-skill merge strategy** in `symlinks.js`:
  - `.claude/skills/` is now a **real directory the user owns**.
  - Each framework skill is symlinked individually:
    `.claude/skills/<name> → ../../.framework/framework/.claude/skills/<name>`.
  - User-authored skills coexist in the same directory and are NEVER
    overwritten or backed up by the framework.
  - Name collisions are logged to `.baldart/skill-conflicts.json` with
    enough context for the user to resolve them.
- **`createSymlink` modes** in `symlinks.js`:
  - `safe`: refuses to overwrite a user-customised file/dir (used on fresh
    installs).
  - `prompt` (default): asks before backing up a customised file/dir.
  - `force`: legacy v2.0.x behaviour (no longer used by default).

### Changed

- `npx baldart add` now runs symlink creation in `safe` mode — first-time
  installs never silently rename user customisations.
- `npx baldart update` no longer issues a generic "Recreate broken symlinks?"
  prompt that triggered the destructive path. Instead it explains exactly
  what will happen (per-item merge, no overwrites without confirmation) and
  prompts in `prompt` mode.
- `verifySymlinks()` now spot-checks `.claude/skills/` for the per-item
  layout and reports a clear actionable warning when the legacy bulk
  symlink is detected (`Run npx baldart update or npx baldart migrate`).
- `createAllSymlinks(opts)` is now `async` because some bulk-symlink
  reconciliations may ask the user for confirmation. All callers updated.

### Fixed

- **Destructive update path on existing repos**: v2.0.0 / v2.1.0
  `npx baldart update` could rename `.claude/skills/` to `.claude/skills.backup`
  on existing installs without explicit user consent. The new merge strategy
  never touches user-owned content under `.claude/skills/`.

### Migration (v1.x / v2.0.x / v2.1.0 → v2.1.1)

After updating BALDART:

```bash
npx baldart update     # pulls v2.1.1 framework
npx baldart migrate    # one-shot: converts skills layout, restores .backup
npx baldart status     # confirm
```

`migrate` is idempotent — safe to re-run.

---

## [2.1.0] - 2026-05-18

Minor release that closes the gap left by v2.0.0: the framework now ships the
**scheduled routines** that make agents like `wiki-curator` and `skill-improver`
actually run on a cadence. Without these schedules the auto-learning loops are
dormant, so every BALDART install is now prompted to configure them.

### Added

- **`framework/routines/` directory** with 6 declarative routine specs:
  - `wiki-review` — nightly @ 02:00 UTC — drives the LLM-wiki auto-learning loop
  - `doc-review` — nightly @ 00:00 UTC — audits doc changes, flags SSOT drift
  - `code-review` — nightly @ 01:00 UTC — reviews last-24h commits
  - `skill-improve` — weekly Sunday @ 02:00 UTC — refines skills/agents
  - `ds-drift` — weekly Monday @ 03:00 UTC — design-system drift check (optional)
  - `full-sweep` — weekly Sunday @ 03:00 UTC — full SSOT audit (optional)
- **3 backend adapters** (`src/utils/routine-adapters/`):
  - `claude-code-cloud` — generates `.claude/routines/<name>.json` for RemoteTrigger
  - `github-actions` — generates `.github/workflows/baldart-<name>.yml`
  - `cron` — generates `scripts/routines/<name>.sh` + prints crontab line
- **`npx baldart routines` command** with 4 subcommands:
  - `list` — shows status for every routine (installed / available / blocked /
    skipped / disabled / unavailable)
  - `install <name>` — interactive install with backend picker
  - `disable <name>` — removes the schedule artifact and updates the lock
  - `doctor` — verifies installed routines still have their artifacts present
- **Post-install / post-update wizard** — `npx baldart add` and
  `npx baldart update` now surface newly-available routines and offer to
  configure them right after a successful install/update. Skipped routines
  are remembered in `.baldart/routines.lock.json` and not re-prompted.
- **`.baldart/routines.lock.json`** — durable record of which routines are
  installed, their backend, since-version, and per-routine status.

### Changed

- `npx baldart add` next-steps box now includes a "Review scheduled routines"
  step pointing to `npx baldart routines list`.
- `npx baldart update` "what changed" box now includes a routines review
  pointer.
- Added `js-yaml@^4.1.0` as a runtime dependency for parsing routine specs.

### Notes

- Routines are **opt-in per backend**: BALDART never installs schedules
  silently. The user picks the backend and reviews the generated artifact
  before committing.
- The wizard skips routines whose `required_artifacts` are missing (e.g.
  `wiki-review` is blocked until `docs/wiki/` exists). Run `npx baldart
  routines install <name>` later once the artifact lands.
- The CLI is non-interactive when stdin is not a TTY — safe to run in CI
  contexts (the wizard simply exits).

---

## [2.0.0] - 2026-05-18

Major release distilled from ~3 months of intensive use of v1.1.0 inside a real
multi-merchant fidelity/loyalty platform. Introduces the **Skills** category, a
new **LLM Wiki Overlay methodology**, the `/codexreview` deep multi-agent code
review command, and a substantial rewrite of every core agent.

### Added

- **Skills category (new)**: `framework/.claude/skills/` directory with 22
  generic, portable skills shipped out-of-the-box:
  - Workflow skills: `new`, `prd`, `prd-add`, `bug`, `simplify`,
    `worktree-manager`, `issue-review`, `context-primer`
  - Code-quality skills: `skill-creator`, `find-skills`,
    `webapp-testing`, `playwright-skill`
  - Design skills: `frontend-design`, `ui-design`, `motion-design`,
    `gamification-design`
  - Product skills: `seo-audit`, `copywriting`, `api-design-principles`
  - Knowledge skills: `doc-writing-for-rag`, `capture` (LLM wiki overlay)
  - Integration: `kie-ai`, `remotion-best-practices`
- **5 new agents**:
  - `security-reviewer` — dedicated AppSec auditor for auth/secrets/multi-tenant/infra
  - `deep-human-insight` — psychological / sociological analysis for B2C UX
  - `email-deliverability-architect` — transactional email design + deliverability
  - `skill-improver` — weekly auto-improvement of skills/agents based on review/QA findings
  - `prd-card-writer` — atomic backlog-card generation from approved PRDs
  - `wiki-curator` — derived LLM wiki overlay maintenance (re-introduced and generalized)
- **`/codexreview` command** — deep multi-agent code review with mandatory
  false-positive validation. Pools findings from `code-reviewer`,
  `security-reviewer`, `api-perf-cost-auditor`, `doc-reviewer`, and
  `plan-auditor` into a single ranked report with dedup and risk weighting.
- **`agents/llm-wiki-methodology.md`** — full methodology for building a
  derived LLM wiki overlay: layered knowledge stack, page types, frontmatter
  discipline, auto-learning loop (RAG telemetry → synthesis candidates →
  `/capture` skill → new wiki pages), adoption checklist.
- **Post-Approval Complexity Gate** (in `coder.md`) — explicit isolation-mode
  decision (heavy/light worktree) made by the orchestrator before spawning
  the coder, eliminating ambiguity about who owns branch creation.
- **High-Risk Path Code Review trigger table** (in `plan-auditor.md`) — five
  generic triggers that force a per-card `/codexreview` BEFORE merge.
- **Specialist Auto-Spawn matrix** (in `plan-auditor.md`) — auto-routes plans
  to specialist auditors (security, api-perf, ui, ml, docs) based on domain
  signals.
- **NFR enforcement framework** in `code-reviewer.md` and `new` skill —
  generic performance / security / accessibility MUST rules with a project-
  adaptable anti-pattern list (the "Never demote" cluster).
- **Findings Schema** — single YAML emit shape across all reviewing agents so
  `/codexreview` can pool and rank findings.
- **YOLO MODE protocol** — explicit instruction that all Task-spawned
  subagents must use `mode: "bypassPermissions"`.

### Changed

- **All 9 core agents heavily rewritten** based on 3 months of production use
  in a real project:
  - `plan-auditor.md` (+22 KB): 4 expert personas (Staff Eng / Tech Lead /
    Security / SRE), High-Risk Path triggers, Specialist Auto-Spawn matrix,
    multi-phase review with risk matrix, anti-patterns checklist, pre-mortem.
  - `coder.md` (+17.6 KB): explicit build/test/lint gates, design-system
    SSOT read protocol, conditional binary-outcome items, reuse analysis,
    branch/worktree safety check, completion-report schema, Linked Skills.
  - `codebase-architect.md` (+15.3 KB): Linking Protocol v1 for canonical-
    source resolution, design-system resolution path, agent memory, retrieval
    protocol consumption.
  - `code-reviewer.md` (+14.3 KB): confidence-based filtering (HIGH/MED/LOW),
    "Never demote" anti-pattern cluster, Findings Schema emission, design-
    system compliance section, scope boundary discipline.
  - `doc-reviewer.md` (+10.6 KB): design-system scope (when project has one),
    SCIP-symbol code refs, drift validator suite integration.
  - `api-perf-cost-auditor.md` (+9.2 KB): project context integration, agent
    memory, Findings Schema, retrieval protocol consumption.
  - `qa-sentinel.md` (+6.5 KB): clear gate-runner mandate (no AC checks, no
    code analysis), command output management, operating modes
    (Quick / Full / Deep), documentation coverage check.
  - `senior-researcher.md` (+2.7 KB): retrieval-optimized report format,
    structured reading notes, agent memory.
  - `REGISTRY.md` (+5.4 KB): expanded to 24 agents, Model Selection Matrix,
    QA Protocol, Domain Ownership map, doc-writing responsibility split.
- **REGISTRY.md** is the **single source of truth** for agent invocation
  routing. `agents/index.md` defers to it.
- **`coder` vs `doc-reviewer` responsibility split** — `coder` writes only
  minimal invariant doc stubs (API index, UI route, collection entry, env
  var, dependency, SSOT registry); `doc-reviewer` owns full doc writing,
  SSOT sync, and drift resolution.
- **Pre-commit vs Pre-PR build separation** (clarified in AGENTS.md): lint +
  typecheck + markdownlint at every commit; full `npm run build` and test
  suite only before opening a PR. This matches multi-hour build realities
  in larger projects.

### Removed

- **`obsidian-sync` agent** removed from the core framework. Mirroring docs to
  an external corpus (Obsidian / Notion / Confluence) is **project-specific**
  and should ship as a project-local agent. The `wiki-curator` agent still
  mentions external-vault sync as an optional coordination step.
- **Standalone `/qa` command file**. QA functionality remains available via
  the `qa-sentinel` agent and the `/new` skill's Phase 3.5; invocation
  patterns documented in REGISTRY.md § QA Protocol.

### Migration Guide (v1.1.0 → v2.0.0)

- Run `npx baldart update` from your project root. The framework files in
  `.framework/` are auto-updated via git subtree.
- The new `.claude/skills/` directory will land in your project after the
  update. Existing skills you authored locally are not touched.
- If you relied on the old `obsidian-sync` agent, move it to your project's
  `.claude/agents/` (it's a project-specific concern). Update any references
  in `AGENTS.md` to make the dependency explicit.
- The new `wiki-curator` agent expects a `docs/wiki/` directory and a
  frontmatter validator. See `framework/agents/llm-wiki-methodology.md` for
  the adoption checklist; the agent degrades gracefully if you don't ship a
  RAG layer.
- Calls to `Task tool` spawning subagents must now pass
  `mode: "bypassPermissions"` (the YOLO MODE protocol). Without this flag the
  subagent will pause for permission prompts and break the orchestration.

---

## [Unreleased]

### Added

### Changed

### Deprecated

### Removed

### Fixed

### Security

---

## Version Format

**MAJOR.MINOR.PATCH**

- **MAJOR**: Breaking changes (incompatible API updates, directory structure changes, removed features)
- **MINOR**: New features (backwards compatible additions)
- **PATCH**: Bug fixes (backwards compatible fixes)

## Change Categories

- **Added**: New features
- **Changed**: Changes in existing functionality
- **Deprecated**: Soon-to-be removed features
- **Removed**: Removed features
- **Fixed**: Bug fixes
- **Security**: Security fixes
