# SpecVerse AI Architecture

This is the implementation-architecture map for SpecVerse's AI layer. Companion to [SPECVERSE-AI.md](./SPECVERSE-AI.md), which is the user-facing config guide. If you want to *use* the AI layer, read SPECVERSE-AI.md first; if you want to *understand or modify* it, read this.

Last reviewed: 2026-06-04 (engines 6.97.12, self 5.21.6, assets 1.25.0).

---

## Scope

**Covers**: how a prompt invocation flows end-to-end; where the artefacts live and who consumes them; the five integration points where the same prompts are used; skill-aware system-prompt assembly; the provider-resolution decision tree; session-resume mechanics on the claude-cli path; the two-stage `create → verify-create` pattern; the eval-harness flow.

**Doesn't cover**: per-provider config knobs (see SPECVERSE-AI.md), how to add a new provider (see [SPECVERSE-EXTENDING.md](./SPECVERSE-EXTENDING.md)), the inference rule engine (see [SPECVERSE-ARCHITECTURE.md](./SPECVERSE-ARCHITECTURE.md) → "Composition Pipelines").

---

## The end-to-end journey of a prompt

```
                                              ┌──────────────────────────────────┐
USER ACTION                                   │  IN-CONVERSATION (no AI call)    │
e.g. `spv ai create "..."`                    │  Workflow prompts shown via MCP  │
       │                                      │  / Skill — Claude Desktop user   │
       ▼                                      │  triggers them by slash command  │
┌──────────────┐                              └──────────────────────────────────┘
│ CLI command  │  bootstrap/cli/commands/ai/*.ts
└──────┬───────┘
       │ instantiates
       ▼
┌──────────────────────┐    @specverse/engines/ai
│ BehaviorAIService    │    src/ai/behavior-ai-service.ts
│   or                 │    src/ai/commands/*.ts (template / spec-analyser / fill / suggest / enhance)
│ commands/template.ts │
└──────┬───────────────┘
       │ 1. loadPrompt(operation)         ──┐
       │ 2. assembleSystem(role+ctxt)       │
       │ 3. assembleUser(template, vars)    │   prompt-runner / behavior-ai-service
       │ 4. resolveModel({timeout})         │   src/ai/{prompt-loader,model-resolver}.ts
       │ 5. generateText({...})           ──┘   uses Vercel AI SDK
       ▼
┌──────────────────────┐
│ Vercel AI SDK        │  ai@^6 (`generateText`)
│ + LanguageModelV3    │
└──────┬───────────────┘
       │ delegates to provider
       ▼
┌─────────────────────────────────────────────────────────────────┐
│                                                                 │
│   claude-cli      anthropic      openai-compatible      stub    │
│   (Max plan,      (API key,      (DeepSeek/Together/    (no LLM,│
│    `claude        ephemeral      Ollama/vLLM/...)        emits  │
│    --print`)      prompt cache)  via openai-compat        prompt│
│                                                          back)  │
└─────────┬───────────────────────────────────────────────────────┘
          │ returns text
          ▼
   harness writes spec / CLI prints / behavior gets injected into generated code
```

The five integration points (right-hand side of the CLI box) all funnel through the same prompt YAMLs and the same `resolveModel`/`generateText` plumbing. Only the trigger and the consumer of the output differ.

---

## The artefacts and where they live

| Artefact | Source of truth | Loaded by | Used at |
|---|---|---|---|
| Prompt YAMLs — 11 as of assets 1.25.0: `create`, `verify-create`, `analyse`, `analyse-action`, `verify-analyse` (frozen — see note below), `behavior`, `attribute`, `realize-owner`, `realize-action`, `manifest`, `app-demo` | `@specverse/assets/prompts/core/standard/default/*.prompt.yaml` (composed from partials into `_composed/current/`; v9/ kept as a frozen baseline) | `prompt-loader.ts` (engines) and `prompt-runner.mjs` (eval harness) — both via `require.resolve('@specverse/assets/package.json')` | CLI / BehaviorAIService / generated MCP / skill installer / eval harness |
| AI guidance | `@specverse/entities/schema/SPECVERSE-SCHEMA-AI.yaml` | `{{aiSchemaPath}}` substitution in prompt templates; also a generated MCP resource | Every create/analyse/verify prompt |
| Prompt partials (DRY) | `@specverse/assets/prompts/core/standard/default/partials/*.md` | `prompt-runner.ts:loadPrompt` applies `inlinePartials()` at LOAD time, expanding `{{> partial-name}}` references before assembleSystem/assembleUser see content | The active prompts share the canonical partials (before-starting, output-format-{dual,single}-block, attribute-conventions, lifecycle-rules, invocation-discipline, verify-output-caveat, section-placement, deployment-canonical, steps-extraction) |
| Minimal reference example | `@specverse/entities/schema/MINIMAL-SYNTAX-REFERENCE.specly` | `{{referenceSchemaPath}}` and `{{referenceExamplePath}}` substitution | Same prompts |
| JSON schema | `@specverse/entities/schema/SPECVERSE-SCHEMA.json` | Parser + ajv validation; also a generated MCP resource | Validity battery, MCP `validate` tool, IDE/LSP |
| Skill | Assembled by `spv skill install` from the prompt YAMLs + `@specverse/entities/schema/*` + cli-reference into `~/.claude/skills/specverse/` (or project-local) | Claude Code at session start (auto-loaded if claude-cli is the active provider) | Any claude-cli interaction |
| MCP server | Generated by `spv realize` (template at `engines/libs/instance-factories/tools/templates/mcp/`); reads prompts at startup via `require.resolve('@specverse/assets/package.json')` | Claude Desktop / agents | Slash commands + resources at edit time |
| Eval harness | `specverse-demo-self/harness/` | `npm run eval:create` / `eval:analyse` | Benchmarking the prompts; runs flow through `prompt-runner.mjs` → `resolveModel` → provider |

The single source of truth is **`@specverse/assets`**. Everything that consumes a prompt resolves it via `require.resolve('@specverse/assets/package.json')` — both the in-tree engines code and the generated MCP server. A prompt edit shipped in `assets@x.y.z` reaches every consumer through normal `npm install`. (Before assets@1.0.0 the prompts lived in `engines/assets/prompts/`; the legacy fallback path is still present in `prompt-loader.ts` for older installs.)

---

## The five integration points

All five load the same prompt YAMLs from `@specverse/assets`. Only the trigger and the consumer differ.

### 1. CLI direct (`spv ai create`, `spv ai template`, etc.)

`bootstrap/cli/commands/ai/*.ts` instantiates one of the command classes in `engines/src/ai/commands/`. Those classes load the appropriate prompt via `prompt-loader.ts`, fill template variables, call `resolveModel()` + `generateText()`, and write the result to stdout or a file.

### 2. Autonomous behaviour generation (`BehaviorAIService`)

When `spv realize` synthesises behaviour-step bodies, `BehaviorAIService` (`engines/src/ai/behavior-ai-service.ts`) is the engine. It loads `behavior.prompt.yaml` once at construction time, then for each behaviour step in the spec it fills the template, resolves a model via the same `resolveModel`, and asks the LLM (or the stub) to produce the function body. The result is injected into the generated code with cache markers so it survives re-realize cycles.

### 3. Generated MCP server

`spv realize` runs the MCP factory (`engines/libs/instance-factories/tools/templates/mcp/`) which emits a Node MCP server into `generated/code/tools/specverse-mcp/`. At startup the generated server walks `@specverse/assets/prompts/core/standard/default/` and registers each YAML as an MCP `prompts/*` endpoint, plus the schema files as resources. Claude Desktop renders these as slash commands. **The MCP server itself does not call any LLM** — it returns the prompt text for Claude Desktop to execute (which is why the resolver picks `stub` when `MCP_SERVER=1`).

### 4. Claude skill (`spv skill install`)

`bootstrap/cli/commands/skill.ts` materialises a skill at `~/.claude/skills/specverse/` (or project-local with `--project`). It reads the same prompt YAMLs, plus the schema files from `@specverse/entities`, plus a static cli-reference, and assembles:

- `SKILL.md` — the entry point
- `reference/{guide,schema,ai-guidance,minimal-example,cli-reference}.md`
- `workflows/{create,verify-create,analyse,verify-analyse,behavior,app-demo}.md` — one per prompt YAML, formatted as a Claude-skill workflow (inputs table + role + instructions + task template)

Claude Code auto-loads any installed skill at session start. **The skill is not the prompts running** — it's a primer that gives Claude Code background context on SpecVerse, plus discoverable workflows it can suggest invoking.

### 5. Eval harness

`specverse-demo-self/harness/run-{create,analyse}.mjs` calls `runPrompt()` from `harness/lib/prompt-runner.mjs`, which is a parallel implementation of the engines path: it loads the same YAML, assembles system + user, calls `resolveModel`, calls `generateText`. Identical plumbing, but instrumented for benchmarking (writes prompts before the LLM call, captures duration + tokens, scores the output).

---

## The three assembly steps

Inside each integration point, the same three steps run before any LLM call:

### Step 1 — `loadPrompt(operation)`

Resolves `@specverse/assets/package.json`, joins to `prompts/core/standard/default/${operation}.prompt.yaml`, parses the YAML. Falls back to `@specverse/engines/assets/prompts/...` if the assets package isn't installed (legacy path; useful for old installs).

The YAML follows one canonical shape (since assets@1.0.0):

```yaml
name: create
version: "9.3.0"
system:
  role:           ...   # always present
  minimalContext: ...   # optional; loaded when skill is detected
  fullContext:    ...   # optional; loaded when skill is NOT detected
  context:        ...   # legacy older shape; still supported
user:
  template:       ...   # has {{var}} placeholders
  variables:
    - { name, type, required, description }
```

### Step 2 — `assembleSystemPrompt()` (skill-aware)

```ts
if (claude-cli is the active provider AND ~/.claude/skills/specverse/ exists) {
  // Skill is going to load — fullContext primer would be redundant
  return [role, minimalContext].join('\n\n');
} else {
  // No skill mechanism available (anthropic/openai-compat/stub)
  return [role, fullContext].join('\n\n');
}
```

This is a real cost saving — `fullContext` is ~2K tokens of "what is SpecVerse, here are the four pillars, here's the canonical example" primer. When the skill is going to deliver that anyway, sending it in the system prompt would double-charge. The detection lives in `engines/src/ai/skill-detection.ts`.

### Step 3 — `resolveModel({ timeout })` and `generateText({ model, system, prompt: user, … })`

Provider resolution (next section) returns a `LanguageModelV3` instance. `generateText` from Vercel AI SDK runs it. The harness uses `providerOptions: { anthropic: { cacheControl: 'ephemeral' } }` for the API path; that's a no-op on claude-cli (which has its own session-resume mechanism).

---

## Provider resolution — the four-mode decision tree

`resolveProviderId()` in `engines/src/ai/model-resolver.ts`:

```
1. SPECVERSE_AI_PROVIDER set explicitly?
   → claude-cli | anthropic | openai-compatible | stub
   (unrecognized values throw; see model-resolver.ts:30)

2. MCP_SERVER=1?
   → stub
   (rationale: the ambient Claude Desktop is the executor; calling a
    separate API would be wasteful and would lose context)

3. Is the `claude` binary available + responds to --version?
   → claude-cli
   (default for Max users; checks ~/.claude/local/claude then PATH)

4. ANTHROPIC_API_KEY set?
   → anthropic

5. Else
   → stub (with diagnostic comment in output)
```

Each provider implements `LanguageModelV3` from `@ai-sdk/provider`. The Vercel AI SDK plumbing is identical across them; only the per-call mechanics differ:

- **claude-cli** — spawns `claude --print` as a subprocess. First call uses `--session-id <uuid>` + `--system-prompt <s>`; subsequent calls on the same model instance use `--resume <uuid>`. No marginal cost on Max but rate-limited by the 5-hour window quota. Lives in `engines/src/ai/providers/claude-cli.ts`.
- **anthropic** — `@ai-sdk/anthropic`. `cacheControl: 'ephemeral'` enables Anthropic's 5-minute prompt cache on the API.
- **openai-compatible** — `@ai-sdk/openai-compatible` with a configurable `baseURL`. Unlocks DeepSeek / Together / Ollama / vLLM. Cheap; no caching unless the endpoint supports it.
- **stub** — emits the prompt text verbatim, prefixed with explanation comments. Used in MCP, in CI smoke tests, and as the diagnostic fallback. Lives in `engines/src/ai/providers/stub.ts`.

---

## Session-resume mechanics (claude-cli path)

Important detail: the claude-cli provider holds session state **inside the `claudeCli()` factory closure**:

```ts
export function claudeCli(options): LanguageModelV3 {
  const sessionId = randomUUID();
  let initialized = false;
  // ...
  doGenerate(opts) {
    if (!initialized) {
      args.push('--session-id', sessionId, '--system-prompt', system);
    } else {
      args.push('--resume', sessionId);   // system prompt re-read from cached session
    }
    // ...
    initialized = true;
  }
}
```

Consequence: **session caching only fires across calls on the *same model instance***. Caller code that builds a fresh model per call (e.g. by calling `resolveModel()` in a loop) defeats the cache and re-tokenises the full system prompt each time.

**The eval harness explicitly hoists one model per run** (`harness/lib/prompt-runner.mjs::buildRunModel`, called once in `run-create.mjs` / `run-analyse.mjs`'s `main()` and threaded through every case). Empirical effect: an 8-case create run that rate-limited at case 2 yesterday completes cleanly today.

`BehaviorAIService` does the same — it constructs one model in `beginGeneration()` and reuses it across all behaviour steps in a realize pass.

CLI commands that issue a single call don't benefit from session-resume (one call doesn't have a cache target), but they don't pay a cost either.

---

## The two-stage `create → verify-create` pattern (new 2026-04-26)

Empirical observation: the `create` prompt consistently misses **branch states** in lifecycles (cancelled, lost, overdue, reopened) — even when the spec mentions them explicitly in invariants/behaviors. Same prompt, same model, same case scored 67% one run and 67% the next; failure was deterministic, not noise.

Adding a second pass — same prompt library, new YAML — fixes this without touching the upstream `create` prompt:

```
NL request
   │
   ▼
[create prompt] ──► first-draft .specly  (linear flow, branch states often missing)
   │
   ▼
[verify-create prompt] ──► final .specly  (checklist applied; branch states recovered)
   │
   ▼
score(final)
```

The verify prompt receives the original requirements + the draft + the schema/reference paths via `{{aiSchemaPath}}`/`{{referenceSchemaPath}}`. Its checklist explicitly enumerates the failure modes from the eval (`cancelled` / `lost` / `overdue` / `reopened`; transitions named after recovery actions imply destination states). Output is the FINAL spec, not a diff.

Mechanically this is just two `runPrompt()` calls on the **same model instance** — second call resumes the cached session, so the system-prompt cost is paid only on the first draft. Net cost on Max: ~2× output tokens (no input-side overhead after pass 1).

**Analyse verification took a different route.** The two-stage LLM shape above is create-only. Analyse was *expected* to mirror it via `verify-analyse.prompt.yaml`, but that monolithic LLM re-pass proved destructive on weak models (it could capture chat prose as the spec) — so engines 6.92→6.93 **retired the `verify-analyse` LLM re-pass and repointed `spv ai analyse --verify` to a deterministic, rule-driven pass** (`engines/src/ai/analyse-verify/`, the analyse-side analog of realize's post-emit-verify): named verifiers (`ANL-STEP-CODE`, `ANL-EVENT-NONDOMAIN`, `ANL-REQUIRE-EXISTENCE`, …) flag findings, then surgically re-fire only the affected action via the `analyse-action` prompt's `{{correction}}` slot. The `verify-analyse.prompt.yaml` file remains in `assets/` but is **frozen** — kept as a deterministic surface for the behaviour cross-check, not run as a second full pass. Net: there is no `create → verify-create`-style LLM second pass for analyse.

Implementation:

- Prompt: `assets/prompts/core/standard/default/verify-create.prompt.yaml`
- Harness wiring: `harness/run-create.mjs::runOneCase` calls `runPrompt({ operation: 'verify-create', model, ... })` after the draft pass; `--skip-verify` falls back to single-pass for ablation studies.
- Future: when verify proves out for analyse too, the same pattern can be added to `BehaviorAIService` and the CLI commands.

---

## Eval harness specifics

The harness lives in **specverse-demo-self** (separate repo so it can install `@specverse/*` from npm like a real consumer). It's the empirical feedback loop for prompt iteration.

```
specverse-demo-self/
├── corpus/
│   ├── create/   *.yaml   — NL request + expected_entities/relationships/lifecycles
│   └── analyse/  *.yaml   — source_setup script + expected_spec_path
├── harness/
│   ├── run-create.mjs     — two-stage create flow + scoring
│   ├── run-analyse.mjs    — analyse flow + round-trip scoring
│   ├── lib/
│   │   ├── prompt-runner.mjs   — buildRunModel + runPrompt; uses @specverse/assets
│   │   ├── corpus.mjs          — loadCreateCorpus / loadAnalyseCorpus
│   │   └── reporter.mjs        — buildRunId, aggregate, printSummary
│   └── scoring/
│       ├── validity-battery.mjs    — parse / schema / infer / realize gates
│       ├── domain-coverage.mjs     — for create; alias-aware
│       └── round-trip.mjs          — for analyse; structural diff
├── runs/<run-id>/<case>/  — per-case artifacts (system-prompt, user-prompt, draft, generated, raw-output)
└── reports/<run-id>.json
```

The run-id encodes everything that affects the result: timestamp + provider + model + prompt-version, e.g. `2026-04-26T07-31-30__create-v9__claude-cli__default`. That's the reproducibility key.

For prompt iteration without re-publishing assets every cycle: the assets package files can be edited in the `specverse-self` workspace and copied into `specverse-demo-self/node_modules/@specverse/assets/prompts/...` for testing. Once changes are validated, publish `@specverse/assets@x.y.z` for the rest of the ecosystem to pick up.

---

## Where to make changes

| Change | Edit here | Effect on consumers |
|---|---|---|
| Add a new prompt | `specverse-self/assets/prompts/core/standard/default/<name>.prompt.yaml` (+ mirror in `v9/`) | Reaches CLI, MCP, skill, eval after `npm install @specverse/assets@<new>` |
| Tune an existing prompt | Same path | Same |
| Add a new provider | `engines/src/ai/providers/<name>.ts`; wire into `model-resolver.ts` | Available everywhere `resolveModel()` is called; ships in next engines release |
| Add a new env-var trigger | `model-resolver.ts::resolveProviderId()` | Same |
| Add a new schema reference variable (`{{newVarPath}}`) | `prompt-runner.mjs::SCHEMA_VAR_PATHS` (eval) + `prompt-loader.ts` (engines path); place file in `entities/schema/`; bump entities | Available in templates after publish ritual |
| Add a new skill workflow | Add the prompt YAML + run `spv skill install --target=...` | Skill is regenerated locally; published when assets ships |
| Add a verify-style second pass | New YAML in `assets/prompts/...` + harness change in `run-{create,analyse}.mjs` | Verify pattern; mirror the `create → verify-create` shape |

---

## See also

- [SPECVERSE-AI.md](./SPECVERSE-AI.md) — user-facing config: provider modes, env vars, quick-start scenarios, skill interaction, FAQ
- [SPECVERSE-ARCHITECTURE.md](./SPECVERSE-ARCHITECTURE.md) — overall system architecture (parser, inference, realize, composition pipelines)
- [SPECVERSE-EXTENDING.md](./SPECVERSE-EXTENDING.md) — adding new entity types / engines / factories
- [docs/plans/2026-04-24-AI-PROVIDER-REPLATFORM.md](../plans/2026-04-24-AI-PROVIDER-REPLATFORM.md) — design rationale for the provider layer
- [docs/plans/2026-04-25-CONTENT-PACKAGE-REFACTOR.md](../plans/2026-04-25-CONTENT-PACKAGE-REFACTOR.md) — why `@specverse/assets` exists; what moved
- [docs/plans/2026-04-25-SPEC-QUALITY-BENCHMARK.md](../plans/2026-04-25-SPEC-QUALITY-BENCHMARK.md) — eval-harness design + roadmap
- [docs/proposals/in-progress/2026-04-26-STRUCTURAL-PREPASS-FOR-ANALYSE.md](../proposals/in-progress/2026-04-26-STRUCTURAL-PREPASS-FOR-ANALYSE.md) — proposal to add a deterministic structural pre-pass to analyse (ties into the integration point #5 here)
- [specverse-demo-self/RUNBOOK.md](../../../specverse-demo-self/RUNBOOK.md) — operational guide for the eval harness
