# SpecVerse — analyse-pipeline extraction decision tree

`spv ai analyse <source>` reads a source codebase and produces a faithful `.specly` mirror plus an `implementation.yaml` manifest. The pipeline is layered: each layer reads cheaper, higher-fidelity signals before falling back to more inferential ones, and each layer's output is forward-logged to a sidecar so downstream stages can audit (and re-run) the extraction without re-paying its cost.

This document describes the extraction decision tree — which adapter fires when, what each is responsible for, what signal it consumes, and what fallback runs when it produces nothing.

---

## Pipeline overview

```
.specly source code
    │
    ▼
┌───────────────────────────────────────────────────────────────────────┐
│ Stage 1: Extract (analyse-prepass)                                    │
│                                                                       │
│  1a. Backend selection         (GitNexus → CodeGraph → grep-only)     │
│  1b. ORM adapters              (typescript-prisma, typescript-drizzle…)│
│  1c. Decorator adapter         (typescript-decorators)                │
│  1d. Interfaces adapter        (typescript-interfaces)                │
│  1e. Express-routes adapter    (router-file → synthetic controllers)  │
│  1f. Component suggestions     (structural-detector → community)      │
│  1g. Candidate-step extraction (per kept method)                      │
│                                                                       │
│  → SpecVerseFacts { entities, relationships, lifecycles,              │
│                      candidateMethods, suggestedComponents, ... }     │
│  → facts/extract-decisions.json (per-class drop-reason audit)         │
│  → facts/extract-stats.json                                           │
│  → facts/source-fingerprint.json (SHA256 over input tree)             │
└───────────────────────────────────────────────────────────────────────┘
    │
    ▼
┌───────────────────────────────────────────────────────────────────────┐
│ Stage 2: Skeleton (deterministic)                                     │
│                                                                       │
│  - emitFaithfulSkeleton(facts) → spec yaml                            │
│  - Single-home entity mapping (each entity → exactly one component)   │
│  - Action stubs only — no bodies (LLM fills via micro-call)           │
│  - Forward-logged: every yaml line traces to a fact reference         │
│                                                                       │
│  → specs/skeleton.specly                                              │
│  → facts/skeleton-provenance.json                                     │
└───────────────────────────────────────────────────────────────────────┘
    │
    ▼
┌───────────────────────────────────────────────────────────────────────┐
│ Stage 3: Microcalls (parallel LLM)                                    │
│                                                                       │
│  - One LLM call per action stub (concurrency-bounded)                 │
│  - Manifest call (separate)                                           │
│  - Surgical retry on validation errors (re-fire only failed stubs)    │
│  - Post-process: yaml-escape, manifest specVersion, events injection  │
│                                                                       │
│  → specs/main.specly (final spec)                                     │
│  → manifests/implementation.yaml                                      │
│  → llm-output/microcall-decisions.json                                │
└───────────────────────────────────────────────────────────────────────┘
    │
    ▼
┌───────────────────────────────────────────────────────────────────────┐
│ Stage 4: Verify + Realize                                             │
│                                                                       │
│  - spv validate (parse + schema + structural invariants)              │
│  - spv infer (rule engine, default expansion)                         │
│  - spv realize all (instance factories generate code)                 │
│                                                                       │
│  → verify/invariants.json + verify/retries.json                       │
│  → realize-decisions.json + realize-coverage.json                     │
│  → REPORT.md (assembled from all sidecars)                            │
└───────────────────────────────────────────────────────────────────────┘
```

The remainder of this document focuses on **Stage 1** — the extraction decision tree that determines what facts the rest of the pipeline operates on.

---

## Stage 1a — Backend selection

The structural backend is whatever can resolve symbols across the source tree. Backends differ in capability (call graphs, file watching, community detection) and external-tool dependency (GitNexus needs a separate binary).

`selectBackend('auto')` resolves to:

| Order | Backend | When it wins | What it provides |
|---|---|---|---|
| 1 | `GitNexus` | Available (`gitnexus` on `$PATH`) | Full call graph, cross-file resolution, Leiden community detection |
| 2 | `grep-only` | Always | Symbol listing via filesystem walk + regex; no graph |

`CodeGraph` was the original auto-default; deprecated in 6.21.1 because its CLI 0.7+ schema change broke the in-process queries. Still selectable via explicit `--backend codegraph`.

---

## Stage 1b — ORM adapters (entity + relationship extraction)

Each ORM adapter detects its own applicability by reading the source tree, then emits entity + relationship facts.

| Adapter | Detects on | Source signal | Fact slots produced |
|---|---|---|---|
| `typescript-prisma` | `**/schema.prisma` exists | Prisma schema DSL | `entities[]`, `relationships[]`, `lifecycles[]` (from `@@enum` field types), `auto`/storage profile |
| `typescript-drizzle` | a `pgTable(...)` declaration (via `drizzle.config` / layout globs) | `pgTable(...)` calls | same — typed columns, `.references()` FKs (+inverse hasMany, cascade), `pgEnum` lifecycles, dialect→storage profile (engines 6.97.0+) |
| `typescript-sequelize` | (planned, deferred) | `sequelize.define(...)` | same |

If an ORM adapter detects entities, the decorator + interfaces adapters still run but skip names already present.

**ORM precedence (engines 6.97.0+).** When a *complete-schema* ORM adapter (prisma / drizzle) declares the entity set, the typescript-interfaces class-walker may **not mint new entities** — it would otherwise promote service/utility CLASSES (scrapers, coordinators, managers) to entities. A suppressed method-bearing class is instead walked as a **service** (its methods → service operations) rather than dropped. So adapter-detected ORM tables win; the class-walker fills only genuine gaps. (Adapter-free codebases keep the original gap-filling behaviour: ORM > decorators > interfaces.)

**Identity preservation (engines 6.97.x).** ORM adapters carry id-generation semantics through to the spec instead of discarding them and re-guessing uuid: `@default(uuid())`/`autoincrement()`/`cuid()`/`now()` and Drizzle `serial`/`.defaultNow()` map to `auto=uuid4`/`auto=autoincrement`/`auto=now`; the source PK type is preserved (not coerced to String). A primary key with **no** generator (from a definitive prisma/drizzle parse) is emitted as **`auto=none`** — the externally-owned-key marker (case 7), so realize requires it rather than inventing a uuid. See `proposals/implemented/2026-06-02-IDENTITY-IN-SPECVERSE-ANALYSIS.md`.

---

## Stage 1c — `typescript-decorators` adapter

For decorator-based codebases (TypeORM, NestJS, MikroORM). Detects `@Entity()`, `@Column()`, `@OneToMany()`, etc.

| Source signal | Fact emitted |
|---|---|
| `@Entity()` on a class | `entities[]` entry with the class name |
| `@Column()` on a field | attribute on the entity |
| `@Column({ enum: PaymentStatus })` + `enum PaymentStatus { ... }` | `lifecycles[]` entry with states from the enum members |
| `@OneToMany`, `@ManyToOne`, `@OneToOne`, `@ManyToMany` | `relationships[]` entry |
| `this.<statusField> = <enumMember>` inside a method body | lifecycle `transition` (`from` inferred from the enclosing `if`/guard, `action` from method name) |

When this adapter fires (decorator-rich codebases), lifecycles are recovered with both **states AND transitions**. Higher fidelity than Stage 1d.

---

## Stage 1d — `typescript-interfaces` adapter (engines 6.27.6+)

For adapter-free codebases that declare types as raw `class` / `interface` / `type` declarations (no decorators, no schema files). Idle-meta-style.

### Phase 1d-1: Entity detection

| Signal | Heuristic |
|---|---|
| Class / interface has an `id` field | **Liberal heuristic (default):** entity |
| Class / interface has `id` AND a timestamp field (`createdAt` / `updatedAt` / `_id`) | **Conservative heuristic** (opt-in via `--entity-shape conservative`): entity |
| Class name ends in `Service` / `Controller` / `Adapter` / `Engine` / etc. | NOT an entity (excluded explicitly to avoid constructor-DI false positives) |

### Phase 1d-2: Relationship inference

For each entity field whose declared type matches another entity:

| Field shape | Relationship |
|---|---|
| `field: T[]` where `T` is an entity | `hasMany` |
| `field: T` where `T` is an entity | `belongsTo` |

### Phase 1d-3: Lifecycle inference (engines 6.27.6+)

The adapter scans for **literal-string union type aliases** like:

```typescript
type PaymentStatus = "pending" | "confirmed" | "failed";
```

When an entity has a non-array field whose declared type matches such an alias:

| Source pattern | Lifecycle emitted |
|---|---|
| `type Status = "a" \| "b" \| "c"` + `interface E { status: Status }` | `{ model: E, field: status, states: [a, b, c], transitions: [] }` |
| `type Status = "a" \| number` (mixed-type union) | NOT emitted (only pure string-literal unions) |
| `tags: Tag[]` (array-typed) | NOT emitted (lifecycles are scalar by convention) |
| `status: Status \| null` | states extracted, `\| null`/`\| undefined` ignored |

Transitions are intentionally NOT inferred at this layer — that needs method-body analysis (which the decorator adapter does have, via `method-body-walker.ts`'s `this.<field> = <enumMember>` detector). On adapter-free codebases the spec ends up with states-only lifecycles; the realize stack treats them as enum-typed status fields without state-transition guards.

---

## Stage 1e — `express-routes` adapter (engines 6.27.1+)

For router-file codebases (Express, Fastify-with-routes-pattern). idle-meta uses this — it has `routes/auth.ts`, `routes/games.v3.ts`, `routes/players.ts` instead of class-based `*Controller`s.

### Detection

A file in `**/routes/**` qualifies when BOTH:
- Contains `import { Router }` from express OR a `Router()` call
- Has at least one `router.<verb>(path, handler)` call where verb ∈ `{get, post, put, patch, delete, all, head, options}`

### Synthesis

Each qualifying route file becomes one synthetic controller in `candidateMethods`:

| Source signal | Fact emitted |
|---|---|
| File `routes/auth.ts` | Controller named `AuthController` |
| File `routes/games.v3.ts` | Controller named `GamesV3Controller` |
| File `routes/user-profile.ts` | Controller named `UserProfileController` |
| File `routes/FooController.ts` | Controller named `FooController` (no double-suffix) |

Each `router.<verb>(path, handler)` becomes one method (action). Action-name derivation, in priority order:

1. Named handler reference (`router.post('/x', registerHandler)`) → strip `Handler` suffix → `register`
2. Single-segment lowercase path (`router.post('/login', ...)`) → `login`
3. Verb + path PascalCase (`router.get('/users/:id', ...)`) → `getUsers` (`:id` parameters skipped)
4. Fallback: `<verb><N>` where N is the handler ordinal

Handler bodies (inline arrow functions) are extracted verbatim and run through `extractCandidateSteps`, which means the candidate-step timeline includes all `validate` / `find` / `create` / `emit` / etc. operations recognised in the body.

Skeleton emitter then routes `*Controller` suffix to `controllers:` block and emits a `model: <X>` reference.

### Smart `model:` resolution

The schema requires every controller to declare `model: <ModelName>` referencing a real entity. For synthesized route-file controllers, the stripped name (`Auth` from `AuthController`) usually isn't a detected entity. Resolution priority:

1. Strip `Controller` suffix and match a detected entity by name (works for `UserController` → `User`)
2. Scan the controller's handler bodies for entity-name references; pick the most-frequent (`AuthController` → `User` because handlers reference `User.findOne`, etc.)
3. Fall back to the first entity assigned to the same component (via `entityHome` mapping)
4. As a last resort, omit the line — schema validator complains `missing required` but the spec doesn't reference a non-existent model

---

## Stage 1f — Component suggestions

Three sources, ordered by **author-stated-intent precedence**:

```
                  ┌──────────────────────────────────┐
                  │ structural detector (highest)    │
                  │ - sub-package.json files         │
                  │ - apps/* / services/* / etc      │
                  └─────────────┬────────────────────┘
                                │ ≥ 2 suggestions?
                                ▼ no
                  ┌──────────────────────────────────┐
                  │ entity-graph Louvain             │
                  │ - facts.entities ≥ 5             │
                  │ - cluster relationship graph     │
                  └─────────────┬────────────────────┘
                                │ enough entities?
                                ▼ no
                  ┌──────────────────────────────────┐
                  │ empty (LLM falls back to its own │
                  │ judgement — high variance)       │
                  └──────────────────────────────────┘
```

Why structural beats community detection: package names like `idle-clicker-mobile` are **author-chosen architectural identifiers**. Louvain on a synthetic single-package graph is just clustering noise.

After structural buckets are chosen, entities are assigned to them by **filePath prefix matching**. Each component's `structural.sourceDir` determines which entities live inside (entities in `apps/editor/src/...` go to the `Editor` component).

---

## Stage 1g — Candidate-step extraction

For every `Controller` / `Service` / `Handler` / `Engine` / `Manager` / `Evaluator` / `Adapter` / `Cache` / `Provider` / `Factory` / `Repository` / `Store` / `Resolver` / `Worker` / `Job` / `Listener` / `Subscriber` class the backend can list, walk each method body and classify it.

### 3-level method-shape filter (engines 6.21.4+)

Each method is evaluated against three resolution levels:

| Level | What it is | Where it's resolved |
|---|---|---|
| 0 | Container-op accessor (`this.<x>.{get,set,has,delete,clear,size,...}`) | Skipped — not a business action |
| 1 | Single-CURVED-shape method (`find / create / update / delete` with no other shape) | Skipped — instance generator handles via convention |
| 2 | Multi-step business action with recognized step phrases | Convention library auto-handles ("Look up X by Y", "Validate X", etc.) |
| 3 | Business action with unconventional step phrases | AI fallback — generated `*.ai.ts` body |

Levels 0 and 1 are dropped from `candidateMethods`. Levels 2 and 3 are kept and fed to the LLM micro-call as candidate-step timelines.

### Drop-reason audit (`facts/extract-decisions.json`)

Every dropped method records its reason in an 8-value enum:

- `level-0-accessor` — `this.<X>.<container-op>` only
- `level-1-curved-create-update-delete` — single CRUD shape
- `level-1-curved-retrieve` — single retrieve shape
- `container-op-or-self-delegation-only` — `this.<helper>(...)` self-delegation
- `no-candidates-classified` — body is opaque (no recognised steps)
- (others)

This audit replaces post-hoc heuristic reconstruction; sidecar tools read the structured record directly.

---

## What runs when — example matrix

| Codebase shape | 1b ORM | 1c Decorators | 1d Interfaces | 1e Routes | Lifecycles recovered? |
|---|---|---|---|---|---|
| Prisma + raw routes (e.g. cal-com-style) | ✓ entities | — | gap-fills | ✓ controllers | from `@@enum` fields |
| NestJS + TypeORM + decorator-rich | — | ✓ everything | redundant | (skipped — class controllers exist) | from enum columns + transition methods |
| Raw TS + Express routes (idle-meta-style) | — | — | ✓ entities | ✓ controllers | from literal-union aliases (states only, no transitions) |
| Raw TS + class controllers + no decorators | — | — | ✓ entities | — (no `routes/`) | from literal-union aliases |
| Adapter-free + no `id` fields | — | — | — | depends | no |

---

## Forward-logging discipline

Every layer records its decisions structurally **at the moment they happen**. No layer reconstructs decisions from scratch by re-walking sources or re-classifying.

Sidecars produced per run:

| Sidecar | Source layer | What it answers |
|---|---|---|
| `facts/extract-decisions.json` | 1g (audit) | Why was class X walked? Why was method Y dropped? |
| `facts/extract-stats.json` | 1g aggregates | How many methods kept vs dropped per category? |
| `facts/source-fingerprint.json` | run identity | Are two runs reproducing the same source state? |
| `facts/skeleton-provenance.json` | Stage 2 | Why did entity X land in component Y? Which fact backs spec yaml line N? |
| `llm-output/microcall-decisions.json` | Stage 3 | What was the prompt hash, latency, parsed body, failure for action Z's micro-call? |
| `verify/invariants.json` + `verify/retries.json` | Stage 4 | Which structural-check invariants passed? What did each retry change? |

Forward-logging is the architectural shift that lets reports regenerate from sidecars deterministically — `scripts/build-report.mjs` reads these and emits markdown without re-paying any extraction cost.

---

## Adding a new adapter

The path follows existing precedent:

1. Create `engines/src/analyse-prepass/adapters/<name>.ts` exporting `extract<Name>(prepass, options): Promise<Facts>`.
2. Wire into `analyse-prepass/index.ts` after the relevant precedence boundary (e.g. ORM adapters before decorators, decorators before interfaces).
3. Merge results into `facts.entities` / `relationships` / `lifecycles` / `candidateMethods` with first-detector-wins by name.
4. Push into `facts.meta.adaptersRun` for visibility.
5. Add tests under `__tests__/<name>-*.test.ts` with an in-memory `StructuralPrepass` stub.
6. Document the precedence + signal in this file's matrix.

---

## See also

- `docs/plans/2026-05-04-ANALYSE-VIA-FAITHFUL-SKELETON-AND-MICROCALLS.md` — the architectural shift this pipeline embodies
- `docs/plans/2026-05-03-PIPELINE-AUDIT-TRAIL-INSTRUMENTATION.md` — forward-logged sidecar specification
- `engines/src/analyse-prepass/adapters/typescript-decorators.ts` — reference implementation of the decorator path
- `engines/src/analyse-prepass/adapters/typescript-interfaces.ts` — reference implementation of the adapter-free path
- `engines/src/analyse-prepass/adapters/express-routes.ts` — reference implementation of the router-file path
- `engines/src/analyse-prepass/structural-detector.ts` — package.json + apps/*/services/* topology walker
