---
name: erp-kit-app-data
description: Scaffold seed data and gql-ingest scenarios for an erp-kit app. Use after backend implementation (step 5) or standalone when the app has deployed resolvers. Triggers on "generate test data", "scaffold scenarios", "ingest data", "create test fixtures".
disable-model-invocation: true
metadata:
  erp-kit-version: "0.59.0"
---

# Scaffold Seed & Ingest Scenarios

Generate seed data (master/reference) and gql-ingest scenarios (application/test data) for an erp-kit app. Seed goes through seedPlugin (direct DB). Scenarios go through GraphQL resolvers via gql-ingest.

## Version Check

Run `npx erp-kit internal measure versions` from the repo root. If `status` is `"violations"`, relay the findings (each states its own fix) and stop; otherwise proceed.

## When to Use

- After `erp-kit-app-5-impl-backend` when the backend is deployed and resolvers work
- Standalone when you have an app with working resolvers and want test data
- When the user asks for test data, scenarios, fixtures, or demo data for their app

## Prerequisites

- App backend is deployed with working GraphQL endpoint
- Resolver docs exist at `<app>/docs/resolver/`
- Business flow docs exist at `<app>/docs/business-flow/` (optional, for scenario generation)
- `@jackchuka/gql-ingest` v4.2+ is installed (supports `outputCapture` and `$ref`)

## Workflow

```
CONTEXT → SEED → SHARED ARTIFACTS → SCENARIOS → VERIFY
```

### Phase 1: Gather Context

1. Collect `APP_PATH` from the user (the app directory, e.g., `apps/my-app`)
2. Collect `ENDPOINT` — the deployed GraphQL endpoint URL
3. Read all resolver docs at `<APP_PATH>/docs/resolver/*.md`
4. Read business flow docs at `<APP_PATH>/docs/business-flow/*.md` (if they exist)
5. Introspect the live GraphQL schema:
   ```bash
   curl -s -X POST $ENDPOINT \
     -H "Content-Type: application/json" \
     -H "Authorization: Bearer $TOKEN" \
     -d '{"query":"{ __schema { mutationType { fields { name args { name type { name kind ofType { name kind ofType { name kind } } } } } } } }"}' \
     | jq '.data.__schema.mutationType.fields'
   ```
6. Build entity list: for each resolver doc, identify the entity name, mutation name, and input type fields from the introspection result
7. Build dependency graph: from resolver input types, identify foreign key fields (fields ending in `Id` that reference other entities)

### Phase 2: Generate Seed

1. Run existing seed generation:
   ```bash
   pnpm erp-kit app generate seed -p $APP_PATH
   ```
2. Validate the generated JSONL files exist at `<APP_PATH>/backend/seed/data/`

### Phase 3: Generate Shared Artifacts

Create the `ingest/` directory structure at `<APP_PATH>/backend/ingest/`. See [directory structure reference](references/directory-structure.md).

#### 3a: config.yaml

Generate `ingest/config.yaml` with entity dependencies derived from Phase 1 analysis. See [directory structure reference](references/directory-structure.md) for the config format.

#### 3b: mutations/

Generate one `.graphql` file per resolver mutation. See [mutation generation reference](references/mutation-generation.md).

> **Parallelize if possible:** Each mutation file is independent. Dispatch one agent per resolver to generate its mutation file if there are many resolvers.

### Phase 4: Generate Scenarios

Generate scenario directories under `ingest/scenarios/`. Always generate a `baseline` scenario first.

#### 4a: baseline

The baseline scenario contains the minimal data needed to make the app functional:

- One record per entity (or the minimum needed for referential integrity)
- **No hardcoded UUIDs** — use `outputCapture` on entities to capture server-generated IDs, and `$ref` in downstream entities to reference them
- Covers all entities that have resolvers

See [scenario scaffolding reference](references/scenario-scaffolding.md) for data generation patterns.

#### 4b: Business flow scenarios

For each business flow doc at `<APP_PATH>/docs/business-flow/*.md`:

1. Read the flow and identify which entities and operations it exercises
2. Create a scenario directory named after the flow (kebab-case)
3. Generate data that exercises the flow end-to-end
4. **Use `$ref` to reference baseline entities** — never hardcode UUIDs from other scenarios
5. Run business flow scenarios together with baseline in a single invocation so the output store is shared

> **Parallelize if possible:** Each scenario is independent. Dispatch one agent per business flow to generate its scenario.

#### 4c: Mapping configs

For each scenario, generate mapping JSON files. See [mapping config reference](references/mapping-config.md).

### Phase 5: Verify

1. Check all JSONL files parse correctly:

   ```bash
   for f in $APP_PATH/backend/ingest/scenarios/*/*/data.jsonl; do
     echo "Checking $f..."
     node -e "require('fs').readFileSync('$f','utf-8').trim().split('\n').forEach((l,i) => { try { JSON.parse(l) } catch(e) { console.error('Line '+(i+1)+': '+e.message); process.exit(1) } })"
   done
   ```

2. Check all entity.json files reference valid mutation and data files:

   ```bash
   for f in $APP_PATH/backend/ingest/scenarios/*/*/entity.json; do
     node -e "
       const f = '$f';
       const m = JSON.parse(require('fs').readFileSync(f,'utf-8'));
       const dir = require('path').dirname(f);
       const data = require('path').resolve(dir, m.dataFile);
       const gql = require('path').resolve(dir, m.graphqlFile);
       if (!require('fs').existsSync(data)) { console.error('Missing: '+data); process.exit(1) }
       if (!require('fs').existsSync(gql)) { console.error('Missing: '+gql); process.exit(1) }
       console.log('OK: '+f);
     "
   done
   ```

3. Run seed:reset then baseline scenario:

   ```bash
   cd $APP_PATH/backend && pnpm run seed:reset
   npx gql-ingest ./ingest/scenarios/baseline/*/entity.json \
     -e $ENDPOINT -c ./ingest/config.yaml -h '{"Authorization": "Bearer $TOKEN"}'
   ```

4. Run baseline + a business flow scenario together (shared output store):
   ```bash
   cd $APP_PATH/backend && pnpm run seed:reset
   npx gql-ingest \
     ./ingest/scenarios/baseline/*/entity.json \
     ./ingest/scenarios/<scenario>/*/entity.json \
     -e $ENDPOINT -c ./ingest/config.yaml \
     -h '{"Authorization": "Bearer $TOKEN"}' \
     -n Entity1,Entity2,Entity3,...
   ```
   Note: use `-n` to filter entities when multiple scenarios have the same entity name (e.g., PurchaseOrder in both baseline and a flow scenario).

## After Completion

Tell the user:

- Seed data is at `<APP_PATH>/backend/seed/data/` — load with `pnpm run seed:reset`
- Scenarios are at `<APP_PATH>/backend/ingest/scenarios/` — run with:

  ```bash
  # Baseline only
  npx gql-ingest ./ingest/scenarios/baseline/*/entity.json \
    -e $ENDPOINT -c ./ingest/config.yaml -h '{"Authorization": "Bearer $TOKEN"}'

  # Baseline + flow scenario (run together for $ref resolution)
  npx gql-ingest \
    ./ingest/scenarios/baseline/*/entity.json \
    ./ingest/scenarios/<flow>/*/entity.json \
    -e $ENDPOINT -c ./ingest/config.yaml \
    -h '{"Authorization": "Bearer $TOKEN"}' \
    -n Entity1,Entity2,...
  ```

- Mutations are at `<APP_PATH>/backend/ingest/mutations/` — shared across all scenarios
