# Factory v1 evidence baseline

Factory v1 was paused after its control plane grew ahead of evidence
for its core path.

## Observed evidence

- About 17,000 TypeScript lines, 30 tool actions, and more than 200
  exports accumulated across routing, contracts, ownership, execution,
  validation, review, approval, recovery, intake, policy, metrics,
  recommendations, and calibration.
- The local ledger contained ten workflows and no successful terminal
  completion.
- Four incomplete workflows retained active path claims; a stale
  Jupiter claim repeatedly blocked later work.
- Observed calls concentrated on preview, create, and status. Real
  remote delivery bypassed Factory through Team Mode, SSH, tmux,
  worktrees, and direct `my-pi` processes.
- Similar cross-machine work was inconsistently classified as an
  incident and a database migration.
- Synthetic tests were extensive, but the calibration baseline had no
  representative correlated terminal outcomes. Synthetic coverage is
  not real completion evidence.

The build order was inverted: intake, learning, and calibration
preceded proof that one task could execute, validate, receive
independent review, terminate, and clean up reliably.

## Falsification

This conclusion is falsified only if the predeclared ten-task
benchmark produces at least eight validated independent reviews
without manual state repair, leaves zero stale ownership/resources,
and shows lower operator burden or lead time than direct Pi plus
`pi-harness`. Passing synthetic fixtures, adding recovery controls, or
manually repairing state does not qualify.
