# Roadmap

## Completed foundations

- 0.2–0.3: Git-local evidence, deterministic packages, host adapters, and
  bounded review controls.
- 0.5: stable requirement IDs, exact-tree evidence, immutable review packages,
  and deterministic completion decisions.
- 0.6: optional delegation, policy-bounded routing, autonomous convergence,
  and claim-scoped public-beta receipts.

## 0.7.0 portable marketplace release gates

- Agent Plugins 1.0.0 and current Agent Skills compliance from pinned sources.
- Five byte-reproducible artifacts from one canonical tree with matching
  inventories and SHA-256 digests.
- Documented, version-stamped behavior for Codex, Claude Code, Grok, OpenCode,
  Pi, and Kimi Code.
- Signed release-evidence payload bound to tag, commit, tree, standards,
  semantic review, host matrix, verification, and artifact digests.
- A reproducible Codex/DeepSWE protocol and a completed historical prespecified
  minimum paired pilot for pre-change commit
  `95dfedf7d396a7b9faa72ced844a28f70bd6bcef`.
- No credentials, obsolete plans, private scratchpads, or untrusted archives in
  the release tree.

Public registry acceptance is external follow-up. A validated submission bundle
is required; universal registry availability is not a tag gate.

## 0.8: Semantic assurance and evidence sufficiency

Semantic assurance will:

- fail closed when semantic inspection is incomplete;
- distinguish literal values from interpolated, computed, inherited, or dynamic
  values;
- maintain stable semantic graph identity across components;
- reject unresolved material ambiguity;
- add adversarial equivalence testing;
- detect shared-domain authority;
- add assurance-report deduplication.

Before expensive evidence-producing work, the workflow will reason through:

```text
intended claim
→ required evidence
→ evidence adequacy
→ required coverage or sample
→ projected time, token, and cash cost
→ execute, redesign the evidence, or narrow the claim
```

The intended claim and whether the work is a pilot or confirmatory study must be
clear before collection begins. Proposed coverage must be capable of supporting
that claim. Initial resource estimates come before a long campaign; after
calibration supplies observed runtime or cost, projected cost and likely
evidentiary value are reconsidered before more runs continue. This planning does
not add approval ceremony to ordinary tests.

### Evaluation continuation

The next comparative campaign for the final v0.7 candidate or a later release
will use the historical pilot's variance, runtime, and cost observations to
choose its run allocation. It will explicitly test whether greater task
diversity, using more unique tasks with fewer repeats, provides more information
per Codex hour than the former 8-task by 3-repeat design.

Neutral comparator work may include Superpowers and GSD under the same frozen
run contract. Other benchmark suites will be added only for a defined evaluation
question. Null and negative results remain publishable. Model-backed host smokes
may expand without broadening claims beyond their receipts.

## 0.9: Optional visual design companion

Explore a lightweight companion for design questions where richer visual
representation materially improves the decision. It may support UI comparisons,
architecture diagrams, state and data-flow views, workflow diagrams, annotated
screenshots, and side-by-side alternatives.

The companion remains optional and integrates with `designing-work`. Text stays
the default when it is sufficient. The design should avoid heavyweight persistent
infrastructure, a persistent browser service, and visual generation without a
specific reason. This is future work, not part of v0.7's lightweight visual
treatment decision.

## 1.0

- Publish evidence-backed quality/economy targets across more than one host.
- Stabilize adapter compatibility and migration policy.
- Complete applicable official-directory review and publication.
