# Benchmark learning notes

This folder holds public-safe benchmark lessons that affect Open Scaffold's product direction.

It is not a scoreboard, raw evidence archive, benchmark specification folder, or adoption-proof folder. Raw run logs, transcripts, screenshots, JSONL files, local machine paths, benchmark-v2 proposals, and benchmark-owned scoring contracts stay out of this repo. Notes here should say what was actually tested, what failed, what remains unproven, and which generic Open Scaffold product lesson follows.

Current notes:

- [`2000m-v1-two-lane-postmortem.md`](2000m-v1-two-lane-postmortem.md) — a local two-lane 2000m v1 run showed no raw-score advantage for Open Scaffold over a naked Codex/GPT-5.5 lane.

Honesty rules:

- A tie or loss is reported as a tie or loss.
- Evidence/recovery value is not smuggled into a raw-score claim.
- Owner-run local experiments are not adoption proof.
- Benchmark design, scorer mechanics, hidden seeds, harnesses, and result schemas belong in the benchmark repo that owns the benchmark.
- Open Scaffold may absorb benchmark findings only as generic workflow improvements, not as benchmark-specific fixtures or contracts.
- Headless protocol drivers are not called playable games unless a real player/viewer exists.
