---
created: 2026-05-06
gap: G08
order: 11
severity: P1
category: bench
estimate: L
issue: "https://github.com/lidge-jun/agbrowse/issues/67"
depends_on: ['G06', 'G02', 'G03', 'G01', 'G11']
---

# G08 — Reference benchmark adapters without score claims (vs WebVoyager/WebArena/VWA/Mind2Web)

> Severity **P1** · Category `bench` · Estimate **L** ·
> Tracking issue [#67](https://github.com/lidge-jun/agbrowse/issues/67) · Depends on **G06, G02, G03, G01, G11**

## GPT Pro evidence

Evidence (competitor side): URL: https://github.com/browser-use/online-mind2web — quote: “300 real-world web navigation tasks.” Browser Use publishes an Online-Mind2Web runner and results; WebVoyager, WebArena, VisualWebArena, and Mind2Web define widely referenced task formats. 
OSU NLP Group
+4
GitHub
+4
GitHub
+4

Evidence (agbrowse side): structure/phase_status.md:258 says Phase 20 is ready for trajectory bundles only; structure/phase_status.md:267-269 forbids leaderboard/competitor benchmark claims. 
GitHub

Why this matters: Bench adapters let agbrowse collect apples-to-apples trajectories before making any score claim. This creates a path to future evaluation without violating the current no-leaderboard rule.
Proposed scope, respecting forbidden list:

web-ai/eval-adapters/webvoyager.mjs — convert WebVoyager JSONL tasks into dry-run/local trajectory jobs.

web-ai/eval-adapters/webarena.mjs — add environment reset/trajectory hooks only; no score publication.

web-ai/eval-adapters/visualwebarena.mjs — require ObservationBundleV1 and fail closed without screenshots/boxes.

web-ai/eval-adapters/mind2web.mjs — map action-sequence tasks to trace replay/evidence format.

bin/agbrowse-eval — add --dry-run, --limit, --write-trajectory, and --no-score defaults.

structure/benchmarks.md — restate fixed model/planner/env/task-set prerequisites before any score claim.
Test surface: test/eval.
cli-jaw mirror impact: none; cli-jaw can consume trajectory bundles later.
Acceptance gate: keep benchmark trajectory gate; add gate:benchmark-adapters-no-score.
Estimate: L, 4+ days.

## Diff-level work breakdown

> Fill in concrete diffs (NEW / MODIFY / DELETE with file:line) once this gap
> reaches the active sprint. Until then, the bullets in **Proposed scope**
> above are the agreed shape; do not implement before the depending gaps
> (G06, G02, G03, G01, G11) ship and `gate:all` stays green.

### NEW files
- _to be filled before implementation_

### MODIFY
- _to be filled before implementation_

### DELETE
- _to be filled before implementation_

## Tests
- `test/eval/...` — _list test files once written_

## Truth-table update
- `structure/CAPABILITY_TRUTH_TABLE.md` — add row or update status when this
  gap reaches `ready` in agbrowse.
- `cli-jaw/structure/CAPABILITY_TRUTH_TABLE.md` — mirror entry per the
  `cli-jaw mirror impact` line above.

## Release gates touched
- Existing: `gate:typecheck`, `gate:tests`, `gate:truth-table-fresh`,
  `gate:mcp-scope-frozen`, `gate:no-experimental-in-readme-ready-section`.
- Added by this gap: see **Acceptance gate** in the GPT Pro evidence block.
