---
created: 2026-05-06
gap: G05
order: 9
severity: P1
category: observation
estimate: M
issue: "https://github.com/lidge-jun/agbrowse/issues/65"
depends_on: ['G02', 'G06', 'G09']
---

# G05 — Schema-bound page extraction (vs Stagehand, AgentQL)

> Severity **P1** · Category `observation` · Estimate **M** ·
> Tracking issue [#65](https://github.com/lidge-jun/agbrowse/issues/65) · Depends on **G02, G06, G09**

## GPT Pro evidence

Evidence (competitor side): URL: https://github.com/tinyfish-io/agentql — quote: “Structured output defined by the shape of your query.” AgentQL also documents SDKs, REST API, natural-language selectors, and browser debugger tooling. 
GitHub

Evidence (agbrowse side): README.md:446-454 provides text, text --format html, get-dom, console, and network; README.md:699-710 provides source audit for completed answers, not active-page schema extraction. 
GitHub
+1

Why this matters: Many browser-agent tasks are extraction tasks, not just click tasks. A schema contract lets tests assert exact output shape and prevents “looks plausible” scraping results.
Proposed scope, respecting forbidden list:

web-ai/extract-schema.mjs — add active-page extraction contract with JSON Schema/Zod-like shape validation.

web-ai/answer-artifact.mjs — add pageExtraction artifact type with URL, selector scope, schema, and validation result.

skills/browser/browser.mjs — add extract --schema <file> --selector <selector> --json.

web-ai/source-audit.mjs — allow extraction provenance from URL, selector, timestamp, and text span evidence.

structure/commands.md — document extract output and fail-closed schema errors.

structure/release_gates.md — add gate:extract-schema-fixtures.
Test surface: test/unit + test/eval.
cli-jaw mirror impact: parity optional until cli-jaw publicly claims page-extraction parity.
Acceptance gate: keep gate:truth-table-fresh; add gate:extract-schema-fixtures.
Estimate: M, 2–3 days.

## Diff-level work breakdown

> Fill in concrete diffs (NEW / MODIFY / DELETE with file:line) once this gap
> reaches the active sprint. Until then, the bullets in **Proposed scope**
> above are the agreed shape; do not implement before the depending gaps
> (G02, G06, G09) ship and `gate:all` stays green.

### NEW files
- _to be filled before implementation_

### MODIFY
- _to be filled before implementation_

### DELETE
- _to be filled before implementation_

## Tests
- `test/unit/...` — _list test files once written_

## Truth-table update
- `structure/CAPABILITY_TRUTH_TABLE.md` — add row or update status when this
  gap reaches `ready` in agbrowse.
- `cli-jaw/structure/CAPABILITY_TRUTH_TABLE.md` — mirror entry per the
  `cli-jaw mirror impact` line above.

## Release gates touched
- Existing: `gate:typecheck`, `gate:tests`, `gate:truth-table-fresh`,
  `gate:mcp-scope-frozen`, `gate:no-experimental-in-readme-ready-section`.
- Added by this gap: see **Acceptance gate** in the GPT Pro evidence block.
