# Discover Harness — Orchestrator

This is the **policy** half of the `discover` harness. The plugin ships the capabilities (11 dimension-collector agents + 5 mode rules + a `tradecraft` keyboard reference); this file drives them. Together they discover a product across every surface it exposes and assemble a structured, provenance-tracked corpus in the resolved corpus root (default `research/<target>/`).

A run is five modes, in order: **Discovery → Ingestion → Evaluation → Cartography → Self-correction.** The main session drives all five — Ingestion and Evaluation fan out parallel sub-agents (dispatch each via the Agent/Task tool with `model: opus`); Cartography re-enters the live product (the browser singleton) to walk + screenshot it and may fan out doc-drafting sub-agents; a sub-agent cannot spawn further sub-agents.

---

## Delegation

The policy for sub-agent dispatch across all five modes — when to fan out, at what tier, and how far to trust the result.

- **Delegation keeps bulk tool output out of the main context.** Wide reads, sweeps, broad research, and audits default to a sub-agent: a sub-agent's raw output dies with it, while the same output read inline is re-read and re-billed on every later turn. The hand-off is self-contained (the sub-agent sees none of the conversation), so it must still cost less than the work; independent delegations run in parallel.
- **Small or judgment work stays inline.** A task finishable in a handful of tool calls is faster done than briefed. Judgment is never delegated by reflex — a second opinion is for a high-stakes call or an explicit request.
- **A delegated result is a report, not verifiable ground truth.**
- **Tiers**: mechanical fan-out runs on the cheapest capable tier; judgment-heavy delegation runs on the standard tier; the strongest tier only on explicit user request.
- **An explicit user directive about tier, cost, or delegation overrides this section.**
- **Ingestion is the sanctioned mechanical case** — each of the eleven dimensions is large and independent, so its collectors are fanned out deliberately as the default, never by reflex.

No dispatch mechanism ⇒ inline.

---

## Access is a capability vector (the front gate)

A target's access is four orthogonal axes, each `{value, confidence, evidence}` — and the vector decides which dimensions are reachable and which extraction method each uses:

| Axis | Values | Unlocks | Case |
| --- | --- | --- | --- |
| **source** | full · partial · none | codebase, packages | open-source repo |
| **runtime** | reachable · gated · none | deployed-client-bundle, api, infra-backend-fingerprint, wire-capture | deployed, no source |
| **auth** | have-login · none | session | logged-in SaaS |
| **presence** | rich · thin · none | docs, website, community | — |

Targets light several axes at once. The full model + grade→method mapping is in `.claude/rules/discovery.md`.

---

## The eleven dimensions → agents

| Dimension / agent |     | Dimension / agent           |
| ----------------- | --- | --------------------------- |
| `codebase`        |     | `session` ⚠                 |
| `docs`            |     | `deployed-client-bundle`    |
| `packages`        |     | `infra-backend-fingerprint` |
| `api` (+MCP)      |     | `wire-capture` ⚠            |
| `website`         |     | `distribution-artifacts`    |
| `community`       |     |                             |

Agent name = dimension name. ⚠ = gated (needs user authorization before dispatch). The catalog — detection signals + strategy + per-dimension artifacts — is in `.claude/rules/discovery.md`.

---

## Step 0 — first run (tooling check)

**Once per folder, before Mode 1.** Capture depth depends on host tooling most machines lack (`gron`, `trafilatura`, `mitmproxy`, `chrome-remote-interface`, …). On the first run in a corpus folder, check the state file `.claude/discover-tooling.json`:

- **It exists** → tooling was already reviewed here — skip straight to Mode 1. **Never re-prompt in the same folder.**
- **It's absent** → run the one-time check:
  1. Read [`.claude/rules/tooling.md`](.claude/rules/tooling.md) — the Tier 1 (clean) and Tier 2 (gated) tool lists + the AVOID list.
  2. **Validate what's already installed before recommending anything** — `command -v <tool>` per recommended tool (`npm ls -g <pkg>` for the Node libraries that ship no binary, e.g. `chrome-remote-interface`). Detection only — nothing is installed at this step.
  3. **Recommend only the missing tools**, grouped **clean** (safe to install freely — offer the Tier-1 one-paste block) vs **gated** (first-party / user-owned / opt-in — surface the command but **never run it for the user**; these need the user in the loop). Never suggest anything on the AVOID list. If nothing is missing, say so in one line.
  4. **Write `.claude/discover-tooling.json`** so the harness never asks again in this folder:
     ```json
     { "checked": "<ISO date>", "present": ["gron", "pandoc"], "missing": ["mitmproxy"], "recommended": true }
     ```

The state file is **host-local** (tools live on the machine, not the repo) — gitignore it in a shared folder; delete it to force a fresh check after installing a batch. This step never blocks a run: a user who declines proceeds to Mode 1 on whatever tooling is present, and a genuinely missing *preferred* capability is re-surfaced at the Mode 2 capability-gap checkpoint.

---

## Mode 1 — Discovery

Cheap, read-only. Scope the target, grade its access, detect which dimensions it exposes.

**0. Resolve the corpus root** — the folder this run writes into. The harness chooses it _intelligently_ rather than hardcoding one: default `research/` at the repo root, but if the project already holds a discovery corpus (any folder with a prior run's `*/00-scope-verdict.md`) under another name (`discovery/`, `analysis/`, …) **reuse it**, and honor a folder the user names. **Confirm the chosen root in one line** before writing. Every `research/<target>/` path in this harness is the _default form_ — read it as `<corpus>/<target>/`, relative to the resolved root.

1. **Scope-feasibility gate** — run the `.claude/rules/discovery.md` Part A (scope gate) algorithm: resolve identity → ambiguity → ethics → suite-size → primitive → reachability. If the verdict isn't `accept`, **halt** (present the narrowing/redirect prompt, or refuse). Write `research/<target>/00-scope-verdict.md`.
2. **Grade the access vector** (the four axes, with evidence) per `discovery` (Part B).
3. **Probe each of the 11 dimensions** (cheap signals only); mark ✅ available / ⚠️ partial / ❌ absent with evidence. An absent dimension is a finding.
4. Write `research/<target>/00-recon-plan.md` — frontmatter carries `scope:` + the `access_grade` vector; the body is the dimension-availability table + the per-dimension collection plan (warm entry points) + the gating/ethics flags.

---

## Mode 2 — Ingestion

Dispatch one collector per available dimension; collect exhaustively before context is lost. Every collector writes `research/<target>/dimensions/<dim>/_summary.md` (with the **provenance frontmatter** — method · confidence · completeness · gaps) + `raw/`, obeying `.claude/rules/ingestion.md`.

- **Run independents in parallel** — dispatch `docs`, `packages`, `api`, `website`, `community`, `deployed-client-bundle`, `infra-backend-fingerprint`, `distribution-artifacts` together (`subagent_type='<dimension>'`, one message, multiple Agent calls). Seed each collector with **warm entry points (URLs / repos / hosts), not claimed facts** — any fact **or inference** carried from Discovery (a price, a tier name, a version — **or a derived hypothesis like "localStorage-key X ⇒ backend is Y"**) is tagged `verify:may-be-stale` so the collector **re-derives it instead of anchoring** on a possibly-stale (or possibly-wrong) value. A Discovery _inference_ is never seeded as a plan fact: label it a hypothesis in `00-recon-plan.md` and tag it `verify:` in the seed, so a collector that inherits it reports it hedged rather than echoing it as observed (a real run seeded an `nustackAnonId ⇒ GraphQL backend` guess that two captures then repeated before the session refuted it). **A sibling collector's _summary_ is a claim, not a source** — when seeding one collector with something another collector reported, tag it `verify:` too, unless the main session has actually read the underlying `raw/` artifact. A returned summary is a compressed, unverified assertion; relaying it into a fresh brief launders it into apparent provenance, and the receiving sub-agent (which cannot see the conversation) has no way to tell a seeded claim from an established fact. (Origin: Emergent — the orchestrator relayed a "documented limitation: the connector cannot deploy" line from one collector's summary into another's brief; it was **not in the corpus**, and only the receiving agent's own verification caught it.) **Every QUANTITATIVE claim in a brief carries its lane.** The rule above covers a *sibling's summary*; it does not cover a number the **orchestrator itself** picked up in passing — from a search result, a page it skimmed, or its own recollection — and then re-emitted as framing. A figure the main session cannot name a `raw/` anchor for is either seeded **`verify:`** or **omitted**; it is never stated flat. A number in a brief reads as established even when nothing established it, and the receiving agent cannot tell the difference. (Origin: indus-sarvam — the orchestrator seeded a "₹98.68 crore IndiaAI funding" figure into the positioning brief as context; the receiving agent traced it to a **third-party user's GitHub issue body**, not to any first-party source, and had to de-launder it. The vendor states only "compute provided under the IndiaAI mission", with no amount.) Only one dimension may drive the browser at a time (see `ingestion` §7 rule 9): the parallel collectors use WebFetch/curl, leaving Chrome to `session`.
- **Codebase fan-out** — dispatch `codebase`; it clones + maps + returns a per-package work-list; then the main session fans out one reader per substantial package.
- **Gated dimensions** — `session` and `wire-capture` (rung ≥2) need explicit user authorization first (see `ingestion` §7). `session` follows `.claude/rules/ingestion.md` §9. When `session` is authorized, **dispatch the non-browser collectors in the background (`run_in_background`) and drive the gated browser session in the main session concurrently** — they don't contend (Chrome belongs to the main session; the collectors use WebFetch/curl), so the whole of Ingestion runs in parallel.
- **Shared seam** — **seed the two shared files `dimensions/_shared/{api-path-catalog.md, feature-flags.md}` (empty, with a header) once, before the collector fan-out**; bundle/session/wire/ distribution then each _append_ a `source:`-tagged section to the same physical file (never their own copy). One owner dimension per static asset; a second mine of the same asset is one source, not corroboration (`ingestion` §6).

Validate each return (`_summary.md` + provenance present). A `blocked`/`partial` return is recorded, not blindly re-run.

**Validate a delegated artifact by READING THE FILE, not by trusting the agent's return.** A sub-agent can report success, deliver one of two assigned artifacts, or leave a **stale prior-run file** untouched — and a stale file is **worse than a missing one, because it scores silently**. For any artifact a later mode reads programmatically (above all `evaluation/feature-coverage.md`, whose `coverage_scorecard` frontmatter Mode 5 scores directly), assert (a) the file's mtime is newer than the dispatch **and** (b) its frontmatter matches this run's facts. If it doesn't, write it yourself. (Origin: Emergent — a sub-agent delivered its IA doc and silently dropped the coverage doc; the untouched Pass-1 file would have scored a full write-side run at `screenshot_coverage: 0` / `write_side_observed: false`, i.e. the *previous* run's numbers.)

- **Capability gap = a checkpoint, not a silent fallback** (`ingestion` §8). If a _preferred_ capture capability is found unavailable mid-run (CDP/debugging port closed, browser MCP unreachable, a gated proxy can't be stood up) and only a **materially weaker** method remains, **pause and confirm** with the user before continuing on the fallback — naming what's unavailable and the coverage lost. (Distinct from a planned method pre-grade, which is recorded, not confirmed.)

---

## Mode 3 — Evaluation

Reconcile the collected dimensions into four cross-dimension rollups under `research/<target>/evaluation/`, plus the run `README.md`, **weighting every claim by how it was obtained**. Follow `.claude/rules/evaluation.md`: read each dimension's provenance frontmatter first; anchor + provenance-tag every claim `(api: raw/endpoint-catalog.md · openapi-verbatim · high)`; weight (max-of-sources, promote one band only on ≥2 _independent_ dimensions, never average); present high→fact, low/single-source→tentative, conflict→flagged. The four rollups: `technology-architecture`, `product-features`, `data-model-api-surface`, `competitive-positioning`.

---

## Mode 4 — Product Cartography & Coverage

Make the product **legible as a product**. After Evaluation reconciles the dimensions, Cartography **re-enters the live product** (read-only nav, the browser singleton) and compiles the navigable **information architecture**, the primary **UX flows**, a **screenshot surface map**, and a **feature-coverage gate** — reconciling every _marketed / documented_ feature against **where it actually lives in the product** and **whether the run walked it**. This catches the blind spot the dimension collectors structurally miss: a feature **hyped on the website but buried three levels deep** — or marketed but **absent from this edition**. Follow `.claude/rules/cartography.md` (the spine: inputs · scorecard · orchestration · ethics), which routes to three artifact references — `cartography-ia.md` · `cartography-flows.md` · `cartography-coverage.md`. Four artifacts:

1. `evaluation/information-architecture.md` — nav tree + per-surface cards, each tagged with **depth** (D0 top-level … D3+ sub-panel) and **promoted | buried**.
2. `evaluation/ux-flows.md` — the primary journeys as mermaid sequence diagrams, traced from the **observed `session` wire** (not source). A write/generate flow on a `write_side_observed: false` run is diagrammed **as inferred**, never observed.
3. `evaluation/feature-coverage.md` — the **claimed-vs-located-vs-walked** matrix + the machine-readable **coverage scorecard** (`feature_location_rate`, `flow_coverage`, `screenshot_coverage`, `ia_nav_complete`) that Mode 5 scores against. A `claimed-but-not-located` feature is **never dropped** — it is dispositioned (deeper-than-looked → re-walk queue · edition/roadmap-gated → recorded gap · over-claim → an evaluation flag).
4. `dimensions/session/captures/screens/` — one **redacted** screenshot per primary surface (+ `_index.md`); the textual surface-map is the fallback only when the runtime can't screenshot.

Main-session-driven (it owns the browser); it **may** dispatch sub-agents (`model: opus`) to draft the three docs from captured material, but the **live walk + screenshots are main-session**. **Read-only / no-cost only** — locating a feature never triggers it; any write/generate stays behind the Pass-2 gate.

---

## Mode 5 — Self-correction (the run corrects itself; the harness improves only on approval)

Mode 5 is a loop, not a one-shot. Follow `.claude/rules/self-correction.md` — it carries the **defect checklist** (A), the **measurement rubric** (B), and the **iterate-or-propose loop** (C). Three stages, driven by the **main session** (not a sub-agent):

1. **Inspect** the finished run (scope verdict, plan, every dimension `_summary.md`, the rollups, **the Cartography IA / UX-flows / feature-coverage docs**) against the **built-in defect checklist** → a concrete list of _behavioral defects_ in the prior steps (mode · defect · severity · evidence anchor · within-run|harness-changing).
2. **Measure** the run against the **rubric** → a `/100` score + vibe band — including the heavily-weighted **experiential / IA coverage** axis fed by the Cartography scorecard. Always — every iteration, including the final one; the self-correcting step ALWAYS evaluates the current run.
3. **Iterate-or-propose**, splitting each defect by one test — _does the fix require editing a rule or an agent?_
   - **No → within-run:** the main session **auto-iterates** (no approval), re-running the cheapest affected Mode 1–4 steps on THIS target and re-measuring each round — including **re-entering the live product (read-only nav) to walk the `located-but-not-walked` queue** and capture missed screenshots. The loop is **capable of ≥5 iterations** and **stops on regression** (a clear score drop → revert to the best prior run), **convergence** (no within-run defects left), or the cap. **Read-only / no-cost only** — any fix needing a Pass-2 / credit-spend / state-change / credential action goes through the explicit-confirmation gate, never the auto-loop.
   - **Yes → harness-changing:** accumulate into `research/<target>/05-self-correction-proposal.md` (the shape is in `self-correction.md`: run measurement · iteration ledger · defect list · within-run fixes applied · the approval-gated harness-change table).

**Approval gate (unchanged):** present the harness-change table + the iteration ledger and STOP — nothing in the harness is applied. The user accepts / rejects / defers per item. Then the **main session** (not a sub-agent) applies accepted items, flips each row to APPLIED, and logs the change. Promote only what the _next, different_ target would also hit; prefer additive edits; structural edits are flagged, never auto-applied. (For a batch, additionally roll up into a cross-target synthesis that promotes only patterns ≥2 _independent_ targets hit — a homogeneous batch's 6/6 is one vote, not six.)

---

## Mirroring applied harness changes (the delivery step)

`self-correction.md` instructs the main session to "mirror to the harness home" and attributes the procedure
to "per ingestion / the orchestrator" — **which never defined one.** It was a dangling reference, and the
cost was measurable: **275 lines of accumulated rules across 4 commits sat only in one working copy**, so
the distributable package a new user installs was missing every lesson the fleet had learned.

**The procedure, defined here.** After applying accepted harness changes and flipping the rows to APPLIED:

1. Resolve the home. Default `~/dev/etna/harness/discover/`; if absent, ask — never guess, never skip.
2. Copy `CLAUDE.md` → `<home>/CLAUDE.md`, `.claude/rules/*.md` → `<home>/rules/`,
   `.claude/agents/*.md` → `<home>/agents/`.
3. **Assert zero drift** (`diff -r`) and report the line delta. A non-zero delta after mirroring is a
   failure, not a warning.
4. If the home is a distributable package (e.g. `packages/cli/data/harnesses/`), note that it ships to new
   users and flag whether a version bump is needed — a rule that exists only in a working copy has not
   shipped.

**A harness change is not "applied" until it is mirrored.** Applying locally and stopping means the next
*different* target — the entire justification for promoting the rule — never sees it.

## Calibration — the only external signal (a cross-run process, not a mode)

Everything in Modes 1–5 is self-assessment: collectors grade their own confidence, Evaluation weights those
self-grades, Mode 5 scores against a checklist this system wrote. **Calibration is the one procedure that
tests whether any of it is true** — it re-derives a prior run's claims **blind** and records what survived.

Governed by [`.claude/rules/calibration.md`](.claude/rules/calibration.md); the fleet metric accumulates in
`research/_calibration/ledger.md`. Two tiers: a **monthly Tier-A claim spot-check** (~15–20 claims, one
session — this is the default, because a metric too expensive to run leaves the sample at n=1 forever) and a
**quarterly Tier-B full blind re-run**. Target is chosen by **rotation, never judgment**.

**Three rules make it evidence rather than bookkeeping:**

1. **Blind.** Re-derive from the live product and write down what you observed **before** opening the prior
   corpus. This deliberately defers `discovery.md` Step 00 for calibration runs — a run that reads the prior
   claim first anchors on it and reports its own suggestibility as a confirmation rate.
2. **`refuted` ≠ `stale`.** "Wrong when made" indicts the method; "true when made, product changed" does
   not. Conflating them makes the harness fix the wrong problem.
3. **Never grade with it.** The moment `refutation_rate` becomes a rubric input, the run acquires an
   interest in the number and the one uncorrupted signal is gone. It measures the **harness**, not the run.

**Current state: n=1**, and the one data point is uncomfortable — 44% of re-tested claims refuted, with the
`high`-confidence band refuting *more* often than `low`. Do not act on it yet; do grow n.

## Ethics (non-negotiable)

For spec-writing, integration planning, and competitive analysis — **not** credential theft, rate-limit bypass, or scraping. Full rules: `.claude/rules/ingestion.md` §7.

- **Never enter credentials.** The user logs in; `session` and `wire-capture` rung 3 are opt-in.
- **Redact every secret before writing `raw/`.** Cookies/tokens/JWT-signatures/API-keys never hit disk.
- **wire-capture** climbs the 3-rung ladder (browser-tap → HAR → mitmproxy); rung 3 is user-run, first-party traffic only. Wireshark is not used (TLS ciphertext).
- **Confirm before any state change or cost.** **No destructive actions, no exfiltration.** Respect TOS/robots.

---

## Output contract

Output lives under the **resolved corpus root** (default `research/`; chosen per Mode 1 step 0):

| Artifact (under `<corpus>/<target>/`) | Written by | Contents |
| --- | --- | --- |
| `00-scope-verdict.md` | Discovery | accept / narrow / redirect / refuse |
| `00-recon-plan.md` | Discovery | `access_grade` + scope + dimension plan |
| `dimensions/<dim>/{_summary.md, raw/}` | Ingestion | per-dimension collection (`session` & `wire-capture` also: `captures/`, `curls.md`) |
| `dimensions/session/captures/screens/` | Cartography | redacted per-surface screenshots (+ `_index.md`) |
| `evaluation/{technology-architecture, product-features, data-model-api-surface, competitive-positioning}.md` | Evaluation | the four rollups |
| `evaluation/{information-architecture, ux-flows, feature-coverage}.md` | Cartography (Mode 4) | the three cartography artifacts |
| `README.md` | Evaluation | run story + headline findings |
| `05-self-correction-proposal.md` | Self-correction (Mode 5) | proposed harness fixes (approval-gated) |
| `00-prior-run-reconciliation.md` | Calibration (when a prior corpus exists) | refuted / corrected / confirmed / stale per claim + the confidence regression (`calibration.md`) |

Cloned repos live in `source/<target>/`. Provenance is the "self-aware" layer: each collector records _how_ it was obtained; Evaluation weights claims by it. Detailed shapes live in the rules.
