# Ingestion Contract — Capture Shape · Provenance · Ethics · Session Playbook

The reference knowledge for **Ingestion** (`discover` harness, Mode 2). Every dimension capture obeys the contract in §1–8 so (a) the **evaluation** mode can rely on a stable shape and weight each claim by _how it was obtained_, and (b) a future agent can reproduce or extend any capture. §9 carries the deepest, gated dimension's method — the logged-in browser-session playbook — which the `session` agent follows.

This contract generalizes the source playbook's three-format contract (`captures/*.md` narrative + `raw/*.json` verbatim + `curls.md`) to _any_ dimension.

> **Contract style = "lighter".** Provenance ("how it was captured") lives in the **`_summary.md` frontmatter**, not a separate `manifest.yaml`. There is no parallel YAML machine-spine — everything stays markdown-with-frontmatter. Evaluation reads the frontmatter to weight claims; self-correction reads it to detect under-delivery.

---

## 1. Folder layout

```
research/<target>/dimensions/
├── _shared/           ← cross-dimension seam artifacts, ONE physical copy (see §6)
│   ├── api-path-catalog.md
│   └── feature-flags.md
└── <dimension>/
    ├── _summary.md    ← decoded narrative + provenance frontmatter (human + machine read this)
    └── raw/           ← verbatim artifacts, DIMENSION-PRIVATE (the reproducer reads these)
```

Richer dimensions inherit the playbook contract (extra files, same `_summary.md` provenance):

- **session**: `dimensions/session/{_summary.md, captures/p*.md, raw/*.json, curls.md}`
- **wire-capture**: `dimensions/wire-capture/{_summary.md, captures/wire-narrative.md, raw/flows.json, curls.md}`

Two layout rules the seam and the "absence is a finding" principle depend on:

- **Shared-seam files live in `dimensions/_shared/`, one physical copy each — never under a single dimension's `raw/`** (§6). Everything else under a dimension's `raw/` is **dimension-private by default**; only the §6 allowlist (`api-path-catalog.md`, `feature-flags.md`) is shared.
- **A folded or absent dimension still leaves a record.** When a dimension is folded into another or is absent, write a one-paragraph `_summary.md` with `status: folded` (naming the fold target) or `status: absent` (+ the cheap evidence). "An absent dimension is a finding" must be **materialized as a file**, never a silent drop.

---

## 2. `_summary.md` — required shape (with provenance frontmatter)

```markdown
---
dimension: <codebase|docs|packages|api|website|community|session|deployed-client-bundle|infra-backend-fingerprint|wire-capture|distribution-artifacts>
target: <slug>
status: complete | partial | blocked
# --- provenance (the "how it was captured" layer evaluation weights on) ---
access_grade_used: <source:full | runtime:reachable | auth:have-login | presence:rich | ...>
method: <the ACTUAL extraction method — see §3 vocabulary>
completeness_pct: <0-100, honest self-estimate of reachable surface captured>
confidence: high | medium | low # how much to TRUST what was captured (see §3)
captured_at: <YYYY-MM-DD>
sources: [<urls / paths actually consulted>]
gaps: [<structured — each is a thing evaluation must NOT treat as covered>]
---

# <Target> — <Dimension> capture

## Method

Exactly what was done, which tools, what was skipped and why. (The prose behind the frontmatter `method`.)

## Findings

Ordered, concrete observations — real values (package names, versions, endpoint paths, file paths, page URLs, prices). Tables liberally. The load-bearing section.

## Inferences

What the findings imply. Mark inference vs. fact.

## Open questions

What couldn't be captured and why. Never fabricate to fill a gap.

## Artifacts

Index of `raw/` files, one line each.
```

Keep `status` and `confidence` honest — they drive the entire evaluation weighting.

---

## 3. Method & confidence vocabulary (drives evaluation weighting)

`method` is one of: `clone-and-map`, `source-map-reassembly`, `registry-metadata`, `ats-api`, `first-party-api-json`, `openapi-verbatim`, `graphql-introspection`, `crawl-clip`, `docs-reconstructed`, `bundle-string-mine`, `dns-ct-fingerprint`, `wire-tap-browser`, `wire-har`, `wire-mitmproxy`, `binary-extract` (asar/apk/crx), `session-playbook`, `sdk-harness`, `inferred`.

**Confidence is a function of method** (evaluation relies on this mapping):

- **high** — verbatim machine artifacts: `openapi-verbatim`, `graphql-introspection`, `registry-metadata`, `ats-api`, `first-party-api-json`, `clone-and-map`, `binary-extract`.

> **`ats-api` — a first-party job-board / ATS JSON API is NOT a crawl.** A response from an official
> applicant-tracking endpoint (`boards-api.greenhouse.io/v1/boards/<slug>/jobs?content=true`,
> `api.ashbyhq.com/posting-api/job-board/<slug>`, `api.lever.co/v0/postings/<slug>?mode=json`) is a
> **verbatim machine artifact published by the company itself** — the same evidence class as
> `registry-metadata`, and it maps to **high**. Grading it `crawl-clip` (medium) forces the strict reader
> to *downgrade a high-quality first-party artifact*, which is the wrong direction. The `careers` marketing
> page, an aggregator listing, or a search-result snippet **is** `crawl-clip`; the ATS API is not.
> (Origin: Emergent — a 433 KB verbatim Greenhouse JSON body carrying a byte-identical declared tech stack
> across two postings was labelled `crawl-clip` with `confidence: high`, a mismatch against this very
> table. The confidence was right and the vocabulary had no word for the artifact.)

> **`first-party-api-json` — the general case of the rule above.** `ats-api` fixed the *hiring* instance
> of a wider gap: **any verbatim JSON response from a first-party or official platform API is a machine
> artifact, not a crawl**, and maps to **high**. Use it for a GitHub org/issues/releases API response
> (`api.github.com/orgs/<org>/repos`, `/search/issues?q=org:<org>`), a status-page API
> (`status.<vendor>.com/api/v2/summary.json`), a changelog/release JSON feed, a public app-store listing
> API, or any comparable official endpoint. **The distinction is the artifact, not the topic:** the
> *rendered page* of the same information is `crawl-clip` (medium); the *API body* is
> `first-party-api-json` (high). Prefer the more specific term where one exists — `registry-metadata` for
> npm/PyPI/crates, `ats-api` for job boards, `openapi-verbatim` for a served spec — and reach for this one
> only when none fits. (Origin: indus-sarvam — the `community` lane mined the GitHub API and, finding no
> correct term, graded itself `registry-metadata`. The evidence class and the `high` confidence were both
> right; the label named the wrong registry. A vocabulary with no word for an artifact does not stop a
> collector using it — it just makes the provenance record lie.)
- **medium** — `crawl-clip`, `docs-reconstructed`, `bundle-string-mine`, `dns-ct-fingerprint`, `wire-har`.
- **low** — `inferred`, `session-playbook` Pass-1-only, `wire-tap-browser` shapes-only, any wire-drift observation without a spec.

---

## 4. `raw/*` — required shape

Structured captures (API responses, registry metadata, schemas) → JSON with a `_meta` header:

```json
{
  "_meta": {
    "source": "...",
    "captured_at": "YYYY-MM-DD",
    "kind": "...",
    "byte_length": 0,
    "redactions": ["field → <REDACTED:...>"]
  },
  "body": {
    /* verbatim, PII + secrets redacted */
  },
  "_decoded": { "field": "meaning + observed values" },
  "_observation": "optional — what's load-bearing"
}
```

Prose captures (doc/marketing pages, structure maps, catalogs) → markdown with a metadata line: `<!-- source: <url-or-path> · captured_at: YYYY-MM-DD · method: <method> -->`

**One file per logical artifact**, named by what it is (`endpoint-catalog.md`, `react-exports.md`, `pricing.md`, `structure-map.md`).

---

## 5. Capture discipline — DUMP VERBATIM, digest alongside

> **This rule was inverted until 2026-07-28.** It used to read *"Size discipline — reduce, don't dump… a
> `raw/` file is readable, not a memory dump,"* and collectors obeyed it: across 21 corpora `raw/` came out
> **640 `.md` to 115 `.json` — 84.8% model-written prose** — and **6 of 12 sessions marked `status: complete`
> held ZERO verbatim envelopes.** The corpus became unreproducible by following its own contract.
>
> The old rule conflated two different problems with one answer:
> **(a) don't put 6 MB through your context** — correct, always.
> **(b) don't put 6 MB on disk** — wrong, and the cause of the damage.
> They have different solutions. **Write to disk without routing through context** (§5.2). Context budget is
> never a reason to discard evidence.

### 5.1 The two layers — never let one substitute for the other

| Layer | Files | Audience | Rule |
| --- | --- | --- | --- |
| **Evidence** | `raw/*`, `captures/*.json` | a *future run* re-deriving the claim | **verbatim**, redacted, complete. Unreadable is fine. |
| **Narrative** | `_summary.md`, `captures/*.md`, `evaluation/*` | a human, and Evaluation | digested, decoded, opinionated. |

**A digest is written IN ADDITION TO the artifact, never INSTEAD OF it.** If you find yourself writing "the
response contained 47 fields including…", stop: dump the response, *then* write that sentence. The sentence
is worth little without the file; the file is worth a lot without the sentence.

### 5.2 Size is a routing problem, not a reason to discard

Never read a large artifact into context in order to save it. Route it to disk directly:

```bash
curl -sS "$URL" -o raw/spec.json                      # never touches context
curl -sS "$URL" | gzip > raw/bundle.js.gz             # compress in transit
```
- **In-page (browser) captures** → the clipboard channel (`tradecraft.md` §1b): scrub → `__CP(obj)` →
  `pbpaste > raw/capture.json`. Unlimited size, never passes through the display layer.
- **Genuinely unbounded streams** (an infinite log, a 700-page crawl): bound it explicitly — first N,
  every Nth, or a time window — and **record the bound in `_meta.sampling`**. A recorded sample is evidence;
  an unrecorded summary is not.
- **Too big to keep** (>50 MB): keep a verbatim head/tail slice + the full `sha256` + the retrieval command
  in `_meta.refetch`, so a later run can reproduce it exactly.

`_meta` gains: `verbatim: true|false`, `sampling: <bound, if any>`, `refetch: <exact command>`.
Set `digested: true` **only** when the verbatim form is also on disk beside it.

### 5.3 THE TRACEABILITY RULE — if you cite it, dump it

**Every value that appears in a `_summary.md`, a rollup, a README, or a provenance anchor MUST be traceable
to a file on disk in the same corpus.** No exceptions.

**The session is not a storage medium.** A collector that reads a response, extracts a number, writes the
number into prose, and never saves the response has produced a claim that **cannot be checked by anyone,
ever** — including the next run, including the same model an hour later. The claim outlives its evidence,
and the corpus silently degrades from research into assertion.

This is the single highest-value rule in this file, because it is the root cause of an entire class of
apparent defects. When a later reader finds a citation pointing at a file that does not contain the claim,
the overwhelmingly likely explanation is **not** fabrication — it is that the evidence existed in the
session, shaped the claim honestly, and was never written down. The citation outlived the data.

**Practical form:** before writing any Finding, ask *"which file on disk would let someone else check
this?"* If the answer is "none," write the file first. If the value came from a tool result you did not
save, save it now — a `raw/` file with `_meta.source` and the raw body is thirty seconds of work and is the
difference between evidence and hearsay.

> (Origin: layer-ai — a rollup promoted a finding on *"five independent dimensions"*, citing
> `infra-backend-fingerprint`; the string it relied on appears **zero** times in that dimension's files. The
> lane was real in-session and never dumped. Also emergent — 26 session envelopes, the only corpus of 12
> where the session lane is genuinely reproducible.)

### 5.4 Dump manifest — the floor, per dimension

Each dimension is `status: complete` only when these exist as **verbatim** files. Anything else is `partial`.

| Dimension | MUST land in `raw/` (verbatim, redacted) |
| --- | --- |
| `api` | the served spec **as served** (`openapi.json` / SDL / introspection result), not a re-typed table |
| `packages` | the registry metadata JSON per package (`npm view --json`, PyPI JSON), + the `.d.ts` / exports surface |
| `deployed-client-bundle` | the entry bundle (gzipped is fine) + any `.map`; the route/flag extracts *in addition* |
| `session` | ≥3 envelopes (identity read · one canonical list read · the primary write) — §8 |
| `wire-capture` | the flow log (HAR/JSON), not only the narrative |
| `infra-backend-fingerprint` | raw `dig`/`crt.sh`/header responses, not only the decoded table |
| `docs` | the source form of each captured page (md/html), not only the distilled section |
| `website` | the raw HTML or extracted text per captured page, + the pricing table as data |
| `community` | the changelog/issue payloads as fetched (JSON where an API exists) |
| `distribution-artifacts` | the manifest(s) verbatim (`manifest.json`, `Info.plist`, `AndroidManifest`) + the extracted file list |
| `codebase` | (the clone **is** the artifact) the structure map + resolved manifests |
| `external-reputation` | the platform's own structured payload (`__NEXT_DATA__`, schema.org block, API JSON) |
| `hiring-intel` | the ATS API response verbatim |

**When a MUST is unobtainable, that is fine — record it in `gaps:` with the reason and mark the dimension
`partial`.** An honest `partial` is worth more than a `complete` resting on prose.

---

### 5.5 Visual capture — screenshot GENEROUSLY (for the human, not the score)

**Any dimension driving a real browser captures screenshots continuously, not one-per-surface.** A capture
session that walks fifteen screens should leave far more than fifteen images.

> **This rule is deliberately NOT tied to a score.** `screenshot_coverage` (`cartography-coverage.md`)
> measures something narrow — did each *primary surface* get one image — and it is a Mode-5 rubric input.
> **This is a different obligation with a different beneficiary: the human who reads the corpus weeks
> later, for a purpose nobody anticipated during the run.** Do not capture to satisfy the metric and stop.
> Satisfying the metric is a floor so low it is nearly always already met; the rule here is *keep going*.
> Extra images cost a few hundred KB and buy a reader something no prose reconstructs — what the product
> **actually looked like**, at that version, in that plan tier, on that date. That is a genuinely
> irreproducible asset: the run can be repeated, the product's past state cannot.

**Capture every surface, AND every materially different state within a surface.** The states worth having,
roughly in order of how often a later reader wants them:

| State | Why a reader wants it |
| --- | --- |
| **Empty** (no data yet, zero-state, onboarding) | the new-user experience — usually the most-designed and least-documented screen in the product |
| **Populated** (real data present) | the actual working UI |
| **A write in progress** (form filled *before* submit, mid-action) | the input schema as the product presents it — often richer than the API's |
| **After the write** (success state, the created entity) | what the product considers a result |
| **Modal / drawer / panel open** | usually a whole surface the nav never reveals |
| **Dropdown / picker expanded** | enumerates the option set — engine lists, model lists, tier lists, permission sets |
| **Error and validation states** | constraints the docs never state |
| **Permission / plan walls** | exactly what is gated, at which tier |
| **Loading / async / in-flight** | reveals the job model and whether work is sync or queued |
| **Hover, tooltip, help text** | inline documentation that exists nowhere else |
| **Settings and config surfaces, fully scrolled** | the configuration surface is where products differ most |

Also capture the **before/after pair** around any state-changing action you were authorized to take — the
delta is the finding, and one image of the result cannot show it.

**Naming and index.** Volume only helps if it stays navigable:

- `captures/screens/<nn>-<surface>[-<state>].png` — e.g. `07-agents-modal-open.png`,
  `12-composer-model-picker-expanded.png`, `18-deploy-confirm-before.png` / `19-deploy-confirm-after.png`.
- Keep `screens/_index.md` current: **file → surface → route → depth → state → one line on what it shows.**
  An unindexed directory of 60 images is an archive nobody opens.
- Sequential numbering encodes the walk order, which is itself information — it preserves the path taken.

**Redaction happens BEFORE the shutter, and it is not optional.** More images means more leak surface, and
a leaked pixel is undetectable afterwards — `grep` cannot audit a PNG, so the §7.0 4c literal sweep has **no
image analogue**. Apply DOM-level redaction (`tradecraft.md` §1c) and re-check before each capture, because
the elements needing redaction change as you navigate.

> **The third-party trap — check the CORNERS of the frame, not just the content.** It is easy to redact the
> account email in the header and ship a screenshot whose bottom-right corner holds a support-chat widget
> displaying **another human's name and message**. Redact any: support/chat widget (Intercom, Crisp,
> Drift, Zendesk), notification toast, presence avatar or "who's online" strip, comment or activity feed,
> shared-workspace member list, and any other tenant's data visible in a list.
>
> (Origin: Emergent — a run redacted the account holder's email and display name at the DOM level, wrote
> *"no PII exists in the pixels — verified by reading each capture"* into `screens/_index.md`, and shipped
> ~20 images each carrying an Intercom widget with a named support agent and her message text. The
> redaction was real; the *frame* was never examined. **Never write a zero-PII assertion you did not verify
> by looking at the image** — and prefer "redacted: <what>" over a blanket claim.)

### 5.6 Reconstructing an ALGORITHM — differential testing, in separated layers

§5.4 says what a *catalog* capture must dump. This is the extra bar for the harder case: a dimension
that reconstructs **behaviour** — a scoring formula, a ranking function, a pricing calculation, a
matcher, a state machine. A word list can be checked by reading it. An algorithm cannot: it is only
`complete` once it has been **run against the live product on inputs chosen to break it**.

**The rule.** A reconstructed algorithm is `partial` until:

1. It is **executable** and committed to the corpus, not merely described in prose.
2. It has been **differentially tested against the live product** on a real corpus plus deliberately
   adversarial inputs — not only on the examples it was derived from.
3. The test **separates the estimator from the feature extraction**, and the per-layer pass rates are
   recorded separately. One aggregate number is not enough.
4. The inputs and the product's own observed outputs are frozen in `raw/` (or beside the
   implementation) as **re-runnable fixtures**, so a later run can repeat the comparison.

**Why the layer split is the load-bearing part.** Almost every reconstruction has two independent
places to be wrong — the *computation* and the *inputs fed to it*. Tested together they are
indistinguishable, and the aggregate score points at whichever one you already suspect:

- **Layer A — the estimator.** Feed the algorithm **the product's own reported feature values** and
  check it reproduces the product's own output. This tests the mathematics alone.
- **Layer B — the feature extraction.** Compare **your** computed features against the product's
  reported features. This tests tokenisation / parsing / counting alone.
- **Layer C — end-to-end.** Both together. C failing while A passes localises the fault to B in one
  step, with no re-derivation of a formula that was already right.

**A residual that survives a correct layer B is a finding about the TARGET, not an error in the
reconstruction** — and it is often the most valuable thing the dimension produces. Chase it; do not
tune it away. And when the target's behaviour is a defect, **do not silently replicate it** to make
the numbers match: implement the correct behaviour by default and put bug-parity behind an explicit,
documented flag. A reconstruction that quietly reproduces a bug teaches the next reader the bug.

> (Origin: hemingway — the run shipped a "validated" reference implementation of the app's readability
> scorer whose **formula was exactly right** and whose **tokenizer agreed with the target on only 41 of
> 62 real inputs**. A single end-to-end comparison read as *"the model is wrong"* and would have sent
> the run back to re-derive a correct formula. Splitting the test localised it immediately: layer A was
> **62/62** from the first run, and `Δletters = 0` in 20 of the 21 layer-B divergences pinned the fault
> to word/sentence segmentation, which was then reverse-engineered to **61/62**. The one residual that
> survived a correct tokenizer was **a defect in the product** — a literal `...` silently deletes the
> preceding text from every statistic — which no static lane could have found and which only appeared
> because adversarial inputs were included.)

---

## 6. The shared-surface seam rule (bundle ↔ session ↔ wire-capture)

`deployed-client-bundle` (static unauth SPA), `session` (authenticated browser runtime), and `wire-capture` (non-browser client) can all surface the same **API path catalog** and **feature flags**. To keep them consumable instead of triplicated **N times** (a real run forked `api-path-catalog.md` up to 7×, and "two copies agree" then mis-reads as corroboration):

- **One physical file, one canonical location, append-only.** The seam files live at `dimensions/_shared/api-path-catalog.md` and `dimensions/_shared/feature-flags.md` — **Ingestion seeds the two empty files (with a header) once, before the collector fan-out.** Every contributing dimension **appends a `source:`-tagged section** (`source: bundle | session | wire | distribution | api`) and **never creates its own copy**; each collector's `_summary.md` Artifacts list **links** the shared file rather than duplicating it. **Shared-seam allowlist:** these two filenames are the _only_ shared/single-physical artifacts — every other `raw/*` is dimension-private.
- Evaluation owns the **API-path diff** (an _adaptive N-way_ diff over whichever runtime sources were actually run — see evaluation template 3) in `data-model-api-surface.md`. A path that appears in `bundle` but never in `session`/`wire` is a dormant/unshipped route — a finding, not noise.
- **The cloned source tree is a shared working copy too (a sequencing race, not just the seam files).** On a `source:full` target, `codebase` clones into `source/<target>/`, and other dimensions (`docs`/`packages`) may read in-repo files from it. A clone-dependent collector dispatched **concurrently** with the clone can read a **mid-checkout tree** (a `.git`-only state) and silently fall back to a weaker source. So either (a) treat the `codebase` clone as a **sequencing barrier** — let it finish before dispatching any collector that reads `source/<target>/` — or (b) require each clone-dependent collector to **fetch its own copy** (GitHub API / its own shallow clone) and never assume the shared clone is ready. (Origin: Twenty — `docs` raced the `codebase` clone, saw a `.git`-only tree, and fell back to the live docs site; `packages` correctly fetched its own manifests and was unaffected.)

> **Same artifact ⇒ ONE source (never independent corroboration).** Two dimensions that string-mine the _same_ static asset (one bundle / one spec / one binary) are **one source** for the promotion rule — never a confidence band-bump — and their divergent counts are a **reported range, not a flagged conflict.** Name **one owner dimension per asset**: when `source:none` and no machine spec is served, `deployed-client-bundle` **owns** the app-API bundle mine, and `api` is re-scoped to the _published/partner_ spec only (or marked `partial` if none is served) — the others **cite the owner by path and do not re-mine**.

> **Published API ≠ the app's own API.** A product's developer/REST/MCP surface (`api`/`docs`/`packages` dims) is frequently _not_ the API its own app runs on (`session`/`bundle` dims) — different protocol, host, and operation set. Always diff the two: operations the app uses but the public API omits, **and** public operations the app never calls, are both findings. The published surface is the integration story; the runtime surface is the real architecture. When a published API exists, Evaluation emits this split as a _named_ output, not inline prose (evaluation template 3).
>
> **Paired inputs make the diff feedable (capture cue).** The split is only emittable if BOTH lanes were captured: when a target has **a published API AND an authed session**, the **pre-nav in-page fetch/XHR tap** (`session` — the app-own paths it actually calls) and the **unauth bundle/route-manifest** (`deployed-client-bundle` — the published paths) are **paired required inputs**. Install the tap _before_ navigation (else on-load app calls are missed) and capture the manifest unauth; a session that only screenshots, or a bundle mine with no session, leaves the diff un-feedable. (Origin: Akeneo — 54 session-observed app-own paths vs a 124-path unauth `/api/rest/v1` manifest, near-disjoint.)
>
> **Host-level dimensions cite, they don't append (seam allowlist).** The shared `api-path-catalog.md` `## source:` allowlist is **bundle / session / wire / distribution** (the path-contributing *runtime* lanes) **plus `api`** (the *published* lane — see below). A host-level dimension (`infra-backend-fingerprint`) that wants to cross-reference the catalog does so **from its own `raw/`** (cite hosts/paths by reference), and does **not** append its own `## source:` section to the shared file. (Origin: Akeneo — infra appended a host-level note to the path catalog; benign but blurs the allowlist.)
>
> **`api` IS on the allowlist — but it is the PUBLISHED lane, and it never corroborates a runtime path.** The published developer/partner spec belongs in the same physical file as the runtime lanes because the **published-vs-app-own diff** (evaluation template 3) is the whole point of the seam, and it is far easier to compute when both sides sit side by side under one header. Two hard constraints ride with the inclusion: (a) an `## source: api` section carries the **spec-declared** surface only — never a path the collector merely inferred; and (b) **`api` is excluded from the promotion rule for app-own paths** — a path appearing in `## source: api` and in `## source: session` is *the published surface and the runtime surface agreeing*, which is a **finding about drift (or its absence)**, not two independent dimensions corroborating one claim. (Origin: indus-sarvam — the `api` lane appended a 70-operation published catalog in violation of the then-allowlist; deleting it would have made the corpus strictly worse, because that section is exactly what made the published-vs-app-own diff feedable. The rule was wrong, not the collector.)

---

## 7. Ethics & redaction (non-negotiable — binds EVERY dimension)

Violating them voids the work. Rules 2–3 are enforced by the **runnable redaction layer in §7.0** — install it inside the tap wrapper so no secret can reach `raw/` in the first place.

1. **Never enter credentials.** The user signs into any logged-in surface beforehand.
2. **Storage: shape, not values.** No raw cookies/tokens/JWT-signatures. Decode JWT _structure_ with identity claims (`sub`,`email`,`phone`,`user_id`,`uid`,`sid`) redacted.
3. **Never save secrets.** Before any `raw/` write, redact every credential-shaped field → `<REDACTED:...>` + log in `_meta.redactions`. Field patterns: `api_key`,`*_secret`,`client_secret`,`webhook_secret`,`signing_*`,`password`,`private_key`,`access_token`/`refresh_token` in bodies, vendor-named secrets. Value patterns: `sk_live_…`/`pk_live_…` (Stripe), `sk-…` (OpenAI), `xox[bp]-…` (Slack), `AKIA…` (AWS), `gh[ps]_…` (GitHub), `AIza…` (Google), `eyJ…` (JWT), opaque base64 ≥32 chars in a credential field. Over-redact when unsure. **The executable scrubbers that enforce this rule are in §7.0 below** (`SCRUB` / `redactSecrets` / `ULTRA`); the full keyboard layer lives in `tradecraft.md`.
4. **Confirm before state changes** — anything that creates state, costs money, or is visible to others (even a draft toggle). State the exact action + estimated cost.
5. **No destructive actions.** Document the absence; move on.
6. **No bypassing safety harnesses** — redact, never evade.
7. **No exfiltration.** Captures stay local under `research/<target>/`.
8. **Respect TOS & robots.** For spec-writing / integration / competitive analysis — not credential theft, rate-limit bypass, or scraping.
9. **One browser-driving dimension at a time.** Chrome MCP is a shared singleton. Only the gated `session`/`wire-capture` dimensions drive the browser; parallel non-browser collectors use WebFetch/curl, or — if a page genuinely needs a real browser — create their **own** tab and never touch the session tab. Never run two browser-driving collectors concurrently.
10. **Verify an absence before asserting it.** A "host is NXDOMAIN / dimension is absent / endpoint 404s everywhere" claim that would _contradict another dimension_ must be re-checked by a **second method** (alternate resolver, direct HTTP probe, second source) before it is written as fact. A single-resolver/single-probe negative is recorded as _tentative_, never as a cross-dimension directive. (An HTTP 4xx/5xx means the host **resolves** — present, not absent.) **Feedback-channel / changelog absence (`community`): a 404 on the marketing host's `/changelog` is NOT proof of "no public feedback loop"** — before asserting it, probe the common alternate homes: `feedback.`/`changelog.`/`updates.`/`roadmap.` subdomains, hosted boards (FeatureOS/Canny/Featurebase), a status-page "announcements" tab, and `<link rel="alternate" type="application/rss+xml">` autodiscovery on the docs/blog. Only after those miss is the closed-feedback cap (evaluation) warranted; a wrong-host 404 alone leaves it _tentative_.
11. **Bounded live-token use (read-only, in-page only).** To fill a Pass-1 capture gap you MAY read the live token from its store and replay a **read-only** request with it — _only_ if the token is used **in-page / in-process only**, is **never written to `raw/`, never returned, never logged**, and the replay changes no state. This is the single permitted touch of a live credential; rules 2–3 (shape-not-values, never-save-secrets) still bind everything written to disk.

### 7.0 Redaction — runnable (the executable layer behind rule 3)

These are **executable guardrails**, not prose policy. Install them **inside the tap/fetch wrapper** so every captured string is scrubbed _before it reaches the harness or disk_ — credentials never reach `raw/`, and the browser MCP's content-block (`[BLOCKED: Cookie/query string data]`) never fires on a legitimate capture. Run `SCRUB(value)` on every string returned from a `javascript_tool` capture; run `redactSecrets(body, redactions)` on every object before any `Write` of `raw/*.json`; fall back to `ULTRA` only when the harness still blocks output.

#### 1. `SCRUB` — the value-shape scrubber (run on every captured string)

Combine into one `SCRUB` function and apply to every string returned from `javascript_tool`. It neutralises PII + provider-credential value shapes regardless of field name (catches credentials embedded in string concatenations like a deep-link URL or a log line):

```js
const SCRUB = (s) =>
  String(s || "")
    .replace(/eyJ[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+/g, "<jwt>")
    .replace(/[A-Za-z0-9+/]{40,}={0,2}/g, "<b64>")
    .replace(/data:image\/[^"]+/g, "<data-image>")
    .replace(/\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g, "<email>")
    .replace(
      /[a-f0-9]{8}-[a-f0-9]{4}-[a-f0-9]{4}-[a-f0-9]{4}-[a-f0-9]{12}/g,
      "<uuid>",
    )
    .replace(/[a-f0-9]{32,}/g, "<hex>")
    .replace(/(user|org|ins|sess|cli)_[A-Za-z0-9]{20,}/g, ".$1id.") // Clerk-style IDs
    .replace(/\+91\d{10}/g, "<phone-in>") // Indian phone numbers
    .replace(/\+1\d{10}/g, "<phone-us>") // US phone numbers
    .replace(/\b\d{10,15}\b/g, "<phone>") // generic numeric phone
    .replace(/[?&]([a-z_]+)=([^&\s"]+)/gi, (_, k, v) => "?" + k + "=<v>");
```

The regex table behind `SCRUB` (PII value shapes + known provider-key prefixes), so a new run can extend it:

| Pattern | Replacement | Catches |
| --- | --- | --- |
| `eyJ[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+\.[A-Za-z0-9_\-]+` | `<jwt>` | JWT tokens (3-segment base64) |
| `[A-Za-z0-9+/]{40,}={0,2}` | `<b64>` | Long base64 |
| `data:image/[^"]+` | `<data-image>` | Inline image data URIs |
| `[a-f0-9]{8}-[a-f0-9]{4}-[a-f0-9]{4}-[a-f0-9]{4}-[a-f0-9]{12}` | `<uuid>` | UUIDs |
| `[a-f0-9]{32,}` | `<hex>` | Long hex strings (hashes) |
| `(user\|org\|ins\|sess\|cli)_[A-Za-z0-9]{20,}` | `.<group>id.` | Clerk-style IDs |
| `\+91\d{10}` | `<phone-in>` | Indian phone numbers |
| `\+1\d{10}` | `<phone-us>` | US phone numbers |
| `\b\d{10,15}\b` | `<phone>` | Generic numeric phone |
| `\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b` | `<email>` | Email addresses |
| `[?&]([a-z_]+)=([^&\s"]+)` | `?$1=<v>` | Query string values |
| `sk_(live\|test)_[A-Za-z0-9]{16,}` | `<REDACTED:STRIPE_SK>` | Stripe secret key |
| `(rk\|pk)_(live\|test)_[A-Za-z0-9]{16,}` | `<REDACTED:STRIPE_KEY>` | Stripe restricted / publishable key |
| `sk-(proj-)?[A-Za-z0-9_\-]{20,}` | `<REDACTED:OPENAI>` | OpenAI API key (incl. project-scoped) |
| `xox[bpoars]-[A-Za-z0-9-]{10,}` | `<REDACTED:SLACK>` | Slack token |
| `\b(AKIA\|ASIA)[A-Z0-9]{16}\b` | `<REDACTED:AWS_KEY_ID>` | AWS access key ID (long-lived + temporary) |
| `gh[pousr]_[A-Za-z0-9]{30,}` | `<REDACTED:GITHUB>` | GitHub PAT / OAuth / refresh / server token |
| `AIza[0-9A-Za-z_\-]{35}` | `<REDACTED:GOOGLE>` | Google API key |

#### 2. `redactSecrets` — the field-name walker (run before every `raw/*.json` write)

`SCRUB` catches known value _shapes_; some providers use opaque base64 tokens with **no recognisable prefix**, so you also need a recursive walker that redacts by **key name** regardless of value shape. Run both — `SCRUB` on captured strings, `redactSecrets` on the parsed object before any `Write`:

```js
const SECRET_KEY_RE =
  /(api[_-]?key|apikey|secret|password|passphrase|pwd|client[_-]?secret|webhook[_-]?secret|signing[_-]?(?:key|secret)|private[_-]?key|service[_-]?account[_-]?key|access[_-]?token|refresh[_-]?token|bearer[_-]?token|auth[_-]?token)$/i;

function redactSecrets(obj, redactions) {
  if (obj === null || typeof obj !== "object") return obj;
  if (Array.isArray(obj)) return obj.map((v) => redactSecrets(v, redactions));
  const out = {};
  for (const [k, v] of Object.entries(obj)) {
    if (typeof v === "string" && SECRET_KEY_RE.test(k) && v.length > 0) {
      redactions.push(`${k} → <REDACTED:SECRET> (was ${v.length} chars)`);
      out[k] = "<REDACTED:SECRET>";
    } else {
      out[k] = redactSecrets(v, redactions);
    }
  }
  return out;
}

// Use:
const redactions = [];
const safeBody = redactSecrets(rawBody, redactions);
// Then write { _meta: { ..., redactions }, body: safeBody, _decoded_fields: {...} }
```

At run time, append the **current session's** own identifiers (your user UUID, email, display name, org slug) as explicit `.replace(...)` lines in a final pass over the serialized output — but **never commit those literal values into this file.** The generic UUID / email / phone / Clerk-ID patterns above already neutralise them; the explicit pass is a belt-and-suspenders step that stays local to the run.

For a logged-in session, also redact: the user's display name, the user's avatar URL, any organisation slug (`org_slug` claim) the user might want hidden, and phone numbers at any masking level. Over-redaction is cheap; a leaked partner API key is not.

#### 3. `ULTRA` — the last-resort scrubber (when the harness still blocks output)

When `SCRUB` isn't enough (e.g. the body has agent-generated content with patterns that match the blocklist), drop to a regex that strips everything except readable ASCII:

```js
const ULTRA = (s) => String(s || "").replace(/[^a-zA-Z0-9 _.,:!?\-/]/g, ".");
```

You lose JSON braces and special characters, but you keep the **structural information** (field names, enum values, lengths) — useful for getting any output at all when the body is irretrievably full of blocked patterns.

#### 4. PII / value-shape coverage (what these patterns neutralise)

- **JWTs** — 3-segment base64 `eyJ…` anywhere (body fields, not just `Authorization`).
- **Phone numbers** — Indian `+91` (10 digits), US `+1` (10 digits), and generic 10–15-digit numeric.
- **Emails** — `local@domain.tld` (and the user's own specific address explicitly).
- **UUIDs** — canonical `8-4-4-4-12` hex (and the user's own specific user UUID explicitly).
- **Clerk-style IDs** — `user_` / `org_` / `ins_` / `sess_` / `cli_` + 20+ chars.
- **Query-string values** — every `?k=v` / `&k=v` value collapsed to `?k=<v>`.
- **Long base64 / long hex** — `≥40`-char base64 and `≥32`-char hex (hashes, opaque tokens).
- **Inline image data URIs** — `data:image/…`.
- **Provider key prefixes** — Stripe (`sk_`/`rk_`/`pk_`), OpenAI (`sk-`/`sk-proj-`), Slack (`xox*`), AWS (`AKIA`/`ASIA`), GitHub (`gh[pousr]_`), Google (`AIza`).

#### 4b. Field-level redaction does NOT protect prose — scrub the VALUE document-wide

`redactSecrets` walks **field names**, so it reliably turns `{"password": "abc123"}` into
`<REDACTED:SECRET>`. It does **nothing** for the same value written into surrounding narrative —
a capture that correctly redacts the field and then explains *"the password is returned in the `abc123`
form"* has **leaked the secret anyway**, and the field-name walker will never catch it.

**Rule: once a value is identified as secret, scrub that literal from EVERY file in the capture, not just
from the JSON field it arrived in.** After redacting, grep the corpus for the literal and confirm zero hits.

**And do not trust `SCRUB` alone to find it afterwards.** `SCRUB`'s generic hex rule is `[a-f0-9]{32,}`,
which is **structurally incapable** of matching a short secret — an 8-char hex password, a 6-digit OTP, a
short numeric PIN, a 4-char confirmation code. A "sweep clean" result from a pattern that cannot match the
shape you are looking for is **not evidence of cleanliness**. Grep for the **specific literal**.

(Origin: Emergent — a per-job code-server password `<8-hex>` was correctly redacted in the `password` JSON
field of three artifacts and then quoted verbatim in the adjacent prose of all three. A later corpus-wide
sweep declared all 175 files clean; the sweep's own regex could not match an 8-char hex value, so the
"clean" result was meaningless. The leak was found only by a cross-run reader that read the prose.)

#### 4c. The literal pre-commit sweep (MANDATORY — a regex sweep is not evidence)

**Before committing a corpus, grep for the LITERAL of every value ever identified as secret for this
target — including values named in PRIOR runs' own defect notes — and assert zero hits.** Record the
literal-grep result in the run's Mode-5 output.

Rule 4b tells you *why* short secrets evade a value-shape regex. This rule is the *verification step*,
and it exists because documenting the lesson turned out not to be enough:

> **(Origin: Emergent, second run.)** The 4b rule above was written FROM the Emergent password leak — and
> the three leaked files were still in the corpus a month later, unremediated. The next run's opening
> sweep found them only because it grepped the **literal 8-char value** rather than a shape. A rule that
> records a lesson without a verification step lets the same leak ship again.

```bash
# For every value ever flagged secret for this target (incl. ones quoted in prior defect notes):
for lit in "$KNOWN_SECRET_LITERAL" ...; do
  grep -rIl -- "$lit" . && echo "LEAK: $lit" || echo "clean: $lit"
done
# Plus the field-shape pass, which catches NEW secrets the literal list doesn't know about:
grep -rIoE '"(password|passwd|secret|access_token|refresh_token|api_key)"[[:space:]]*:[[:space:]]*"[^"<][^"]{2,}"' . | grep -v REDACTED
```

A sweep that reports "clean" using a pattern structurally incapable of matching the target shape is
**not a clean result** — it is an absence of evidence. Say which patterns you ran.

#### 5. Maintenance loop

When a run uncovers a new key or value shape (a regional BSP / telephony token, a niche analytics SDK key, an internal-tool API, an opaquely-named credential field), **add the regex to the `SCRUB` table here** (and the field name to `SECRET_KEY_RE` if it's name-based) — the layer only stays a real guardrail if every new shape is folded back in.

### 7.1 wire-capture & active-interception ethics (the strictest tier)

The `wire-capture` dimension climbs a **3-rung ladder, lowest privilege first**, and records the chosen rung in `_meta`:

1. **Browser network tap (default, unauth).** Chrome `read_network_requests` + the in-page fetch/XHR interceptor on the **public** surface only. Capture request/response _shapes_; redact in the wrapper before disk. No login. → `method: wire-tap-browser`.
2. **HAR export.** Export DevTools HAR of the unauth SPA's XHRs, digest to an endpoint catalog, scrub `Authorization`/`Cookie`/secret-shaped values. → `method: wire-har`.
3. **mitmproxy / intercepting proxy — GATED, opt-in, user-run.** ONLY first-party traffic the **user owns** (their own desktop/mobile/CLI client), ONLY with explicit chat authorization, ONLY TOS-checked. The harness NEVER stands up a proxy against a third party's infra unattended; the _user_ installs the CA and routes their client — the agent only ingests exported flows. → `method: wire-mitmproxy`. Same explicit-authorization loop as session Pass-2.

> **Rung 3+ — the first-party SDK-harness lane (the deepest method for a closed runtime that PUBLISHES an SDK).** When the product ships an official client SDK (npm / PyPI / Maven / …) and the _user_ can provision credentials (free tier / trial), the deepest no-source method is not to intercept the _vendor's_ app but to **scaffold a minimal throw-away app with the product's OWN SDK + the user's keys, run it, and instrument the client you control**: inject fetch/WS/EventSource taps at `document_start` (so there is **no interceptor-timing problem** — you wrap the constructors before any connection opens), observe every request/response shape, and — with the user routing it through _their own_ mitmproxy — the fully-decoded wire. Because you **own** the client, there is no third-party-traffic concern and no missed on-load socket. **Gated exactly like rung 3 / session Pass-2:** the _user_ provisions the account + keys and runs any proxy; the **agent never self-signs-up and never enters credentials**, only scaffolds + instruments, and confirms any cost/quota first. → `method: sdk-harness` (a first-party runtime lane; confidence = what is actually observed, typically high since you control the client). Especially powerful for `runtime:reachable + source:none` targets whose published SDK is the only public mirror of the runtime — and it independently corroborates the `api`/`packages` lanes.

> **Never fold a browser-rung into a non-browser (curl/WebFetch) collector.** Rung 1 _is_ a browser tap, and the browser is a singleton owned by `session`/`wire-capture` — routing it to a curl collector produces zero `source: wire` rows (a real run did exactly this). When `wire-capture` is **folded**, the rung-1 tap is the **in-page interceptor that `session` already installs**: `session` appends its observed paths to `dimensions/_shared/api-path-catalog.md` as `source: wire`. A folded `wire-capture` still leaves a one-paragraph `_summary.md` (`status: folded`, naming the fold target) per §1 — "folded" is recorded, never silently dropped.

> Wireshark / packet capture is explicitly **not** used: for HTTPS it yields only TLS ciphertext, and a packet sweep of a third party is the bright line this harness will not cross.

### 7.2 distribution-artifacts ethics

Downloading a public binary (CRX/`.asar`/`.ipa`/`.apk`) and reading its manifest/strings is fine. But binary **decompilation** can cross from "inspect" to "decompile in violation of EULA": the agent reads the artifact's own license/EULA first and records a restrictive one as an Open question rather than silently decompiling.

---

## 8. Completeness check

A capture is complete when: `_summary.md` exists with full provenance frontmatter + honest `status`; **every Finding is backed by a `raw/` artifact — an inline value is NOT sufficient (§5.3)**; the dimension's **§5.4 dump manifest is satisfied verbatim**; `raw/` artifacts carry `_meta` (with `verbatim:` and, if bounded, `sampling:`); redaction pass run; `gaps` recorded (never fabricated).

> **Algorithm reconstructions have an extra bar (§5.6):** if the dimension reconstructed *behaviour*
> rather than a catalog, it is `partial` until the implementation is executable, differentially tested
> against the live product with the **estimator and feature extraction scored as separate layers**, and
> the inputs + observed outputs are frozen as re-runnable fixtures. Record the per-layer pass rates.

> **Self-check before writing `status: complete`:** count the verbatim artifacts in `raw/`. If that count is **zero**, the dimension is `partial` no matter how good the prose is — you have written an essay, not a capture. If a Finding names a number, a path, a version, or a field, open the file that contains it; if no such file exists, write it now (§5.3).

**Summary-first + bounded, visible fan-out (never go silent long enough to look dead).** Write a `status: partial` `_summary.md` with whatever is mined so far **before** any long-tail fetch loop (downloading many chunks, crawling many pages, deep pagination), then enrich it in place and flip to `complete`. **Bound** every fan-out — cap it by count / total bytes / wall-time (e.g. ≤40 items or ≤90 s) and record the remainder as a `gap` rather than fetching exhaustively; emit a progress note per batch. A collector that fetches everything before writing anything can run many times longer than its siblings and read as hung to the orchestrator, forcing a needless salvage. Partial-but-finalized beats complete-but-silent.

**A single reading of a mutable numeric field is a sample, not a schema.** Before describing any observed
limit / quota / cap / budget / ceiling as **fixed**, either re-read it at **≥2 points in the entity's
lifecycle** (e.g. at creation and again mid-run) or record it as `observed-at-T` rather than as a structural
property. A first reading that merely *looks* bounded is an observation, not a contract — promoting it to a
structural fact is an inference-stated-as-fact defect, and it is one the same run usually refutes on its own.
(Origin: Emergent — `GET /budget/{job_id}` returned `max_budget: 25.0` on job acceptance and was written up
as "a hard per-job credit ceiling"; the *same field on the same job* later read `100.87`, because it tracks
available account credit. The cap was never real.)

**And when a re-read moves, look for the SIBLING FIELD that is the real cap.** Do not merely re-label a
moving value `observed-at-T` — **search the write payload for a client-declared field that states the
limit outright.** A moving server-side figure and a static request field are usually two different
things, and the request field is the contract. (Origin: Emergent, second run — `budget.max_budget` read
`25.0`, then `100.89` on the same job, while the submit payload carried `per_instance_cost_limit: 25`
all along. The moving number was account credit; the static one was the actual per-job cap.)

**Capability gap = a checkpoint, not a silent fallback.** When a _preferred_ capture capability is found **unavailable at runtime** — a CDP/remote-debugging port is closed, Chrome / the browser MCP is unreachable, a gated proxy or tool can't be stood up — such that you would proceed on a **materially weaker** method (e.g. source-only or a digest instead of a live wire tap), **pause and confirm with the user before continuing.** State plainly: what is unavailable, the degraded method you would fall back to, and the coverage that would be lost. This is distinct from a _planned_ method pre-grade (e.g. `openapi-verbatim` → `docs-reconstructed` when no spec is served), which is **recorded, not confirmed** — the checkpoint is for an _infrastructure/capability_ gap discovered mid-run, where silently degrading hides a real coverage loss behind a confident-looking capture. (A real run found the CDP port closed and silently fell back to a browser-MCP tap; it worked, but the degradation should have been surfaced first.)


**A `session` capture is not `complete` until the verbatim envelopes exist, not just the decoded prose.**
The §9.5 contract requires `raw/*.json` (with `_meta` + redacted `body` + `_decoded_fields`) so a later run
can *reproduce* the capture, not merely read about it. A session lane may therefore only be marked
`status: complete` once **at least** these three exist as redacted `raw/*.json`:

1. the **identity/account** read (`/me`, `/account`, `/credits`, or the tenant equivalent),
2. one **canonical list** read (the entity taxonomy), and
3. the **primary write** request+response envelope.

Rich `captures/*.md` narrative does **not** substitute — prose can record a *schema* while losing the
*envelope* (status codes, headers-shape, error shape, field ordering) that the reproducer needs. If an
envelope genuinely could not be captured, mark the lane `partial` and record it in `gaps:` rather than
calling it complete. (Origin: Emergent — the session lane produced excellent decoded captures and a full
shared path catalog, but zero `raw/*.json`, so the endpoint envelopes are unreproducible.)

---

## 9. The session dimension — logged-in browser-session playbook

The method for the **session** dimension: discovering the _authenticated runtime_ of a product by driving it through a browser (Chrome MCP) and instrumenting it. It produces a section-by-section understanding of a product's wire protocol, auth, and state model **without source-code access**. The `session` agent follows this section.

This is the deepest and most sensitive dimension. The ethics in §7 above bind hardest here.

### 9.1 Prerequisites (verify before anything else — all three non-negotiable)

| # | Precondition | How to verify |
| --- | --- | --- |
| 1 | Target is **pre-logged-in** in Chrome | open the target URL in the connected tab → an authenticated surface (a dashboard, not a sign-in page). The agent never enters credentials. |
| 2 | **Claude Chrome extension** installed | `mcp__claude-in-chrome__tabs_context_mcp` returns live tabs |
| 3 | **Chrome CDP** connected | the same call returns stable tab IDs |

If any fails, return `blocked` naming the missing precondition.

### 9.2 The three-pass strategy

A single linear sweep wastes context on things you can read for free and wastes credits on actions whose response shape you could infer. Use three passes.

**Pass 1 — read-only mining (no money, no state change).** If the user already has data (a project, conversation, account, order), mine it first. Hit every list/state/history endpoint with the existing token; capture verbatim bodies. This yields ~70% of the spec for free. Mine, by role: identity (`/me`, `/account`, `/workspaces`), list (the canonical taxonomy), single-resource state (server-injected vs client-sent fields), history/event-log (a whole trajectory in one call), runtime/preview state, per-resource billing, account billing, feature flags/cohort, **the account's `entitlements`/`permissions` arrays** (the single richest data-model artifact — `entitlements`/`features` enumerate every gated feature _including ones your plan hides_; `permissions` are CRUD-verbs-per-entity, so the verb set enumerates the domain entity set — `X_CREATE/VIEW/EDIT/DELETE` ⇒ entity `X`; mine these first), integration catalog (`/mcp/`, `/integrations/`), and schema/capabilities (`/openapi.json`, `/.well-known/*`, `/page-metadata`).

> **Mine the cohort / entitlement / flags endpoint FIRST — it is routinely the single richest artifact of a run.** Modern SaaS SPAs resolve the whole tenant's feature surface in **one authenticated call** (`/user/cohort`, `/entitlements`, `/flags`, `/bootstrap`, `/me/features`, `/config`). One request returns the **complete flag key-set with this tenant's resolved values** — richer than an exhaustive bundle string-mine, and it names capabilities that have **no public marketing footprint at all**. Read the **key set** as the durable finding and the **values as tenant-scoped**: a `false` means *not enabled for this tenant*, **never** "the feature does not exist". (Origin: Emergent — a single `POST /user/cohort` returned **72 flags**, including whole unshipped surfaces and a sibling product's feature family, in one call.)

**Pass 2 — fresh-run validation (with explicit cost/state-change confirmation).** Only after Pass 1, and only with explicit user OK stating the exact action + cheapest mode + estimated cost. Catches what Pass 1 can't: the submit/write request body, live state-machine transitions, the completion event, server-side resolution of shorthand, ID/slug-generation, the real iteration endpoint, and the full realtime event catalog. **Legitimate to skip** if the user opts out — leave write-side shapes as open questions; never fabricate them.

> **Record the write-side decision as a machine flag, not re-derived prose.** A Pass-1-only run leaves the _entire_ write/generate/launch surface unobserved — feature _existence/gating_ is solid, but feature _execution behavior_ rests on marketing/docs and must never be promoted to fact (this caveat otherwise gets re-hedged across 3–4 rollups every cost-gated run). So the session `_summary.md` frontmatter carries `write_side_observed: <true|false>` and `pass2: <not-applicable | offered-pending | declined-by-user | deferred | authorized-declined | done>` (this flag is the load-bearing part — it lets Evaluation render the write-side gap _once_).

> **An AUTHORIZED Pass-2 that is never exercised is a DEFECT, not a default.** When the user has granted
> Pass-2 (explicitly, or by saying the run may spend credits / test it), the session lane must do one of
> exactly two things before Evaluation:
>
> 1. **Exercise the cheapest write action** that closes the gap — and there is almost always a cheap one:
>    a simulated/dry run, a draft save, a version bump, a single smallest-unit generation. Price it, name
>    it, run it, capture the request+response envelope.
> 2. **Record `pass2: authorized-declined` with the reason** (the user wrapped the run early; the cheapest
>    action still mutated shared state; no non-destructive write exists), plus the exact actions and costs
>    that would close it.
>
> `pass2: deferred` is **not** a valid terminal state for an authorized run — it is the state that lets an
> authorization quietly expire. Mode 5 treats an authorized-but-unexercised Pass-2 as a **medium Ingestion
> defect** and charges it, because a run that holds permission to observe the write side and ships without
> it has left its single largest blind spot open by inattention rather than by policy. (Origin:
> indus-sarvam — the user granted credit spend mid-run; the orchestrator acknowledged it, planned around
> it, asked which action to take, and shipped `write_side_observed: false` when no reply arrived. The
> whole write surface, the deploy transition and the live-call realtime plane went unobserved with 100
> credits unspent.) When it would help a later run resume, optionally add `captures/pass2-followup.md` listing the **exact cheapest-mode write actions + est. cost** that would close the gap. Discovery seeds the `pass2:` decision in `00-recon-plan.md` for any `auth:have-login` target.

**Pass 3 — synthesis.** Walk the captures, pull each finding into the dimension's `_summary.md`, note the cleverest design choices, defaults imposed, performance characteristics, and open questions.

### 9.3 The P0 → P6 phases

The same phases apply whether the target is an AI builder, analytics dashboard, CRM, or collab tool:

| Phase | Focus |
| --- | --- |
| **P0 Surface** | routes, bundles, storage shape, the API host(s), the auth tell; build the **coverage checklist** from any route/op catalog a sibling declared (§9.4) |
| **P1 Auth** | token type + issuer + storage location + refresh path; tenancy model; **disambiguate the transport by omission** (replay one authed read with each credential mechanism dropped) — and never test auth with a schema-only field |
| **P2 Core action** | the primary write — the submit/create body and its server resolution |
| **P3 Streaming** | the realtime transport (WS / SSE / WebChannel / Pusher / none) + event catalog; **fingerprint binary frames by signature, never log them as opaque** (§9.4) |
| **P4 Runtime/state** | sandbox/preview/build/deploy state; the live result surface |
| **P5 Iteration** | the follow-up/edit endpoint (often surprisingly named) + any auto-iteration loop |
| **P6 Synthesis** | reconcile into the dimension's `_summary.md` + `captures/` + `curls.md` |

### 9.4 Runtime taps (Pass-1 instrumentation, run in-page)

> **The runnable taps + the ~27-entry Chrome-MCP gotcha catalogue + the dead-end recovery table + the NuStack read-side recipe + the fingerprint tables live in `tradecraft.md`** — paste them, do not re-derive. The prose below is the _why_; that file is the _keyboard layer_. The §7.0 redaction scrubbers (`SCRUB`/`redactSecrets`/`ULTRA`) install inside these tap wrappers.

> **Install the in-page interceptor FIRST, then move section-to-section WITHOUT a full reload.** CDP `read_network_requests` is unreliable for an already-loaded single-page app: it starts tracking only when first called and silently misses requests already in flight or served from cache (seen on a real run where only a third-party analytics beacon was captured while the app's own API calls went unseen). Treat it as a _supplement_; the in-page interceptor is the primary capture. A full-page `navigate` wipes the interceptor, so traverse the app using whatever it exposes — **clicking in-app nav works on any stack**; or call the client router if one is reachable (`history.pushState`+`popstate`, React Router / Vue Router / Angular Router, or a globally-exposed router such as `window.next.router` on Next.js). The principle is framework-agnostic; the specific call is just whatever that app ships.

- **WebSocket / EventSource / fetch+XHR interceptors** — wrap the constructors/methods to log URLs, frames, and bodies (shapes, not secrets — redact in the wrapper before anything reaches disk). Add a `SCRUB(value)` step inside the fetch wrapper so credentials never reach the harness. **Never log a binary frame as opaque `<binary>`.** Any realtime/streaming/collab layer (token streams, CRDT/OT edits, presence, telemetry, game or market state) may ship `ArrayBuffer`/`Blob` frames; for each _distinct_ frame capture its **byte-length + a hex dump of the first ~16 bytes** — enough to fingerprint the wire format (protobuf, MessagePack, CBOR, Yjs/CRDT, gRPC-web, or a text/topic-prefixed channel protocol) without recording content. **Decode the signature as ASCII first** — many "binary" protocols are framed text (STOMP, Phoenix, Centrifugo, a `+topic` channel prefix), so a legible prefix names the protocol in one step before a binary-codec guess. The frame _signature_ is the finding; the payload is not.
  - **Realtime-on-load capture (the connection that opens before you can wrap it).** The "install the interceptor FIRST" rule assumes you are _already on_ the target surface; it does **not** cover a realtime connection that fires during the **initial load of a page you `navigate` to** (or a lazy-loaded/`iframe`d live demo). A post-load interceptor **cannot** wrap a `WebSocket`/`EventSource` constructed before injection, and an `offline`→`online`/`focus` reconnect-trigger may spawn no new connection on such a demo — so frame capture comes back empty (a real run got 0 frames on a collab-demo this way). When that happens: (a) capture the **CDP-visible auth/token-mint fetch** that precedes the socket (it names the realtime host/handshake); (b) **the technique that works — client-side-navigate within the SPA** (a router `<Link>` click, not a full reload): the in-page wrapper **survives client-side nav**, and the next route **mounts a fresh room → a new socket through your wrapper** (this captures the handshake + frames where `offline`/`online` reconnect-triggers don't — a real run decoded Liveblocks' v8 JSON protocol this way after a 0-frame first pass); (c) if a raw **CDP `Network.webSocketFrameReceived`** tap is reachable, it captures frames protocol-level with no injection-timing problem at all (the browser MCP's HTTP-only `read_network_requests` does **not** surface WS frames — that gap is what forces (b)); and (d) fall back to the `codebase`/`distribution` lanes for the wire format — record it as a _folded_ gap, not a silent miss, since the protocol is then sourced independently rather than live-observed. **Decode captured frames against any in-repo protocol enums** (message-code / op-code tables) so the live wire _validates_ the source read.
- **Storage shape inspector** — enumerate localStorage / sessionStorage / cookies / IndexedDB **key names + lengths**, never raw values. **Treat this census as an independent vendor-corroboration lane, not just an auth probe.** Key *names* alone fingerprint the embedded vendor stack (see the `tradecraft.md` §2.1 prefix table) and are derived from what the app **actually stores** — so they corroborate (or refute) a CSP/bundle-declared vendor list from a genuinely different source. Run it early and diff it against the CSP. (Origin: Emergent — a 40-key census confirmed Supabase, Pusher, PostHog, New Relic, VWO and Razorpay *and* surfaced Intercom, Criteo, Reddit Pixel, Bing UET and Commission Junction, none of which the CSP named.)
- **JWT decoder** — decode the token _structure_ (alg, claim names, exp) with identity claims (`sub`, `email`, `phone`, `user_id`, `uid`, `sid`) redacted.
- **Auth-channel disambiguation.** Token _location_ is itself a tell: not in localStorage ⇒ an httpOnly cookie **or** an SDK store (IndexedDB — Firebase/Auth0/Supabase/Cognito keep tokens there, which is why a "no JWT in localStorage" result is a finding, not a dead end). Determine the _required_ transport by **omission**: replay one harmless authed read — using a field that needs the authed context (`viewer`/ `me`), **never** a schema-only field (GraphQL `__typename`, a health route) that resolves unauthenticated — with each candidate credential dropped (cookie-only, then header-only); whichever response comes back unauthorized names the channel the app actually requires.
- **MANDATORY realtime library tap — a `window.WebSocket` wrapper is NOT sufficient evidence of absence.** If a realtime SDK global is present (`window.Pusher` / `Ably` / `io` / `Centrifuge` / `Phoenix` / `Socket`) **or** a `pusher*` / `ably-*` / `*Transport*` / `socket*` key appears in the storage census, the **library-specific event-bus tap is REQUIRED before ANY realtime claim — positive or negative.** These libraries open their socket through `window.WebSocket` **at app boot**, before any wrapper you install post-load; a constructor wrapper is therefore *structurally blind* to them (gotcha #13/§2.12), and "my tap saw no frames" is **not** evidence the channel is idle. Install the tap (`p.bind_global(...)` for Pusher; the same pattern via `Ably.Realtime.instances` / `io.managers` / `Centrifuge` handles) and read the live handshake (`state`, `cluster`, `channels`) — see `tradecraft.md` Tap toolkit item 4. **Reporting "vendor X carries no traffic" without library-tap evidence is a named Ingestion defect.** (Origin: Emergent — a post-boot `window.WebSocket` wrapper saw zero Pusher frames and the run published "Pusher does not carry build streaming" as a refuted hypothesis; Pusher was live the entire time — `connected`, `cluster: us2`, channel `user-{uuid}`, carrying `traj-update`/`terminal-state`/`streak-update`.)
- **Third-party-SDK fingerprint** (cross-checks `infra-backend-fingerprint`). Storage key names + the network host list name the embedded vendor stack — auth (Firebase/Auth0), flags (LaunchDarkly/Statsig), support (Intercom/Zendesk/Ada), analytics (Segment/Amplitude/Datadog), payments (Stripe). Record them; they corroborate the sub-processor list _independently_ (a different dimension → promotes confidence).
- **Drive op coverage from any catalog a sibling declared.** If `deployed-client-bundle` (or `api`/`docs`) recovered a declared **route / endpoint / operation catalog** — a REST path table, a GraphQL operation set, an RPC/tool method list — consume it as a **coverage checklist**: walk the routes/actions that trigger the un-observed entries so the declared-vs-observed diff (§6) reflects a real walk, not just the landing page.
- **Server-side-read (RSC) check — don't mis-record "no read API."** In a modern Next.js App-Router / React-Server-Components app (and similar SSR/streaming stacks), the initial **read** data is fetched _server-side_ with the session and streamed — so the client wire shows only mutations, asset/media fetches, telemetry, and token-mints, and **no client-side data-read XHRs fire on load even though the UI is fully populated.** Before recording "read API not observed," check for the server-side path: a document/`?_rsc=` response with `text/x-component`, embedded `__next_f`/flight payloads, `_next/data` JSON, or server actions. When reads are server-resolved, say so — and note that a sibling bundle's app-own _path count_ then **overstates the client read surface** (fold this into the published-vs-app-own diff, §6, rather than reporting a phantom "unused routes" gap). (Writes can be off-main-thread too — a generate POST issued from a **Web Worker** won't appear in a main-thread fetch tap; corroborate via CDP / Performance resource timing.)

### 9.5 The capture contract (richer than a standard dimension capture)

Under `research/<target>/dimensions/session/`:

- **`captures/p*.md`** — per-phase narrative for the human: frontmatter (title, summary, scope, status), **Method** (taps run, actions taken), **Findings** (concrete values — URLs, fields, status codes), **Inferences**, **Diff vs reference**, **Open questions**. Decoded, opinionated; don't paste raw bodies here.
- **`raw/*.json`** — verbatim endpoint captures for the reproducer: a `_meta` header (endpoint, auth-shape, captured_at, status, byte_length, redactions) + `body` (PII + secrets redacted) + `_decoded_fields`. Digest bodies > 50 KB into a structured summary.
- **`curls.md`** — reproducible curl per endpoint, the auth model documented, a variables block, and an endpoint-catalog table.
- **Visual / surface map** — `captures/screens/` (**preferred when available**): one annotated screenshot per primary surface (each major nav section), redact visible PII before saving. **When screenshot capture is unavailable** (e.g. the runtime lacks an image tool, or the app holds a persistent connection that never reaches page-idle), a **textual `captures/surface-map.md` is the accepted substitute** and satisfies this visual-map requirement — and the **visual-coverage gap is recorded in the `gaps:` frontmatter** so Evaluation reads a _known constraint_, not a silent miss. (The per-surface screenshot set is then **completed in Mode 4 — Cartography** (`cartography.md` §D), which re-enters the live product to capture each surface alongside its IA depth; the session-pass surface-map is the seed, the Cartography screens are the finished deliverable, and `screenshot_coverage` is the measured gap.)
- **`captures/pass2-followup.md`** (optional, when Pass 2 is deferred/declined and it would help a later run resume) — the exact cheapest-mode write/generate actions + estimated cost that would close the write-side gap (§9.2). The `write_side_observed`/`pass2` frontmatter flags are mandatory; this file is the convenience.
- **`_summary.md`** — the standard dimension bridge file (provenance frontmatter per §2) pointing into the above, so Evaluation treats session like any other dimension. `method: session-playbook`, `confidence: low` for Pass-1-only/inferred. **Session frontmatter additionally carries `write_side_observed: <true|false>` and `pass2: <not-applicable | offered-pending | declined-by-user | deferred | done>`** (§9.2) so Evaluation renders the write-side gap once, machine-readably.

### 9.6 Session ethics (binds hardest here — see §7 above)

Never enter credentials; storage shape not values; redact every secret **inside the tap wrapper** before it reaches disk; **confirm before any Pass-2 state change or cost** (exact action + exact cost); no destructive actions; no bypassing the harness (redact, never evade); captures stay local; respect TOS. The provenance/redaction rules in §1–8 apply to everything written here.
