# Hosting, verification, and the REVIEW REQUIRED report

## Run directory layout

```
<run-dir>/
  site/        ← the deployable folder. Deploy THIS path, never its parent.
    index.html
    assets/    css/js exactly as the source named them — do not rename
    media/     re-hosted external media, grouped by source host
  reports/     ← run-report.json, substitution-report.json, REVIEW-REQUIRED.md
```

`reports/` must never reach a public host: it contains the user's rights declaration and an itemized list of the original owner's liabilities. Deploying the run directory instead of `site/` publishes both.

New client media goes in `site/img/brand/` with descriptive kebab-case names (`owner-headshot.jpg`), never `IMG_4821.jpg`. Re-hosted media keeps its source filename, hashes included.

## Host profiles

| Host | Deploy | Per-file cap | Oversized media | 404 behavior |
|---|---|---|---|---|
| **Cloudflare Pages** (default) | `wrangler pages deploy <run-dir>/site --project-name=<slug>` | **25 MB** | R2 or Stream | serves 200 + HTML unless `_redirects` has `/* /404.html 404` — **add it** |
| **Vercel** | `vercel deploy --prod` (from `site/`) | ~100 MB, plan-dependent | Vercel Blob | configure in `vercel.json` |
| **Netlify** | `netlify deploy --prod --dir=<run-dir>/site` | ~100 MB | Netlify LFS / external | `_redirects`, same syntax as Cloudflare |
| **GHL-native** | hand off to Page Studio (`compose_page` / `compose_website`) | GHL media library | GHL media library | GHL-managed |
| **Folder only** | zip `site/` and hand it over | n/a | advise per their host | n/a |

Create or confirm the project **before** deploying — Cloudflare Pages fails opaquely when the project doesn't exist.

## Oversized media

`mirror.mjs` records anything over the cap in `run-report.json` under `oversized` and does **not** download it. Route those files explicitly:

```bash
wrangler r2 bucket create <client-slug>-media
wrangler r2 object put <client-slug>-media/<file> --file ./<file>
# then bind media.<clientdomain> to the bucket and rewrite the refs to that host
```

Never leave oversized media hot-linked to the original owner's CDN in a delivered clone. It breaks the day they delete the file, and until then the original owner pays for the bandwidth.

## Verification

```bash
node scripts/verify.mjs --url "<deployed-url>" --dir <run-dir>
```

1. **Content-type, not status code.** Static hosts return 200 with the page shell for a missing asset. An asset is present only if the served content-type matches its extension. (Proof: requesting `/reports/REVIEW-REQUIRED.md` on a correctly-deployed clone returns `200 text/html` — the file is not there at all.)
2. **One retry before calling anything broken.** A CDN rollout or a burst rate-limit can answer one request with HTML. Verified live: an unretried pass reported 9 broken assets, a clean re-run reported 2, and the truth was 0. A flaky verifier is worse than a strict one.
3. **Inherited breakage is separated out.** A reference the *source* page was already serving badly is reported but does not fail the run — redeploying cannot fix a file the original site does not have either.
4. **Old tokens at zero, new tokens above zero**, checked against `substitution-report.json`, with asset paths masked so a filename never counts as a brand hit — and checked in the **rendered** text too, which is where a tag-split brand mark hides.
5. Exit code 0 = pass, 5 = fail. A failing verify means fix and redeploy — do not show the user the URL first.

Then the part no script can do: open the deployed URL in a browser and confirm every `<img>` reports `naturalWidth > 0`, no `<video>` has an error, and `document.body.innerText` contains **zero** occurrences of the old brand. Report counts. "Looks right" is not a verification.

**Force lazy images to load before you measure.** A page with `loading="lazy"` images reports `naturalWidth === 0` for everything below the fold, and you will report a pile of broken images that are perfectly fine. Set `loading = 'eager'`, fill any `data-src`, scroll to the bottom, wait, and only count images where `img.complete` is true. Live-caught 2026-07-30 on a WordPress/Elementor page with 43 lazy images: the "broken" list was entirely lazy-loading, and every file served 200 with the right content-type.

The rendered-text check earns its place twice over: it catches tag-split brand marks, and it is the only way to see text a third-party widget injects at runtime. A cloned page keeps the original owner's chat widget live until you repoint it — it loads their business identity from their account, and nothing in the files can be edited to change that.

## Walking the user through REVIEW REQUIRED

Lead with consequences, in this order:

1. **MUST REPOINT first.** Every item here routes something real — a lead, a booking, a payment, an ad conversion — to the original owner. Read out the form endpoints, payment links, booking embeds and pixel IDs by name and ask what each should become. A cloned page that still posts to someone else's CRM is the worst failure mode this skill has.
2. **The regulated-claims flag next.** FDA/medical/efficacy language, "clinically proven," "board certified." These were substantiated (or not) by a different business, and the exposure follows whoever publishes them. Ask directly whether the new business can stand behind each one.
3. **Other people's content.** Names, street addresses, testimonials, before/after photos, staff headshots. Never delete these silently — deleting a testimonial block breaks the layout and hides the problem. Propose replacements and let the user decide.
4. **Ratings, counts, tenure, awards, pricing.** Numbers that describe the original business. Each needs a true number for the new one or it comes out.
5. **Excluded and unresolved.** Oversized files, assets the origin didn't actually serve, third-party scripts left external, and any token with no replacement value.

The report is deliberately over-inclusive — a quoted line of marketing copy shows up alongside a real customer testimonial. Say so. Over-flagging costs a minute of reading; under-flagging ships someone else's customer's face on a page selling a health service.

Close with the report's own line: *everything in it describes the original business, not the client, and each item needs replacing or removing before the page takes real traffic.*

## Runtime-loaded assets (the render→repair loop)

A file-level check cannot see a URL that JavaScript constructs. The pattern that bites hardest:

1. A page-builder publishes an asset base path in inline config.
2. At runtime it concatenates chunk filenames onto that base.
3. Those filenames appear in no attribute, stylesheet or string literal, so the mirror never fetches them.
4. The rebrand rewrites the base path's host, so every runtime URL resolves against a domain that does not exist.

The result is a page that passes every static check and renders wrong. Live-caught 2026-07-30: six Elementor handler chunks failed, `section-frontend-handlers` among them, so no section background was painted — the hero was blank and its white headline invisible on white.

`mirror.mjs` now prevents step 4 by making same-origin asset URLs root-relative. Step 3 still needs the loop:

```
render the clone → collect failed same-origin requests → repair.mjs → render again
```

Repeat until nothing fails. One pass commonly reveals the next, because code restored in pass one runs in pass two and asks for its own dependencies.

## What "done" means

A clone run is done when all of these are true, and not before:

- the render→repair loop has run to zero failed requests
- the deployed URL is live and `verify.mjs` exits 0
- a browser check confirms zero broken images and zero video errors
- `REVIEW-REQUIRED.md` has been shown to the user and the MUST REPOINT items have been answered
- `reports/` is not on the live host
