---
name: clone-site
description: Clone an existing live website or landing page from its URL and rebrand it for a new client — same layout, images, video and CSS, with new company name, contact details, colors and copy. Starts by asking what rights you have to the page and records your answer; with rights it makes a true byte-for-byte clone, without them it extracts the page's real design system and rebuilds on that. Every run ends with a REVIEW REQUIRED report of everything still belonging to the original owner — lead, booking and payment destinations, other people's names and testimonials, and inherited regulated claims. Triggers on clone this site, clone this page, copy this website, duplicate this funnel, rebuild this landing page for, make this site for my client, mirror this URL, clone URL, copy this landing page, steal this layout, make me one like this.
compatibility: Claude Code, Claude Cowork, Claude.ai
---

# Clone Site

A clone is a **copy** operation, never a generation operation. You never retype a page from your understanding of it — the scripts in `scripts/` copy the bytes, you build the substitution map, and you verify the result. That distinction is the whole skill: models paraphrase, and a paraphrased clone is the thing that comes back "not even close."

One run = **rights → acquire → inventory → rebrand → host → verify → REVIEW REQUIRED report.**

**The scripts live at `~/.claude/skills/clone-site/scripts/`** (installed with GHL Command). Call them with that path — dependency-free Node, nothing to install. Work in a run directory outside the user's repo unless they name one.

## Hard rules

1. **Ask for rights first, record the answer, then proceed.** Step 0 below. You are not a rights verifier; the user's declaration is the user's responsibility, and recording it makes that explicit.
2. **Coach, never block.** "No rights" routes to the style lane — never to a refusal, and never to a from-memory redraw.
3. **Copy, never regenerate.** If a fetch fails, say so. Do not reconstruct a page you could not download.
4. **Never silently fix ownership content.** Names, addresses, testimonials, likenesses and claims belonging to the original owner go into the report for a human decision. Flag and propose; never quietly delete or invent replacements.
5. **No credentials or payment identifiers are ever carried into a live clone.** `pk_live_*`, payment links, form endpoints and webhook URLs are MUST REPOINT findings.
6. **Verify by content-type and by rendering, never by status code.** Static hosts serve 200 + HTML for a missing asset.
7. **Never deploy the `reports/` folder.** It contains the user's rights declaration and a list of the original owner's liabilities. Deploy `site/` only.

## STEP 0 — Rights (before any bytes are copied)

Ask this first, in these words, before fetching anything:

> **Before I copy this page — who owns it, or what permission do you have to copy it?**
> 1. **I own it**
> 2. **My client owns it and authorized the migration**
> 3. **I have written permission from the owner**
> 4. **None of these** — I found the page and want something like it
>
> I'm not checking this, and I'm not going to argue with your answer — it goes in the run report so it's on the record, and it decides which of two ways I build this.

Map the answer to `own` / `client-authorized` / `written-permission` / `none` and carry it as `--rights`. `mirror.mjs` refuses to run without it, so this step cannot be skipped by accident.

Do not lecture. Do not repeat the rights topic later in the run — it is asked once, recorded once, and the audit at the end carries the practical consequences.

| Answer | Lane |
|---|---|
| own / client-authorized / written-permission | **Full clone** — byte-for-byte mirror + rebrand |
| none | **Style clone** — extract the real design system, rebuild with original copy and licensed media |

If the user has no rights but insists on a byte-for-byte copy, tell them once, plainly, what that exposes them to (copyright in the layout and copy, other people's testimonials and likenesses, trademark, and inherited advertising claims), and that you'll do the style lane instead — which in practice looks closer to the original than a redraw does. Read `references/rights-and-lanes.md` for the full wording and the edge cases.

---

# LANE A — Full clone (rights declared)

## STEP 1 — Scope and ask (one message, defaults offered)

1. **Source URL — and how much of the site?** Ask this outright; never assume a default:
   > *"Just this one page, or the whole site? If it's the whole site, how deep should I follow the navigation — one level (the main menu), two levels (menu plus the pages those link to), or everything I can find?"*

   A one-page mirror inherits the original's entire menu, so every nav item 404s. For a real client migration the answer is almost always the whole site. Give them the trade-off in their terms: a 36-page contractor site at depth 2 came to **2,104 files and 935 MB**, and the audit then listed **1,598 images with the old brand baked into the artwork**. Run the crawl, then show them the discovered page list and the size before going further.
2. **New client facts** — company name, phone, email, address, domain, owner/staff names.
3. **Host** — default **Cloudflare Pages**; also Vercel, Netlify, GHL-native (→ Page Studio), or "just give me the folder."
4. **Brand changes** — keep the original colors/fonts, or new hex/font values?
5. **Media** — re-host everything (default and recommended) or keep hot-links (fragile, and it bills the original owner's bandwidth).

If they give a URL and nothing else, run step 2 first and show them the inventory — the substitution questions are much easier to answer once they can see what is actually in the page.

## STEP 2 — Mirror (deterministic)

```bash
# one page
node scripts/mirror.mjs --url "<URL>" --out <run-dir> --rights <declaration> --declared "<their words>"

# the whole site — follows the site's own navigation
node scripts/mirror.mjs --url "<URL>" --out <run-dir> --rights <declaration> \
  --crawl --depth 2 --max-pages 40
```

**Answer the single-page-or-whole-site question with `--crawl`, not with a caveat.** A one-page mirror keeps the original's full navigation, so every menu item 404s the moment the client clicks it — the first thing they will do. Crawling writes each page at its own path (`/contact/` → `contact/index.html`) and rewrites `<a>` links to your own pages as root-relative so the nav works locally and after deploy. `<link rel=canonical>` and `og:url` keep their host on purpose, so the rebrand can point them at the client's real domain.

Depth 2 with a 40-page cap covered a 36-page contractor site in one pass. Assets are shared across pages, so the marginal cost per page is small — but the total is not: that site came to 2,104 files and 935 MB.

Creates `<run-dir>/site/` (deployable) and `<run-dir>/reports/run-report.json`. It collects assets from **four** places — miss any one and images break:

1. HTML `src` / `href` / `poster` / `srcset`
2. `url(...)` inside every downloaded stylesheet
3. Full CDN URLs baked into JS bundles
4. **Root-relative paths inside JS bundles** (`"/logo.png"`) — the one everyone misses; it caused 6 broken images in the proving run

It also does two things that decide whether the clone renders at all:

- **URLs inside escaped JSON attributes** (`data-settings="{&quot;url&quot;:&quot;https:\/\/host\/hero.jpg&quot;}"`) are found and localised. Page-builder background slideshows and galleries live here and are invisible to a src/href scan.
- **Same-origin asset URLs are made root-relative**, including inside inline JS config. A page-builder publishes an asset base path and concatenates chunk filenames onto it at runtime; left absolute, the rebrand points every runtime-built URL at a domain that does not exist.

External **media** is re-hosted locally; third-party **scripts** (analytics, chat, payment SDKs) are deliberately left external so the audit can see and flag them. Files over 25 MB are recorded, not downloaded — route those to object storage (`references/hosting-verify-and-audit.md`).

Show the user the inventory before rebranding: file count, size, breakdown by type, anything excluded, anything oversized.

## STEP 3 — Rebrand (substitution, not rewriting)

Write a `facts.json` (schema in `references/facts-and-substitution.md`), then **dry run first**:

```bash
node scripts/substitute.mjs --dir <run-dir> --facts facts.json          # counts only
node scripts/substitute.mjs --dir <run-dir> --facts facts.json --apply  # after they approve
```

The script derives the tokens that leak in practice — the bare brand word, every phone format, the address as a unit, domain/www/email variants — protects asset filenames from being rewritten, and after `--apply` runs a zero-check that no replaced token survives.

Two warnings decide whether the rebrand is actually finished:

- **NEAR MISS** — a token that never matches exactly but appears loosely (different case, spacing or punctuation). The user's fact is slightly wrong. Fix it and re-run the dry run. This is how the old brand survives a "complete" list.
- **SPLIT ACROSS TAGS** — the old brand is visible on screen but broken up by markup in the file, almost always a two-tone brand mark like `GHL <span style="color:#D4AF37;">Command</span>`. **No string replacement can reach these.** You must edit them by hand, keeping the styling and changing only the words. The script prints the exact markup. The zero-check counts them, so `--apply` reports FAILED until they are fixed — a literal-only check would have said "clean" while the live page still showed the original brand in the header and the footer.

Never edit page copy to "improve" it during a clone. That is a different request; offer it afterwards.

## STEP 3b — Render and repair (never skip; a clone that does not render is not a clone)

Static scanning cannot see a URL that JavaScript builds at runtime. Webpack lazy chunks are the usual case: the filename exists in no attribute, stylesheet or literal string, so it is never mirrored.

Serve the clone locally and load it in a browser:

```bash
python3 -m http.server 8899 --bind 127.0.0.1   # from <run-dir>/site
```

In the page, list every same-origin resource that failed:

```js
const bad=[];
for (const e of performance.getEntriesByType('resource')) {
  if (!e.name.startsWith(location.origin)) continue;
  const r = await fetch(e.name, {method:'HEAD'});
  if (!r.ok) bad.push(e.name.replace(location.origin,''));
}
bad
```

Then fetch those from the original and re-render:

```bash
node scripts/repair.mjs --dir <run-dir> --origin "<original-site-url>" --urls failed.txt
```

**Repeat until nothing fails** — one repair pass often reveals the next, because code that was missing then runs and asks for more.

Why this is mandatory: on a real page six missing Elementor chunks meant `section-frontend-handlers` never loaded, so no section background was painted. The hero rendered blank, white headline text sat on white, and the clone looked nothing like the original — while every file-level check reported clean.

## STEP 4 — Pre-launch audit (mandatory, never skipped)

```bash
node scripts/audit.mjs --dir <run-dir>
```

Writes `reports/REVIEW-REQUIRED.md`: MUST REPOINT (leads/bookings/money), other people's content, inherited and regulated claims, plus everything excluded. **Show it to the user before deploying**, lead with the MUST REPOINT list and the regulated-claims flag, and never auto-delete a finding. `references/hosting-verify-and-audit.md` explains how to walk them through it.

## STEP 5 — Host

Deploy **`<run-dir>/site`** — never the run directory itself.

```bash
wrangler pages deploy <run-dir>/site --project-name=<slug>   # Cloudflare Pages (default)
```

Create/confirm the project first (Pages fails opaquely if it doesn't exist), and add a `_redirects` 404 rule so missing assets stop returning 200. Host table, per-file caps and the R2 commands for oversized media: `references/hosting-verify-and-audit.md`.

**GHL-native seam:** if the clone should live *inside* GHL, stop after step 3 and hand the content plus the asset manifest to Page Studio (`compose_page` / `compose_website`). Do not write a second GHL page writer here.

## STEP 6 — Verify the deployed URL

```bash
node scripts/verify.mjs --url "<deployed-url>" --dir <run-dir>
```

Checks every asset by content-type (not status code), retries once before calling anything broken (a single transient answer during a CDN rollout is not a broken asset), separates references the **source page was already serving badly** from ones the clone caused, and confirms old brand tokens are at zero — in the file *and* in the rendered text — with the new ones present.

Then open the URL in a browser and check three things, because no fetch-based check can see them:

```js
// Lazy images report naturalWidth 0 until they load — force them first, or you
// will report broken images that are perfectly fine (live-caught 2026-07-30).
document.querySelectorAll('img[loading="lazy"]').forEach(i => { i.loading = 'eager'; });
document.querySelectorAll('img[data-src]').forEach(i => { if (!i.src) i.src = i.dataset.src; });
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 4000));
window.scrollTo(0, 0);

[...document.querySelectorAll('img')].filter(i => i.complete && i.naturalWidth === 0)  // must be empty
[...document.querySelectorAll('video')].filter(v => v.error)                           // must be empty
document.body.innerText.match(/<old brand>/g)                                          // must be null
```

That last one is not redundant. Rendered text is where a tag-split brand mark shows up, and it is the only check that sees text injected at runtime by third-party widgets. Report the numbers, not an impression.

Finish with: the live URL, the verification table, and the REVIEW REQUIRED report.

---

# LANE B — Style clone (no rights, or "I just want the look")

The failure this lane exists to prevent: refusing the copy, then redrawing the page from memory and producing something drastically worse. Do not do that. Extract the actual design system:

```bash
node scripts/extract-design.mjs --url "<URL>" --out <dir>
```

Produces `DESIGN-SYSTEM.md` + `design-tokens.json`: palette ranked by frequency, the author's own CSS variables, font families and the real type scale, spacing rhythm, radii, shadows, breakpoints, container widths, and the section skeleton as **shape** (role, heading word counts, media and CTA counts) — deliberately never the words.

**Extraction is step 1 of 5. The deliverable is a working page, never a markdown file.** Do not stop here — a design-system document is not something the user can look at next to the original.

### B1 — Read the extracted artefacts, not your instincts

Open `design-tokens.json`. Use `declaredPalette` first when it exists (the site names its own colours — no guessing), then `brandAccents` with their source files. Take the type scale, spacing, radii, container width and grid **verbatim**.

### B2 — Lay the page out from `structure.sections`

The skeleton gives you section order, role, heading levels **with word counts**, and image/button/list counts. Build that. If it says the hero is `h1:3w · h5:17w · h2:5w · h6:25w` centred, build exactly that shape — matching the rhythm is what makes it read like the same designer's work. Do not substitute your own preferred layout; that is the from-memory redraw this lane exists to prevent.

If `spaShell` is true the structure is client-rendered and absent from the source; open the page in a browser and record the section order by eye first.

### B3 — Write original copy

Every headline, paragraph and CTA fresh. Match length and function, never wording. Filling a "17-word subhead" slot by pasting the original's 17-word subhead is copying — it will be caught in B5.

### B4 — Supply real media (this step is not optional)

A style-lane page with no imagery reads as a downgrade no matter how good the tokens are. Three legitimate sources: the client's own photographs, licensed stock, or **generated originals**. Never the source site's files. Label any slot you cannot fill yet, and say plainly in the handover that placeholders must be replaced before launch.

### B5 — Leak check, then audit

```bash
# compare against the ORIGINAL page, never against a rebranded clone
```

Fetch the original's rendered text and confirm **zero shared 5-, 6- and 8-word sequences** with your rebuild, and zero shared image filenames. Comparing against a rebranded clone is meaningless — the new brand name shows up as "shared" and hides real leaks.

Then run `audit.mjs` on the finished page. A fresh build can still inherit a regulated claim you typed from memory.

---

## Anti-patterns

- Rewriting page copy "to make it better" mid-clone.
- Trusting HTTP 200 as proof an asset exists.
- Scanning only the HTML for assets.
- Replacing a city without the street address.
- Auto-deleting testimonials — it destroys the layout and hides the liability instead of surfacing it.
- Deploying with the original owner's CDN still serving the media.
- Shipping `reports/` to the live host.
- Repeating the rights warning after step 0. Ask once, record, move on.
