# Deploy playbooks, exe.dev VM operations

Operational playbook for services on **exe.dev**, shared Linux VMs reached via
`ssh <host>.exe.xyz`. This reference captures the platform contract (which
isn't obvious from docs), the standard layout used across services, and the
two-three commands that solve 90% of day-to-day tasks.

## When to use

- Deploying a new static site or API to a fresh exe.dev VM.
- Updating an existing deployed site (push new `dist/`).
- Rotating API keys or other secrets on the VM.
- Diagnosing a down service, the **"Port 8000 unbound."** error page, 502s, or failed health checks.
- Adding a new systemd service / Caddy site to a VM already in service.

## When NOT to use

- DNS config, exe.dev owns `<host>.exe.xyz` subdomains; nothing to do locally.
- TLS, the exe.dev front-door terminates HTTPS. Never bind :443 on the VM.
- Multi-VM orchestration / auto-scaling, out of scope, single VM per service.
- Anything requiring root on the host kernel: these are shared VMs, not full boxes.

## Platform contract (the non-obvious bits)

These facts are easy to get wrong because the docs are thin, learned the hard way.

- **Your app must bind `:8000`** (plain HTTP). The exe.dev edge proxy terminates TLS on `<host>.exe.xyz` and forwards to VM `:8000`. If nothing is listening, exe.dev serves:

  > **Port 8000 unbound.** `ssh <host>.exe.xyz sudo systemctl enable --now nginx`
  >
  > (The nginx suggestion is just their canonical example, Caddy, Node, anything that binds :8000 works.)

- **Default VM user: `exedev`**, uid 1000, member of `sudo` and `docker`. Use this for service processes; don't create a dedicated service user unless you have a real isolation requirement.
- **What's preinstalled:** `git`, `rsync`, `docker`. **Not preinstalled:** `caddy`, `node`, `npm`, `nginx`. Use apt + NodeSource for node.
- **`127.0.0.1:9999` runs `shelley`**, exe.dev's internal agent. Localhost-only. Don't kill it, don't bind anything to 9999.
- **Disk:** 25 GB on `/`; no separate `/srv` mount. Fits a typical static bundle + node modules comfortably.
- **Host SSH key:** currently RSA-2048 only on new VMs (as of 2026-04). Verify the fingerprint in the exe.dev console on first connect; then `ssh-keyscan >> ~/.ssh/known_hosts`.

## Standard layout (convention we use)

| Path                          | Purpose                                         | Owner             |
| ----------------------------- | ----------------------------------------------- | ----------------- |
| `/srv/<app>/dist/`            | Static bundle served by Caddy                   | `exedev:exedev`   |
| `/srv/<app>/server.js`        | Optional runtime (e.g. API proxy)               | `exedev:exedev`   |
| `/etc/caddy/Caddyfile`        | Single-site; binds `:8000`                      | `root:root`       |
| `/etc/systemd/system/<app>.service` | systemd unit for server.js               | `root:root`       |
| `/etc/<app>.env`              | Secrets loaded via `EnvironmentFile=`           | `root:root` 0600  |

**Never** put secrets in `/srv/<app>/`, that's the webroot.

## Playbook: fresh provisioning

```sh
# 1. SSH in; verify host key first via exe.dev console.
ssh <host>.exe.xyz

# 2. Install runtime.
sudo apt-get update -qq
sudo apt-get install -y caddy
curl -fsSL https://deb.nodesource.com/setup_20.x | sudo -E bash -
sudo apt-get install -y nodejs

# 3. Filesystem.
sudo mkdir -p /srv/<app>
sudo chown exedev:exedev /srv/<app>
```

Then from your **local** machine (this repo's VM artifacts live in `deploy/`:
`deploy/Caddyfile`, `deploy/adia-ui.service`, `deploy/adia-ui.env.example`):

```sh
# 4. Push artifacts.
rsync -az --delete dist/ <host>.exe.xyz:/srv/<app>/dist/
rsync -a server.js <host>.exe.xyz:/srv/<app>/server.js   # if applicable

# 5. Push Caddyfile + systemd unit.
scp deploy/Caddyfile <host>.exe.xyz:/tmp/
scp deploy/<app>.service <host>.exe.xyz:/tmp/
```

Back on the VM:

```sh
# 6. Activate Caddy.
sudo mv /etc/caddy/Caddyfile /etc/caddy/Caddyfile.default.bak
sudo mv /tmp/Caddyfile /etc/caddy/Caddyfile
sudo systemctl reload caddy

# 7. Activate service.
sudo mv /tmp/<app>.service /etc/systemd/system/<app>.service
sudo systemctl daemon-reload
sudo systemctl enable --now <app>

# 8. Seed secrets, do this by hand; never via agent transcript.
sudo install -m 0600 -o root -g root /dev/null /etc/<app>.env
sudo vim /etc/<app>.env
sudo systemctl restart <app>
```

## One-time CI setup for `ui-kit.exe.xyz`

- **Repo secret `SITE_DEPLOY_SSH_KEY`**, done. An ed25519 keypair generated
  *by a human*, never by the agent (Hard gate 1). Public half goes in the
  VM's `~exedev/.ssh/authorized_keys`; private half goes in
  Settings → Secrets and variables → Actions, pasted directly, it should
  never appear in an agent's Bash context or a commit.
- **Environment `production-site` reviewer gate, CONFIGURED** (verified
  live 2026-08-11, `gh api repos/<org>/<repo>/environments`:
  `protection_rules` carries `required_reviewers`). The `deploy` job in
  `deploy-site.yml` therefore blocks on a human approval after its dry-run
  job, the delete-adjudication gate this skill's hardened-deploy design
  assumes. Changing the reviewer set is operator-only (repo Settings →
  Environments → `production-site`), no agent can configure it. Re-check
  the API output before trusting this line; it drifts with repo settings.

## Playbook: deploy an update

**Push a tag matching `site-v*`** (or trigger `.github/workflows/deploy-site.yml`
via `workflow_dispatch`). That workflow automates every step below: build,
dry-run with a delete summary posted to the job, snapshot, real rsync,
fixture-file + headless-render verify, auto-rollback on failure. This is
the PR-first, tag-triggered path, **do not run `npm run deploy:site` from
a local shell**; that command still exists only as a fallback for when CI
itself is unavailable.

`npm run deploy:site` under the hood: `npm run build:site && rsync -az --delete dist/ <host>.exe.xyz:/srv/<app>/dist/`. Caddy picks up static changes without a reload.

> ⚠️ **This is a `--delete` deploy, treat it as destructive whichever path runs it.**
> `build:site` was, until 2026-07-11, **entirely manual** and not wired into any
> pipeline, `dist/` silently drifted behind `main` until someone remembered to run
> it, and the rsync **deletes** everything on the server that isn't in the fresh
> `dist/`. A real ui-kit run (2026-06-08, run by hand) sent ~160 real content deltas
> and **deleted 3,572 files**. The CI workflow closes the drift problem (tag-triggered,
> so a deploy happens deliberately, not "whenever someone remembers") and keeps every
> safety step from the sequence below, it doesn't remove the risk, it enforces the
> discipline that used to depend on whoever ran the command remembering it.

### Hardened `--delete` deploy sequence (what `deploy-site.yml` automates)

1. **Build from a clean, fully-merged `main`**, never a feature branch. `dist/` ships
   verbatim; whatever's missing on the server gets deleted.
   - **In a FRESH worktree, build `@adia-ai/llm` FIRST:**
     ```sh
     npm run build -w @adia-ai/llm   # BEFORE build:site
     ```
     The `llm` package **compiles its JS at publish time** and its outputs are
     **gitignored**, so a fresh worktree (or any tree that hasn't published llm
     locally) has no `packages/llm/core/index.js`, `build:site` copies nothing, and
     **`/packages/llm/core/index.js` 404s on the deployed site → component registration
     breaks on every `/site/components/*` page** (the docs components reference it). Found
     **live 2026-06-09**; the **0-delete dry-run proved it had never been deployed** (the
     file was absent on the server, so there was nothing to delete, not a regression, a
     standing gap across every prior deploy). Build llm, then `build:site`, then the
     dry-run.
   - **`build:site` copies packages but does NOT rebuild their dist bundles**, after
     component `.css`/`.js` source changes, rebuild first (`npm run build -w
     @adia-ai/llm`, then `npm run build:bundles`) or the deployed bundles are stale.
   - **Package registration is one manifest, `scripts/build/site-package-registry.mjs`
     (ADR-0062)**, every `@adia-ai/*` package the site ships is one registry entry, from
     which the dist copy, BOTH importmaps (one shared `renderImportMapBlock()`: the
     `site/index.html` copy is a generated artifact, `npm run build:site-importmap`),
     the `dist/node_modules` symlinks, and the CI build order all derive. The old
     failure class (per-package `copyX()` hand-edits; local Vite works, prod 404s, v0.3.0 llm/a2ui-runtime, v0.8.27 persona+agent) is gated mechanically:
     `check:site-packages-registered` scans every shipped source root for an
     unregistered bare `@adia-ai/*` import and fails naming file:line, alongside
     `check:site-importmap-fresh` and `check:deploy-workflow-build-order`, all three
     in `npm run check` and early in `deploy-site.yml`. After any package add/rename,
     add the registry entry; the gates say the rest.
2. **Dry-run first, and adjudicate every delete, BEFORE the real rsync, never after:**
   ```sh
   rsync -azni --delete --exclude='packages/gen-ui/a2ui/corpus/feedback/' \
     dist/ <host>.exe.xyz:/srv/<app>/dist/   # -n simulates · -i itemizes
   ```
   - **Exclude server-side runtime-written paths.** Some files exist ONLY on prod, written by the running service at runtime, never present in a local build, so
     `--delete` wipes them on every deploy. Known class on ui-kit:
     `packages/gen-ui/a2ui/corpus/feedback/*.jsonl` (the gen-UI canvas training-feedback log).
     Found **live 2026-06-10** (the deploy deleted the day's feedback log; restored from
     the pre-deploy snapshot). Carry the same `--exclude` list on BOTH the dry-run and
     the real rsync; when a new runtime-written path appears, add it here.
   - `*deleting` lines = files removed from prod. Bucket **every one** into a known-safe
     class; **abort if any served-content delete is unexplained.** Safe classes seen on
     ui-kit: gallery review artifacts (`apps/genui/.../review/cycle-*/`), stale
     `packages/gen-ui/a2ui/retrieval/` + eval reports, **content-hash-rotated** CodeMirror
     chunks (`code/{chunk,dist}-<hash>.js`, old hash deleted, new hash sent = rotation,
     not loss), restructured `packages/llm/core/*.js` dist copies, `node_modules/`
     symlink-farm dirs, and refactor-orphaned app files. None are served HTML.
   - **Don't panic at the send count.** A fresh local build never mtime-aligns with the
     remote, so `rsync -a` flags ~every file as a send (`<f..t` = mtime-only touch). Only
     `<f+++` (new) and `<f.s.` / `<fcst` (content) are real deltas, the 2026-06-08 run
     itemized 12,540 "sends" of which only ~160 carried real content.
3. **Snapshot prod FIRST, the only safety net for `--delete`:**
   ```sh
   ssh <host>.exe.xyz 'cp -al /srv/<app>/dist /srv/<app>/dist.bak-<date>'   # hardlink: instant, reversible
   ssh <host>.exe.xyz 'find /srv/<app>/dist.bak-<date> -type f | wc -l'     # confirm it materialized
   ```
   `cp -al` is a hardlink farm, instant, ~0 extra disk, and a true point-in-time
   snapshot because rsync replaces inodes (writes a temp file + renames) rather than
   mutating in place. Restore with `rm -rf dist && mv dist.bak-<date> dist`.
4. **Deploy** (the dry-run, minus `-n`, SAME excludes):
   ```sh
   rsync -az --delete --exclude='packages/gen-ui/a2ui/corpus/feedback/' \
     dist/ <host>.exe.xyz:/srv/<app>/dist/
   ```
5. **Verify the FILE, not the route.** A SPA returns `200` + the app shell for *any*
   route even when stale, `curl https://<host>.exe.xyz/` proves nothing. Curl a
   **fixture file** that only exists in the new build, then **render-check the real
   screens**:
   ```sh
   curl -s https://<host>.exe.xyz/.../scenarios/manifest.json | grep -c <new-scenario>          # >0
   curl -s -o /dev/null -w '%{http_code}' https://<host>.exe.xyz/.../<scenario>/manifest.json   # 200
   ```
   File-presence is necessary but not sufficient, SPAs render client-side from the
   fixture, so only a headless-Chromium render proves the page composes (and didn't fall
   back to a default scenario). Run Playwright with `env -u NODE_OPTIONS` (a cmux
   `--require` in `NODE_OPTIONS` crashes node from the repo cwd).
6. **Rollback if verify fails**, restore the snapshot to the exact pre-deploy state:
   ```sh
   ssh <host>.exe.xyz 'rm -rf /srv/<app>/dist && mv /srv/<app>/dist.bak-<date> /srv/<app>/dist'
   ```
   Keep the snapshot until the deploy is confirmed good, then prune it.

> **`build:site` DOES include the app's components** (a pre-deploy worry, disproven
> 2026-06-08). AdiaUI apps are Light-DOM / no-bundle, components live in
> `app/.../src/components/` (the **source** tree), which `build:site` copies wholesale,
> so the SPA is fully renderable in `dist/`. Presence ≠ render, though; step 5's
> render-check is still the proof.
>
> **The one exception, `@adia-ai/llm` (2026-06-09):** llm is NOT a source-tree
> component; it **builds its JS at publish** (gitignored outputs), so `build:site` copies
> nothing for it in a fresh worktree and `/packages/llm/core/index.js` 404s → component
> registration breaks site-wide. Run `npm run build -w @adia-ai/llm` **before** `build:site`
> (step 1 above). This is exactly why step 5 render-checks a **`/site/components/*`** page,
> not just a fixture file: the 404 is invisible to file-presence and to the SPA shell.

If `server.js` changed:

```sh
rsync -a server.js <host>.exe.xyz:/srv/<app>/server.js
ssh <host>.exe.xyz 'sudo systemctl restart <app>'
```

## Playbook: rotate secrets

```sh
ssh <host>.exe.xyz 'sudo vim /etc/<app>.env && sudo systemctl restart <app>'
```

No rebuild, no redeploy. The static bundle never sees keys.

## Playbook: diagnose

When the site is down or returning 502, start here, everything is read-only:

```sh
ssh <host>.exe.xyz '
  echo "=== caddy ==="
  systemctl is-active caddy; systemctl status caddy --no-pager | head -10
  echo "=== app service ==="
  systemctl is-active <app>; journalctl -u <app> -n 30 --no-pager
  echo "=== listening ports ==="
  sudo ss -tlnp | grep -E ":(8000|3456|80|443)"
  echo "=== caddy syntax check ==="
  sudo caddy validate --config /etc/caddy/Caddyfile
  echo "=== disk ==="
  df -h /
'
```

Common failures:

| Symptom                         | Likely cause                                   | Fix                                               |
| ------------------------------- | ---------------------------------------------- | ------------------------------------------------- |
| "Port 8000 unbound" page        | Caddy not listening on `:8000`                 | Check Caddyfile header is `:8000` not `:80`       |
| 502 on `/api/*`                 | `<app>` service down or wrong port             | `systemctl status <app>`; tail journalctl         |
| 500 from `/api/llm/*`           | Missing/invalid API key                        | Check `/etc/<app>.env`, restart service           |
| Static assets 404               | `rsync --delete` ran with wrong source         | Rebuild local `dist/`, push again                 |
| Site serves a stale build (new fixtures 404, SPA shows a default scenario) | Deploy is tag-triggered CI now (`deploy-site.yml`, since the 2026-ci/tag-triggered-site-deploy cut), a merge to `main` alone deploys nothing. Either no `site-v*` tag was pushed for this merge, or its workflow run is sitting in the `production-site` environment's reviewer-approval queue, or the run failed a step | Check the Actions tab for a run against the expected tag first. If none exists, push a `site-v*` tag (or re-run via `workflow_dispatch`); if one exists and is pending, approve it; if it failed, fix the failing step. Only fall back to the manual `--delete` sequence if CI itself is unavailable, verify the fixture **file**, not the route |
| Caddy won't reload              | Syntax error in Caddyfile                      | `sudo caddy validate --config /etc/caddy/Caddyfile` |

## Invariants

- **Secrets never flow through agent context.** The agent should `sudo vim` *or* pause and let the human seed `/etc/<app>.env`. Never `echo "sk-..." > /etc/<app>.env` in a Bash call.
- **Caddy binds `:8000`**, not `:80` or `:443`.
- **`--delete` on rsync is scoped to the webroot only** (`/srv/<app>/dist/`). Never rsync-delete against `/srv/<app>/` or the VM's home.
- **Snapshot the webroot before any `--delete` deploy.** `cp -al /srv/<app>/dist /srv/<app>/dist.bak-<date>` is the only rollback for the 3,500+ files a fresh `--delete` prunes. Dry-run and adjudicate every delete before you send.
- **Verify a fixture FILE, never the route.** `curl /` returns `200` even when the site is stale (the SPA serves its shell for any path); proof of a live deploy is a fixture file `200` plus a browser render-check.
- **Back up before overwriting system files.** `mv Caddyfile Caddyfile.default.bak`, not `rm`.

## Current deployments

- `ui-kit.exe.xyz` → AdiaUI docs + demos (repo: `gen-ui-kit`, artifacts: `deploy/`, webroot: `/srv/adia-ui/dist/`, secrets: `/etc/adia-ui.env`, service: `adia-ui.service`). Also serves the embedded-app HCC demo at `/apps/embedded-app/app`; updated by pushing a `site-v*` tag (`.github/workflows/deploy-site.yml`, `--delete`, the hardened sequence above, automated). `npm run deploy:site` remains as a manual fallback only.

Add new hosts here as they come online.

## Hand-off from package-release

Release engineering (`package-release`) builds and publishes artifacts; this
playbook owns the deploy step that pushes them to the VM.

## Deploy Record, a filled example

The schema lives in SKILL.md's own "The Deploy Record" section; this is a
worked example (a real cut, `gh run view 29586391343`):

```text
Deploy Record
tag / run id:      site-v4 (workflow run 29586391343, 2026-07-17T14:03:15Z)
dry-run deletes:   see the run's dry-run job log for the class breakdown
fixture verified:  pass, CI's post-deploy verify step, run marked success
render verified:   pass, CI's post-deploy verify step, run marked success
snapshot:          CI pre-deploy hardlink step (deploy-site.yml)
rollback state:    not-needed
verdict:           shipped
```

This example cites the run URL rather than restating its log inline: the
record's job is to point at the evidence, not transcribe it; re-derive the
dry-run/fixture/render lines from `gh run view <id> --log` if the detail is
ever needed, don't assume this filled example's prose stays current with a
run that already happened.
