# Self-Hosted Runner Provisioning — <PROJECT_NAME>

> **Self-contained runbook** (not a thin stub). Unlike the other templates in
> this directory, there is no canonical `docs/runbooks/` counterpart to link —
> this file IS the process. It provisions a **persistent** (non-ephemeral)
> GitHub Actions runner on a macOS host using the mandrel-platform runner kit
> (`templates/runner/`), which ships the two job hooks — job-start hygiene
> (`job-cleanup.sh`) and job-end reaping (`job-completed.sh`) — and the
> per-runner `.env` (`.env.example`).
>
> Placeholder convention: `<UPPER_SNAKE>` between angle brackets — search for
> `<` after copying to find everything that still needs a value.

---

## ⚠️ Read first: the shared-`$HOME` concurrency hazard

Every runner on a host typically runs as the **same OS user**, so anything
resolved against `$HOME` is **shared across all co-resident runners** — it is
*not* scoped to "this runner", no matter what a comment claims. The two
hazards this kit exists to close:

1. **`~/setup-pnpm` is shared.** `pnpm/action-setup`'s `dest` input defaults
   to `~/setup-pnpm`. With N runners on one host, N concurrent jobs race on
   one pnpm shim install. Worse, a cleanup hook that `pkill`s processes
   matching `~/setup-pnpm` or `rm -rf`s it at job start will **destroy a pnpm
   install a concurrent runner is mid-flight on**. The fix is ordering:
   **first** make the pnpm shim location per-runner (workflow-side, see
   [pnpm scoping](#3-pnpm-scoping-workflow-side-prerequisite) below), and only
   **then** is a reap/delete hook safe — and even then it must target only the
   runner-scoped path. The kit's `job-cleanup.sh` never touches
   `~/setup-pnpm`.
2. **The default tool cache is shared.** Without a runner-scoped
   `RUNNER_TOOL_CACHE`, toolchain actions extract into a host-shared cache
   and co-resident runners race on it. The kit's `.env` scopes it to the
   runner's own `_work/_tool`.

Do not roll the hook out to a host until every workflow that job's runner
serves installs pnpm to a runner-scoped `dest`. Rolling out the hook without
the pnpm-scoping prerequisite reintroduces the exact corruption it guards
against.

## Host Values

| Value | Setting |
|-------|---------|
| Runner host | `<RUNNER_HOST>` |
| Runner OS user | `<RUNNER_USER>` |
| Runner root dir | `<RUNNER_DIR>` (e.g. `~/Development/github-runners/<REPO>`) |
| Repository | `<OWNER>/<REPO>` |
| Runner name | `<REPO>-runner` (or `<REPO>-runner-<N>` for a pool) |
| Labels | `self-hosted, macOS, ARM64, <REPO>-runner` |
| Runner version | `<RUNNER_VERSION>` (latest from [actions/runner releases](https://github.com/actions/runner/releases)) |

## 1. Download and unpack the runner

One directory per runner — never share a runner root between registrations.

```bash
mkdir -p <RUNNER_DIR> && cd <RUNNER_DIR>
curl -o actions-runner-osx-arm64-<RUNNER_VERSION>.tar.gz -L \
  https://github.com/actions/runner/releases/download/v<RUNNER_VERSION>/actions-runner-osx-arm64-<RUNNER_VERSION>.tar.gz
# Verify the SHA-256 against the checksum published on the release page
# before unpacking — same download-and-verify posture as the platform's
# pinned gitleaks/actionlint installs.
shasum -a 256 actions-runner-osx-arm64-<RUNNER_VERSION>.tar.gz
tar xzf actions-runner-osx-arm64-<RUNNER_VERSION>.tar.gz
```

## 2. Register with `config.sh` (repo-level)

Registration is **repo-level** (the fleet's standing model), not org-level.
Mint a short-lived registration token via the repo UI
(*Settings → Actions → Runners → New self-hosted runner*) or:

```bash
gh api -X POST repos/<OWNER>/<REPO>/actions/runners/registration-token --jq .token
```

Then configure:

```bash
cd <RUNNER_DIR>
./config.sh \
  --url https://github.com/<OWNER>/<REPO> \
  --token <REGISTRATION_TOKEN> \
  --name <REPO>-runner \
  --labels self-hosted,macOS,ARM64,<REPO>-runner \
  --work _work \
  --unattended
```

- `--labels` — the four-label contract the fleet's workflows target
  (`self-hosted, macOS, ARM64, <REPO>-runner`). The `<REPO>-runner` label is
  the routing key; keep it unique per repo.
- `--work _work` — keeps the work tree inside `<RUNNER_DIR>`, which is what
  makes every path in the hygiene kit runner-scoped.
- Registration tokens expire after ~1 hour; mint a fresh one per runner.

## 3. pnpm scoping (workflow-side prerequisite)

Before installing the hook, confirm every workflow this runner serves
installs the pnpm shim to a **runner-scoped** destination:

- Workflows using the platform's `setup-toolchain` composite action (all
  `pr-quality.yml` tiers) are already safe: it defaults `pnpm/action-setup`'s
  `dest` to `${{ runner.temp }}/pnpm`, i.e. `<RUNNER_DIR>/_work/_temp/pnpm`,
  unique per runner. `pr-quality.yml` also exposes a `pnpm-dest` input for
  explicit overrides.
- Workflows calling `pnpm/action-setup` directly MUST pass
  `dest: ${{ runner.temp }}/pnpm`. The action's default (`~/setup-pnpm`) is
  host-shared and unsafe under runner concurrency (see the hazard header).

There is **no runner-side override** for the pnpm `dest` — it is a workflow
input — which is why this step is a rollout gate, not an `.env` line.

## 4. Install the hygiene kit (hooks + `.env`)

Copy the kit from the platform payload into the runner root:

```bash
cp node_modules/mandrel-platform/templates/runner/job-cleanup.sh <RUNNER_DIR>/job-cleanup.sh
chmod +x <RUNNER_DIR>/job-cleanup.sh
cp node_modules/mandrel-platform/templates/runner/job-completed.sh <RUNNER_DIR>/job-completed.sh
chmod +x <RUNNER_DIR>/job-completed.sh
cp node_modules/mandrel-platform/templates/runner/check-runner-env-drift.sh <RUNNER_DIR>/check-runner-env-drift.sh
chmod +x <RUNNER_DIR>/check-runner-env-drift.sh
cp node_modules/mandrel-platform/templates/runner/.env.example <RUNNER_DIR>/.env
```

Then edit `<RUNNER_DIR>/.env` and replace every `<RUNNER_DIR>` placeholder
with the runner root's absolute path. The resulting file wires:

- `ACTIONS_RUNNER_HOOK_JOB_STARTED=<RUNNER_DIR>/job-cleanup.sh` — the
  job-start hook. It reaps orphaned pnpm/node processes parented to **this**
  runner's work tree, clears stale runner-scoped pnpm installs, and removes
  leftover tool-download dirs from `<RUNNER_DIR>/_work/_temp`. It never fails
  a job (always exits 0) and never touches another runner's state.

  **Every path it reads is runner-scoped, and that is load-bearing** (issue
  #343). The hook runs inside the *job's* clock, so its cost is charged to
  `Set up runner` and counts against the job's own `timeout-minutes`. An
  earlier version swept the host-shared OS temp root; on a host where that
  directory had grown to ~840k entries, `Set up runner` reached 5m29s and
  jobs were killed before their first real step — surfacing as `cancelled`
  on unrelated diffs. If you add a sweep to this hook, root it at
  `_work/_temp`, never at `$TMPDIR`.
- `ACTIONS_RUNNER_HOOK_JOB_COMPLETED=<RUNNER_DIR>/job-completed.sh` — the
  job-end hook. After the last step of every job it terminates whatever of
  **this** job's process tree is still alive — SIGTERM, a bounded grace, then
  SIGKILL — matching processes whose command line resolves inside
  `<RUNNER_DIR>/_work/` plus their descendants. It never signals itself, its
  own ancestors, or the runner's `Runner.Worker`/`Runner.Listener`, never
  fails a job (always exits 0), and does nothing at all when the job left
  nothing behind.
- `RUNNER_TOOL_CACHE=<RUNNER_DIR>/_work/_tool` and
  `AGENT_TOOLSDIRECTORY=<RUNNER_DIR>/_work/_tool` — runner-scoped tool cache
  (two env names, one dir; some actions read the legacy name).
- `LANG=en_US.UTF-8`.

None of the shipped scripts needs per-runner editing: each derives its paths
from its own location, so the same files work verbatim on every runner.

### Why both hooks

They cover opposite ends of the same job and neither substitutes for the
other:

| Hook | Runs | Defends against |
|------|------|-----------------|
| `job-cleanup.sh` (`..._JOB_STARTED`) | before the first step | the **previous** job's leftovers — orphaned pnpm/node processes and stale `_work/_temp` artifacts already on the runner |
| `job-completed.sh` (`..._JOB_COMPLETED`) | after the last step | **this** job's own survivors reaching the **next** job |

The started hook cannot close the second case: it runs before the new job's
processes exist, so once a job is minutes in, nothing it did can help. A
**cancelled** job is where survivors are most likely — the runner terminates
the step it is executing, not everything that step forked — so a coalesced
push (`concurrency: cancel-in-progress`) is the routine way a runner ends up
hosting a previous job's vitest forks or dev server. The observed symptom is
the next job on that runner exiting 143 (SIGTERM) mid-run, with no
cancellation request in the runner's `Worker_*.log` and every concurrent job
on the pool's other runners passing.

The completed hook runs on the **job's** clock, like the started one, so it is
held to the same cost rule: its whole input is one `ps` snapshot, and it
sleeps only while waiting out the grace period of a tree it actually
signalled. A job with nothing to reap pays a few milliseconds.

### Confirm the pool is uniform (`check-runner-env-drift.sh`)

Run this **after provisioning each runner**, and again whenever two runners
behave differently on the same job. It walks the pool and names the runners
missing any of the five mandated keys:

```bash
cd <RUNNER_DIR>
./check-runner-env-drift.sh                       # pool root = this dir's parent
./check-runner-env-drift.sh --pool-root <POOL_ROOT>
```

The pool root is the directory holding one subdirectory per runner (§1); a
child directory counts as a runner iff it contains `config.sh`. The checker is
read-only — it never writes into a runner root and never touches a service.

| Exit | Meaning | Operator response |
|------|---------|-------------------|
| `0` | No drift. Every mandated key is set on every runner, or unset on every runner. | None. |
| `1` | **Drift** — a key is set on some runners but not all. | Copy `.env.example` onto each runner the report names, substitute its `<RUNNER_DIR>`, then `./svc.sh stop && ./svc.sh start` on those runners so they reload `.env`. |
| `2` | Usage error — bad flag, or a pool root holding no runner directories. | Re-check `--pool-root`. |

A key absent from **every** runner is reported as a uniform gap and does *not*
exit non-zero, so a fleet that has deliberately not adopted a key is not a
standing alarm.

**Why this check exists at all:** the GitHub runners API reports a runner's
name, labels and online status — it cannot see `<RUNNER_DIR>/.env`. So partial
provisioning never surfaces as a configuration fault; it surfaces as an
unattributable behavioural difference between two runs of the same job. In
issue #343, 16 of 19 runners on one host carried the hook, and the resulting
`Set up runner` spread (5m29s against 54s) took far longer to attribute than
reading nineteen `.env` files would have.

Do **not** wire this into `ACTIONS_RUNNER_HOOK_JOB_STARTED`. That hook runs
inside the job's clock, where every read is billed to `Set up runner` and
counts against the job's `timeout-minutes` — the exact cost model that
made #343 a job-killer. This is an operator-run tool.

## 5. Install as a launchd service (`svc.sh`)

```bash
cd <RUNNER_DIR>
./svc.sh install   # generates the launchd plist for the current user
./svc.sh start
./svc.sh status    # expect: Started · running
```

Verify end-to-end: push a trivial workflow run targeting
`runs-on: [self-hosted, macOS, ARM64, <REPO>-runner]` and confirm (a) the job
is picked up, (b) the job log shows the `Set up runner` hook phase running
`job-cleanup.sh` before the first step, and (c) the job log shows
`job-completed.sh` running after the last step (it prints either what it
reaped or `nothing to reap`).

The runner loads `.env` at service start — after any `.env` change, restart:

```bash
./svc.sh stop && ./svc.sh start
```

## 6. Update / rotation guidance

- **Runner version updates.** Persistent runners self-update by default when
  GitHub releases a new runner version; no action needed. If a runner is
  pinned or the self-update wedges, stop the service, download/unpack the new
  tarball over `<RUNNER_DIR>` (config and `.env` survive), and restart via
  `svc.sh`.
- **Kit updates.** The hooks and `.env.example` are versioned in
  mandrel-platform. On a platform release that touches `templates/runner/`,
  re-copy `job-cleanup.sh`, `job-completed.sh` and
  `check-runner-env-drift.sh` (all verbatim — they are parameterized) and
  diff `.env.example` against the live `.env`,
  then `./svc.sh stop && ./svc.sh start`. There is no `mandrel sync`
  equivalent for a runner host's filesystem — this is an operator-applied
  step. Re-run `./check-runner-env-drift.sh` afterwards: a kit update applied
  to some runners and not others is exactly the drift it reports.
- **Token/registration rotation.** Registration tokens are one-shot at
  config time; nothing persists to rotate. To move a runner between repos or
  rename it: `./svc.sh stop && ./svc.sh uninstall && ./config.sh remove
  --token <REMOVAL_TOKEN>`, then re-register (§2) and reinstall the service
  (§5). Mint the removal token via
  `gh api -X POST repos/<OWNER>/<REPO>/actions/runners/remove-token --jq .token`.
- **Decommission.** Same removal sequence, then delete `<RUNNER_DIR>`.
  Confirm the runner disappeared from *Settings → Actions → Runners*.

## Project-Specific Notes

<!-- Record host quirks: co-resident runner inventory for this host, Xcode /
     toolchain versions the workloads assume, monitoring hooks, etc. -->
