# Rollbridge

Rollbridge is a Node.js process supervisor and local traffic switcher for zero-downtime deploys.

Nginx points at one stable Rollbridge proxy port. Deploy tooling asks Rollbridge to start a new release, health-check it, and switch new traffic to it. Retirement then continues asynchronously: HTTP/WebSocket connections and retained process generations drain independently after the deploy command returns.

> **Required jobs-generation contract:** a release-scoped background-jobs
> runtime is its own `background-jobs-main` plus worker pool. Rollbridge starts a
> complete candidate generation before activation. After activation the retired
> main stops schedules, new dispatch, and new worker handoffs, but remains with
> its old workers to supervise every handoff it already made until all jobs and
> workers settle. Old workers never move to the new main. Generations may overlap
> for hours on separate ports, and neither deploy completion nor HTTP drain
> completion waits for or kills them. See [Background-job worker
> deployment](docs/workers.md) and the [Velocious deployment
> guide](docs/velocious.md).

This is the required architecture, not evidence that every released runtime or
consumer config already implements durable recovery and release-reference
reporting; verify source and config before claiming compliance.

## Install

```bash
npm install rollbridge
```

For local development in this repository:

```bash
npm install
npm run all-checks
```

## Config

A Rollbridge config is a JavaScript module that `export default`s a config
object. It can also export a function (sync or async) that returns the object,
which is handy for computing values from the environment. Write it in your
project's module system — `export default` for ESM (`"type": "module"`) or
`module.exports` for CommonJS.

```js
// rollbridge.js
export default {
  application: "ticket-server",

  control: {
    path: "/tmp/rollbridge-ticket-server.sock"
  },

  statePath: "/var/lib/rollbridge/ticket-server.state.json",
  ownerRecovery: {reconnectGraceMs: 30000},

  proxy: {
    host: "127.0.0.1",
    port: 8182,
    upstreamHost: "127.0.0.1",
    healthPath: "/ping",
    healthTimeoutMs: 30000,
    drainTimeoutMs: 60000,
    forceStopTimeoutMs: 10000
  },

  processes: [
    {
      id: "beacon",
      policy: "service",
      cwd: "{{releasePath}}",
      command: "env VELOCIOUS_BEACON_PORT={{port}} npx velocious beacon",
      port: 7330
    },
    {
      id: "background-jobs-worker",
      policy: "companion",
      cwd: "{{releasePath}}",
      env: {VELOCIOUS_BACKGROUND_JOBS_PORT: "{{ports.background-jobs-main}}"},
      command: "npx velocious background-jobs-worker",
      nonBlockingDrain: true,
      gracefulStopMs: "indefinite",
      outputLines: 200
    },
    {
      id: "background-jobs-main",
      policy: "service",
      deployStrategy: "handoff",
      cwd: "{{releasePath}}",
      env: {VELOCIOUS_BACKGROUND_JOBS_LIFECYCLE_SOCKET: "{{releasePath}}/tmp/background-jobs-main.sock"},
      command: "env VELOCIOUS_BACKGROUND_JOBS_PORT={{port}} npx velocious background-jobs-main",
      lifecycle: {
        activateCommand: 'npx velocious background-jobs:activate --generation "$ROLLBRIDGE_RELEASE_ID" --socket "$VELOCIOUS_BACKGROUND_JOBS_LIFECYCLE_SOCKET"',
        activateTimeoutMs: 60000,
        quietCommand: 'npx velocious background-jobs:retire --generation "$ROLLBRIDGE_RELEASE_ID" --socket "$VELOCIOUS_BACKGROUND_JOBS_LIFECYCLE_SOCKET"'
      },
      port: {from: 7331, to: 7399}
    },
    {
      id: "web",
      policy: "proxied",
      cwd: "{{releasePath}}",
      command: "npx velocious server --host 127.0.0.1 --port {{port}}",
      port: {from: 18182, to: 18299},
      health: {path: "/ping", timeoutMs: 30000}
    }
  ]
}
```

The lifecycle socket path is illustrative; set it to the reviewed release-local
Velocious socket used by jobs-main.

Each process retains its most recent stdout/stderr lines and reports them in
`status`. Set `outputLines` (a positive integer, default 50) per process to keep
more or fewer lines for chatty or quiet processes.

Set `control.mode` to an octal permission string (for example `"660"`) to
chmod the control socket after it binds. This restricts which users can send
control commands — useful when several deploy users share a group. When unset,
the socket keeps the default permissions from the daemon's umask. Pair it with
`control.owner` and `control.group` (a numeric id or a user/group name) to
`chown` the socket to a shared deploy group; names resolve via
`/etc/passwd`/`/etc/group`, and the daemon must run as a user allowed to chown it.

Set the proxied process's `health.startDelayMs` (default `0`) to wait that long
after the process starts before the first health probe — like a readiness
probe's initial delay, useful for apps with a known boot time. The delay runs
before the `health.timeoutMs` window begins.

Set a process's `restart` policy to control automatic restarts after a crash.
`restart.maxRestarts` caps how many restarts are allowed within `restart.windowMs`
before Rollbridge gives up and leaves the process `failed` (`maxRestarts: 0`
disables restarts entirely), while `restart.backoffFactor` — with an optional
`restart.maxDelayMs` cap — backs off the `restartDelayMs` delay on each successive
restart. With no `restart` block, a crashed process keeps restarting after
`restartDelayMs`, as before. See [`docs/config.md`](docs/config.md#processesrestart).

```js
restart: {maxRestarts: 5, windowMs: 60000, backoffFactor: 2, maxDelayMs: 30000}
```

Set a process's `memory` policy to supervise its resident memory (RSS) and
gracefully restart it when it grows too large. `memory.limitBytes` is the RSS
limit (measured across the whole process group, not just the wrapper);
`memory.warnBytes` logs a warning before the limit; `memory.checkIntervalMs`
(default `5000`) sets how often RSS is sampled. A memory restart is reported in
`status` and recorded in `events` (a `process started` with `reason: "memory"`).
See [`docs/config.md`](docs/config.md#processesmemory).

```js
memory: {limitBytes: 536870912, warnBytes: 402653184, checkIntervalMs: 5000}
```

Set a process's `stopSignal` (default `"SIGTERM"`) to the signal it quiets on.
This is a process-level stop mechanism, not the primary deploy-completion
mechanism for a release-scoped jobs generation. A retired generation may remain
for hours, and its old jobs-main must stay available until its workers finish.
For example, a generic worker that drains on `SIGINT`:

```js
{id: "worker", policy: "companion", command: "…", stopSignal: "SIGINT", gracefulStopMs: 60000}
```

Set `replicas` on a port-less `companion` to run a pool of identical workers.
Each instance runs as `<id>#<index>` (`worker#0`, `worker#1`, …) — visible in
`status` and targetable by `rollbridge restart` (base id for all, `worker#0` for
one) — and gets `{{replicaIndex}}`/`{{replicaCount}}` and
`ROLLBRIDGE_REPLICA_INDEX`/`_COUNT` so each instance can pick a distinct shard or
queue. See [`docs/config.md`](docs/config.md#processesreplicas).

```js
{id: "worker", policy: "companion", command: "npx velocious background-jobs-worker", replicas: 4}
```

For generic workers that quiesce or drain via a command, set a `lifecycle` block —
Rollbridge runs `quietCommand`, then drains (`drainCommand`/`drainTimeoutMs`),
then `stopCommand`/`stopSignal`, then `SIGKILL` after `gracefulStopMs` when
gracefully stopping the process. Each hook is bounded so it can't wedge a stop.

Set `nonBlockingDrain: true` on a jobs worker companion so it stops accepting new
handoffs as soon as its release retires, independently of the proxied connection
drain. Its release-scoped handoff jobs-main remains running to supervise existing
handoffs and exits only after the worker pool has drained.

For a coordinator that starts quiescent, give the handoff jobs-main paired
`lifecycle.activateCommand` and `quietCommand` hooks. After candidate health,
Rollbridge waits for old retirement, waits for candidate activation, and commits
the active proxy target without an awaited boundary. The durable transition is
generation-scoped and resumable; failures remain visible and block unrelated
deploys. Post-commit singleton replacement is also journaled and must complete
before an exact retry reports success. If the active coordinator restarts,
Rollbridge restores its active role with the same bounded, generation-scoped
activation command before reporting it running. Set `activateTimeoutMs` when the
activation acknowledgement can exceed its 30-second default. Omit `activateCommand` to
preserve the existing hook-free ordering.

See [`docs/workers.md`](docs/workers.md) for the full release-generation
deployment pattern: a handoff `background-jobs-main`, its companion worker pool,
independent quiescence, and durable supervision while retained generations drain.

Set `releaseRetention` to bound how many stopped (drained) releases the daemon
keeps in memory and reports in `status`. `keep` (default `10`) retains the most
recent stopped releases; `maxAgeMs` (default `0`, disabled) also prunes stopped
releases older than that many milliseconds. The active and draining releases are
never pruned. This is Rollbridge's own release records — your deploy tool still
owns cleaning up on-disk release directories.

```js
releaseRetention: {keep: 5, maxAgeMs: 86400000}
```

Set `statePath` to have the daemon persist sanitized recovery state to a file
(active/draining releases, process pids, counters, sanitized recent events).
Commands, environment mappings, child command lines, and captured process output
are deliberately excluded; those diagnostics remain available through the live
status/log/event APIs. Without `ownerRecovery`, the next startup reads any
leftover file and reports managed processes still alive from a daemon that
didn't shut down cleanly — advisory orphan detection. In that mode, after a crash run
`rollbridge recover` to list those leftovers and `rollbridge recover --force` to
stop them before restarting the daemon. A clean `shutdown` removes the file. See
[`docs/config.md`](docs/config.md#statepath).

```js
statePath: "/var/lib/rollbridge/ticket-server.state.json"
```

For durable daemon process recovery and atomic owner replacement, opt into
`ownerRecovery`. Rollbridge
then runs a private local process guardian which remains the OS supervisor for
managed processes if the control daemon exits unexpectedly. A replacement using
the exact same normalized config/runtime reconnects within `reconnectGraceMs`,
reconstructs active and draining generations and their ports, and fences
concurrent replacements. If no replacement reconnects during that grace, the
guardian restarts the exact accepted daemon command and environment itself; the
recovery definition is kept only in the guardian's private authenticated state
and is refreshed atomically during a package/runtime replacement. A restart
attempt which cannot claim ownership and publish ready listeners within the
accepted startup timeout is terminated with its process group and retried with a
nonzero backoff. `ensure-daemon` can also prepare a requested
config/control-socket/package/runtime owner, prove it healthy, and atomically
transfer guardian authority while every retained generation keeps its exact
release reference and drains asynchronously. The old `statePath` is the durable
transaction anchor and cannot change during this handoff. The `0600` state file
contains the guardian capability; protect its directory accordingly.
Prepared transactions fence owner mutations and compare a monotonic guardian
state revision at staging. Existing HTTP/WebSocket connections remain owned by
the retired listener process, while their counts transfer to the new daemon so
later deploys continue to honor the original drain boundary.
If that new daemon itself exits while a retired listener still owns connections,
recovery conservatively retains the last authenticated transferred count until
the configured drain timeout; it never guesses that the older sockets closed.
An intermediate guardian which supports atomic owner replacement but predates
daemon recovery cannot be hot-upgraded because it is the existing processes' OS
supervisor. The upgrade fails before handoff and requires one explicit clean
`shutdown` followed by `ensure-daemon`; subsequent package/runtime replacements
remain atomic. A genuinely pre-replacement guardian still uses the separately
documented one-time disruptive compatibility bridge below.

There is one explicit compatibility boundary: the first upgrade from a genuine
pre-owner-replacement Rollbridge guardian and daemon cannot share its listeners
on supported Node 20. The same boundary applies to a retained transition guardian
that supports prepare/stage but predates the retired-owner commit command. After
an explicit capability probe and authentication of the guardian, exact daemon
PID/socket, runtime authority, and durable owned-process state, `ensure-daemon`
performs a one-time **disruptive** bridge. Existing proxy/control connections may close,
the retained processes keep their exact PIDs under guardian supervision, and
`status.ownerTransition` reports `mode: "legacy-first-upgrade"` with
`disruptive: true`. The bridge requires the existing config identity unchanged;
apply config/socket changes in a subsequent invocation, which uses the atomic
protocol. Unknown commands, auth/transport failures, malformed responses, and
identity mismatches fail closed without entering this bridge. Every replacement
after this protocol upgrade remains candidate-first and atomic.

```js
statePath: "/var/lib/rollbridge/ticket-server.state.json",
ownerRecovery: {reconnectGraceMs: 30000}
```

During the first migration from an old supervisor, set `legacyTakeover` and run
`rollbridge predeploy-cleanup --release-path <path>` before `rollbridge deploy`.
Rollbridge will only stop configured legacy processes when no reusable active
Rollbridge release is running.

```js
legacyTakeover: {
  screens: ["ticket-server"],
  processes: [
    {name: "legacy web", includes: ["/home/dev/ticket-server/", "velocious server", "--port 8082"]}
  ]
}
```

A function export receives no arguments and lets you build the config at load
time:

```js
// rollbridge.js
export default () => ({
  application: process.env.APP_NAME || "ticket-server",
  control: {path: `/tmp/rollbridge-${process.env.APP_NAME || "ticket-server"}.sock`},
  proxy: {host: "127.0.0.1", port: 8182},
  processes: [
    {id: "web", policy: "proxied", cwd: "{{releasePath}}", command: "npx velocious server --port {{port}}", port: {from: 18182, to: 18299}}
  ]
})
```

### Template variables

A process `command`, `cwd`, and `env` values support `{{...}}` placeholders
rendered when the process starts:

- `{{releasePath}}`, `{{releaseId}}`, `{{revision}}`, `{{application}}`, `{{processId}}`
- `{{port}}` — the port allocated to this process; `{{ports.<id>}}` — another process's allocated port
- `{{proxy.host}}`, `{{proxy.port}}`, `{{proxy.upstreamHost}}`
- `{{env.<NAME>}}` — a variable from the daemon's own environment, e.g. `{{env.HOME}}`

Referencing a placeholder with no value (including an unset `{{env.<NAME>}}`)
fails the process start with a clear error, so typos surface immediately.

Configuration examples live in `examples/`, including
`examples/tensorbuzz.com.js` for the current TensorBuzz backend deployment; see
[`docs/tensorbuzz-runbook.md`](docs/tensorbuzz-runbook.md) for the matching
production runbook (ports, deploy ordering, rollback constraints, and day-to-day
operations).

See [`docs/velocious.md`](docs/velocious.md) for a Velocious deployment guide —
how Beacon, background-jobs-main, background-jobs-worker, and the web process map
to Rollbridge policies, with startup ordering and deploy behavior.

See [`docs/config.md`](docs/config.md) for the full config reference — every
field, its default, validation rules, template variables, and the environment
variables Rollbridge injects.

## Process Policies

Every process declares a `policy` that controls its lifecycle. Pick one per
process:

| You need… | Use |
| --- | --- |
| The process that receives external HTTP/WebSocket traffic | `proxied` |
| A per-release helper tied to the release lifecycle | `companion` |
| Exactly one instance, never overlapping across deploys | `singleton` |
| A long-lived shared broker that survives deploys | `service` |

### `proxied`

The web/API process — exactly one per config. Rollbridge forwards HTTP and
WebSocket traffic to the active release's proxied process and tracks open
connections so they can be drained on the next deploy. It must define a `port`
range, is health-checked before traffic switches to a new release, and is
auto-restarted while its release is active.

```js
{
  id: "web",
  policy: "proxied",
  cwd: "{{releasePath}}",
  command: "npx velocious server --host 127.0.0.1 --port {{port}}",
  port: {from: 18182, to: 18299},
  health: {path: "/ping", timeoutMs: 30000}
}
```

### `companion`

A release-scoped helper (for example a background worker bound to one release).
It starts **before** the proxied process in the same release, so release-local
dependencies are ready before the health check, and it is auto-restarted while
its release is active. Each release gets its own companions; a release's
companions normally stop when that release is drained and retired after a newer
release takes over. A jobs worker configured with `nonBlockingDrain` instead
quiesces at retirement and drains independently of HTTP/WebSocket connections.

```js
{
  id: "background-jobs-worker",
  policy: "companion",
  cwd: "{{releasePath}}",
  command: "npx velocious background-jobs-worker",
  nonBlockingDrain: true,
  gracefulStopMs: "indefinite"
}
```

### `singleton`

A one-at-a-time helper for duplicate-unsafe schedulers or job dispatchers. After
a new release becomes active, Rollbridge stops the old singleton and then starts
the new one, so two copies never run at once. Use it when running the old and
new copies simultaneously during a deploy would be unsafe.

```js
{
  id: "scheduler",
  policy: "singleton",
  cwd: "{{releasePath}}",
  command: "npx velocious scheduler"
}
```

### `service`

A service can be daemon-wide and persistent, or release-scoped with
`deployStrategy: "handoff"`. Velocious Beacon is normally persistent on a stable
port. `background-jobs-main` is not: it must be a handoff service with one port
per release so old workers keep their old coordinator while new workers use the
candidate coordinator.

```js
{
  id: "background-jobs-main",
  policy: "service",
  deployStrategy: "handoff",
  cwd: "{{releasePath}}",
  command: "npx velocious background-jobs-main",
  lifecycle: {
    activateCommand: 'npx velocious background-jobs:activate --generation "$ROLLBRIDGE_RELEASE_ID" --socket "$VELOCIOUS_BACKGROUND_JOBS_LIFECYCLE_SOCKET"',
    quietCommand: 'npx velocious background-jobs:retire --generation "$ROLLBRIDGE_RELEASE_ID" --socket "$VELOCIOUS_BACKGROUND_JOBS_LIFECYCLE_SOCKET"'
  },
  port: {from: 7331, to: 7399}
}
```

### Deploy ordering

On `rollbridge deploy`, the required ordering is:

1. starts any missing persistent service and the candidate's handoff services;
2. starts the new release's `companion`s, then its `proxied` process, and
   health-checks the proxied process;
3. when `activateCommand` is configured, retires and acknowledges the previous
   jobs generation, then activates and acknowledges the candidate;
4. synchronously switches new traffic to the new release (hook-free configs keep
   the existing switch-then-retire behavior);
5. replaces `singleton`s (stops the old one, then starts the new one);
6. returns success without waiting for the previous generation or its independent
   HTTP/WebSocket drain; Rollbridge supervises all retained drains in the
   background and reaps each generation only after its handoffs and workers end.

If the new release fails to start or health-check, the previous release stays
active and any service started during this deploy is rolled back.

`status.releaseReferences` lists the id and path of every active or draining
release, plus a stopped release that still owns a persistent service definition
or singleton, or remains part of an unresolved generation transition. A
reference disappears only when that release has stopped and no longer owns
runtime state. With
`ownerRecovery`, those references and generations survive both same-authority
daemon recovery and guardian-fenced incompatible
config/control-socket/package/runtime replacement through `ensure-daemon`.

## Commands

`--config` is optional for every command. When omitted, Rollbridge looks for
`rollbridge.js` in the current directory. The examples below pass `--config`
explicitly, but `rollbridge validate` (or any command) works with no flag when a
`rollbridge.js` is present.

For machine-readable output, `deploy`, `status`, `stop`, `shutdown`, and
`ensure-daemon` already print JSON, and `validate`, `doctor`, and `logs` accept
a `--json` flag that switches their output to JSON (with the same exit codes),
so deploy tooling can parse results.

See [`docs/cli.md`](docs/cli.md) for the full per-command reference (every
option, default, output shape, and exit code).

Validate a config without starting the daemon:

```bash
rollbridge validate --config rollbridge.js
```

`validate` reports every config error at once with an example fix and exits
non-zero when issues are found, so deploy tooling can gate on it. It checks
required fields and types, duplicate process IDs, port ranges, that exactly one
process is `proxied`, and that the proxied process defines a port range. Example
output for a misconfigured file:

```text
Found 2 configuration issues in rollbridge.js:

1. Config must define exactly one proxied process; found 0
   Fix: Mark exactly one process with policy: proxied so Rollbridge knows where to forward traffic.

2. Duplicate process id: web
   Fix: Give each process a unique id; "web" is used more than once.
```

Check the environment before starting the daemon:

```bash
rollbridge doctor --config rollbridge.js
```

`doctor` validates the config and then probes the runtime environment, exiting
non-zero if any check fails (so deploy tooling can gate on it):

```text
✓ config: valid: 4 processes, proxy on 127.0.0.1:8182
✓ control socket: no daemon running; /tmp/rollbridge-ticket-server.sock is free to bind
✓ control socket directory: /tmp is writable
✓ proxy port: 127.0.0.1:8182 is available

All checks passed.
```

A free control socket, a writable socket directory, and a bindable proxy port
pass. Because `rollbridge daemon` cannot bind a socket or port that is already
taken, doctor fails the relevant check when a Rollbridge daemon (or any other
process) is already listening on the control socket or holding the proxy port —
so a green `doctor` means a fresh daemon can actually start.

Start the daemon:

```bash
rollbridge daemon --config rollbridge.js
```

Start the daemon and bootstrap an exact prepared release before leaving it in
the foreground (for example from a boot-time service manager):

```bash
rollbridge daemon --config /srv/ticket-server/rollbridge.js \
  --release-path /srv/ticket-server/releases/20260813090000/ticket-server \
  --release-id 20260813090000 --revision abc123 \
  --boot-attestation sha256:0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef
```

External supervisors that need to replace a foreground owner without waiting
for retained generations to drain add `--takeover-owner`. The candidate starts
and health-checks the exact release first. Only then does it retire the accepted
owner's proxy/control listeners. Candidate bootstrap failure leaves the accepted
owner untouched. Current takeover does not preserve retained jobs generations
or transfer listeners atomically, so it is not a zero-downtime package, config,
or socket upgrade mechanism. Use `ownerRecovery` plus `ensure-daemon` for the
guardian-fenced atomic replacement contract. This is opt-in; ordinary daemon
bootstrap and `shutdown` keep their existing behavior.

The four bootstrap inputs are all-or-nothing and use absolute config/release
paths. Rollbridge binds its proxy, activates the release through the normal
deploy path, then exposes the control socket and stays foreground. A failed
activation stops only processes started by that attempt and exits non-zero;
without `ownerRecovery`, persisted processes from a previous daemon are reported
as advisory orphans and are never recovered or killed implicitly. With
`ownerRecovery`, bootstrap reconnects only to the matching guardian and
reconstructs its provenanced generations before exposing control.

External supervisors may add `--boot-attestation` with exactly `sha256:` plus
64 lowercase hexadecimal characters. After successful activation, `rollbridge
status` echoes the opaque, non-secret value under `bootstrap.attestation`
alongside the exact release id/path/revision. Listener-only and detached ensured
daemons omit `bootstrap`, so a supervisor can distinguish a newly accepted
foreground owner from a stale daemon without Rollbridge interpreting the token.
See [`docs/cli.md`](docs/cli.md#daemon) for the complete contract.

Start the daemon only when it is not already running:

```bash
rollbridge ensure-daemon --config rollbridge.js --daemon-log-path log/rollbridge.log --daemon-pid-path tmp/pids/rollbridge.pid
```

Deploy a prepared release:

```bash
rollbridge deploy --config rollbridge.js --release-path /home/dev/ticket-server/releases/20260521073000/ticket-server --revision abc123
```

Deploy and start the daemon first when needed:

```bash
rollbridge deploy --ensure-daemon --config rollbridge.js --release-path /home/dev/ticket-server/releases/20260521073000/ticket-server --revision abc123
```

Inspect state:

```bash
rollbridge status --config rollbridge.js
rollbridge status --no-logs --config rollbridge.js
```

`status` reports each managed process's `state`, `pid`, recent `logs`, last
`exitCode`/`exitSignal`, and — per process — its automatic-restart count
(`restarts`), last start time (`startedAt`), current `uptimeMs` while running,
and why it last started (`lastStartReason`: `deploy`, `crash`, `manual`, or
`memory`). The same reason appears on each `process started` entry in
`rollbridge events`. For memory-supervised processes it also reports current
`rssBytes`, `memoryRestarts`, `lastMemoryRestartAt`, and `children` (the sampled
process tree — each group member's `pid`, `command`, and `rssBytes`).

For machine lifecycle attestations that do not need captured process output,
pass `--no-logs`. It returns the same status projection while omitting only each
release, service, and singleton process's `logs` array.

Print the recent captured stdout/stderr per process (a one-shot snapshot of the
retained `outputLines`, not a live stream):

```bash
rollbridge logs --config rollbridge.js
rollbridge logs --config rollbridge.js --process web
```

Print the daemon's recent structured event history — deploys, traffic switches,
release stops, process crashes/restarts, and failed commands (the most recent
1000 events, in memory):

```bash
rollbridge events --config rollbridge.js
rollbridge events --config rollbridge.js --limit 20
```

Stop the active release:

```bash
rollbridge stop --config rollbridge.js
```

Roll back to a previous release — re-starts it, health-checks it, and switches
traffic back (defaults to the most recently retired release; a failed rollback
leaves the current release active). Rollback manages processes only, not
database migrations:

```bash
rollbridge rollback --config rollbridge.js                  # the previous release
rollbridge rollback --config rollbridge.js --release-id v3
```

Restart non-proxied processes in place — all of them, one by id, or a policy
group (the proxied process is never restarted; use `deploy` for that):

```bash
rollbridge restart --config rollbridge.js                      # all non-proxied processes
rollbridge restart --config rollbridge.js --process background-jobs-worker
rollbridge restart --config rollbridge.js --policy companion
```

Shut down the daemon and managed processes:

```bash
rollbridge shutdown --config rollbridge.js
```

A successful shutdown response is emitted only after the targeted control
endpoint has stopped accepting connections and been removed, owned processes
and the proxy have stopped, and persistent state cleanup has finished. It is
therefore safe to start or ensure a replacement daemon immediately, without a
delay or retry loop. Cleanup failure or an already-missing daemon exits non-zero.

Prepare a first Rollbridge deploy by recovering Rollbridge-managed orphans and
stopping configured legacy processes:

```bash
rollbridge predeploy-cleanup --config rollbridge.js --release-path /srv/app/current
```

Enable shell completion (bash or zsh) for command names and option flags:

```bash
source <(rollbridge completion bash)   # add to ~/.bashrc
source <(rollbridge completion zsh)    # add to ~/.zshrc
```

## Nginx

Nginx should proxy to Rollbridge, not directly to Velocious:

```nginx
location / {
  proxy_pass http://127.0.0.1:8182;
  proxy_http_version 1.1;
  proxy_set_header Upgrade $http_upgrade;
  proxy_set_header Connection "upgrade";
  proxy_set_header Host $host;
  proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
  proxy_set_header X-Forwarded-Proto $scheme;
}
```

See [`docs/nginx.md`](docs/nginx.md) for the full guide — WebSocket upgrade
headers, timeouts for long-lived connections, forwarded headers, and common
failure modes (502/503, dropped WebSockets).

## Running under systemd

Run the long-lived daemon as a systemd service so it starts on boot and is
restarted if it crashes. A ready-to-edit unit lives at
`examples/rollbridge.service`:

```bash
sudo cp examples/rollbridge.service /etc/systemd/system/rollbridge.service
# edit User/Group, WorkingDirectory, the ExecStart path, and --config
sudo systemctl daemon-reload
sudo systemctl enable --now rollbridge
sudo systemctl status rollbridge
```

The unit runs `rollbridge daemon --config <stable-config>` in the foreground,
so its output goes to the journal (`journalctl -u rollbridge`). Key directives:

- `KillMode=mixed` / `KillSignal=SIGTERM`: Rollbridge stops its own managed
  child process groups on `SIGTERM`, so systemd signals only the daemon and
  lets it shut down gracefully before escalating to `SIGKILL`.
- `TimeoutStopSec`: give the daemon time to stop its managed processes; size it
  above the largest process `gracefulStopMs` (the daemon `SIGKILL`s stragglers
  after that). Note that `systemctl stop`/reboot stops processes but does **not**
  drain HTTP/WebSocket connections — connection draining happens only during
  `rollbridge deploy` release transitions.

The daemon is long-lived and survives deploys. **Deploy with
`rollbridge deploy` (or `rollbridge deploy --ensure-daemon`), not
`systemctl restart`** — pointing `--config` at a stable, daemon-wide file while
release paths are passed per deploy. The daemon reloads compatible process and
lifecycle config before each deploy, so updated graceful-stop deadlines govern
the release being retired without interrupting the stable proxy. Listener and
process-topology changes still require a daemon restart; see
[`docs/config.md`](docs/config.md#config-reloads). Use `command -v rollbridge`
to find the absolute CLI path for `ExecStart`.

See [`docs/logging.md`](docs/logging.md) for where the daemon's JSON logs go
(stdout / journald / the `--daemon-log-path` file) and how to rotate them — the
daemon holds its log file open, so logrotate needs `copytruncate`.

## Deployment Notes

Run migrations before `rollbridge deploy`, and keep migrations backwards-compatible while old and new web and jobs generations overlap. Velocious Beacon may be a persistent fixed-port `service`; configure `background-jobs-main` as a release-scoped handoff service on a port range. A normal deploy may leave several retired generations draining concurrently.

See [`docs/deploy-recipes.md`](docs/deploy-recipes.md) for ready-to-use shell, CI, and Capistrano recipes that drive Rollbridge through its CLI, and [`docs/troubleshooting.md`](docs/troubleshooting.md) for diagnosing health-check failures, port conflicts, stale sockets, crash loops, and stuck draining releases.

## Releasing

Maintainers can publish a patch release from the latest default branch:

```bash
npm run release:patch
```

The release script owns the package version bump, lockfile update, default-branch
commit, push, and npm publish. Do not run `npm version` manually before running
it.

See [`docs/releasing.md`](docs/releasing.md) for the maintainer release checklist
— the pre-flight checks before `npm run release:patch` and what to verify after.

## License

Rollbridge is released under the [MIT License](LICENSE).
