# Shared multi-computer render queue

## Problem and boundary

`LaneQueue` originally stored leases in a local SQLite database. That is correct
for several processes on one computer but cannot coordinate two computers: each
machine sees an empty database and both can submit to the same one-slot
Higgsfield Unlimited account.

The shared coordinator is the authority for production provider-account lanes.
It coordinates leases and audit receipts only. Prompts, images, videos, browser
profiles and provider credentials remain on the originating computer.

## Coordination atom

One Cloudflare Durable Object exists per exact `route:account` lane, for example:

```text
seedance-direct:main-unlimited
```

Every client deterministically addresses the same object by this name. The
object's single-threaded execution and strongly consistent SQLite storage make
acquire, deduplication, promotion, heartbeat, expiry and release serializable.

## State machine

```text
queued -> held -> done
             \-> expired -> done
```

- `queued`: request is waiting under project-round-robin fairness.
- `held`: request owns provider capacity until release or lease expiry.
- `done`: terminal, retained in a bounded history window for audit and ETA.

A lane can also be **paused** (`lane_cooldowns`, one row per lane): while a
pause is active nothing is promoted from `queued` to `held` on any computer,
`acquire` answers `queued` with the pause attached and an ETA that includes it,
and the alarm fires when it ends. See "Provider human check" below.

Provider work has a separate append-only receipt trail:

```text
lease-acquired -> references-applied -> provider-submitted -> provider-terminal
```

A receipt cannot be replaced with different content. Replaying the same
content hash is idempotent; a different hash for the same receipt identity is a
conflict.

The coordinator checks the hash it is sent. It recomputes `sha256:<hex>` over the
payload's canonical JSON (keys sorted by code unit, not by locale, so the CLI and
the Worker produce the same bytes) and stores the verdict: the receipt response
and each `recentReceipts` entry carry `contentHashVerified` — `true` when the sent
hash matched, `false` when it did not, `null` for a receipt stored before hashes
were checked. A mismatch is stored and flagged, with a `receipt-hash-mismatch`
warning in the control room; it is not refused, because a receipt that reads
"unverified" is worth more than one that was dropped. `vclaw video lane receipt
--content-hash <value>` still accepts an override, and an override that does not
match its payload is exactly that case.

### Reading the trail

```bash
npm run shared-queue:status -- --verbose      # or: vclaw video lane status --route <id> --verbose
```

`--verbose` adds `recentReceipts` (newest first, up to 50 per render queue) to the
status answer: each one's id, ticket, phase, machine, payload, content hash,
`contentHashVerified` and `createdAt` (fractional epoch **seconds**, not
milliseconds — multiply by 1000 before handing it to `new Date`). Without the
flag the answer is unchanged, so the reply a driver reads in a loop while it
waits for a turn stays short. The flag saves no network: the coordinator returns
those rows either way and the client already parses the whole body, so what
`--verbose` decides is only what gets printed. On the local single-computer queue
the field is simply absent — there are no receipts there.

Alongside it comes `unreadableReceipts`: how many rows the coordinator sent that
this client could not read. `0` says the trail is complete. A malformed row is
skipped rather than failing the whole status read, but it is never skipped
*silently* — on a surface whose job is evidence, a short list must not be
mistakable for a step that left no receipt. Read a non-zero count as a CLI older
than whatever wrote the receipt, never as lost work.

One case where `--verbose` carries nothing and is not telling you the trail is
empty: a route with no measured limit. An unenforced lane grants `unlimited-…`
tickets the coordinator never holds, so it stores no receipts for them and the
status answer has no receipt keys at all — the same shape as the local queue.

`contentHashVerified` is `true` when the coordinator recomputed the payload hash
and it matched, `false` when it did not, and `null` in two cases that look the
same from here: a coordinator too old to check, and a receipt written **before**
this coordinator was upgraded to 0.4.0 — the verified column was added by
migration, so rows stored earlier hold no verdict. Receipts written by the first
release that recorded them pre-date the check and read `null` on a current
coordinator.

The Cinema console does not show receipts yet. That is a deliberate deferral, not
an oversight: the console is a read-only surface built from on-disk artifacts, and
showing receipts means adding a coordinator network call to it, which deserves its
own decision. (The console already models the render queue — `CinemaConsoleTask`
carries `laneId`, `lanePosition` and `laneEtaSeconds`; only `ticketId` is not yet
projected.)

## HTTP contract

All `/v1/*` endpoints require `Authorization: Bearer <shared secret>`. The
public `/healthz` endpoint exposes no queue data.

| Endpoint | Method | Purpose |
|---|---:|---|
| `/v1/lanes/acquire` | POST | Atomically deduplicate, enqueue and fill capacity |
| `/v1/lanes/status` | POST | Reap expiry, promote capacity and return holder/waiters |
| `/v1/lanes/heartbeat` | POST | Extend one exact held ticket |
| `/v1/lanes/release` | POST | Finish one exact ticket and promote the next waiter |
| `/v1/lanes/receipt` | POST | Append an immutable prompt/reference/provider receipt |
| `/v1/lanes/history` | POST | The render log: done tickets newest first, with the outcome reported on release (`limit`, `project`) |
| `/v1/lanes/cooldown` | POST | Pause this render queue for every computer (`reasonCode: provider-human-check`, `retryAfterSeconds` 30–7200); only the owner of a HELD ticket may |
| `/v1/dashboard` | GET | Legacy authenticated redirect to the control room |

Every mutation carries `protocolVersion`, exact `lane`, `machineId`, exact
`ticketId` where applicable, and stable request identity (`project`, `scene`,
`promptHash`) on acquisition.

## One claim per driver — two sessions on one machine

The ticket's `holder` is the machine id. Two Claude sessions on one Mac share
`VCLAW_MACHINE_ID`, so on 2026-09-04 the coordinator could not tell the two
drivers apart: session B's `release` ended session A's held ticket under A's
in-flight render, and a probe from the same machine was granted the slot.

Every driver now mints a **claim** (`VCLAW_LANE_CLAIM`, a UUID by default;
`lane_render.py` mints one per process and exports it into every lane call) and
sends it on `heartbeat` and `release`. When a request names a claim, the
coordinator matches it against the ticket's own `claimId`; a request without
one keeps the machine-only rule. That is deliberate, not only legacy: a hand
`vclaw video lane release --ticket <id>` from the same machine names no claim,
so a SIGKILLed driver's ticket can still be unwedged before its TTL. The
dashboard still groups by machine, and every held or queued row on the status reply and the control
room now carries the ticket's `claimId` beside the holder, so two drivers on one
Mac read as `render-macbook · lane-render-58628-…` and `render-macbook ·
lane-render-71204-…`, not as one holder.

A driver reads the heartbeat's answer. `refreshed: false` three times in a row
means the coordinator no longer holds the ticket for this driver: the poll
stops, the window is parked as `needs: rerender:lease-lost` (the provider job
keeps running; resume reconciles it from the live-job record) and nothing is
resubmitted on a slot someone else may hold.

`skills/rap-avatar-mv/scripts/lane_trace.py` is the receipt trail for one
lease: it holds a ticket and logs the heartbeat answer and the queue's
`held`/`queued` lists every 30 s. Run it only on a quiet lane, alone and then
with a second copy under a different claim, and read the two trails together.

## The render log

Every ticket that reaches `done` is a row in the render log — one row per draw of
one window — and the driver says what happened as it gives the slot back:
`release` carries `outcome` (`completed`, `failed`, `timeout`, `lease-lost`,
`not-submitted`, `slot-busy`, `abandoned`), the provider job id and a short
note; a lease the coordinator reaps is logged `expired`; a release with no
outcome (an older client, a hand release) leaves it null. `/v1/lanes/history`
returns the log newest first (`limit` up to the retained 5000, `project` to
filter), and the control room's **Render log** panel lists it with a project
filter and outcome counts. It is kept out of `/v1/lanes/status`, which stays
small enough for `lane status` and `lane_trace.py` to print whole.

## Provider human check — one driver's wall pauses the render queue for everyone

A Cloudflare Turnstile in front of a submit is anti-bot **rate limiting keyed on
the account's cadence** (three back-to-back submits tripped it on 2026-08-07),
not a verdict on the prompt. Every other driver on the same `route:account` lane
is about to hit the same wall and burn a draw on it. So the driver that sees it
reports a cooldown **on the ticket it still holds** — the coordinator refuses
one from a queued or released ticket, because a ticket that never spoke to the
provider has no wall to report — and then releases as usual:

```bash
vclaw video lane cooldown --route seedance-direct --ticket <ticketId> \
  --reason provider-human-check [--retry-after 900]
vclaw video lane release --route seedance-direct --ticket <ticketId> --outcome not-submitted
```

This is automatic wherever a ticket is still held when the wall is seen
(`src/video/provider-human-check.ts`): `execute-status` reports it when a poll
returns a Turnstile in its `issues` — the shape the free Seedance engine uses,
since its submit always mints a job id — and `produce`/`execute` report it when
an adapter throws one at submit. The rap lane's `lane_render.py` does the same
before it releases. So does the Cinema production queue (`cinema-work`): a
failed lookup, a submit that throws and `failCinemaTask` each report it on the
held ticket first and release second — a submit that throws keeps its ticket,
because whether the provider accepted the job is still unknown. The machine that paused the queue prints one line saying
so. The pause length is
`VCLAW_PROVIDER_HUMAN_CHECK_COOLDOWN_SECONDS` (default 900, clamped to 30–7200).
A longer report replaces a shorter one; a shorter one never trims an active
pause. `lane status` shows the pause with who reported it; a queued `lane await`
prints "paused — provider-human-check reported by <machine>" once, so "position
1 forever" reads as a pause rather than a wedged lane. The control room lists
`provider-cooldown-started` and `provider-cooldown-ended` incidents.

What it does **not** do: clear the challenged driver's own fingerprint (nothing
does; a challenged headless fingerprint stays challenged), and it does not fire
on Google Flow's reCAPTCHA 403, which the useapi Flow API solves itself
(`captchaRetry`, `flow-captcha.ts`) — a bare "captcha" trigger would pause the
whole `veo-useapi` lane for fifteen minutes on a condition the auto-solver
handles. The reporting driver's own retry wait (`TURNSTILE_WAIT_S`, 120 s) now
queues behind its own lane pause: its next `await` waits the pause out, which
is the point — the wall is on the account's cadence, its own included.

A coordinator deployed before this endpoint existed answers the cooldown call
with 404; the client surfaces that as a warning and the render report is the
signal to you as before. Redeploy with the readiness gate below to enforce
it across computers.

## Fail-closed client policy

Configure both computers with:

```text
VCLAW_SHARED_QUEUE_URL=https://<worker>.<account>.workers.dev
VCLAW_SHARED_QUEUE_TOKEN=<same secret on both computers>
VCLAW_MACHINE_ID=<unique stable machine name>
VCLAW_LANE_ACCOUNT=main-unlimited
VCLAW_SHARED_QUEUE_REQUIRED=1
```

When `VCLAW_SHARED_QUEUE_REQUIRED=1`, missing configuration, coordinator
timeouts, authentication failures and malformed responses are pre-provider
errors. VideoClaw never falls back to local SQLite. Local SQLite remains the
explicit development default when the required flag is absent. Never place
`lanes.db` in Dropbox, iCloud Drive, NFS or another synced folder; SQLite WAL is
not a cross-machine consensus protocol.

## Reference and provider receipts

Evidence policy belongs to each job, not to the shared lane. Before a
reference-sensitive render is accepted, the client records the required slot
order, prompt and reference hashes, provider media IDs, provider job ID,
execution mode, wallet bracket, machine, project, scene and attempt identity.
A text-only job can explicitly declare `requiredSlots: []`; image-to-video,
extend, edit, audio and multi-reference jobs declare only the inputs they
actually require. A project that needs a stricter rule — say, exactly three
ordered image slots with non-empty provider media IDs — states it in its own
contract. A rule like that is never inherited by another project.

### What `produce` writes for you

On the shared queue, a direct render (`vclaw video produce`, and the
`execute-status` that follows it) writes four receipts on the ticket it holds, in
order: `lease-acquired` (project, route, the request hash the ticket was granted
for, scene indices), `references-applied` (the sha256 of each reference file,
measured before the submit), `provider-submitted` (the provider job id) and
`provider-terminal` (how it ended: `completed`, `failed`, or `submit-failed`).
Each is written while the ticket is still held — the coordinator checks that the
machine writing a receipt is the ticket's holder, though not that the ticket is
still live — and each has the id `<ticket>:<phase>:<hash prefix>`. So the same
facts are the same receipt however often they are written, and different facts
are a new line in the trail rather than a rewrite the coordinator would read as
tampering.

They are evidence about a render, never a condition of it: a receipt the
coordinator refuses, or cannot be reached for, is dropped and the render carries
on. They hold hashes, ids and counts, never a prompt, a file path, a token or an
account name. The local single-computer queue has no receipts, so there nothing
is written. Not yet: the coordinator stores each receipt's content hash as sent
and does not recompute it, `lane status` does not show receipts, and the Cinema
queue worker does not write them.

## Security

- Store the token with `wrangler secret put SHARED_QUEUE_TOKEN` and in each
  computer's local secret store; never commit it.
- The Worker compares bearer-token bytes with constant-time Web Crypto.
- Request bodies are size-bounded and schema-validated.
- Status and dashboard endpoints are authenticated because ticket IDs are lease
  capabilities.
- Provider credentials and browser sessions never reach Cloudflare.

## Deployment readiness gate

Deployment does not resume generation. Coordinator readiness requires service
tests and a Wrangler dry-run, both computers configured with distinct machine
IDs, racing same-request clients converging on one ticket, distinct requests
respecting capacity one, heartbeat/expiry recovery, and immutable receipt
conflict checks. Render readiness is evaluated separately against the selected
job's own contract and human approval. One project's evidence cannot approve
another project.

The first three of those — the service typecheck, its tests and the Wrangler
dry-run — are one command, `npm run shared-queue:verify`, and CI runs the same
three on every pull request that changes `services/shared-lane-coordinator/`
(other pull requests skip it). The tests pass on code that does not typecheck,
so a green test run alone is not the gate: run the whole command. It needs
Node 22 (Wrangler 4 refuses Node 20) and no Cloudflare login.

## Deploy once

Wrangler is scoped to `services/shared-lane-coordinator/`:

```bash
cd services/shared-lane-coordinator
npm ci            # a fresh checkout has no node_modules here; the typecheck fails without it
npm run typecheck
npm test
npm run deploy:dry
npx wrangler secret put SHARED_QUEUE_TOKEN
npm run deploy
```

Use a randomly generated secret with at least 32 bytes. `wrangler secret put`
reads it without committing it to `wrangler.jsonc`. The first deployment creates
the Durable Object class through migration `v1`; later schema changes must add a
new migration tag instead of changing `v1`.

## Configure both computers

Run once on each Mac, with a different stable machine ID:

```bash
scripts/setup-shared-queue-machine.sh \
  https://videoclaw-shared-lane-coordinator.<account>.workers.dev \
  studio-mac-mini
```

Use `render-macbook` (or another distinct label) on the other computer. The
helper stores the token in the macOS Keychain and non-secret settings in
`~/.config/videoclaw/shared-queue.env` with mode `0600`.

Run VideoClaw through the safe wrapper:

```bash
scripts/with-shared-queue.sh node dist/cli/vclaw.js video lane status \
  --route seedance-direct --account main-unlimited
```

The result must include `"coordinator":"cloudflare-durable-object"`. Any
timeout, bad token or malformed response exits before provider submission.

## Read-only control room

The deployed control room shows live leases, queued requests, recently
observed computers, immutable provider/reference receipts, coordinator
incidents, coordinator evidence and the selected job's scoped evidence. It
cannot acquire, release, submit, approve, upload or generate anything.

On an enrolled Mac, open it without putting the bearer token in a URL or shell
history:

```bash
npm run shared-queue:dashboard
```

The helper copies the existing Keychain credential to the clipboard and opens
the dashboard. Paste it into **Unlock ledger**. The browser keeps the token in
that tab's `sessionStorage`; **Lock** removes it immediately.

Machine presence means “the coordinator recently observed a signed queue
operation”, not merely “the computer is switched on”. A machine is shown as
online for three minutes after its last signed action. Coordinator proof (two
recent machines and a safely deduplicated collision) is displayed separately
from job proof (stable execution identity, that exact ticket's reference policy
and that exact ticket's human approval). Missing evidence is reported as
unknown or incomplete; it is never borrowed from another scene, attempt or
project.

## Two-computer acceptance drill

This drill spends no credits and calls no video provider.

Run the automated single-machine engine portion on a dedicated remote lane:

```bash
npm run shared-queue:self-test
```

The script refuses the `main-unlimited` account, verifies that the remote
coordinator is enforced, exercises capacity-one serialization, same-request
deduplication, heartbeat, atomic promotion, receipt replay and conflicting
receipt rejection, and releases every test ticket. It never invokes a provider
adapter, uploads media or submits a generation. Override its isolated account
only with `VCLAW_QUEUE_SELF_TEST_ACCOUNT`; production is always rejected.

Then perform the genuinely cross-computer portion:

1. On computer A, acquire a long-lived ticket with prompt hash
   `shared-queue-drill-a`. It must be `granted`.
2. On computer B, acquire a distinct request `shared-queue-drill-b`. It must be
   `queued` at position 1.
3. On both computers at nearly the same time, acquire the same request
   `shared-queue-race`. They must report one canonical ticket and only its
   owning process may be granted.
4. Release A's exact ticket. B's distinct request must become held.
5. Record a test receipt, replay it unchanged (idempotent), then try the same
   receipt ID with different content (HTTP 409).
6. Stop heartbeats on a short test lease and verify the alarm promotes its
   waiter after expiry.

Release every drill ticket. Do not use production prompts or references during
the drill. Each project remains governed by its own execution approval and
reference contract: a project whose contract demands proof of its references
stays paused until its own real `references-applied` receipt provides it.

## Rollback

For local development only, omit `VCLAW_SHARED_QUEUE_REQUIRED` and all shared
queue variables; VideoClaw then uses local SQLite. Never do this on either Mac
that can reach the shared Higgsfield Unlimited account. Production rollback is
to stop submissions, repair/redeploy the coordinator, and verify the acceptance
drill—not to set `VCLAW_LANE_DISABLE=1`.
