/** * Per-HQ-root operation gate for long-running operations (`sync`, `rescue`, * `reindex`). * * Contract: * - Normal `sync` processes publish distinct shared worker memberships. This * lets workers mutate independent paths concurrently. `rescue` (and a * format-changing migration) take the same per-root gate exclusively: its * queued writer intent stops later shared admissions, then it waits for the * already-admitted worker PIDs to leave. Different HQ roots are independent. * - `reindex` uses a SEPARATE per-root scope ("reindex"), so it is guarded * against other reindexes but is NOT blocked by a held sync/rescue lock. * This is deliberate: a watch-mode `hq-sync-runner` holds the "operation" * lock across its entire lifetime, and a shared lock would starve a * standalone `hq reindex` indefinitely (feedback_ed98d810). reindex only * READS the source trees to rebuild derived artifacts (skill wrappers, * overlay mirrors, the workers registry) and is idempotent + re-run after * every sync pass, so decoupling it from the sync writer lock is safe. * See {@link lockPathFor} for how the scope partitions the lock file. * - The push watcher / watch+event-push runner is EXEMPT: it never calls in * here, so it neither takes the lock nor is blocked by it (its targeted * in-process push passes are likewise lock-free). * * ## Where the lock lives — and why * * `/locks/operation-.lock`, plus short-lived * admission, writer-intent, and per-PID membership siblings, where * `stateDir = $HQ_STATE_DIR || ~/.hq`. This is deliberately NOT inside the HQ * root: * - It must never round-trip to the cloud. A lock is machine-local, per-run * state; syncing it to S3 (and thence to other machines/roots) would be a * correctness bug. `~/.hq` is the established machine-local state dir * (journals already live there) and is never synced. * - `rescue` repairs a possibly-broken HQ root; a lock that depends on the * root being healthy is exactly backwards. `~/.hq` is independent of the * root's health. * - Keying the filename by a hash of the *canonical* root path makes the * lock per-root and prevents leakage across roots, while keeping the path * short and filesystem-safe. * * ## Atomicity, liveness, takeover * * - Every gate record is fully written to a same-directory temp file, fsynced, * and published with create-if-absent semantics. The short admission record * serializes only membership transitions, never a sync pass or network I/O. * - Gate records carry `{ pid, command, startedAt, hqRoot, mode }`. On an * existing record we test its PID with `process.kill(pid, 0)`: * * ESRCH → the holder is gone (crashed / killed -9 / stale file) → * reclaim the lock IMMEDIATELY (a dead holder never makes us * wait). * * EPERM → the PID exists but is owned by another user → treat as ALIVE * (conservative: wait rather than risk two concurrent ops). * * success → alive → WAIT for the holder to release, then acquire (see * "Waiting" below). The fast-refusal path is still reachable * via an explicit timeout / `wait: false`. * * ## Waiting for a live holder (default behavior) * * When a LIVE holder owns the lock, acquisition WAITS by default: it polls * (~2s) and acquires the instant the holder releases, rather than refusing * fast. A single status line is written to stderr the first time we start * waiting ("Waiting for (pid N) to finish…"), never per-poll. * This is what an interactive `sync` / `rescue` / `reindex` invocation wants — * queue behind the running op instead of erroring out. * * A bounded escape exists for scripts that must not block forever: * - `timeoutSec` option, or the `HQ_OP_LOCK_TIMEOUT` env var (seconds). * The option wins over the env. After the bound elapses we throw * {@link OperationLockedError} (exit 17) with the same clear refusal * message as before. * - `timeoutSec === 0` (or `HQ_OP_LOCK_TIMEOUT=0`, or `wait: false`) → do * not wait at all; refuse immediately. This is the pre-wait behavior. * - absent / negative / unparseable → INFINITE wait (the documented * default). * Stale-PID takeover is unconditional and happens BEFORE any wait — a dead * holder is reclaimed at once regardless of the wait config. * * Ordering / scope caveats: * - Wait-capable acquisitions publish a per-process waiter intent before * trying the mutex. Long-lived/watch callers can opt to defer to those * intents, preventing repeated short sync passes from starving an * already-waiting rescue at each release/reacquire boundary. * - This is a CROSS-PROCESS mutex keyed on the holder's PID. In-process * concurrent acquire is unsupported (the real consumers — sync / rescue / * reindex — are separate processes), but a live same-PID holder is still * treated as busy rather than reclaimed. * - When several distinct processes wait on the same lock, the next one to * win the O_EXCL race after a free acquires. Order is best-effort, NOT * FIFO — do not depend on arrival order. * - PID reuse is an inherent, un-eliminable race for any PID-based scheme: if * the original holder crashed and the OS later handed its PID to an * unrelated process, we conservatively read that as "still held" and * refuse. We accept that false-busy over the far worse false-free, and * record `startedAt`/`command` so an operator can diagnose a wedged lock. * * ## Release * * - Normal exit: the `with*` wrappers release in a `finally`. * - Signals (SIGINT/SIGTERM): a one-time handler releases every held lock, * then re-raises the default disposition so exit status is unchanged. * - Hard crash (SIGKILL / power loss): nothing runs, but a later gate * contender reclaims dead exclusive owners and dead shared memberships. * - `process.on("exit")`: a final best-effort synchronous unlink. * * ## Escape hatch * * `HQ_DISABLE_OP_LOCK=1` makes acquisition a no-op (returns a handle whose * release does nothing). For emergencies and for callers that manage * exclusion themselves; documented, off by default. */ /** Process exit code used when an operation is refused because the lock is held. */ export declare const OPERATION_LOCKED_EXIT = 17; /** * Process exit code used when an operation is refused because the lock's state * directory is not writable (a permission-class fs error, not a live holder). */ export declare const OPERATION_LOCK_UNWRITABLE_EXIT = 18; export interface LockInfo { pid: number; command: string; /** ISO-8601 acquisition time. */ startedAt: string; /** Canonical HQ root the lock guards (diagnostic only). */ hqRoot: string; /** Gate role. Omitted payloads are legacy exclusive mutex holders. */ mode?: OperationLockMode; } /** * Normal sync processes hold shared membership. Repair and migration callers * hold exclusive ownership. The explicit option is primarily for callers whose * command is not named `sync`; the historical API continues to classify the * `sync` command as shared and every other command as exclusive. */ export type OperationLockMode = "shared" | "exclusive"; /** Thrown by `acquireOperationLock` when a LIVE holder owns the lock. */ export declare class OperationLockedError extends Error { readonly holder: LockInfo; readonly attempted: string; constructor(holder: LockInfo, attempted: string); } /** * Thrown by `acquireOperationLock*` when the per-root lock CANNOT BE CREATED * because its state directory is not writable — a permission-class fs error * (`EPERM` / `EACCES` / `EROFS`) on the lock dir or its temp file, NOT a live * holder. Distinct from {@link OperationLockedError}: nothing is holding the * lock; this process simply cannot write there — e.g. `~/.hq` owned by another * user (created by a past `sudo` run), a macOS privacy/security restriction * (which surfaces as `EPERM: operation not permitted`), or a read-only volume. * Callers turn this into a clear, actionable message + clean exit rather than a * raw uncaught `fs` crash (HQ-CLI-2). */ export declare class OperationLockUnwritableError extends Error { readonly lockDir: string; readonly cause: NodeJS.ErrnoException; constructor(lockDir: string, cause: NodeJS.ErrnoException); } export interface LockHandle { /** Absolute path of the lock file. */ readonly path: string; /** The info written for this holder. */ readonly info: LockInfo; /** Idempotently release the lock iff this process still owns it. */ release(): void; } /** Default poll interval while waiting on a live holder. */ export declare const DEFAULT_LOCK_POLL_MS = 2000; /** Options controlling how `acquireOperationLock*` behaves against a LIVE holder. */ export interface AcquireOptions { /** * When a LIVE holder owns the lock: `true` (default) → WAIT-poll until it * frees, then acquire; `false` → refuse immediately with * {@link OperationLockedError}. A `timeoutSec` of 0 is equivalent to * `wait: false`. */ wait?: boolean; /** * Bounded wait, in seconds, before giving up and throwing * {@link OperationLockedError} (exit 17). Precedence: this option > the * `HQ_OP_LOCK_TIMEOUT` env var > infinite. `0` → do not wait at all (refuse * immediately). Negative / non-finite → treated as absent (infinite wait). * Fractional values are honored (used by tests); the CLI flags accept whole * seconds. */ timeoutSec?: number; /** Poll interval in ms while waiting. Defaults to {@link DEFAULT_LOCK_POLL_MS}. */ pollIntervalMs?: number; /** * Invoked exactly ONCE, the first time we begin waiting on a live holder. * Defaults to a single "Waiting for …" line on stderr. Pass a custom hook * (or a no-op) to redirect/silence the status line. */ onWaitStart?: (holder: LockInfo, attempted: string) => void; /** * Yield acquisition to wait-capable callers that already published a waiter * intent for this lock. Watch mode enables this so its next short pass cannot * beat a rescue that was already waiting when the previous pass released. * Deferred callers do not publish their own intent, so periodic watch work * cannot queue ahead of interactive operations. */ deferToWaiters?: boolean; /** * Lock scope — partitions the per-root mutex into independent lock files (see * {@link lockPathFor}). Defaults to {@link DEFAULT_LOCK_SCOPE} ("operation"), * which `sync`/`rescue` share. `reindex` passes "reindex" so it is guarded * against other reindexes without being blocked by a held sync/rescue lock. */ scope?: string; /** Override the command-derived gate role for this acquisition. */ mode?: OperationLockMode; } /** * Default lock scope. `sync` and `rescue` share this one so they stay mutually * exclusive with each other, keyed only by the root. */ export declare const DEFAULT_LOCK_SCOPE = "operation"; /** * Absolute lock path for a given HQ root and `scope`. Exported for tests. * * The `scope` prefixes the lock filename so callers can partition the mutex: * `sync`/`rescue` use {@link DEFAULT_LOCK_SCOPE} ("operation"); `reindex` uses * its own "reindex" scope so a long-lived watch-mode sync-runner (which holds * the "operation" lock across its whole lifetime) can never starve a standalone * `hq reindex`. Different scopes hash to different lock files and never block * one another; the same scope is a real cross-process mutex. */ export declare function lockPathFor(hqRoot: string, scope?: string): string; /** * If `err` is a permission-class fs error against the lock dir, rethrow it as an * actionable {@link OperationLockUnwritableError}; otherwise rethrow it * unchanged. Always throws (return type `never`). * * Exported for unit testing: ESM module namespaces can't be spied, so the * classification decision is verified directly here rather than by mocking * `fs.openSync` (see operation-lock.test.ts, HQ-CLI-2). */ export declare function rethrowLockCreateError(err: unknown, lockDir: string): never; /** Acquire the operation gate synchronously. Watch-loop probes always pass a * zero timeout; all wait-capable parent paths use the async counterpart below. */ export declare function acquireOperationLock(hqRoot: string, command: string, opts?: AcquireOptions): LockHandle; /** Async gate acquisition. This is the only wait-capable path used by the * parent watch loop, so maintenance waits never block watcher heartbeats. */ export declare function acquireOperationLockAsync(hqRoot: string, command: string, opts?: AcquireOptions): Promise; /** Run `fn` while holding the per-root lock for `command` (async). */ export declare function withOperationLock(hqRoot: string, command: string, fn: () => Promise, opts?: AcquireOptions): Promise; /** Run `fn` while holding the per-root lock for `command` (synchronous). */ export declare function withOperationLockSync(hqRoot: string, command: string, fn: () => T, opts?: AcquireOptions): T; //# sourceMappingURL=operation-lock.d.ts.map