/** * Sustained Latency Demotion Subsystem. * * Charter & Invariants: * - Demotion Only: Latency measurements are used exclusively to demote candidate usability; * they never promote, drop, or re-sort deployment order. * - Two-Tier Signal Precedence: Per-token latency (ms/token p95 over ≥ minSamples request samples) * is the primary signal. Absolute probe latency (p95 over probe samples only) acts as fallback. * - Unmeasured Invariance: Deployments without qualifying sample history evaluate to Infinity * (no opinion), ensuring unmeasured models are never penalized or demoted. * - Self-Healing Rolling Window: No persistent cooldowns are registered; demotion state is * dynamically derived from recent ping samples and clears automatically upon performance recovery. * - Fault Tolerance: Evaluator functions are fail-safe and catch internal errors to yield null, * preventing latency estimation faults from failing live proxy requests. * * ⚠ **This deliberately reverses a rationale recorded in place, by OWNER DECISION (2026-08-30).** * `server.ts` argues against re-ranking on stability, in these words: a second ranking pass * "means neither decides the order", and "live health then PROMOTES on evidence that is often a * single request's latency". That objection was shown to the owner with this option's cost stated, * and the owner chose this option anyway. So a later reader must NOT treat this module as drift * and "restore" the old behaviour as a regression fix. The objection was also right about the * failure modes, and every bound below is a direct answer to one of them: * * - **Demote only.** Never promote, never drop, never re-sort. The objection is about a competing * ranking PASS; a one-way term folded into `targetUsability` is not one. Fitness still decides * the order, exactly as it does for quota. * - **Never act on one request's latency.** Every figure is a p95 over at least `minSamples` * qualifying samples — the objection's own worst case, ruled out by construction. * - **Unmeasured has NO effect whatsoever.** Both statistics answer `Infinity` when nothing * qualifies, and `Infinity` is an UNMEASURED deployment, never an infinitely slow one. Treating * it as slow would demote every never-probed member, which is the "unknown stays null, never 0" * invariant broken in the most damaging possible direction. * * WHY IT EXISTS. Measured 2026-08-30: an offload lane read as "stalled" was in fact paying the * relay's own candidate walk — single requests took 120-123 s across 2-6 attempts. Five top-ranked * `pool/medium` members carried an OPEN breaker, and `nim/deepseek-ai/deepseek-v4-flash` was * breaker-CLOSED with a p95 of 70364 ms. Because health banding read breaker state and nothing * else, that healthy-but-glacial member was walked AHEAD of every cooling one. Evidence: * `docs/backlog.md`. * * ⚠ **WHICH DATASET — the answer changed once already, so it is stated plainly.** This reads the * PROBE dataset (`probe-cache.json`, through the injected `readPings` seam), which is what * `llm-relay candidates` displays and which SURVIVES A RESTART. It briefly read the BREAKER's * pings instead; those are request-path only, in memory only, and never written by `PingLoop`, so * the term went inert after every restart and could disagree with the surface an operator reads. * Owner decision, same day: use the probe dataset, and EXPAND it to carry request latency too * (`probe-cache.ts` `recordRequestSample`). * * ⚠ **TWO STATISTICS, because absolute latency alone cannot compare a probe with a generation.** * A probe asks for one token; a real request may generate hundreds, amortising the same fixed * overhead. So: * - **per-token** (`getP95MsPerToken`) is the primary signal, over REQUEST samples only. It is * what the owner asked for, and the only figure that is fair across sample kinds. * - **absolute** (`getP95`) is the fallback, over measurable **PROBE** samples only. It still * catches a deployment that is slow before it emits anything, and it works before any request * sample exists. * Per-token is tested FIRST, so a deployment with real traffic is judged on the better evidence * rather than on whichever ceiling happens to trip first. * * ⚠ **The absolute fallback reads PROBE samples ONLY, and that split is load-bearing * (2026-08-30).** It briefly read every measurable sample, request samples included, which put a * probe-calibrated ceiling in front of generation data — the exact comparison the paragraph above * says cannot be made. Measured live: `nim/nvidia/nemotron-3-ultra-550b-a55b`, which had served 59 * of this machine's 62 successful requests, answered one request with 632 tokens in 34863 ms. That * is **55.2 ms/token** against a 250 ceiling — healthy by the primary signal — yet it pushed the * mixed absolute p95 to 34863 and DEMOTED the deployment. Probe-only, the same deployment reads * 23478 ms and is not demoted. The per-token guard could not save it, because per-token engages * only at `minSamples` REQUEST samples and it had three. So a deployment demoted itself by * succeeding, inside a window every deployment passes through on its way to being measured. * ⚠ A request sample with NO token count therefore reaches NEITHER statistic. That is deliberate: * it is a generation of unknown length, so it is not normalisable and not what `p95Ms` describes. * Evidence: `docs/history/latency-demotion-regression-2026-08-30.md`. * * ⚠ **NO breaker cooldown is registered, and that is the design, not an omission.** Quota demotion * can register one because its evidence STATES a `resetsAt`. Latency states no reset, and this * relay never invents a cooldown duration — so instead the term is re-resolved from the rolling * sample window on every request, which means it lifts BY ITSELF as soon as the measurement * recovers, with no expiry anybody had to guess. The cost, stated: a latency demotion is invisible * to `/candidates`' cooldown column and to the dashboard Cooldowns panel, because those read * breaker state. The response header is its surface. * * Pure and bounded: no IO and no clock of its own. The factory wraps everything in try/catch — * this runs on the request path, and a routing hint must never be able to fail a request. It logs * nothing: routing is not an error stream. */ import type { LatencyDemotionConfig } from "./config.js"; import type { ResolvedAttempt } from "./resolved-attempt.js"; import { type PingRecord } from "./ping/metrics.js"; /** * Tunable defaults. These are TUNABLES, not provider facts, which is the distinction the * "a guess must never be labelled a measurement" invariant draws — it forbids inventing an * unpublished provider limit, price or context ceiling, and explicitly permits a tunable default. * * ⚠ Both are nonetheless CALIBRATED against this machine's own traffic rather than picked, and the * measurement is recorded here so it can be re-run: * * - **`msPerToken` = 250.** Over 68 real requests (2026-08-30) the population ran p50 40.4, * p75 70.5, p90 292.0, p95 967.1 ms/token. Per deployment the separation was clean: a healthy * `nemotron-3-ultra` at a median 36.3 and `minimax-m3` at 57.3, against `gemini-3.6-flash` at * 687.8. 250 sits about 3.5x above the healthy band and well under the bad one, so it demotes * the deployment that was costing whole requests and leaves the ones that were working. * - **`p95Ms` = 30000.** The fallback ceiling, from the same window: the member that actually * SERVED had an absolute p95 of 23478 ms, the one that burned the walk 70364 ms. */ export declare const DEFAULT_LATENCY_MS_PER_TOKEN = 250; export declare const DEFAULT_LATENCY_P95_MS = 30000; /** * Minimum qualifying samples before latency may demote anything, applied to EACH statistic against * its own sample set. Five is the smallest count for which a p95 is not simply "the worst of a * handful", and it is the direct answer to the recorded objection about acting on a single * request's latency. */ export declare const DEFAULT_LATENCY_MIN_SAMPLES = 5; /** One sustained-latency verdict — the smallest honest statement of "why this cell stepped aside". */ export interface LatencyDemotion { /** Which statistic crossed its ceiling. Per-token wins when it has enough evidence. */ readonly basis: "per-token" | "absolute"; /** The measured figure: ms/token for `per-token`, ms for `absolute`. Finite by construction. */ readonly measured: number; /** The ceiling it exceeded, in the same unit. */ readonly threshold: number; /** Qualifying samples behind `measured`. Never below `minSamples`. */ readonly samples: number; } export type LatencyDemotionFn = (attempt: ResolvedAttempt, now: number) => LatencyDemotion | null; /** * How this module reads latency samples. * * ⚠ A plain function, NOT a `PingLoop`. Keeping the seam narrow is what lets this module stay pure * and testable, and it means nothing here can reach into probe scheduling or quota state. The * server passes `PingLoop.getModelPings`; the suite passes an array. */ export type PingReader = (provider: string, model: string) => readonly PingRecord[]; export interface LatencyDemotionDeps { readonly readPings: PingReader; /** * ⚠ The SHAPE is owned by `config.ts` (`LatencyDemotionConfig`) and imported, never re-declared * here. Two hand-written copies of one settings object is a defect class this codebase has hit * repeatedly, and the copies always drift. `config.ts` also normalizes the boolean shorthand * away, so this module never has to decide what `false` means. */ readonly settings?: LatencyDemotionConfig | undefined; } export declare function resolveLatencyDemotion(deps: LatencyDemotionDeps, attempt: ResolvedAttempt): LatencyDemotion | null; /** * Build the request-path evaluator. The wrapper is the safety seam: NOTHING inside may throw into * the request path, and a failure degrades to "no opinion" — the pre-2026-08-30 behaviour — rather * than to a refused request. Deliberately silent: routing hints are not log-worthy events. */ export declare function createLatencyDemotionFn(deps: LatencyDemotionDeps): LatencyDemotionFn; /** * `" (p95 687.8ms/token > 250ms/token over 8 request samples)"`, or the absolute form. * Bounded, metadata only — a rate and a count, never a prompt, a credential or an id. */ export declare function latencyDemotionLabel(spec: string, demotion: LatencyDemotion): string;