/** * Two judges reaching one account of what is wrong. * * One AI proposes, a second challenges what it filed, and what comes out is a * single set of findings each carrying who agreed and who did not. The order * matters and is not symmetric: the proposer names each defect, and the * challenger rules on those names rather than inventing its own, which is what * keeps one defect one row in the backlog. * * The fan-out lives HERE and not inside `judgeBatch`, deliberately. A dialogue * has to be a property of the call, because one caller must never inherit it: * `skills replay` grades an amendment by re-judging a frozen set of * screenshots, and the gate is already fighting the variance of one model * answering differently twice. Adding a second model's variance to that would * make the gate reject good amendments. Replay calls `judgeBatch` directly, so * it opts out by construction rather than by remembering to pass a flag. * * What a disagreement does, and does not, do: * * it is kept a disputed finding is still filed and still open. Nothing an * AI saw is discarded because another disagreed; both accounts * are recorded and a person reads them. * it is not a a dispute changes no status and no severity. Letting one AI * verdict close or downgrade another's finding would hand a single * vendor a veto over the backlog. * it is bounded one challenge round, and a second only when the challenger * added findings of its own for the proposer to rule on. There * is no convergence loop: residual disagreement is the OUTPUT, * not a failure to retry away. * * A challenge that fails is not fatal and not silent. The proposal stands * alone, the batch is reported as undialogued so the caller can decline to * cache it, and the reason is recorded. Caching a verdict as though two judges * had agreed when only one ever spoke is the one outcome worth avoiding. */ import { type JudgeBatchResult, type JudgeContext } from "./engine.js"; import type { ShotRecord } from "../types.js"; /** The AIs in a dialogue, and what each judges with. */ export interface Roster { proposer: { ai: string; model: string; }; /** Absent when only one AI is configured, which is the single-judge pipeline. */ challenger?: { ai: string; model: string; }; } /** One dialogue's outcome: a batch result, plus what the second judge did. */ export interface DialogueResult extends JudgeBatchResult { /** * False when a challenger was configured and could not be reached or could * not be understood. The caller leaves such a batch out of the cache: the * verdict is one judge's, and a cache entry keyed to a two-judge identity * would durably record a dialogue that never happened. */ dialogued: boolean; /** Why there was no dialogue, when one was expected. */ undialoguedReason?: string; } /** * Judge a batch with one AI, then have the other rule on what it filed. * * With no challenger this is `judgeBatch` and nothing else, which is what every * existing project gets until it configures a second AI. */ export declare function judgeWithDialogue(skillText: string, project: string, shots: ShotRecord[], evidenceDir: string, roster: Roster, ctx?: JudgeContext & { challengeSkill?: string; }): Promise;