/** * panel-grader — multi-judge LLM grader with consensus aggregation. * * Runs the same rubric prompt against N judge models in parallel, then * combines their normalized scores with a configurable aggregation * strategy (majority / unanimous / mean / median / min). * * Re-uses the same tool schema, system message, and per-criterion * structure as `prompt-grader` (via `buildRubricJudge` and `buildUserMessage`) * so reporters and downstream tooling see one consistent "judge response" * shape per judge. * * Design notes: * - Judge-error strictness: any judge that fails (non-zero exit / crash / * timeout / malformed response) forces the whole panel to fail, regardless * of the surviving judges' consensus. The surviving judges still produce the * reported `score`, but `passed` is forced `false` and the failed judges are * surfaced in `metadata.failed_judges`. * - Concurrency: all judges fan out together via `Promise.allSettled`. The * panel grader intentionally does not expose a per-invocation concurrency * knob — LLM rate-limit / throughput control belongs at a global layer * (e.g. on the `LlmClient` or runner) so it applies uniformly across every * LLM-using grader rather than only within a single panel call. * - Disagreement: the spread (`max − min`) of judge scores is exposed in * `metadata.disagreement` as a raw primitive. The grader intentionally * does not flag it against a threshold or affect `passed`/`score` — * downstream tooling (reporters, dashboards, alerts) can apply its own * policy on top of the raw value. */ import type { Grader, GraderInput, GraderMetadata, GraderResult } from "../types.js"; import type { LlmClient } from "./types.js"; export declare class PanelGrader implements Grader { metadata: GraderMetadata; private client; constructor(client: LlmClient); /** * `panel "did it use bicep"`, else `panel gpt-5,claude@high` — the rubric is * the clearest discriminator; the judge lineup is the fallback since `models` * is the only required config. * * A judge's `reasoning_effort` is part of that lineup: two panels differing * only by effort run genuinely different judges, so omitting it would derive * one name for both and drop them back to positional `#n` labels. Uses the * same `model@effort` spelling as {@link judgeLabels}. */ defaultName(config: Record): string; grade(input: GraderInput): Promise; } //# sourceMappingURL=panel-grader.d.ts.map