/** * Provider-truth-driven request repair for OpenAI-compatible endpoints. * * Model families disagree about which chat-completions parameters they accept: * OpenAI's reasoning models (o-series, gpt-5 except gpt-5-chat) reject * `max_tokens`, `frequency_penalty` and `presence_penalty` outright, and reject * any non-default `temperature` / `top_p`. Encoding that taxonomy here would * duplicate knowledge we do not own and would rot on every model release — and * it cannot be complete anyway, because an Azure deployment name is chosen by * whoever created it ("prod-llm" tells us nothing about the model behind it). * * So we let the provider tell us. A 400 names the offending parameter exactly: * * { error: { code: "unsupported_parameter", param: "max_tokens", * message: "... Use 'max_completion_tokens' instead." } } * { error: { code: "unsupported_value", param: "temperature", * message: "... Only the default (1) value is supported." } } * * This middleware removes the named parameter, retries once per rejection, and * REMEMBERS the verdict per model for the rest of the process, so only the first * call to a given deployment pays the extra round-trip. * * The two codes are remembered at DIFFERENT scopes, because they say different * things. `unsupported_parameter` means the deployment does not know the parameter * at all, so it is dropped from every later request. `unsupported_value` means only * the value sent was refused — the parameter itself is fine. Remembering the latter * as a parameter-level drop is actively harmful: one call sending * `reasoning_effort: "none"` to a deployment that only knows low/medium/high would * otherwise strip `reasoning_effort` from EVERY later request, silently discarding * the `"low"` latency win the triage/retrieval/citation-audit nodes depend on. So a * value rejection is memoised against the (parameter, value) pair, and a later * request carrying a different value for the same parameter is still sent. * * A rejected parameter is not always a top-level key: the Responses API nests * `reasoning.effort` inside a `reasoning` object, and the provider reports the * dotted path (`param: "reasoning.effort"`) rather than the leaf name. Every * body access below goes through the dotted-path helpers, of which a flat * name is simply the one-segment case. And some rejected VALUES are not * simply unsupported but renamed by generation — see `VALUE_SUCCESSORS`. */ /** Test seam — clears everything learned so far. */ export declare function resetLearnedUnsupportedParams(): void; /** * Wraps `fetch` so a rejected chat-completions parameter repairs itself. * * @param modelKey - Identifies the deployment the verdicts belong to. Two * deployments of the same model name on different endpoints get separate * entries, so one provider's rejection never suppresses a parameter another * provider accepts. * @param inner - The fetch to wrap. Composes with other middleware — passing * `openRouterEscalatingFetch(...)` keeps its provider pinning intact. */ export declare function unsupportedParamFetch(modelKey: string, inner?: typeof fetch): typeof fetch; //# sourceMappingURL=unsupported-param-fetch.d.ts.map