/** * Response repair for providers that deliver the answer in `reasoning`. * * A reasoning-capable model served over an OpenAI-compatible endpoint splits its * output into two fields: the hidden trace in `message.reasoning`, the answer in * `message.content`. That split is done by the SERVING STACK, not the model — a * vLLM reasoning parser looking for the boundary token the model was trained to * emit. When the parser is misconfigured for the checkpoint it is hosting, the * boundary is never found, EVERYTHING is classified as reasoning, and `content` * comes back null. * * Captured verbatim from OpenRouter → Parasail → minimax/minimax-m3 on * 2026-08-21, answering a `response_format: json_schema` request: * * { "finish_reason": "stop", * "message": { "content": null, * "reasoning": "{ \"where\": \"A rain-soaked harbour...\" }" } } * * The payload was complete and schema-conforming. It was simply in the wrong * field, so LangChain — which reads `message.content` and DISCARDS `reasoning` * entirely (it survives in neither `additional_kwargs` nor `response_metadata`) * — surfaced an empty string, and the whole salvage ladder in `LLMService.call` * failed with "No content" on a response that had been generated and billed. * That is why this repair lives in the fetch middleware rather than as another * rung of that ladder: by the time the ladder runs, the evidence is gone. * * The rule is deliberately narrow. Reasoning is only promoted when `content` is * absent or blank AND the reasoning text is itself a JSON document. Prose * reasoning is left exactly where it is: promoting a chain of thought into the * answer would leak deliberation into narration, which is far worse than the * failure being repaired. A response that already carries content or tool calls * is never touched, so a healthy provider is unaffected. */ /** * A `fetch` middleware that repairs reasoning-only responses. Wraps an inner * fetch (the OpenRouter pin, the unsupported-parameter repair) rather than * replacing it, and forwards `input`/`init` untouched — it only ever inspects * the response. * * THE BODY IS READ EXACTLY ONCE, AND NEVER THROUGH `clone()`. * * `clone()` tees the body stream: the read that follows pulls through BOTH * branches, so the original is no longer pristine the moment the clone is read. * That is fine until the read FAILS. The OpenAI SDK bounds the whole fetch call * — this middleware included — with `AI_REQUEST_TIMEOUT_MS`, and a generation * that outruns the budget aborts the stream while we are mid-read. Returning * "the original, untouched" then handed the caller a DISTURBED response, whose * `.text()` threw undici's `Body is unusable: Body has already been read` and * buried the real AbortError. Observed live 2026-08-31 on game creation, where * an 88s turn crossed the 120s default (the payload was in `content` all along * — this middleware had nothing to repair and still broke the call). * * So: one read, and a read failure PROPAGATES. The caller owns the timeout and * must see its own error. Everything downstream of the read is handed on as a * rebuilt Response — a repair is no longer the only path that constructs one, * because the single read has consumed the original either way. */ export declare function reasoningContentFetch(innerFetch?: typeof fetch): typeof fetch; //# sourceMappingURL=reasoning-content-fetch.d.ts.map