import { ConfigService } from "@nestjs/config"; import { TokenUsageRecorderInterface } from "../../../common/tokens"; import { BaseConfigInterface } from "../../../config/interfaces"; import { TokenUsageService } from "../../../foundations/tokenusage/services/tokenusage.service"; import { type TranscodeOptions } from "./audio/ffmpeg-transcode"; import { LLMCallDumper } from "./llm-call-dumper.service"; import { ModelService } from "./model.service"; /** * Parameters for AudioLLMService.call. Engine-agnostic: the service decides * how to use `prompt` based on whether audio.directUrl is configured. */ export interface AudioCallParams { audioPath: string; /** * Free-form prompt. Caller sizes/shapes it for the configured engine: * - audio.directUrl unset → used verbatim as the chat-LLM system prompt * - audio.directUrl set → passed as the /audio/transcriptions `prompt` * parameter; the upstream API typically truncates at ~224 tokens and * treats it as vocabulary biasing (no instruction-following). */ prompt: string; temperature?: number; /** * Optional in-pass audio cleanup applied during the universal transcode * (high-pass, silence trim). Omit for the plain resample. See TranscodeOptions. */ transcode?: TranscodeOptions; /** * Cost-attribution category written to the usage record. Free-form so each * application owns its own vocabulary; defaults to "transcription". */ tokenUsageType?: string; /** * Opt-in cost attribution, exactly like `LLMCallParams`: BOTH of these must * be present or no usage record is written at all. The package stays * domain-agnostic — the caller decides what the transcription is billed * against (narr8: the Session the recording belongs to). */ relationshipId?: string; relationshipType?: string; } export interface TranscriptionResult { text: string; /** * Tokens the engine reported. {0, 0} when it reports none — which is common on * the direct `/audio/transcriptions` path, where a duration-priced model * (whisper-large-v3-turbo) charges by the second and reports no tokens at all. */ tokenUsage: { input: number; output: number; }; /** Duration of the audio actually sent (from ffmpeg's Duration line). */ audioSeconds: number; /** * The price the PROVIDER invoiced for this request, when it reports one * (OpenRouter's `/audio/transcriptions` returns `usage.cost`). Absent * otherwise. A measured figure rather than an estimate, so `persistUsage` * bills it in preference to both the token and the duration clock. * * NEVER ASSUME IT IS THERE. Reporting a cost is an OpenRouter courtesy, not * part of the OpenAI-style contract this path speaks: an endpoint on Azure, * Vertex, OpenAI itself or a self-hosted Whisper reports tokens or a duration * at best, and often nothing at all. Every such call falls through to the * manual clocks in {@link AudioLLMService.persistUsage}, which is why those * clocks remain load-bearing rather than legacy. * * CURRENCY IS THE PROVIDER'S. OpenRouter reports USD; `credits.creditCost` * is nominally euro-denominated. The gap is the same one the configured * `*CostPer1MTokens` rates already carry (they are copied off provider * pricing pages in USD too), so nothing is converted here — a deployment * that needs euro-exact accounting must price its credits accordingly. */ providerCost?: number; /** * Audio duration the PROVIDER measured and billed, when it reports one * (OpenAI-style endpoints return `usage: {type: "duration", seconds}`). * Preferred over the locally measured `audioSeconds` by the duration clock, * because it is the number the invoice was calculated from. Absent when the * endpoint reports no usage at all. */ providerSeconds?: number; } /** * Parameters for {@link AudioLLMService.diarize}. No `prompt` and no * `temperature`: a diarizing STT endpoint takes neither, and no `transcode` * because the caller already produced the file that goes on the wire. */ export interface DiarizeCallParams { audioPath: string; /** * Cost-attribution category written to the usage record. Free-form so each * application owns its own vocabulary; defaults to "transcription". */ tokenUsageType?: string; /** * Opt-in cost attribution, exactly like {@link AudioCallParams}: BOTH of * these must be present or no usage record is written at all. */ relationshipId?: string; relationshipType?: string; } /** One diarized speech segment. Times are SECONDS from the start of the file sent. */ export interface DiarizedSegment { start: number; end: number; text: string; /** Provider speaker index within THIS request; -1 when the provider sent none. */ speaker: number; } export interface DiarizationResult extends TranscriptionResult { segments: DiarizedSegment[]; language?: string; } /** * Audio transcription facade. One env var (AUDIO_DIRECT_URL) flips dispatch * between two unrelated backends. The ffmpeg transcode to 16 kHz mono mp3 runs * for **both** backends (universal pre-normalisation): * * - The recorder writes hand-rolled OGG/Opus framing; OpenAI's audio chat * models (gpt-audio*, gpt-4o-audio-preview) reject OGG outright in * `input_audio` content parts, and even compliant OGG occasionally trips * up STT endpoints. Transcoding once at the boundary makes every backend * accept the same bytes. * * - chat-LLM (AUDIO_DIRECT_URL unset): the chat model configured via * ModelService.getAudioLLM receives an `input_audio` content part with the * transcoded MP3 and a system prompt. Plain `invoke` — no structured * output (OpenAI's gpt-audio* family rejects `response_format: json_schema`, * and the response.content is already plain text per the system prompt). * Provider routing (openrouter / vertex / azure / requesty / llamacpp) * lives in ModelService — this service does not branch on it. * * - direct (AUDIO_DIRECT_URL set): the transcoded MP3 is POSTed as * multipart to AUDIO_DIRECT_URL using AUDIO_API_KEY as Bearer auth. Any * OpenAI-style transcription endpoint works (api.openai.com, Groq, * self-hosted Whisper, ...). No provider whitelist. * * No retry layer — BullMQ's job-level retry handles transient failures. */ export declare class AudioLLMService { private readonly modelService; private readonly config; private readonly dumper; private readonly tokenUsageService?; private readonly tokenUsageRecorder?; private readonly logger; constructor(modelService: ModelService, config: ConfigService, dumper: LLMCallDumper, tokenUsageService?: TokenUsageService, tokenUsageRecorder?: TokenUsageRecorderInterface); call(params: AudioCallParams): Promise; /** * Speaker-diarized transcription of ONE audio file through the DIARIZE tier * (`ai.audioDiarize`, env `AUDIO_*_DIARIZE`). Direct JSON endpoint only — * there is no chat-model diarization. No transcode: the caller already * produced a 16 kHz mono mp3 sized under the provider's payload cap, and the * universal transcode would re-encode to WAV and blow through it. */ diarize(params: DiarizeCallParams): Promise; /** * True once the "direct STT reports no tokens" warning has been emitted by * this instance. The audio service is a singleton, so this throttles the * warning to once per process instead of once per utterance. */ private directPathUnbilledWarned; /** * Makes an unbillable engine LOUD instead of silent. * * Some direct `/audio/transcriptions` engines report no usage of their own — * no tokens, no cost. The zero-token rule in {@link persistUsage} then writes * no usage record, so pointing an application at such an engine silently * turns transcription into completely unbilled work. Inventing a token count * would be a lie, so the honest fix is a warning that names the consequence * and the one setting that fixes it. * * NOT every direct engine is in that position: OpenRouter returns a full * `usage` block (tokens AND its own `cost`), which {@link parseDirectUsage} * now reads, so those calls bill exactly and must not be warned about. The * warning fires only when the engine reported NOTHING billable and no * per-minute rate is configured either. * * Emitted once per process — never in the hot per-utterance path beyond a * boolean check. */ private warnIfDirectPathUnbilled; /** * Records transcription usage. Attribution is opt-in — same contract as * `LLMService.persistUsage`: no relationship, no record, so the package stays * domain-agnostic. Cost comes from the `audio` config block rather than * `computeCost()`, which only knows the text tiers (`configForWeight` never * looks at `ai.audio`), hence `costOverride`. * * THREE PRICING CLOCKS, in priority order: * * 0. PROVIDER COST, when the engine invoices the call itself. OpenRouter's * `/audio/transcriptions` returns `usage.cost` — what it actually charged, * not what we estimate it charged — so it outranks both estimates below. * It is also the ONLY correct clock for a duration-priced model such as * whisper-large-v3-turbo, which reports a real cost alongside ZERO tokens: * the token clock prices that call at nothing and the zero-token rule then * drops the record entirely, which is exactly how a full session of 665 * utterances once billed nobody. * 1. TOKENS, when the engine reports counts but no cost. The chat path * (AUDIO_DIRECT_URL unset) counts the audio itself as input tokens, so * this is the truer measure wherever a provider cost is unavailable. * 2. DURATION, when the engine reports neither and `ai.audio.costPerMinute` * is set. A pure estimate, and the last line of defence: without it an * engine that reports nothing records NOTHING, silently turning * transcription into unbilled work. * * They are never combined — billing one call by two clocks would charge it * twice. A call that reports no cost, no tokens and no priced duration still * records nothing, because `recordTokenUsage` floors every row at * `minCreditsPerRecord` and such a row would invent a charge for work nobody * can measure. That case is what {@link warnIfDirectPathUnbilled} shouts * about. * * FLOOR-EXEMPT (`applyMinimum: false`), same rationale as EmbedderService: * transcription is per-utterance, not per-request. A session averages a few * hundred segments and can exceed a thousand, each truly worth a fraction of * a credit, so applying the per-record floor to every one of them would bill * a large multiple of the real cost. The floor exists to stop sub-cent REAL * usage rounding to nothing on a handful of records, not to price a thousand * of them. * * Writes through the application-provided `TOKEN_USAGE_RECORDER` when one is * bound, falling back to the module-local `TokenUsageService` otherwise (see * the token's docblock for why package code must use this seam). * * Never throws: a persistence failure logs a warning and the transcription * result stands. */ private persistUsage; private callChat; /** * Dump the full upstream error for diagnostics. LangChain wraps the openai * SDK's APIError which carries the upstream body in `.error` (or * `.response.data` depending on transport); without printing it we only see * the generic "400 Provider returned error" wrapper from OpenRouter, hiding * the real OpenAI rejection underneath. */ private dumpUpstreamError; private callDirect; } //# sourceMappingURL=audio.llm.service.d.ts.map