import { ConfigService } from "@nestjs/config"; import { TokenUsageRecorderInterface } from "../../../common/tokens"; import { BaseConfigInterface } from "../../../config/interfaces"; import { TokenUsageService } from "../../../foundations/tokenusage/services/tokenusage.service"; import { ModelService } from "../../llm/services/model.service"; /** * Opt-in cost attribution for an embedding call. Without both * `relationshipId` and `relationshipType` no usage record is written — the * package stays domain-agnostic and the caller decides what the usage is * billed against. */ export interface EmbedderAttribution { relationshipId: string; relationshipType: string; tokenUsageType?: string; } export declare class EmbedderService { private readonly modelService; private readonly config; private readonly tokenUsageService?; private readonly tokenUsageRecorder?; private readonly logger; constructor(modelService: ModelService, config: ConfigService, tokenUsageService?: TokenUsageService, tokenUsageRecorder?: TokenUsageRecorderInterface); /** * Runs ONE embedding invocation, with a single failover retry on a transient * provider failure (429/5xx/socket-level — classified by exactly the rule * `LLMService` uses). * * ONE retry only, and only when there is somewhere to go: the failed * connection is put into cooldown and the embedder is rebuilt, so * `getEmbedder()` resolves the NEXT healthy link of the chain. A chain of one * (no DB-backed embedder connection — just the `.env` block) has nowhere to * fail over to, so the rejection propagates immediately, exactly as it does * today. Everything beyond that single re-resolution belongs to the BullMQ * job retry that owns these calls (spec § 2, "Other modalities"). */ private withCandidateFailover; /** * Puts the embedder connection currently in service into its cooldown window * and returns it — or undefined when there is nothing to fail over to: no * candidate-aware `ModelService`, a chain of one (`.env` only), or a registry * that threw. Undefined means "behave exactly as before failover existed". */ private failCurrentEmbedderCandidate; vectoriseText(params: { text: string; attribution?: EmbedderAttribution; }): Promise; vectoriseTextBatch(texts: string[], attribution?: EmbedderAttribution): Promise; /** * Embeds ONE text, slicing it first when it exceeds the provider's per-input * cap: every slice goes through the same (rate-limited) embedder and the * slice vectors are mean-pooled, so the caller always gets exactly one vector * for one text and no oversized request ever reaches the provider. */ private embedGuarded; /** * Sequential on purpose: it mirrors RateLimitedEmbedder.embedDocuments, which * walks its sub-batches one at a time (rate-limited-embedder.ts:77) so a * single large input never raids the shared token bucket in one burst. */ private embedSlices; private warnOversizedInput; /** * Slices `text` into pieces that each measure at most MAX_EMBED_INPUT_TOKENS * with the SAME tokenizer used for billing. A text that already fits is * returned as-is (single element), and the slices always concatenate back to * the original text — nothing is dropped. */ private splitByTokenBudget; /** Longest measured-within-budget prefix of `text` (never empty). */ private takeBudgetedHead; /** * Pulls the cut back to the nearest whitespace within the last fifth of the * window, so slices break between words. A stretch with no whitespace at all * (a long unbroken blob) is cut mid-word rather than left oversized. */ private snapToWordBoundary; /** * Element-wise mean of the slice vectors, L2-normalized: the standard * long-input reduction. Normalising keeps the pooled vector comparable with * single-call vectors under cosine similarity, which is what every Neo4j * vector index in this package searches with. */ private meanPool; /** Lazily-built tokenizer for the configured embedder model; null after a failed init. */ private tokenizer; /** * Embedding providers return vectors only — no usage figures — so tokens are * counted LOCALLY with the embedder model's own tokenizer (js-tiktoken). * For OpenAI-family models (text-embedding-3-large on Azure — a360ai's * embedder) this count is exactly what the provider bills. Models unknown to * tiktoken fall back to the rate limiter's chars-per-token heuristic. * One record per service invocation (a batch is one record), attribution * opt-in exactly like LLMService.persistUsage, floor-exempt (applyMinimum * false): sub-cent embedding calls must never floor to 0.1 credits. */ private persistUsage; /** * Records what a FAILED embedding call already burned — but ONLY when the * rejection is evidence the provider actually did work. * * A failure is not automatically a free call: a batch that dies on the * provider's LAST internal sub-batch has still been charged for every chunk * the provider served, and `RateLimitedEmbedder` surfaces that as a single * rejection. Recording nothing there understates real spend. * * BUT the counts here are computed LOCALLY with tiktoken, never read off a * provider response, so they exist even when the provider was never reached. * Billing every rejection would therefore charge the customer for OUR * outages, inverting the very rule this path claims parity with: * `LLMService.persistUsageOnFailure` bills provider-REPORTED tokens, which * are 0 when nothing was served. The damage would be real — * `RateLimitedEmbedder` rejects with `EmbedderBucketStarvedError` before any * HTTP request when our own token bucket cannot grant within `maxWaitMs`, and * these calls sit inside BullMQ jobs with `attempts: 3`, so one transient * local failure could bill the same batch up to four times. * * So the local count is only trusted when {@link providerServedWork} says the * rejection carries positive evidence of provider-side work. See that method * for the signal and for the residual over-billing it knowingly accepts. * * ZERO-TOKEN RULE, as everywhere else on this path: an operation that burned * nothing records NOTHING — a 0-token row would assert a call that never * happened. * * Never throws: it sits in a catch block and must not mask the original error. */ private persistUsageOnFailure; /** * Does this rejection carry positive evidence that the provider did billable * work? Nothing is billed unless it does. * * THE SIGNAL: a SERVER-side HTTP status (5xx) reported by the provider. That * is the only class of rejection in which the provider accepted the request * and failed during or after processing it — the case the failure-billing * exists for (a multi-sub-batch call whose last sub-batch 500s, after earlier * sub-batches were served and charged). * * Everything else records NOTHING, because in every other case the provider * demonstrably served none of this input: * - NO status at all — raised locally before or without any HTTP exchange: * `EmbedderBucketStarvedError` (our own bucket refusing to grant), DNS * failures, connection refused, aborts and client-side timeouts. This is * the "our outage must not charge the customer" case. * - A 4xx — the provider REFUSED the request rather than serving it: 401/403 * auth, 400/422 malformed, 413 too large, 429 rate-limited (including the * retry-exhausted path). Providers do not charge for a request they * rejected, so neither do we. * * RESIDUAL OVER-BILLING, KNOWINGLY ACCEPTED: on a 5xx this bills the WHOLE * submitted input, including sub-batches the provider never got to. * `RateLimitedEmbedder.embedDocuments` splits the input into sequential * sub-batches and is all-or-nothing to its caller — it returns vectors or it * throws, never a partial result — so this service cannot know how many * sub-batches were served. Over-billing is bounded by the batch size and only * on provider-side faults; the alternative (recording nothing) silently * under-bills every genuinely-served chunk. Narrowing this needs * `RateLimitedEmbedder` to report served-sub-batch counts, which would change * its published surface. */ private providerServedWork; /** * Token count for a SINGLE text, through the very tokenizer billing uses — so * the oversize guard slices by the same measure the provider charges by. */ private countText; private countTokens; } //# sourceMappingURL=embedder.service.d.ts.map