/** * Cross-environment executor adapter for WebLLM Chat Completions compatible endpoints. * * @module @nhtio/adk/batteries/llm/webllm_chat_completions/adapter * * @remarks * Cross-environment LLM adapter for the WebLLM Chat Completions wire shape. Chat Completions was * chosen as the ADK's reference adapter because it is the de-facto interchange format for the * majority of OpenAI-compatible gateways (vLLM, Together, Groq, Fireworks, Ollama, Azure OpenAI, * OpenRouter, DeepSeek, Mistral La Plateforme, and most self-hosted deployments). Its tool-call * synthetic-history shape (`role: 'assistant', tool_calls: [...]` followed by `role: 'tool'` with * `tool_call_id`) is the lowest-common-denominator that every conformant gateway accepts. * * The adapter is built around three pluggable layers: * * 1. **Translation helpers** — the thirteen swappable functions exported from `./helpers` turn * ADK primitives ({@link @nhtio/adk!Tokenizable}, {@link @nhtio/adk!Memory}, {@link @nhtio/adk!Message}, {@link @nhtio/adk!Thought}, * {@link @nhtio/adk!ToolCall}, {@link @nhtio/adk!Tool}, {@link @nhtio/adk!ArtifactTool}, {@link @nhtio/adk!SpooledArtifact}) into Chat * Completions wire shapes. Consumers override individual helpers via `options.helpers.*` to * customise envelope formats, bucket ordering, thought surfacing, or JSON Schema generation * without forking the adapter. * 2. **Three-layer options merging** — constructor baseline, per-`executor()` overrides, and * per-iteration `ctx.stash.webLLMChatCompletions` overrides combine with key-by-key * precedence for `helpers` and wholesale replacement for everything else. * The merged shape is re-validated on every iteration so a malformed stash override * fails loud, not silently. * 3. **WebLLM engine invocation** — accepts a preloaded `engine` or lazy `createEngine` factory. * The resolved request body is passed directly to WebLLM's OpenAI-compatible chat API. * * Per-iteration flow (steps 1–9 of the plan): * 1. Merge constructor / executor / stash options and re-validate. * 2. Resolve helpers, falling back to bundled `default*` for each unset field. * 3. Forge artifact-query tools by walking `ctx.turnToolCalls`, collecting unique * `SpooledArtifact` constructors, calling `.forgeTools(ctx)` on each, and merging the * results with `ctx.tools`. * 4. Pre-render every persisted tool-call result into the prompt-ready string the timeline will * use, cached by `tc.id`. * 5. When `tokenEncoding !== null`, sum the token weight of every persisted bucket and throw * {@link @nhtio/adk/batteries!E_WEBLLM_CHAT_COMPLETIONS_CONTEXT_OVERFLOW} when the total exceeds `contextWindow`. * 6. Build the request body via `buildChatCompletionsHistory`; carry vendor-opaque reasoning * blocks through the `_adk_reasoning_payloads` side-channel. * 7. Resolve or lazily create a WebLLM engine and call `engine.chat.completions.create(body)`. * 8. Streaming path: consume WebLLM's async chunk iterable; surface deltas through * `helpers.reportMessage` / `reportThought` / `reportToolCall`; assemble tool-call deltas via * the accumulator; persist `Message` / `Thought` / `ToolCall` records on stream end. * 9. Non-streaming path: consume the returned Chat Completion object; same persistence + * tool-execution loop. */ import type { DispatchExecutorFn } from "../../../dispatch_runner"; import type { WebLLMChatCompletionsAdapterOptions, WebLLMEngine } from "./types"; /** * Opinionated cross-environment LLM adapter for the WebLLM Chat Completions wire shape. * * @remarks * Construction validates options eagerly via {@link @nhtio/adk/batteries!validateOptions} and throws * {@link @nhtio/adk/batteries!E_INVALID_WEBLLM_CHAT_COMPLETIONS_OPTIONS} on failure — config bugs fail loud, not at * dispatch time. The returned instance is reusable: call {@link WebLLMChatCompletionsAdapter.executor} * once per `DispatchRunner` configuration to obtain an {@link @nhtio/adk!DispatchExecutorFn} bound to the * baseline plus optional executor-scope overrides. * * Per-iteration overrides live on the active {@link @nhtio/adk!DispatchContext}'s * `stash.webLLMChatCompletions` slot and take highest precedence — they merge into the * executor-scope shape on every iteration. `helpers` merge key-by-key across all three layers; * every other field is replaced wholesale at the highest layer that * sets it. */ export declare class WebLLMChatCompletionsAdapter { #private; /** * Customary key for per-iteration overrides on `ctx.stash`. The adapter reads * `ctx.stash.get(WebLLMChatCompletionsAdapter.STASH_KEY, {})` at the start of every * iteration and merges the value into the resolved options shape. */ static readonly STASH_KEY: "webLLMChatCompletions"; /** * Whether the runtime can host a WebLLM engine — i.e. WebGPU (`navigator.gpu`) is present. * * @returns `true` when WebGPU is available in the current environment. */ static isAvailable(): boolean; /** * @param options - Constructor-baseline options. Re-validated on every iteration after * per-dispatch and per-iteration overrides are layered in. * @throws {@link @nhtio/adk/batteries!E_INVALID_WEBLLM_CHAT_COMPLETIONS_OPTIONS} when `options` does not satisfy * {@link @nhtio/adk/batteries!webLLMChatCompletionsOptionsSchema}. */ constructor(options: unknown); /** * Eagerly loads (and caches) the engine so the first dispatch does not pay the model-load cost. * * @param overrides - Optional option overrides layered over the constructor baseline. * @returns The resolved {@link WebLLMEngine}. */ preload(overrides?: Partial): Promise; /** Drops the cached engine and any in-flight load so the next dispatch re-resolves it. */ reset(): void; /** * Instance-level availability check, honouring an injected * {@link WebLLMChatCompletionsAdapterOptions.isWebGPUAvailable} override. * * @returns `true` when a WebLLM engine can run in the current environment. */ isAvailable(): boolean; /** * Returns an {@link @nhtio/adk!DispatchExecutorFn} bound to this adapter's baseline plus optional * executor-scope overrides. The returned function is reusable across iterations — every * iteration re-merges with `ctx.stash[STASH_KEY]` and re-validates the result. * * @param overrides - Optional executor-scope overrides. Higher precedence than the baseline, * lower precedence than `ctx.stash[STASH_KEY]`. * @returns An {@link @nhtio/adk!DispatchExecutorFn} suitable for `DispatchRunner`. */ executor(overrides?: Partial): DispatchExecutorFn; /** * Returns `true` when `value` is an {@link WebLLMChatCompletionsAdapter} instance. * * @param value - The value to test. * @returns `true` when `value` is an `WebLLMChatCompletionsAdapter` instance. */ static isWebLLMChatCompletionsAdapter(value: unknown): value is WebLLMChatCompletionsAdapter; }