/** * transformers.js (ONNX, dual-environment) Caption (image-to-text) specialist adapter battery. * * @module @nhtio/adk/batteries/specialists/caption/transformers_js/adapter * * @remarks * Image-captioning battery backed by transformers.js's `image-to-text` pipeline (the documented * reference model is `Xenova/vit-gpt2-image-captioning`). **Environment-neutral** — runs in Node (via * `onnxruntime-node`) and the browser (via `onnxruntime-web` / WebGPU), auto-selected by the package; * there is no WebGPU requirement. * * Same shape as the transformers.js Embeddings battery: eager constructor validation, a required * `model` (no default), a lazily-imported/single-flight peer, `preload()` / `reset()` / `dispose()`, * and the shared lifecycle hooks. * * **Image input:** {@link @nhtio/adk/batteries/specialists/_shared!toBytes} normalizes the accepted * {@link @nhtio/adk/batteries/specialists/_shared!SpecialistImageInput} forms (bytes / bytes+mime / * media-like) to plain bytes + an optional MIME type. Those bytes become a plain `Blob` — the * `image-to-text` pipeline's `ImageInput` union directly accepts `Blob` (verified against the * installed `@huggingface/transformers` 4.2.0 type declarations, alongside `string | RawImage | URL | * HTMLCanvasElement | OffscreenCanvas`), so building a `Blob` needs no peer import at all — `Blob` is a * cross-env global (Node 18+ and every browser). This keeps the adapter's hot path peer-free until the * pipeline itself is resolved, and means fake-pipeline unit tests never load `@huggingface/transformers`. * * `@huggingface/transformers` is an optional peer dependency, imported lazily (only inside * {@link makeDefaultCreatePipeline}, i.e. only when no `pipeline`/`createPipeline` override is supplied). */ import type { SpecialistImageInput } from "../../_shared/index"; import type { DescribeOptions, DescribeResult } from "./types"; /** * Caption (image-to-text) adapter for transformers.js's `image-to-text` pipeline. * * @remarks * Reusable: construct once, call {@link TransformersJsCaptionAdapter.describe} as many times as * needed. The pipeline is resolved lazily on first use (or via {@link preload}) and cached with * single-flight semantics so concurrent calls share one load. */ export declare class TransformersJsCaptionAdapter { #private; /** * Whether this battery is available. transformers.js is environment-neutral (Node + browser), so * this is `true` whenever the runtime can import the peer — there is no WebGPU requirement. */ static isAvailable(): boolean; /** * @param options - Constructor options. Validated eagerly. * @throws {@link @nhtio/adk/batteries!E_INVALID_TRANSFORMERS_JS_CAPTION_OPTIONS} when invalid. */ constructor(options: unknown); /** Instance availability probe (honours an injected `isAvailable`). */ isAvailable(): boolean; /** Eagerly loads (and caches) the pipeline so the first `describe` call is fast. Idempotent. */ preload(): Promise; /** Drops the cached pipeline and in-flight load so the next call reloads. */ reset(): void; /** * Release the loaded model's ONNX sessions + GPU/wasm buffers, then drop the cached pipeline. * * @remarks * `reset()` only nulls the JS reference; the native ONNX Runtime sessions and WebGPU/wasm device memory * stay alive until GC. `ImageToTextPipeline` extends `Pipeline`, which exposes `dispose()` — this * awaits it so the memory is reclaimed between loads, swallows a disposal error (teardown must not * throw), and finishes with `reset()`. Idempotent. */ dispose(): Promise; /** * Generates a caption for an image. * * @param input - The image in any {@link @nhtio/adk/batteries/specialists/_shared!SpecialistImageInput} * form (bytes / bytes+mime / media-like). * @param opts - Per-call options (`maxNewTokens`, forwarded as `max_new_tokens`; omitted when unset). * @returns The normalized `{ text }` caption result. * @throws {@link @nhtio/adk/batteries!E_TRANSFORMERS_JS_CAPTION_ENGINE_ERROR} when the call fails or * the pipeline returns no usable caption text (an empty/missing caption is treated as an engine * failure, not a valid empty result — a captioner that produces nothing didn't do its job). */ describe(input: SpecialistImageInput, opts?: DescribeOptions): Promise; }