---
name: llm
load-when: wiring @adia-ai/llm — chat/streamChat signatures, StreamChunk handling, the proxy security model
load-size: ~1.2k tokens
required-for: [llm-wiring]
---

# `@adia-ai/llm` — the app-side LLM client

Chat / streaming / AI features for a consumer app. Snapshot verified against **@adia-ai/llm 0.8.53 (lockstep — see `packages/llm/core/package.json`)**. LLM client surfaces drift — confirm precise field names against the installed version and lean on the **smart-proxy contract** (stable) rather than memorized fields. (Version pin re-verified against `packages/llm/core/package.json` 2026-08-26 — still 0.8.53; re-verify on a MINOR cut.)

## Import & core API

```js
import { chat, streamChat, createClient } from '@adia-ai/llm';
import { MODELS, DEFAULT_MODEL } from '@adia-ai/llm/models';
```

- `chat(opts) → Promise<{ text, usage, stopReason }>` — non-streaming.
- `streamChat(opts) → AsyncGenerator<StreamChunk>` — streaming; `for await` the chunks.
- `createClient(defaults) → { chat, stream }` — a client with baked-in defaults.

`ChatOpts` (the fields that matter): `model`, `messages: [{role, content}]`, `provider?` (auto-detected from the model name), `system?`, `maxTokens?`, `temperature?`, `signal?` (AbortSignal), `proxyUrl?`, plus Anthropic extras `thinking?` / `cache?`.

`StreamChunk` is a tagged union — branch on `chunk.type`, handle every branch:

```js
for await (const chunk of streamChat(opts)) {
  if      (chunk.type === 'text')     append(chunk.text);        // also .snapshot (full text so far)
  else if (chunk.type === 'thinking') showThinking(chunk.text);  // Anthropic extended thinking
  else if (chunk.type === 'done')     finalize(chunk.usage, chunk.stopReason);
  else if (chunk.type === 'error')    fail(chunk.error);         // render it — never drop
}
```

## Providers

Anthropic, OpenAI, and Gemini, **auto-detected from the model name** (`claude*` → anthropic, `gpt*`/`o1*` → openai, `gemini*` → gemini — the provider key is `'gemini'`, and an unknown name throws). Override with `provider`. Keys are read server-side from `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GEMINI_API_KEY`.

## The proxy / security model — read this first

**Never ship a provider API key to the browser in production.** Two proxy shapes, auto-selected from the `proxyUrl` (a URL matching `/api/llm/<provider>/…` ⇒ passthrough; anything else ⇒ smart):

- **Smart proxy (production):** point `proxyUrl` at your own same-origin endpoint (e.g. `/api/chat`). The browser sends a provider-neutral body (`{provider, model, messages, system?, maxTokens?, temperature?}`); **your server** holds the real key, reformats per provider, and pipes the SSE bytes back. The browser never sees a key.

  ```js
  for await (const c of streamChat({ proxyUrl: '/api/chat', model, messages })) { … }
  ```

  Reference implementation (monorepo only — `server.js` is not in the npm-published files): `packages/llm/core/server.js` (`POST /api/chat`).

- **Passthrough proxy (Vite dev ONLY):** a `proxyUrl` matching `/api/llm/<provider>/…` (e.g. a Vite dev proxy). The browser sends the real upstream body **and the real key in headers** — anyone with DevTools can read it. **Never deploy this shape.**

When authoring: default to the smart-proxy pattern; use passthrough only behind a Vite dev proxy, and say so explicitly.

## Chat surface wiring

`<chat-shell proxy-url="/api/chat" model="…">` auto-sends on submit via `streamChat` and renders the stream for you; without `proxy-url` (or a dev-only `apiKey` property) it only emits `submit` for you to handle. The full surface — cluster roster, props/events/methods, registration gotchas — lives in [`shell-chat.md`](shell-chat.md); this file owns only the client/proxy contract.

## Not in the package (don't assume)

Tool/function-calling chunks, structured-output/JSON-schema modes, built-in retry/backoff, and client-side token estimation are **not** surfaced by `@adia-ai/llm` — handle them in your own server layer. `createAdapter` from `@adia-ai/llm/bridge` is the A2UI generation pipeline's internal LLM adapter; an app rarely calls it directly.
