<div align="center">

# ⚡ pi-neuralwatt-provider

**Models + energy tracking via [Neuralwatt](https://neuralwatt.com)**

_Kimi, GLM, Qwen, DeepSeek — with real-time ⚡ energy/cost per session for [pi](https://github.com/earendil-works/pi-coding-agent)._

[![pi extension](https://img.shields.io/badge/pi-extension-blueviolet)](https://github.com/earendil-works/pi-coding-agent)
[![license](https://img.shields.io/badge/license-MIT-blue)](./LICENSE)

</div>

## Pi 1.0 compatibility

Tested with Pi **1.0.0** and the **1.0.2 bundled host**. Host-provided packages remain wildcard peers; development uses exact SDK pins.
Run `bun run test:pi` for offline manifest, catalog, lifecycle and streaming checks. To test an installed host, set `PI1_HOST_PACKAGE` to its package directory; add `PI1_HOST_ENTRY=bundle` for its bundled CLI runtime. The probe stubs all network requests and never uses a live provider endpoint.

---

![Energy Reporting Status Widget](assets/screenshot.jpg)

## Features

- **OpenAI-compatible API** - Uses Neuralwatt's `/v1/chat/completions` endpoint
- **Reasoning models** - Support for thinking models with `reasoning_effort` parameter
- **Vision models** - Image input support on Kimi K2.5, K2.6, and Devstral
- **Tool use** - Function calling support, plus opt-in Neuralwatt hosted tools with per-request budgets
- **Classifiers** - Clef Flash via pi core’s classifier registry, separate from chat models
- **Image history** - Separate per-turn and whole-conversation limits; no premature 20-image history eviction
- **Streaming** - Real-time token streaming
- **Fast variants** - Optimized "Fast" versions of popular models for quicker responses
- **Energy reporting** - Displays energy consumption (⚡J/mWh/Wh/kWh) and actual billed cost ($) in a dedicated status widget below the editor, tracked per-session
- **Quota display** - Shows subscription plan, kWh allocation, and credits remaining from your Neuralwatt account, right-aligned in the status widget
- **Configurable display** - Energy and quota can each be shown in the below-editor widget, the built-in status bar, or turned off entirely via a config file

## Available Models

| Model | Context | Vision | Reasoning | Input $/M | Cache Read $/M | Output $/M |
|-------|---------|--------|-----------|-----------|-----------------|------------|
| Clef Flash (classifier; preview access) | 262K | ✅ | ❌ | $0.18 | $0.02 | — |
| DeepSeek V4 Flash | 1.0M | ❌ | ✅ | $0.14 | $0.03 | $0.28 |
| DeepSeek V4 Flash (0731 Canary) | 1.0M | ❌ | ✅ | $0.14 | $0.03 | $0.28 |
| DeepSeek V4 Flash (flex) | 1.0M | ❌ | ✅ | $0.09 | $0.02 | $0.18 |
| DeepSeek V4 Flash (Speed) | 1.0M | ❌ | ✅ | $0.14 | $0.03 | $0.28 |
| DeepSeek V4.1 Flash | 1.0M | ✅ | ✅ | $0.15 | $0.01 | $0.60 |
| DeepSeek V4.1 Flash (flex) | 1.0M | ✅ | ✅ | $0.10 | $0.01 | $0.39 |
| DeepSeek V4.1 Flash (Speed) | 1.0M | ✅ | ✅ | $0.15 | $0.01 | $0.60 |
| Gemma 4 31B | 262K | ✅ | ✅ | $0.14 | $0.01 | $0.42 |
| GLM 5.3 | 1.0M | ❌ | ✅ | $1.45 | $0.14 | $4.50 |
| GLM 5.3 (flex) | 1.0M | ❌ | ✅ | $0.94 | $0.09 | $2.92 |
| GLM-5.3 Flash | 1.0M | ✅ | ✅ | $0.15 | $0.03 | $0.50 |
| GLM-5.3 Flash (flex) | 1.0M | ✅ | ✅ | $0.10 | $0.02 | $0.33 |
| Kimi K2.7 Code | 262K | ✅ | ✅ | $0.95 | $0.10 | $4.00 |
| Kimi K2.7 Code (flex) | 262K | ✅ | ✅ | $0.62 | $0.06 | $2.60 |
| Kimi K2.7 Code Fast | 262K | ✅ | ✅ | $0.95 | $0.10 | $4.00 |
| Kimi K3 | 1.0M | ✅ | ✅ | $3.00 | $0.30 | $15.00 |
| Kimi K3 (flex) | 1.0M | ✅ | ✅ | $1.95 | $0.20 | $9.75 |
| Kimi K3 Fast | 1.0M | ✅ | ❌ | $3.00 | $0.30 | $15.00 |
| MiMo-V2.6-Pro | 1.0M | ✅ | ✅ | $0.87 | $0.04 | $1.74 |
| Neuralwatt Flash | 1.0M | ✅ | ✅ | $0.15 | $0.01 | $0.60 |
| Neuralwatt Flash (flex) | 1.0M | ✅ | ✅ | $0.10 | $0.01 | $0.39 |
| Neuralwatt Large | 1.0M | ✅ | ✅ | $3.00 | $0.30 | $15.00 |
| Neuralwatt Large (flex) | 1.0M | ✅ | ✅ | $1.95 | $0.20 | $9.75 |
| Neuralwatt Small | 262K | ✅ | ✅ | $0.45 | $0.25 | $3.20 |
| Neuralwatt Small (flex) | 262K | ✅ | ✅ | $0.29 | $0.16 | $2.08 |
| Qwen 3.8 27B | 262K | ✅ | ✅ | $0.45 | $0.25 | $3.20 |
| Qwen 3.8 27B (flex) | 262K | ✅ | ✅ | $0.29 | $0.16 | $2.08 |
| Qwen3.6 35B | 262K | ✅ | ✅ | $0.29 | $0.03 | $1.15 |
| Qwen3.6 35B (flex) | 262K | ✅ | ✅ | $0.19 | $0.02 | $0.75 |
| Qwen3.6 35B Fast | 262K | ✅ | ❌ | $0.29 | $0.03 | $1.15 |
| DeepSeek V4-Pro | 1.0M | ❌ | ✅ | $1.00 | $0.10 | $3.00 |
| GLM-5 Long (MCR 1M) | 1.0M | ❌ | ✅ | $1.10 | — | $3.60 |
| GLM-5.1 Fast Long (MCR 1M) | 1.0M | ❌ | ❌ | $1.10 | — | $3.60 |
| Kimi K2.5 Long (MCR 1M) | 1.0M | ✅ | ✅ | $0.52 | — | $2.59 |

## Classifiers and newer API capabilities

Clef Flash is available through **pi core's classifier registry**, not `/model`:

```js
const clef = await models.getModelOfType("classifier", "neuralwatt", "clef-flash");
// In pi codemode: models.classify(clef, { state, questions })
```

Preview access is required. [Capability guide](docs/capabilities.md) covers classifier examples, hosted-tool opt-ins and budgets, sampling/JSON output, image limits, speed/flex lanes, limitations, and live probes. No hosted tool is opted into automatically.

## Authentication

The Neuralwatt API key can be configured in multiple ways (resolved in this order):

1. **`auth.json`** (recommended) — Add to `~/.pi/agent/auth.json`:
   ```json
   { "neuralwatt": { "type": "api_key", "key": "your-api-key" } }
   ```
   The `key` field supports literal values, env var names, and shell commands (prefix with `!`). See [pi's auth file docs](https://github.com/badlogic/pi-mono) for details.
2. **Runtime override** — Use the `--api-key` CLI flag
3. **Environment variable** — Set `NEURALWATT_API_KEY`

Get your API key from [neuralwatt.com](https://neuralwatt.com).

## Installation

### Option 1: Using `pi install` (Recommended)

Install from npm:

```bash
pi install npm:pi-neuralwatt-provider
```

Or install directly from GitHub:

```bash
pi install https://github.com/monotykamary/pi-neuralwatt-provider
```

### Option 2: With npm

Install from npm:

```bash
npm install npm:pi-neuralwatt-provider
```

### Option 3: Manual Clone

Then authenticate and run pi:
```bash
# Recommended: add to auth.json
# See Authentication section below

# Or set as environment variable
export NEURALWATT_API_KEY=your-api-key-here

pi
```

1. Clone this repository:
   ```bash
   git clone git@github.com:monotykamary/pi-neuralwatt-provider.git
   cd pi-neuralwatt-provider
   ```

2. Configure your Neuralwatt API key:
   ```bash
   # Recommended: add to auth.json
   # See Authentication section below

   # Or set as environment variable
   export NEURALWATT_API_KEY=your-api-key-here
   ```

3. Run pi with the extension:
   ```bash
   pi -e /path/to/pi-neuralwatt-provider
   ```

## Environment Variables

| Variable | Required | Description |
|----------|----------|-------------|
| `NEURALWATT_API_KEY` | No | Your Neuralwatt API key (fallback if not in auth.json) |

## Configuration

### Compat Settings

Neuralwatt's API provides compatibility and capability metadata (pricing, reasoning, vision, `developer_role`, `reasoning_effort`, `max_images`, native reasoning levels in `metadata.reasoning`) directly in the `/v1/models` response. The `update-models.js` script reads these and writes them into `models.json`. Only genuinely incorrect API data — or deliberate deviations from the derived thinking-level maps — needs a manual override in `patch.json`.

Currently configured compat settings:

- **`supportsDeveloperRole: false`** — All models. vLLM doesn't support the `developer` role; pi sends system prompts as `system` messages instead.
- **`supportsReasoningEffort: true`** — GLM-5.2, Kimi K3, DeepSeek V4 Flash. Sends the `reasoning_effort` parameter (maps pi's `/reasoning` levels onto each model's native efforts — e.g. GLM-5.2's `high`/`max`/`none` — via the API-derived `thinkingLevelMap`).
- **`requiresReasoningContentOnAssistantMessages: true`** — Kimi K2.6/K2.7 reasoning variants. Pi-ai replays the model's prior-turn `reasoning` field on every assistant message so the model can continue its chain-of-thought across turns. All Neuralwatt reasoning models get this Layer-A replay automatically (the gateway aliases `reasoning` ↔ `reasoning_content`); this flag adds an empty `reasoning_content` scaffold for turns with no thinking block.
- **`chatTemplateKwargs`** — Raw `chat_template_kwargs` merged into every request via pi-ai's `onPayload` hook, mirroring vLLM's request field of the same name. Used to opt reasoning models into **full-history reasoning preservation** (vLLM's Jinja templates otherwise trim older assistant reasoning in alternating chat). The flags are template-level and family-specific — NOT a generic boolean:
  - **Kimi K2.6 / K2.7** → `{ "preserve_thinking": true }` — keeps the full reasoning history across turns (doc-backed; behavioral E2E: 0/6 → 6/6 recall).
  - **GLM-5.2 family** → `{ "clear_thinking": false }` — stops the template clearing older reasoning (functional: 1/4 → 4/4 recall, confirmed family-wide).
  - **GLM-5.1 / Qwen3.x / non-reasoning `-fast`** — no kwarg (their templates expose no flag; they rely on Layer-A replay only).

  These are injected alongside `reasoning_effort` (NOT via `thinkingFormat: "chat-template"`, which would displace the OpenAI `reasoning_effort` path) so thinking-level control and full-history preservation coexist.

### Custom Stream Handler

This extension registers a custom `streamSimple` provider (`api: "neuralwatt"`) that wraps pi-ai's built-in `streamOpenAICompletions`. A per-request fetch wrapper tees the HTTP response body so the OpenAI SDK handles all standard chunk parsing (text, thinking, tool calls, usage) while the extension reads the tee for Neuralwatt's SSE comment lines (`: energy {...}`, `: cost {...}`) that the SDK discards. Per-request wrappers allow concurrent main-agent and helper-model calls to settle in any order without corrupting `globalThis.fetch`.

`X-NW-Conversation-ID` is attached in Pi's `before_provider_headers` lifecycle for agent requests rather than registered as a provider-wide auth header. Raw helper streams therefore cannot accidentally inherit and replace the main agent's Neuralwatt cache lineage.

### Pi Configuration

Add to your pi configuration for automatic loading:

```json
{
  "extensions": [
    "/path/to/pi-neuralwatt-provider"
  ]
}
```

## Usage

Once loaded, select a model with:

```
/model neuralwatt kimi-k2.5
```

Or use `/models` to browse all available Neuralwatt models.

### Reasoning Effort

For reasoning models, control thinking depth:

```
/reasoning high
```

The selectable levels are per-model: each model's `metadata.reasoning` block (native efforts, aliases, `mandatory`/`default_enabled`) is compiled into a `thinkingLevelMap` at sync time and at runtime live-refresh, so `/model` only offers levels the model actually supports (e.g. GLM-5.2 offers `off`/`high`/`max`; a mandatory-reasoning model hides `off`). Selecting a hidden alias level isn't needed — pi clamps up to the native level the alias resolves to.

Full-history reasoning preservation is **on by default** for Kimi K2.6/K2.7 and the GLM-5.2 family (see [Compat Settings](#compat-settings)). Override it per-model via [Model Overrides](#model-overrides).

## Display Configuration

Energy and quota are independently configurable. Create `~/.pi/agent/extensions/neuralwatt.json`:

```json
{
  "energy": "widget",
  "quota": "widget",
  "mcr": "widget",
  "carbon": "widget"
}
```

The file is auto-populated with defaults on first run.

| Key | Values | Default | Description |
|-----|--------|---------|-------------|
| `energy` | `"widget"`, `"statusbar"`, `"off"` | `"widget"` | Energy/cost display mode |
| `quota` | `"widget"`, `"statusbar"`, `"off"` | `"widget"` | Quota display mode |
| `mcr` | `"widget"`, `"statusbar"`, `"off"` | `"widget"` | MCR (context-reuse) display mode |
| `carbon` | `"widget"`, `"statusbar"`, `"off"` | `"widget"` | Carbon (session CO₂ + fleet grid/region badge) display mode |
| `hideOnOtherProvider` | `true`, `false` | `true` | Hide all Neuralwatt display when a non-Neuralwatt model is active |
| `baseUrl` | Any `http(s)` URL | `https://api.neuralwatt.com/v1` | Override the API URL for all requests (chat, `/models`, `/quota`). For use with a proxy such as Headroom |
| `api` | `"chat-completions"`, `"responses"` | `"chat-completions"` | Generation API surface. `"responses"` opts into the staged `/v1/responses` rollout — see below |
| `storeResponses` | `true`, `false` | `true` | Responses-surface retention. Only read when `api` is `"responses"` — see below |
| `glyphs` | `"auto"`, `"unicode"`, `"ascii"` | `"auto"` | Footer glyph set; `"auto"` degrades to ASCII on legacy terminals (mintty/Cygwin) — see below |

The display appears as soon as a Neuralwatt model is selected: the quota line renders when the session-start prefetch lands, the energy line joins it once a turn records data, and `hideOnOtherProvider` (default `true`) clears everything the moment the active model belongs to another provider.

**Display modes:**

- **`"widget"`** — Shown in the dedicated below-editor status line. Energy on the left, quota on the right, padded to terminal width.
- **`"statusbar"`** — Shown in the built-in pi status bar. When both are set to `"statusbar"`, they're combined with a ` | ` separator: `⚡X J $Y | plan ● kWh ∙ $bal`.
- **`"off"`** — Hidden entirely. For `"quota": "off"`, the `/v1/quota` API fetch is also skipped (saving a network round-trip). Energy data is still parsed from the SSE stream and persisted to the session even when `"off"`.

**Glyphs on legacy terminals.** Older mintty/Cygwin builds measure emoji and ambiguous-width codepoints (the energy bolt, the carbon leaf, the quota dot, the region flag) with cell-width tables that disagree with pi’s. The status bar absorbs that difference, but the widget line is padded to the terminal width, so one over-wide glyph can wrap it physically while pi counts a single row — the footer then desyncs and leaves ghost rows behind. `"auto"` (the default) swaps the glyph set for ASCII equivalents on detected legacy terminals:

| Widget | `"unicode"` (elsewhere) | `"ascii"` (auto on mintty/Cygwin) |
|--------|---------------------------|-------------------------------------|
| Energy | ⚡5.68 mWh | *5.68 mWh |
| Carbon | 🌱1.24 g CO₂ | ~1.24 g CO2 |
| Region | 🇺🇸 PJM 416 | PJM 416 |
| Quota | ● 25.0/33.0 kWh ∙ $12.34 | o 25.0/33.0 kWh - $12.34 |
| Flex | flex −82% | flex -82% |

The widget also never paints the terminal's last column, and clamps an explicit `"unicode"` to ASCII on legacy terminals (one notice per session); the status bar is not edge-padded and always honors the exact choice.

`mcr` and `carbon` follow the same three modes. `carbon` adds two segments: **session CO₂** (`🌱X g CO₂`, on the energy line — cumulative, like energy) and a **fleet grid/region badge** (on the quota line — the latest request's electricity grid, e.g. `🇺🇸 PJM 416`). The badge compresses flag → intensity → balancing-authority tag as the terminal narrows, and a `~` marks intensities from a fallback carbon source. The badge also renders **standalone** (on its own) when `quota` is `off`, so the fleet location still shows.

#### API surface: `chat-completions` vs `responses`

By default the extension generates through `/v1/chat/completions`. Setting `"api": "responses"` switches generation to the [Responses API](https://portal.neuralwatt.com/docs/api/responses) (`/v1/responses`), Neuralwatt's staged-rollout successor surface:

```json
{
  "api": "responses"
}
```

What changes on `responses`:

- **Retention defaults to `store: true`.** Neuralwatt is ZDR (no training on customer data), so retention isn't a privacy exposure, and `store: true` preserves a reasoning model's thinking across turns server-side. Stored turns are account-scoped and expire after 24h. Set `"storeResponses": false` to retain nothing (reasoning continuity is then lost between turns). Note pi-ai's underlying Responses client hardcodes `store: false`; this extension overrides it to follow `storeResponses`.
- **Reasoning still works.** Thinking levels map onto the Responses `reasoning.effort` parameter. The `reasoning.encrypted_content` include pi-ai adds is stripped (Neuralwatt lists it as unsupported).
- **Token usage and billing are unchanged** — Responses downgrades to the chat pipeline internally, so usage, rate limits, and prefix caching match chat completions.

> **Telemetry caveat:** flex works on `/v1/responses` (verified live), but current streams still lack the `: energy` / `: cost` / `: mcr-session` SSE comments. Energy/carbon/billed-cost reporting and flex queue telemetry are unavailable for these requests; token usage and the independently fetched quota remain available. Keep `chat-completions` (the default) if you rely on the energy widget.

**Example — custom quota footer:** If you use your own unified quota footer extension, disable the built-in quota display to avoid duplication:

```json
{
  "energy": "widget",
  "quota": "off"
}
```

### Model Overrides

`modelOverrides` lets you override compat flags and other model properties per model id, **on top of** `patch.json` + `custom-models.json`, without editing the extension. Keyed by model id; `compat` and `thinkingLevelMap` are deep-merged (toggle one flag without redeclaring the rest), scalars are replaced. Applied at session start, so edits take effect on the next `pi` session.

```jsonc
{
  "energy": "widget",
  "quota": "widget",
  "mcr": "widget",
  "carbon": "widget",
  "hideOnOtherProvider": true,
  "modelOverrides": {
    // Disable full-history reasoning for kimi-k2.6 (e.g. to save tokens):
    "kimi-k2.6":      { "compat": { "chatTemplateKwargs": { "preserve_thinking": false } } },
    // Override a single thinking level without redeclaring the map:
    "glm-5.2":        { "thinkingLevelMap": { "high": "max" }, "compat": { "chatTemplateKwargs": { "clear_thinking": true } } },
    // Force a smaller image cap:
    "kimi-k2.7-code": { "vision": { "maxImagesPerRequest": 4 } },
    // Evict over-cap images in bigger batches (default: quarter of the cap,
    // min 2; 1 = exact per-turn FIFO). Bigger batches mean fewer cold prefills:
    // each eviction invalidates the server-side prefix cache at the dropped
    // image, so batching buys several image additions per cache rewrite.
    "kimi-k3": { "vision": { "evictionHysteresis": 8 } }
  }
}
```

Supported config override fields are `compat`, `thinkingLevelMap`, `vision`, and `samplingParams`. Use `patch.json` for other catalog overrides. `vision.maxImagesPerTurn` limits the current user turn; `vision.maxImagesPerRequest` remains the whole-conversation ceiling. Sampling defaults are merged per key; request-level sampling parameters win. See [Compat Settings](#compat-settings) for the catalog of compat flags and what `chatTemplateKwargs` values mean per family.

### Settings UI

`/neuralwatt-settings` opens an interactive settings panel (mirrors pi core's `/settings` — bordered `SettingsList`, Esc to go back) to configure Neuralwatt without editing JSON by hand:

- **Preserved thinking** (nested submenu, one row per model) — toggles `clear_thinking` (GLM-5.2 family) / `preserve_thinking` (Kimi K2.6/K2.7) in `modelOverrides` between **Preserve Thinking** (keep full reasoning history across turns; the default, `clear_thinking: false`) and **Clear Thinking** (let the template drop older reasoning; saves tokens, but can degrade multi-turn recall / cause overthinking).
- **Energy / Quota / MCR / Carbon display** (`widget` / `statusbar` / `off`) and **Hide on other provider** — the same fields as [Display Configuration](#display-configuration), editable live.

Changes write to `~/.pi/agent/extensions/neuralwatt.json` (raw read-modify-write, so unrelated fields survive), refresh the in-memory config, and re-register the provider, so they take effect immediately — no restart needed.

When you switch to — or start pi on — a Neuralwatt model that carries a preserved-thinking flag (e.g. the GLM-5.2 family, GLM-5.1, Kimi K2.6/K2.7), an info notification reports the state and how to change it, e.g. `Preserved thinking ON for glm-5.2 (clear_thinking: false) — suited for coding, but not for prose. Open /neuralwatt-settings to change.` (OFF reads `... reasoning trimmed each turn (lighter; better for prose) ...`). It's an ordinary info notification (not a warning), so it doesn't paint bright yellow.

## Energy Reporting

Neuralwatt provides real-time energy consumption data with every API response. This extension captures it and displays a running total in a dedicated status widget between the editor and the pi footer:

| Segment | Meaning |
|---------|----------|
| `⚡5.68mWh` | Cumulative session energy consumption (auto-scaled: J → mWh → Wh → kWh) |
| `$0.003952` | Cumulative session actual billed cost from Neuralwatt |
| `🌱1.24 g CO₂` | Cumulative session CO₂ emissions (auto-scaled: mg → g → kg); on the energy line when `carbon` is on |
| `pro` | Your Neuralwatt subscription plan |
| `●` | Subscription status indicator (● = active, ⊘ = past due/paused) |
| `31.7/33.0 kWh` | kWh remaining / kWh included in your plan |
| `∙ $64.55` | Credits remaining on your account |
| `🔑 .../.../mo` | Key allowance usage (if set on your API key) |
| `🇺🇸 PJM 416` | Fleet grid/region badge (latest request's electricity grid + its carbon intensity, g/kWh); on the quota line when `carbon` is on. A `~` marks fallback intensities |

The energy and cost data comes from Neuralwatt's SSE stream comments (`: energy` and `: cost`), which the standard OpenAI SDK discards. This extension uses a custom stream handler that parses raw SSE to capture them.

Energy is measured directly from GPU hardware using NVIDIA's NVML. For concurrent requests, Neuralwatt uses token-weighted attribution to fairly calculate your share. See [Neuralwatt's energy methodology](https://portal.neuralwatt.com/docs/energy-methodology) for details.

The same `: energy` comment carries the electricity grid the GPU node drew from (`grid_id`), that grid's carbon intensity, and the resulting CO₂e. The fleet routes across multiple grids, so `grid_id` is latest-wins (the "current" grid) while session CO₂ accumulates like energy. `grid_id` is either a bare ISO country code (`FI`) or an EIA/Electricity-Maps-style `CC-SUBREGION-BA` code (`US-MIDA-PJM`); the badge parses it generically (country flag via regional indicators, balancing-authority tag as the short form), so any new grid renders without a code change.

### Persistence

Energy, cost, carbon, and grid data are persisted per-request as custom session entries. On session resume or tree navigation, the totals are rebuilt by replaying all events in the current branch — CO₂ accumulates like energy, while `grid_id`/intensity are latest-wins. This means:

- **Session resume** — Energy/cost/carbon totals (and the latest grid) are restored when you continue a session
- **Branching** — Navigating to a different point in the session tree shows the correct totals for that branch
- **Forking** — Forked sessions carry their energy (and carbon) history forward

### Per-turn energy event

After every Neuralwatt turn (in the `turn_end` handler, once the SSE tee has drained), the extension emits a `neuralwatt:turn-energy` event on pi's shared event bus so other extensions can surface the energy-billed cost without re-parsing the session. The payload:

| Field           | Type     | Description                                                                         |
| --------------- | -------- | ----------------------------------------------------------------------------------- |
| `costUsd`       | `number` | Actual billed cost for this request (USD)                                           |
| `energyJoules` | `number` | Energy consumed for this request (Joules)                                           |
| `turnIndex`     | `number \| null` | pi's turn index for correlation. `null` if the event didn't carry one.       |

The event is only emitted for turns with Neuralwatt activity (the `pending*` state is per-request), so non-Neuralwatt turns never produce a spurious zero-cost signal. Consumers should correlate on `turnIndex` and treat a missing/`null` index defensively.


## API Documentation

- Neuralwatt API: `https://api.neuralwatt.com/v1`
- Models endpoint: `https://api.neuralwatt.com/v1/models`
- Chat completions: `https://api.neuralwatt.com/v1/chat/completions`

## License

MIT
