# Changelog

All notable changes to `pi-sarvam-provider` are documented here.

The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [0.1.5] - 2026-08-11

### Changed

- **Compatible with the Oh My Pi (omp) fork of pi.** omp's bundled `pi-ai` deleted the
  legacy api-registry (`registerBuiltInApiProviders` / `getApiProvider`), so the static
  import of those helpers from `@earendil-works/pi-ai/compat` failed to link when the
  extension loaded on omp. The registry lookup is now a dynamic-import probe: on upstream
  pi it finds the base `openai-completions` `streamSimple` exactly as before, so retry
  and payload normalization still wrap Sarvam traffic; on omp the probe degrades to
  `undefined` and the provider registers with `streamSimple: undefined`, letting omp's
  native openai-completions dispatch take over. The extension now installs and registers
  on both.

## [0.1.4] - 2026-08-08

### Fixed

- **One-sided Claude-style edits no longer corrupt files.** The edit shim previously
  folded `old_string`/`new_string` into `oldText`/`newText` with `?? ""`, so a model
  sending only `new_string` fabricated `oldText: ""` (the edit engine rejects it — a
  wasted turn) and a model sending only `old_string` fabricated `newText: ""` and
  silently **deleted** the matched text from the file. Both fields are now required,
  non-empty; a one-sided call is rejected by schema validation and the model retries.
- **Network-level failures are now retried.** pi-ai's `retryProviderRequest` only
  retries errors carrying an HTTP status, and the ladder's transient filter missed
  TCP/transport failures, so `fetch failed`, `ECONNRESET`, `socket hang up`, SDK
  timeouts, and dropped connections fell through both retry layers. The filter now
  covers them.
- **Aborted turns are labelled `aborted`.** The wrapper's abort error event used
  `reason`/`stopReason: "error"`; pi's event protocol distinguishes cancellations, and
  session metadata now records them correctly.
- **An explicit `retry.provider.maxRetries: 0` in pi's settings now disables the
  `Retry-After` layer.** The previous `|| PROVIDER_RETRIES` treated an intentional `0`
  as "unset" and forced 2 retries. pi's default is *unset* (not `0`), so the `??`
  check keeps the default at 2 while honouring explicit values including `0`.
- **`Stream Errors` in the debug metrics now counts terminal HTTP error events**, not
  just thrown exceptions (the common failure path was under-counted). Aborts are still
  excluded.
- **A hung gateway can no longer stall pi startup.** The `/v1/models` discovery
  fetch has a 10-second timeout; previously an unresponsive host blocked `activate()`
  indefinitely, leaving the session with no provider.
- **Model discovery retries transient failures.** `/v1/models` is subject to the
  same intermittent gateway `403`s as completions (verified live: identical requests
  returned both `200` and `403`). Discovery now retries transient/network failures
  on a short ladder (0.5 s / 1.5 s) while still failing fast on an
  `invalid_api_key_error` body, so a flaky gateway no longer silently removes the
  provider for the whole launch.
- **Cache robustness.** Clock fields are type-checked (a corrupt entry with a string
  `timestamp`/`ttl` compared as `NaN` and served forever), model ids are validated, and
  mode 600 is enforced with `chmod` on every write rather than only at file creation.
- **Degenerate model fields are sanitized.** A `context_window` of `0`/`NaN`/non-
  numeric registered a broken window (pi compacts every turn or never); both
  `context_window` and `max_tokens` now fall back to safe values.
- **Tool shims reject non-object arguments** instead of spreading a string into
  index-keyed garbage; tool-call metrics moved out of `prepareArguments` (which runs
  for every provider) into a provider-gated `tool_call` handler.
- **Image-only messages no longer collapse to empty content.** The payload
  normalizer now emits a `[non-text content omitted]` placeholder instead of an empty
  string some OpenAI-compatible endpoints reject.
- **Size-guard hardening.** `tool_calls[].function` is only spread when it is an
  object (a string-typed `function` previously produced garbage fields).

### Changed

- **Pure compatibility logic extracted into `src/sarvam.ts`** (path remapping, retry
  classification, model-field sanitization, payload normalization, error-event
  construction, abort-aware sleep) so it can be tested without pi or a network.
- **New test suite** (`tests/sarvam.test.ts`, `node:test`, no new dependencies —
  `pnpm test`, requires Node ≥ 22.6 for type stripping). 30 tests cover the remapping,
  retry classification, size-guard stages (stub → drop-at-user-boundary → hard cap),
  idempotence, and abort semantics.
- **Docs aligned with verified behaviour:** the tool shims are registered globally
  (they pass non-Claude-style arguments through untouched); when two installed copies
  collide the first-registered one wins silently; `retry.provider.maxRetries`
  precedence is stated precisely; and the discovery endpoint is now known to answer
  `403` (not `200`) for an invalid key — verified against the live API during testing.

## [0.1.3] - 2026-08-01

### Fixed

- **Auto-compaction no longer fails against Sarvam.** pi supplies the `onPayload` callback
  that fires `before_provider_request` only on the normal turn path; compaction and branch
  summarization build their own request options (`createSummarizationOptions` returns just
  `{maxTokens, signal, apiKey, headers, env}`), so no hook fired and the summarization
  request reached Sarvam with pi-native array content. Sarvam rejected it with
  `body.messages.1.user.content : Input should be a valid string` — index 1 being the
  summarization prompt — surfacing as `Auto-compaction failed`. Because compaction is what
  frees up context, this was terminal: once the window filled, every subsequent turn failed.
  The payload normalization is now also attached as an `onPayload` on the options passed
  down inside the provider's stream wrapper, which every Sarvam request funnels through, so
  both paths are covered. Any callback pi does supply still runs first, and the transform is
  idempotent, so double application is a no-op.

### Changed

- **`Retry-After` is now honoured on rate limits.** pi-ai's `retryProviderRequest` already
  implements this correctly — `retry-after-ms` and `retry-after`, both delta-seconds and
  HTTP-date, capped at `maxRetryDelayMs` — but it is inert unless `maxRetries > 0`, and pi
  ships `retry.provider.maxRetries: 0`. So by default a 429 fell through to this extension's
  fixed 1s/3s/8s ladder, which burned all three attempts in ~12s against a wait the gateway
  had already specified, sending two requests certain to be rejected. The provider options
  now request 2 retries for Sarvam traffic, enabling pi's implementation rather than adding a
  second one. Covers 408/409/429/5xx; pi-ai does not retry 403, so Sarvam's transient gateway
  403s remain the existing ladder's job. Set `SARVAM_PROVIDER_RETRIES=0` to opt out; an
  explicit `retry.provider.maxRetries` above 0 in pi's settings takes precedence.
- **Model discovery cache moved from memory to disk**
  (`$XDG_CACHE_HOME/pi-sarvam-provider/models.json`, mode 600, 5 minute TTL). Discovery runs
  once per process during activate, so the previous in-memory cache could never serve a hit —
  each run missed, wrote, and exited without reading. The disk cache skips the `/v1/models`
  round-trip on the next pi launch. Entries are invalidated by base URL, a 12-char hash of
  the API key (so rotating the key invalidates rather than serving models it cannot access),
  and TTL. The API key itself is never written to disk. Cache I/O failures are non-fatal:
  a missing, corrupt, stale, or unreadable cache just fetches fresh.
- **The `SARVAM_DEBUG` metrics summary now prints on process exit** instead of at the end of
  activate, where every counter was still zero. It is suppressed entirely when no requests
  were made.

### Removed

- The rate-limiting logic in the model cache. It could never fire — the delay was only ever
  set by a cache write, which happened after the sole cache read — and one request per
  process start needs no throttle.
- `getModelDiscoveryMetrics`, `getProviderMetrics`, and `clearAllMetrics`, which were defined
  but never called.

## [0.1.2] - 2026-08-01

### Changed

- **Relicensed from Apache-2.0 to MIT.** The `license` field and `LICENSE` file both change;
  the copyright holder is unchanged. Note that 0.1.0 and 0.1.1 were published under
  Apache-2.0 and remain so — this applies from 0.1.2 onward.
- Added `repository`, `homepage`, and `bugs` metadata now that the GitHub repository exists,
  so the npm page links back to the source.
- Corrected the release-link account in this file (`rocknarayan` → `nrynss`); the earlier
  links pointed at a GitHub user that does not own this repository. The 0.1.0 and 0.1.1 link
  definitions were dropped rather than repointed: this repository's history begins at 0.1.2,
  so there are no tags for them to resolve to.
- Replaced the stale "Before you publish" section of the README with a Development section,
  including the local-testing conflict caveat (an installed copy and a locally loaded copy
  register the same tool names, and pi rejects the duplicate). A leftover copy of the old
  section survived under Install, still referencing `npm pack` and a `0.1.0` tarball; it is
  now removed.
- **Rewrote the README Install section around `pi install npm:pi-sarvam-provider`.** The
  section previously only documented hand-editing the `packages` array in
  `~/.pi/agent/settings.json`, which is now shown as the alternative. Added the related
  package commands (`pi list`, `pi update`, `pi remove`, version pinning, `-l` for
  project-local installs) and a note on persisting `SARVAM_API_KEY` in a shell config.

## [0.1.1] - 2026-07-30

### Fixed

- **Extension no longer fails `tsc`** — the api-registry helpers (`getApiProvider`,
  `registerBuiltInApiProviders`) are now imported from `@earendil-works/pi-ai/compat`
  instead of the package root. pi injects the compat barrel for the bare specifier at
  runtime, but the root's *types* are the newer narrow surface and declare none of them,
  so the old import ran correctly yet failed typecheck. The `/compat` subpath is the same
  module with matching types. Note this entrypoint is documented upstream as temporary;
  it will need a `createProvider` migration when pi-ai removes it.
- **Model discovery failures are handled** — a non-2xx response, a non-JSON body, an empty
  model list, or an unreachable host now log a warning and skip provider registration.
  Previously any of these threw out of the extension entry point (`models.data` is
  `undefined` on an error body, so `.map` raised a `TypeError`), surfacing as an opaque
  stack trace. Note Sarvam's `/v1/models` answers `200` even for an invalid key, so a bad
  key is not caught here — it first surfaces on the completion request.
- **Invalid API keys fail fast** — Sarvam returns `403` for both transient gateway blips
  and a permanently invalid key, so the retry filter treated an auth failure as retryable
  and burned the full backoff ladder before failing anyway (measured: 16s, now 3.6s). The
  body's error code now distinguishes the two.
- **Oversized tool-call arguments are trimmed** — the last-resort stage of the 256 KB guard
  capped message content only. A recent `write` call carrying a whole file body lives in
  `tool_calls[].function.arguments`, which the earlier stage stubs only for *older*
  messages, so such a request stayed oversized and still hit the 403. Oversized arguments
  are now blanked to `{}` rather than sliced, since a truncated JSON string is invalid.
- **No more orphaned `tool` messages** — when dropping the oldest turns found no user-message
  boundary, the fallback sliced by message count and could start the conversation on a
  `tool` message whose `tool_call` had been dropped, which the endpoint rejects. It now
  falls back to the last user turn (shrunk by the hard cap) or keeps nothing.
- **Aborting during retry backoff is immediate** — the backoff `sleep` now races the timer
  against the abort signal instead of always running to completion (up to 8s late).
- **`streamSimple` delegates to the base `streamSimple`**, not `stream`, matching the options
  shape pi passes to the function it is registered as.

### Changed

- Bumped the `@earendil-works/pi-ai` and `@earendil-works/pi-coding-agent` devDependencies
  from `^0.80.2` to `^0.83.0`, matching the version pi actually loads at runtime. While they
  disagreed, `tsc` and pi were checking different copies of the package — the root cause of
  the compat-entrypoint mismatch above.
- Added `pnpm-workspace.yaml` declaring two transitive dev-only install scripts as not
  built, so `pnpm install` exits 0 and `pnpm run typecheck` is not blocked by pnpm's
  dep-status check.

## [0.1.0] - 2026-06-28

First release. Registers Sarvam AI as a pi model provider and adds the compatibility
shims its OpenAI-compatible endpoint needs. All behaviour is scoped to the `sarvam`
provider; other providers are untouched.

### Added

- **Provider registration** — discovers models from `https://api.sarvam.ai/v1/models`
  and registers them with reasoning enabled, thinking-level mapping, and text input.
  Requires the `SARVAM_API_KEY` environment variable.
- **`developer` → `system` role** — sets `compat.supportsDeveloperRole: false` so pi
  stops sending the `developer` role (which Sarvam rejects) for reasoning models, plus a
  payload-level remap as a safety net.
- **Array content → string** — flattens pi's array-of-parts message content to the plain
  string Sarvam requires.
- **Tool-argument compatibility** — remaps Claude-style tool arguments
  (`file_path` → `path`, `old_string`/`new_string` → `edits[{oldText,newText}]`) via
  `prepareArguments`, *before* schema validation, composed with pi's own edit-argument
  recovery (so JSON-string `edits` keep working).
- **Windows path sanitisation** — strips a spurious leading separator before a drive
  letter (`/E:/work` → `E:\work`) to prevent `path.resolve` from doubling the drive
  (`E:\E:\work…`) and failing `mkdir`.
- **Transient-error retry** — wraps the provider stream and retries with backoff
  (1 s / 3 s / 8 s) when the first stream event is a transient error (403 / 429 / 5xx /
  gateway), which is safe because such failures occur at connect time before any content
  streams. Non-Sarvam traffic passes through untouched.
- **256 KB request-size guard** — Sarvam's gateway rejects request bodies ≥ 256 KB with a
  `403` (verified: 255 KB → 200, 262 KB → 403). The guard keeps the outgoing body under the
  limit, escalating only as needed:
  1. stub the content (and tool-call arguments) of older messages, preserving the system
     prompt and the most recent turns;
  2. when the message *count* alone is too large (per-message JSON + tool-call ids), drop
     the oldest turns — cutting at a user-message boundary so no tool message is orphaned
     and no tool_call is left dangling;
  3. as a last resort, hard-cap any remaining oversized content.

  This is non-destructive: only the outgoing request is trimmed; the session history on
  disk is left intact.
- **Editing guidance** — appends concrete exact-match editing rules to the system prompt to
  help smaller Sarvam models land edits reliably.

[0.1.5]: https://github.com/nrynss/pi-sarvam-provider/releases/tag/v0.1.5
[0.1.4]: https://github.com/nrynss/pi-sarvam-provider/releases/tag/v0.1.4
[0.1.3]: https://github.com/nrynss/pi-sarvam-provider/releases/tag/v0.1.3
[0.1.2]: https://github.com/nrynss/pi-sarvam-provider/releases/tag/v0.1.2
