# Changelog

## 0.1.1 - 2026-08-19

- Replace the DeepSeek-only reasoning switch with automatic provider-specific hard-off enforcement for every `thinking: off` trial.

## 0.1.0 - 2026-08-19

- Add generic caller-defined agent profiles and all-pairs comparison reports.
- Export the built-in `vanillaAgentProfile`.
- Remove the hard-coded vanilla/IDE evaluation model.
- Add scripted and real-model filesystem-tools-vs-Bash executable examples.
- Add profile-specific model, thinking, skill, and system-prompt settings with global fallbacks.
- Add Pi skills for writing and inspecting evaluations, plus a local installation helper.
- Add complete reproducible CLI and direct-API examples to the tool-profile comparison.
- Extract the paired Pi evaluation runner into a standalone package.
- Add deterministic schedules, trace metrics, reports, and the `pi-eval` CLI.
