# Recorded evaluation results

## `deepseek-official` / `deepseek-v4-flash` — 2026-08-14

The first frozen Rewrite evaluation used corpus digest `52613993596f4fb5e4298ba5618e2fffedba19f68e01358055756d2981740dc8` and prompt version `query-rewrite-v1`.

| Metric | Result | Gate |
| --- | ---: | ---: |
| Exact Recall@1 | 100.00% | 100% — pass |
| Cross-language raw BM25 Recall@5 | 33.93% | comparison |
| Cross-language Rewrite Recall@5 | 71.43% | at least 80% — fail |
| Cross-language uplift | +37.50 points | at least +20 — pass |
| Same-language raw BM25 Recall@5 | 91.07% | comparison |
| Same-language Rewrite Recall@5 | 92.86% | at least 90% — pass |
| Valid Rewrite rate | 54.46% | diagnostic |

All 51 failed Rewrite calls were rejected as `invalid-response`; there were no provider-route or transport failures. The 112 calls recorded mean latency 1490 ms, p50 1500 ms, and p95 1984 ms. Reported usage was 4,197 uncached input tokens, 13,824 cache-read tokens, 11,488 output tokens, and 10,830 reasoning tokens.

The release gate failed. This run therefore does not support publishing `0.1.0-alpha.1` or claiming at least 80% cross-language Recall@5.

Local artifact digests:

```text
rewrite-responses.json  e92bb0718ad46dba3cec5de051729d06eaa67b9c8f9fcfadbf67e3cbb96d5d0b
rewrite-report.json     abd1120ee95ad042f8f72fe9347eae2aa2305d1ff0ef8e33adff5eb7c5ae199f
```

The full generated files remain ignored because they contain per-query model output. They may be attached to a review or release as external evidence. Do not tune the prompt or parser against this frozen run and then reuse the same corpus as an unseen release gate; use a separate development set and a new held-out frozen set.
