# Evaluation report - MedSeek v0.1.2

Generated: 2026-08-24 - by `pnpm eval` over the fully
synthetic annotated corpus in `eval/corpus`. No real PHI was used. Gates live in
`tests/eval.spec.ts` and run on every CI pass; regenerate this file when the engine,
templates, or corpus change.

## De-identification assist (Safe Harbor pattern detector)

9 synthetic clinical notes, 49 gold identifier spans across
18 categories.

| Category | TP | FP | FN | Recall | Precision | Flagged low-confidence |
|---|---|---|---|---|---|---|
| account_number | 2 | 0 | 0 | 100% | 100% | 0 |
| age_over_89 | 1 | 0 | 0 | 100% | 100% | 0 |
| biometric | 2 | 0 | 0 | 100% | 100% | 2 |
| date | 8 | 0 | 0 | 100% | 100% | 0 |
| device_id | 2 | 0 | 0 | 100% | 100% | 0 |
| email | 1 | 0 | 0 | 100% | 100% | 0 |
| fax | 1 | 0 | 0 | 100% | 100% | 0 |
| geographic | 2 | 0 | 0 | 100% | 100% | 0 |
| health_plan_id | 1 | 0 | 0 | 100% | 100% | 0 |
| ip_address | 1 | 0 | 0 | 100% | 100% | 0 |
| license_number | 1 | 0 | 0 | 100% | 100% | 0 |
| mrn | 3 | 0 | 0 | 100% | 100% | 0 |
| name | 9 | 0 | 1 | 90% | 100% | 4 |
| other_id | 1 | 0 | 1 | 50% | 100% | 1 |
| phone | 7 | 0 | 0 | 100% | 100% | 0 |
| ssn | 1 | 0 | 0 | 100% | 100% | 0 |
| url | 3 | 0 | 0 | 100% | 100% | 0 |
| vehicle_id | 1 | 0 | 0 | 100% | 100% | 0 |

**Micro-average: recall 95.9%, precision 100.0%.**

Known hard misses kept deliberately in the corpus (they are why the tool says
"assist, not certification"): a bare patient name directly after a section heading
with no cue label, and letter-prefixed reference codes (`PA-2026-884100`) that carry
no structural signature.

## Completeness scanning

5 synthetic notes with known omitted sections:

| Metric | Value |
|---|---|
| Genuinely absent required fields | 6 |
| Detected by scan | 6 |
| Missed by scan | 0 |
| False "not found" on present fields | 0 |
| Detection rate | 100% |
| False-alarm rate | 0% |

## CI gates (tests/eval.spec.ts)

- De-id micro recall >= 0.85; precision >= 0.80
- Structured identifiers (SSN, email, URL, IP, MRN, phone, fax): recall = 100% on labeled forms
- Completeness detection rate >= 75%; false-alarm rate <= 50%

## Limitations

- The corpus is synthetic English text written to exercise every detector rule;
  it does not measure performance on real charts, other languages, or OCR noise.
- Regex detection fails silently in the dangerous direction. The tool's output
  always carries a residual-risk note and lists low-confidence spans for review.
- Readability scoring uses published formulas (Flesch-Kincaid, SMOG) and is
  deterministic; its formulas are asserted directly in unit tests.
