# User session report — v1.14.0

> **Provenance, stated first:** the learner was a persona agent ("Sam") in an isolated
> context with a pinned, deliberately patchy knowledge state (two confident misbeliefs,
> genuine blanks), answering tersely and without performing competence. The tutor (this
> session) followed the release-tree skills verbatim; the assessor and architect ran the
> release-tree specs by path (installed plugin is 1.1.1 — the §5.5 trap, dodged). This
> certifies the FLOW under the shipped prose. What it cannot certify is the felt
> experience of a real human across real days — the protocol's own note says
> `ENGRAM_TODAY` cannot fake the feel of returning, and that remains true here.

```
topic: espresso dial-in (why shots are randomly sour/bitter)    mode: standard
real minutes: ~35 (architect ~6 of them)    nodes encoded: 2 (+1 pretest credit)    reviews cleared: 1
```

## WHAT WORKED

- **The contrast-first opening (P18) did its job twice, visibly.** On
  `extraction-continuum`, the three-case set made Sam's own tea-bag misconception
  collide with the data BEFORE any teaching: *"this data kills what I told you earlier…
  I don't love it, but I can't argue with the list."* He ranked two candidate rules, and
  one dialogic question ("why did your rule 2 break on case 3?") got him to derive the
  dissolution-order mechanism himself. The resolve then only had to name what he built.
- **Attempt-built resolution felt natural to run** — quoting his rules back and asking
  where each broke is a better teaching surface than exposition, and the grammar's
  wording made it the obvious move rather than a remembered obligation.
- **The guided analogy pair (P19) produced a genuinely portable schema** — Sam's
  alignment sentence ("one axis, a window in the middle, the two misses are opposite
  directions producing different failures — so the failure tells you the direction")
  was scored 2/2 by the blind assessor, on the receipt as `alignment_quality: 2`, grade
  unmoved.
- **The misconception loop closed the way diSessa-lite says it should**: the tea-bag
  model was bounded ("right for tea, wrongly borrowed"), not mocked; the
  pressure-does-the-extracting belief was re-scoped ("pressure moves the water,
  dissolution extracts — the job title was wrong"). Sam repeated both boundaries back
  unprompted at the review.
- **The ambient return surface was exactly right**: due count + his plan verbatim + the
  87%-and-falling decay line, once each, nothing repeated by the tutor.
- **The blind assessor rounded down correctly** — `good` not `easy` on a complete
  production because the "why acids first" clause was hedged, and its feedback_line
  (replace "something about size" with the real driver) is what Sam actually fixed by
  the review, where he produced it cold.
- Engine plumbing was invisible to the learner: sid round-trip, stash-before-grade,
  zero engine errors surfaced in the dialogue.

## WHAT CONFUSED ME

- **The confidence-pick pipelining is awkward in chat.** Collecting the band for probe
  N in the same message that poses probe N+1 keeps the no-feedback rule but reads
  slightly bureaucratic three times in a row. (Platform picker would hide this; the
  chat-fallback shape is the grammar's documented degrade and it worked.)
- **The dose-cap tension needed an improvised sentence.** After the review, durability
  says ~17 days but the next interval is 3 (relearning dose, by design). The payload
  carries `schedule_policy`, but no skill line tells the tutor what to SAY when the
  growth line ("holds ~2 weeks") and the next-due date ("3 days") land in the same
  breath. I improvised "early reviews are front-loaded on purpose." Worth one scripted
  sentence in `/review`'s close guidance.

## WHAT ANNOYED ME — and what I would have quit over

- Nothing quit-worthy surfaced. The architect wait (~6 min) is exactly the moment the
  skill's mandatory latency line exists for; with the line said, it was fine.
- Minor: the CONNECT analogy and the node's `transfer_probe` can collide. The architect
  put the best interest-analogy (load balancer, for the developer) into `channeling`'s
  transfer_probe, so using it at CONNECT would have leaked the future maturity probe. I
  skipped the pair on that node. **Suggestion:** one line in the architect spec —
  *CONNECT-analogy clothing and transfer_probe clothing must differ.*

## WHAT IT TOLD ME vs WHAT WAS TRUE

Every number checked against the state, run side by side (§4.8.1 by hand):

- `rate` (review): s 3.71 → 17.22, interval 3 (dose-capped, labeled in payload) — real.
- `momentum`: 1 review / 1 recalled / +13.5 days stability in window — matches receipts.
- `retention`: "measured over 1 retrieval — none yet at the 30-day mark" — refuses to
  flatter, and **prepends the grader-unaudited stamp** ("[grader unaudited — QWK
  unknown]") to its own read. True: this sandbox's grader has no audit.
- `decay`: "3 concepts encoded, none due yet… 2.4 of 3 expected to survive 30 days" —
  consistent with retention (different populations, both labeled).
- `adherence.loop_closure`: rate 1.0, "the loop is closing" — true here, Sam returned.
- `calibration`: n=1 → "insufficient-data", no verdict from one datum.
- `doctor` ok: true. No command contradicted any other on this state.
- The pretest credit, both encode receipts, and the review receipt all carry the right
  kinds; `stats.transfer` is honestly empty (nothing mature).

**The three v1.14 moves, statused:** contrast-first — fired twice, exactly at its gate;
guided analogy pair — fired once, scored on the receipt; **concept discrimination drill
— did NOT fire, correctly**: the due item's `contrasts_with` was empty (its authored
confusable pair, grind↔ratio, is still un-encoded, and the engine drops `new`
siblings — the pre-exposure guard doing its job). The drill therefore remains verified
by selftest and live-drive only, not by this user session. First real firing will come
deeper into an arc.

## WOULD A STRANGER GET THROUGH THIS?  **yes** —

zero dead ends, no raw engine output ever reached the learner, the one long wait was
pre-announced, every number shown was true and none flattered. The learner left with
the thing the product promises: *"the thing that changed isn't the shots, it's that
when one goes wrong I know which way to move."*

## VERDICT: **ship** — with the simulation caveat stated above, and two follow-ups
filed, neither blocking: the dose-cap sentence for `/review`'s close, and the
analogy/transfer-probe clothing rule for the architect. The real-days retention feel
should be dogfooded by the founder between this release and the next, per the
protocol's standing note.
