# Coding Interview Transcripts Without Losing the Signal

Between raw transcripts and themes sits the unglamorous step everyone botches: coding. Done badly it launders your priors through participant quotes. This is the lightweight discipline that keeps the participants' reality intact.

## First pass: tag, don't interpret

Read each transcript tagging *what kind of thing* was said, in the participant's frame:

- `[goal]` what they were trying to accomplish
- `[workaround]` what they actually do today (the gold — workarounds are unpaid feature specs)
- `[pain]` friction in their words, with intensity signals ("annoying" ≠ "I nearly cancelled")
- `[belief]` how they think it works (right or wrong — wrong beliefs are UX findings)
- `[quote]` verbatim keeper — emotion, specificity, or surprise

Resist merging into your product's vocabulary at this stage. If they say "the syncing thing eats my edits", the code is their phrase, not "conflict resolution UX".

## Second pass: cluster across participants

Lay same-type tags side by side across transcripts. A cluster is real when tags from *different* participants describe the same underlying thing in different words — that triangulation is the whole method. Three tags from one talkative participant = one data point, not three.

## The intensity dimension

Frequency isn't severity. Tag each pain with the strongest evidence attached:
- 💬 mentioned · 😤 emotional language · 🔁 built a workaround · 💸 caused churn/spend behaviour
A pain that 3 people built workarounds for beats one that 8 people mentioned mildly. Report both numbers.

## What kills coding integrity

- Coding only the parts that support the roadmap (do a full pass before looking at your hypothesis list)
- Treating the screener-articulate as representative — note who's over-quoted
- Losing the question context — an answer to "what frustrates you about X?" is prompted; unprompted mentions are worth 3× (mark them `[unprompted]`)
