# Worked Example — Study Notes Synthesizer

A gold-standard run of this skill: messy Research Methods materials synthesized by dependency, with a confusable pair given the side-by-side treatment and a gap flagged at the top instead of discovered in the exam.

## The brief

Jonas pastes materials for "PSY 210: Research Methods" — final in 9 days, format: 40 MC questions + 2 short essays. The pile: eleven lectures of bullet notes (in teaching order: history of methods, ethics, measurement, experimental design ×2, quasi-experiments, surveys, validity ×2, statistics review, replication crisis), slides for eight of the eleven, and one past paper. His declared weak spot: "I always mix up validity types, and I don't really get what confounds are versus just… other variables." The syllabus lists "single-case and small-N designs" as a topic; his notes contain one bullet on it and no slides.

## The output

### Study Guide: PSY 210 Research Methods — final in 9 days
**Start here (your declared weak spots):** the Validity Family cluster and the confound box below — do those self-tests first, today.
**⚠ Gaps in your materials:** *Single-case and small-N designs* is on the syllabus but your notes contain one bullet and no slides — not enough to study from. Get the textbook chapter or a classmate's notes before treating this guide as complete. (Also: lectures 9–11 have no slides; the notes look thin for the replication crisis — one essay prompt could live there.)

#### Cluster: Measurement (everything else builds on this)
**Plainly:** before you can test anything, you have to turn an idea ("stress") into something countable ("cortisol level" or "self-report score 1–10") — and that translation can be done well or badly.
**Formally:** operationalization — defining a construct via measurable indicators. **Builds on:** nothing; it's the foundation. **Confused with:** validity (validity asks whether the translation was any *good*).

#### Cluster: The Validity Family (your weak spot #1)
**Plainly:** four different questions people compress into "is the study valid":
| Type | The question it asks | Classic threat |
|---|---|---|
| Internal | Did X really cause Y *in this study*? | confounds, selection |
| External | Does it generalize beyond this study? | unrepresentative samples, lab artificiality |
| Construct | Did we measure what we claimed to measure? | bad operationalization |
| Statistical | Are the numbers strong enough to believe? | low power, p-hacking |
**Formally:** [definitions as per lecture 8]. **Builds on:** measurement. **Confused with:** *reliability* — reliability is consistency (same result twice), validity is truth (right result at all). A bathroom scale that's always 3kg heavy is reliable and invalid.

#### Cluster: Confounds vs. everything else that varies (your weak spot #2)
**Plainly:** lots of things vary in a study; a variable is only a *confound* if it moves **with your independent variable** and could affect the outcome — that's what breaks the causal claim. Noise that varies randomly just adds fuzz; a confound adds a *rival explanation*.
**Formally:** a confound covaries with the IV and plausibly affects the DV. **Builds on:** experimental design. **Confused with:** extraneous variables (vary, but not systematically with the IV).
##### Test yourself
1. [recall] Define a confound in one sentence including the word "covaries."
2. [apply] Coffee drinkers show higher anxiety than non-drinkers. Name two plausible confounds.
3. [why X not Y] Why does random assignment handle confounds but *not* fix bad operationalization?

#### Cluster: Design choices (experiments → quasi → surveys)
**Plainly:** the designs are a trade: experiments buy internal validity with control; quasi-experiments accept known threats when assignment isn't possible; surveys buy breadth and give up causation entirely. **Builds on:** validity family — each design is a different answer to which validity you'll protect. *(Self-tests at three depths, as above.)*

#### Cluster: Replication crisis (thin notes — likely essay territory)
**Plainly:** many famous findings don't repeat; causes include p-hacking, publication bias, small samples — i.e., statistical-validity failures at field scale. **Builds on:** statistical validity. *(Flagged above: reinforce from the textbook.)*

#### One-page cram sheet
1. Operationalization = construct → countable. 2. Reliability = consistent; validity = true. 3. Scale-3kg-heavy example. 4. Four validities: internal/external/construct/statistical + one threat each (that's 8 lines). 5. Confound = covaries with IV. 6. Random assignment kills confounds, nothing else. 7. Quasi = no random assignment, name the threat. 8. Surveys ≠ causation. 9. Replication crisis = statistical validity at scale. *(17 items total — answers to all self-tests on the last page.)*

## Why it's shaped this way

- **Clusters follow dependency, not the eleven-lecture chronology** — measurement comes first because validity is unintelligible without it, even though it was lecture 3; the teaching order (history first) is how it was taught, not how it's structured.
- **Both declared weak spots got dedicated clusters at the top** — the guide leads with validity and confounds because the student said so, and the start-here line routes him there on day one of nine.
- **Plain-language before formal, every time** — "turn an idea into something countable" precedes "operationalization," per the explain-then-anchor rule; fluent jargon masquerading as understanding is the failure this skill exists to prevent.
- **Confusable pairs get side-by-side treatment** — validity/reliability (with the 3kg-scale anchor) and confound/extraneous are exactly the distinctions a 40-question MC exam farms.
- **The gap is a finding at the top, not a footnote** — single-case designs is syllabus-listed and note-absent; per the framework, wrong confidence is worse than known ignorance, and the guide refuses to invent material to fill it.
- **Self-tests span the three depths and are checkable against the guide** — the confound apply-question is answerable from the cluster above it; no question outruns the notes.
- **The cram sheet stops at 17 items** — under the 20 cap, telegraphic, highest-yield only; a cram sheet with everything is a guide with nothing.
