# Problem Substantiveness — Extraction Plan

- **Sub-criterion:** 2.1 (Product Thinking)
- **Weight:** 12.5 pts
- **Source quote:** _"Is the problem clearly stated, non-trivial, with defined user/use-case?"_

This recipe scores whether the problem is substantively framed — a crisp, non-trivial problem statement with a defined user and use-case, versus a toy-shaped placeholder.

---

## Extraction recipe

LLM-primary. The evaluator reads the top of the repo (README, any PRD/SPEC doc, sample data, one or two handler files) and answers: **Is this a substantive, well-defined problem, or a toy-shaped placeholder?**

### Signal 1 — Problem framing (LLM interpretive)

Feed the LLM the following context:

```bash
# README (truncated to first ~500 lines)
head -500 "$REPO/README.md" > /tmp/prob-sub-readme.md

# PRD / SPEC / DESIGN docs if present
find "$REPO" -maxdepth 3 -iname "PRD*.md" -o -iname "SPEC*.md" -o -iname "DESIGN*.md" \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' 2>/dev/null \
  | xargs -I{} head -200 {} 2>/dev/null > /tmp/prob-sub-docs.md

# Sample data files — check provenance
find "$REPO" -maxdepth 4 \( -iname '*.csv' -o -iname '*.json' -o -iname 'sample*' -o -iname 'fixture*' -o -iname 'seed*' \) \
  -not -path '*/node_modules/*' -not -path '*/.venv/*' 2>/dev/null | head -10 > /tmp/prob-sub-data-files.txt

# Top-level folder structure (reveals user/use-case intent)
tree -L 2 -I 'node_modules|.venv|.git|dist|build' "$REPO" 2>/dev/null | head -80 > /tmp/prob-sub-tree.txt
```

Write the LLM read to `raw/problem-substantiveness-llm-read.md`. The LLM emits:

```yaml
detected_domain: "internal dev tooling" | "customer-facing SaaS" | "infra automation" | "ops observability" | "data pipeline" | "agent/LLM tooling" | ...
problem_clarity_score: 0-5          # 5 = crisp problem statement with user + constraint + outcome named; 0 = no statement
target_user_inferred: "internal engineers" | "support agents" | "end customers" | "unclear" | ...
sample_data_provenance: "synthetic" | "redacted-real" | "production" | "none" | "unclear"
toy_pattern_detected: true | false  # hello-world, generic TODO app, calculator, tic-tac-toe, etc.
one_sentence: "<LLM's one-line verdict on whether this is a substantive problem>"
```

### Signal 2 — Proxies JSON

Aggregate the LLM output into `raw/problem-substantiveness-proxies.json`:

```json
{
  "detected_domain": "...",
  "problem_clarity_score": 0-5,
  "target_user_inferred": "...",
  "sample_data_provenance": "...",
  "toy_pattern_detected": true|false
}
```

---

## Banding (absolute)

| Band | Criteria |
| --- | --- |
| **5** | Problem statement is crisp: user + constraint + outcome named. Domain is non-trivial (enterprise ops / a non-trivial real-world domain / novel agent workflow). Sample data is production-shaped. `problem_clarity_score: 5`, `toy_pattern_detected: false`. |
| **4** | Problem statement is substantive with one minor gap (user inferable but not named; outcome fuzzy). `problem_clarity_score: 4`, `toy_pattern_detected: false`. |
| **3** | Problem is stated but under-specified: missing user definition OR missing success criterion. `problem_clarity_score: 3`. |
| **2** | Problem statement is vague or generic. README is mostly boilerplate. `problem_clarity_score: 2`. |
| **1** | No problem statement, or README is an unmodified starter/template README. `problem_clarity_score ≤ 1`. |
| **0** | Empty README + no PRD/SPEC + toy pattern detected. |

**Toy-pattern auto-floor:** if `toy_pattern_detected == true` (hello-world, generic TODO, calculator, tic-tac-toe, etc.), cap the band at **2** regardless of other signals — a clear problem statement doesn't rescue a toy.

---

## Raw dumps

- `raw/problem-substantiveness-llm-read.md` — LLM's full interpretive read with axis scores.
- `raw/problem-substantiveness-proxies.json` — aggregated JSON of signals (also embedded in `scorecard.yaml.problem_substantiveness.data`).

---

## Not in scope

- Judging whether the problem matches a specific organization's real priorities needs operational context — out of scope; this scores whether the problem is substantively framed, not whether it's the right problem for any given org to be solving.
- Judging the quality of the implementation — that's 1.1–1.4 + 2.4. This recipe only looks at framing.
