---
name: frame-ml-problem
description: Frame an ML problem before any modeling — define the prediction target, the correct evaluation metric, a baseline, and the success threshold.
keywords: ml, problem framing, metric, baseline, target, success criteria
---

# Frame ML Problem

> This skill **inherits and extends** `superpowers:brainstorming` for collaborative Q&A.
> First, invoke `superpowers:brainstorming` to explore intent and ambiguities with the developer,
> then apply the structured framing steps below.

---

## When to use this skill

- Gate 1 of the ML workflow — always the first step before any modeling
- Any new ML ticket where the prediction target or success criterion is not yet explicit
- Whenever the evaluation metric has not been pinned down or justified against the business goal

---

## Steps

### 1. Read the ticket context

Read `.aiflow/context/current.json`:
- `description` — what the stakeholder wants the model to do
- `comments` — additional context, constraints, or prior attempts
- `context.files` — related notebooks, data samples, or prior model artifacts

### 2. State the prediction target precisely

Define what is being predicted, at what granularity (row level, user level, session level, etc.), and crucially **when the input features are available** relative to the prediction moment. Guard against accidentally incorporating future information — any feature computed from or derived after the target event is a candidate for leakage.

### 3. Classify the problem type

Explicitly label the task:
- **Classification** (binary or multi-class): predict a category
- **Regression**: predict a continuous value
- **Ranking / Recommendation**: order items by relevance
- **Generation / Seq-to-seq**: produce text, code, or structured output
- **Clustering / Anomaly detection**: unsupervised grouping or outlier detection

State the type before choosing any model family or metric.

### 4. Choose the evaluation metric and justify it

Select the primary metric and tie it explicitly to the business goal:
- Imbalanced classification: prefer F1, PR-AUC, or recall at fixed precision over plain accuracy
- Regression: choose MAE (robust to outliers) vs RMSE (penalises large errors) deliberately
- Ranking: use MAP, NDCG, or MRR depending on position sensitivity
- Note the difference between the **offline proxy metric** (what can be measured from historical data) and the **online KPI** (what the business ultimately cares about); acknowledge any gap

**State the metric before any modeling begins. It must not change after training.**

### 5. Define a baseline and success threshold

- **Baseline:** the simplest defensible reference — majority-class prediction, median value, current rule-based system, or the model already in production
- **Success threshold:** the quantitative improvement over baseline that is meaningful to the business (e.g. "+3 pp F1 on the held-out set" or "<10 ms p99 latency with no accuracy regression")
- A threshold of "as high as possible" is not acceptable — make it concrete

### 6. List constraints

Document non-negotiable constraints that shape model choice and architecture:
- **Latency / throughput:** real-time vs batch; maximum inference time
- **Memory / hardware:** CPU-only, GPU budget, edge deployment
- **Interpretability:** regulatory or stakeholder requirements for explanations
- **Data volume:** row count, feature count, labelling cost
- **Compliance:** PII handling, data residency, auditability requirements

### 7. Ask one clarifying question at a time when ambiguous

If any of the above cannot be answered from the ticket, ask exactly **one question**, wait for a reply, and continue. Do not bundle multiple questions.

---

## Completion Checklist

- [ ] Ticket context read (description + comments + related files)
- [ ] Prediction target defined with the exact availability time of input features
- [ ] Problem type classified (classification / regression / ranking / generation / clustering)
- [ ] Evaluation metric chosen and justified against the business goal
- [ ] Offline proxy metric vs online KPI gap acknowledged
- [ ] Baseline defined (majority class, heuristic, or current production model)
- [ ] Success threshold quantified (not "as high as possible")
- [ ] Constraints listed (latency, memory, interpretability, data volume, compliance)
- [ ] Findings written to `AK-Docs/04.Coding/01.Requirements/[functionId]/[ticketId].md` (see `custom/rules/ml-conventions.md` — do NOT use the legacy `plan/[ticket-id]/` path)

Output language: auto-detect from the ticket/task input — see `custom/rules/output-language.md` (Vietnamese input → Vietnamese output; otherwise English).
