---
name: design-experiment
description: Design a leakage-safe validation strategy and experiment plan with ablations and tracking before training begins.
keywords: experiment, validation, cross-validation, ablation, tracking, splits
---

# Design Experiment

## When to use this skill

- Gate 2, after the problem framing document (`ml-problem.md`) has been APPROVED
- Any time a new round of experiments is planned after a framing change
- Before writing any model training code — the plan must exist first

---

## Steps

### 1. Choose the validation strategy

Select the CV scheme that matches the data structure and justify the choice explicitly:
- **Stratified k-fold:** use when the target is imbalanced; preserves class proportions in every fold
- **Group k-fold / GroupShuffleSplit:** use when rows belong to named entities (users, patients, stores) and the task requires generalising to unseen entities — random splits would let the same entity appear in train and validation
- **TimeSeriesSplit:** use for temporal data; always respect chronological order; never shuffle before splitting

State the number of folds (or holdout fraction) and why that choice is appropriate given the dataset size and variance of the metric.

### 2. Define the leakage-safe split protocol

Specify the exact sequence of operations to prevent any information from the validation or test set contaminating the training fold:
- **Split before fit:** the train/validation boundary is established before any preprocessing step is fitted
- **Pipeline inside CV:** all transformations (scaling, encoding, imputation, feature selection) are wrapped in a `Pipeline` (sklearn) or equivalent so they are fit on each training fold and applied to the validation fold — never fit on the combined data
- **Test set quarantine:** the held-out test set is touched only at Gate 4 evaluation; it is never used to tune hyperparameters or select features

### 3. Enumerate candidate approaches and feature sets

List experiments in ascending order of complexity:
1. **Baseline** (from `frame-ml-problem`): the simplest reference (majority class, median, or current rule)
2. **Simple model** (e.g. logistic regression, linear regression, decision tree) with a minimal feature set
3. **Intermediate model** (e.g. gradient-boosted trees) with engineered features
4. **Complex model** (e.g. neural network, fine-tuned transformer) if warranted by the problem

For each approach, list the feature set explicitly — which raw columns, which engineered features. Separating feature sets from model families makes it possible to attribute performance differences cleanly.

### 4. Define the ablation plan

Describe what each experiment isolates. The rule is **one change at a time**:
- Vary the model family while holding the feature set constant
- Add a feature group while holding the model constant
- Vary a key hyperparameter while holding everything else fixed
- Document the hypothesis for each ablation (what do you expect and why)

A plan with no ablations is not acceptable — it produces no actionable insight.

### 5. Set up experiment tracking

Define what every tracked run must log before any experiment is executed:
- **Parameters:** all hyperparameters, including their defaults; the feature set identifier; the validation scheme
- **Metrics:** primary metric (from Gate 1) per fold + mean ± std; secondary metrics if relevant
- **Reproducibility metadata:** random seed, data version (hash or DVC tag), git commit SHA
- Agree on the tracking tool (MLflow, wandb, or simple CSV) and the run-naming convention from `ml-conventions.md` (`[ticket-id]_[approach]_[yyyymmdd-n]`)

### 6. Define the stopping and decision rule

State, before any experiment runs, how you will decide which approach moves forward:
- Which metric on which split is the decision criterion
- What constitutes a meaningful improvement over the baseline (link to the success threshold from Gate 1)
- The budget: maximum number of experiments or wall-clock time before a decision is forced

An open-ended experiment loop with no stopping rule is out of scope for a gated workflow.

### 7. Output the experiment plan

Write all of the above to `AK-Docs/04.Coding/02.Plans/[functionId]/[ticketId].md` (see `custom/rules/ml-conventions.md` — do NOT use the legacy `plan/[ticket-id]/` path). The document must be self-contained: a developer who did not attend the framing discussion should be able to reproduce the full experiment sequence from the plan alone.

Output language: auto-detect from the ticket/task input — see `custom/rules/output-language.md` (Vietnamese input → Vietnamese output; otherwise English).

---

## Completion Checklist

- [ ] Validation scheme chosen (stratified k-fold / group k-fold / time-series split) and justified
- [ ] Leakage-safe split protocol defined (split before fit; transforms inside pipeline; test set quarantined)
- [ ] Candidate approaches listed in ascending complexity order
- [ ] Feature sets enumerated per approach
- [ ] Ablation plan written (one change at a time; hypothesis stated for each ablation)
- [ ] Experiment tracking configured (params, metrics, seed, data version, git commit)
- [ ] Run-naming convention agreed and follows `ml-conventions.md`
- [ ] Stopping and decision rule stated (metric, threshold, experiment budget)
- [ ] Full plan written to `AK-Docs/04.Coding/02.Plans/[functionId]/[ticketId].md` (see `custom/rules/ml-conventions.md` — do NOT use the legacy `plan/[ticket-id]/` path)
