# Experiments

One workflow, M models, N repetitions, one table. This is the layer the
[`model` knob](workflows.md) exists for: comparing models on the same work
requires the model to be an argument, and comparing them honestly requires the
run to be repeated.

```typescript
import { experiment, experimentTable, loop } from "@ai-for-dev/combo";

const report = await experiment({
	models: ["ilaas/gemma-4-31b", "ilaas/gpt-oss-120b"],
	repetitions: 2,
	run: async (cell) => {
		const result = await loop({ ...cell.options, steps: [coder, reviewer], input, until: lgtm });
		return { ok: result.ok, converged: result.converged, iterations: result.iterations };
	},
});

console.log(experimentTable(report).join("\n"));
```

That is the shape of [`examples/12-experiment.ts`](#running-one), and this is
the table it printed on those two models:

```
| model | runs | ok | converged | iterations | usage | mean wall | mean $ |
| --- | --- | --- | --- | --- | --- | --- | --- |
| ilaas/gemma-4-31b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 368.6s ↑30k ↓21k | 184.3s | not reported |
| ilaas/gpt-oss-120b | 2 | 2/2 | 2/2 | 1×2 | 4 turns 47.9s ↑53k ↓5.8k | 24.0s | not reported |
```

`1×2` is two runs of one iteration each. The provider reports no cost, so the
usage has no `$` figure and `mean $` says `not reported`: a missing figure is
not a zero.

## An experiment is a function, not a combinator

It returns no `Result` and composes with nothing. It is a harness placed *above*
a workflow, and one that could be nested inside a workflow would be measuring
itself. Everything else in this library is a combinator precisely because it can
be nested; this one is deliberately not.

Which also means a flow needs no special support: `runFlow` takes the same
`model`, `signal`, `timeoutMs`, `spawn` and `onEvent`, and the cell's directory
is its run directory.

```typescript
run: async (cell) => {
	const result = await runFlow(checked, input, { ...cell.options, runDir: cell.dir });
	return { ok: result.ok };
},
```

## The contract: spread `cell.options`

```typescript
type ExperimentCell = {
	model: string;        // this cell's model
	repetition: number;   // 1-based, matches rep-<n>/ on disk
	dir: string;          // this cell's directory, absolute, already created
	options: WorkflowOptions;
};
```

`cell.options` carries the cell's `model` and `exportDir`, the experiment's
`signal`, `timeoutMs`, `cwd` and `spawn`, and an `onEvent` combining the cell's
private collector with your own listener. **Spreading it is the contract**: a
callback that rebuilds those by hand puts its subagents on the wrong model, in
the wrong directory, and measures nothing.

What the callback returns becomes the table's columns:

```typescript
type ExperimentOutcome = { ok: boolean; error?: string }
	& Record<string, string | number | boolean | undefined>;
```

Flag columns are the union of the outcome keys actually seen - `converged`,
`approved`, `rounds`, whatever this study compares - so there is nothing to
configure. `ok` has its own column and `error` never becomes one: a column of
distinct sentences compares nothing.

## On disk

```
runs/2026-09-24_00-45-16/
├── experiment.json                machine-readable, every cell
├── experiment.md                  the table, plus the failures named under it
├── ilaas-gemma-4-31b/
│   ├── rep-1/                     pi's transcripts per subagent
│   │   ├── usage.json             time and tokens, attributed
│   │   └── events.jsonl           the cell's whole event stream, in order
│   └── rep-2/
└── ilaas-gpt-oss-120b/
    └── …
```

Each cell writes the same [`usage.json`](export.md) a single run writes, from
its own collector. Measurement is reused, never reinvented.

`events.jsonl` is the [record reporter](display.md#keeping-the-stream), wired
into every cell with no way to turn it off: a cell whose stream was not kept can
only be re-run, and a matrix is expensive.

## The rules

- **Sequential by default.** `concurrency` defaults to 1: two cells racing for
  the same machine measure the contention, not the models. Raise it when the
  providers are remote and the wall time matters more than the precision.
- **Model-major order.** Every repetition of the first model, then the second -
  so a matrix interrupted halfway holds finished models rather than a fragment
  of each.
- **A failed cell stays in the report**, with its usage: it spent tokens before
  it broke, and dropping it would quietly turn "two models out of three
  answered" into a clean comparison of the survivors. A callback that *throws*
  is a failed cell too, not a crashed experiment.
- **Sums are stored, means are displayed.** `experiment.json` carries totals;
  the mean wall and mean cost are computed when the table is rendered, never
  written down. Averaging averages is how a study starts lying about itself.
- **An abort stops launching new cells** and the partial report is still
  written, with `error: "aborted"` at the top.

## Running one

```bash
node examples/12-experiment.ts <modelA> <modelB>
```

Two repetitions of the same loop per model, the table printed at the end. Several
providers report no cost, and some report no tokens either - see
[Measurements](measurements.md) - so a usage with no `$` figure, and a
`mean $` of `not reported`, mean what they say, never "free". Give the matrix a `timeoutMs` it can live with: a cell lost to a
turn that would not end is a cell missing from the comparison.

## Reference

- [`measure/experiment`](../reference/api/measure/experiment.md) - `experiment`, `ExperimentCell`, `ExperimentOptions`.
- [`measure/report`](../reference/api/measure/report.md) - the report, the table, the writes.
