---
name: measure
description: Use when the user asks whether a prompt change improved anything, wants to compare two versions, asks if a suite passed, wants the cost of a version, or needs to see what was actually sent to the model.
---

# Measuring a change

Building and running a prompt is half the loop. This file is the other half: deciding whether the
edit made things better, and seeing why when it did not.

## The loop

1. Build or edit the canvas / workflow / agent.
2. Run the **whole** suite against the version.
3. Review the results attribute by attribute — see `test-case-review.md`.
4. **Measure** the version.
5. Compare against the previous version.
6. Promote only if it actually improved.

Steps 4 and 5 are what stop "I changed the prompt and it feels better" from being the whole story.

## Run the whole suite, not a slice

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const cases = await tela.runSuiteAndWait('PROMPT_ID', 'VERSION_ID')
console.log(cases.length + ' cases settled')
"
```

`runSuiteAndWait` runs every test case on the prompt against that version and waits for all of them.
Comparing two versions is only meaningful when both ran the same suite, so do not hand-pick ids
between runs. Narrow deliberately with `filters` when you mean to:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
await tela.runAllTestCases('VERSION_ID', { filters: { tagIds: ['TAG_ID'] } })
"
```

## The scoreboard for one version

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const s = await tela.getVersionStats('VERSION_ID')
console.log(JSON.stringify(s, null, 2))
"
```

```json
{
  "stats":         { "total": 4, "good": 3, "bad": 1, "pending": 0 },
  "reviewedStats": { "total": 4, "good": 3, "bad": 1 },
  "testCases":     { "total": 2, "completed": 2, "pendingReview": 0, "partiallyReviewed": 0 }
}
```

`stats` counts **attributes**, not test cases. `pending` is the number nobody has judged — while it
is above zero the version has not really been measured, and any score you quote is drawn from a
partial sample.

`getGenerationStats(versionId)` counts something **different**: the overall thumb on each generation
(`generation.feedback`), not the per-attribute verdicts. Nothing in this skill writes that field, so
it will usually disagree with `getVersionStats`. Use `getVersionStats` for quality; reach for
`getGenerationStats` only when you specifically want the whole-generation signal.

## Comparing versions

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const comparison = await tela.compareVersions(['V1_ID', 'V2_ID'], {
  titles: { V1_ID: 'v1 vague', V2_ID: 'v2 explicit' },
})
console.log(tela.formatVersionComparison(comparison))
"
```

```
version                              good  bad  pending  score
v1 vague                                2    2        0    50%
v2 explicit                             4    0        0   100%

verdict: better (+50 points)
```

The **first** id is the baseline and the **last** is the candidate. `score` is
`good / (good + bad)`.

The endpoint answers for one version at a time — passing several ids to the API returns a single
aggregate, not a breakdown — so `compareVersions` fetches each and assembles the table itself.

`verdict` is `inconclusive` unless **both** sides are fully reviewed. A score computed over a
partially reviewed version describes only the attributes someone happened to look at: one good result
and ninety-nine unjudged ones scores 100%, which is not a fact about quality. `coverage` on each row
says how much was actually reviewed.

Report inconclusive as inconclusive — do not fill the silence with a guess. The fix is to finish
reviewing, then compare again, and `formatVersionComparison` names which side is incomplete so you
can tell the user exactly what is left.

## Before you ship: the readiness gate

**Do not promote a version, and do not create Workstation tasks in bulk, on a suite that has not been
reviewed.** Production tasks spend real money producing output nobody has checked, and "every run
returned valid JSON" is not a quality signal — it says the shape parsed, not that the content is
right.

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const readiness = await tela.getVersionReadiness('VERSION_ID')
console.log(tela.formatVersionReadiness(readiness))
if (!readiness.ready) process.exit(1)
"
```

```
Not ready to ship:
- 42 attribute(s) are still unreviewed
```

Run it before `promoteCanvasVersion`, before `upsertWorkstation`, and before a batch of `createTask`.
When it reports blockers, say them to the user and stop — do not proceed and mention the caveat
afterwards.

`minScore` defaults to 1, meaning every reviewed attribute must be good. Lower it deliberately and
say what you lowered it to; do not lower it to make a version pass.

Readiness also blocks on test cases that **produced no result** — never run, still running, or
failed. They contribute no attributes, so a suite where two of five cases never ran would otherwise
show zero pending and read as fully reviewed.

## Changing the model

"Cheaper without losing quality" is two questions, and answering only the cost half is the most
common way to get this wrong. Cost is easy to measure and quality is not, so the pull is to report
the cheap answer and hedge on the rest.

The order that works:

1. **List the real catalog.** Do not answer from memory — see `models.md`. The models available to a
   canvas and to an agent are different lists with differently shaped ids.
2. **Fork a candidate version** with the new model. Never edit the promoted one.
3. **Run the same suite** on both, with `runSuiteAndWait`. A partial re-run makes the comparison void.
4. **Review the attributes on both.** Without this there is no quality axis and the comparison is
   about nothing.
5. **Compare**, and report cost and score together.
6. **Promote only if the candidate holds quality**, and only after `getVersionReadiness` passes.

`getUsageCost` and `getUsageCostByModel` read the usage service, which is the billing source of
truth.

**Never reimplement the credit formula, and never sum `creditsUsed` off completion runs.** Both
produce a number that looks like a measurement and is not. If the user asks what a different model
would cost, run the candidate and read the usage service afterwards. Provider list prices are a fair
basis for a rough projection, as long as you say that is what it is.

## Diagnosis: what was actually sent

**Test case runs cannot be diagnosed this way.** A canvas test case generation records its inputs and
its output but not the prompt the model received — `resolvedInput`, `promptInput`, and
`inputMessages` all come back empty — and a test case run does **not** create a completion run
either. Verified: a canvas whose only executions were test cases reports zero completion runs on
every version.

So for a test case, what the model was actually given is not retrievable through the API. Reconstruct
it from the version's message plus the generation's `variables`, and say that is what you did rather
than presenting it as the real payload.

Completion runs cover **Workstation tasks and direct API completions**. For those, they are the full
picture:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const { data } = await tela.listCompletionRuns({ promptVersionId: 'VERSION_ID', limit: 5 })
for (const run of data) {
  console.log(run.status, run.creditsUsed, 'credits')
  console.log('  sent  :', JSON.stringify(run.rawInput).slice(0, 200))
  console.log('  got   :', JSON.stringify(run.rawOutput).slice(0, 200))
  console.log('  model :', run.metadata?.promptVersion?.modelConfigurations?.model)
}
"
```

A generation cannot be joined to a run by id: `generation.metadata.executionId` is a Trigger.dev run
id, not a completion run. Match on `promptVersionId` and timestamp.

Note the route lives under `/v1/`, unlike most of the API.

For workflows, step-level outputs come from `getWorkflowRun` and **do** cover test case runs. For
agents, use `getAgentSessionThread` / `getAgentSessionTimeline`. Canvas test cases are the one blind
spot.

## Cost

**Cost comes from the usage service, never from a completion run.** `creditsUsed` and
`usage.cost` on a run are execution-time artifacts: they are not reconciled and are not what the
workspace is billed. Summing them produces a number that looks authoritative and is not.

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const end = new Date().toISOString()
const start = new Date(Date.now() - 30 * 864e5).toISOString()

console.log(await tela.getUsageCost({ canvasId: 'CANVAS_ID', start, end }))
console.log(tela.formatModelUsage(await tela.getUsageCostByModel({ canvasId: 'CANVAS_ID', start, end })))
"
```

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
console.log(await tela.getCurrentUsage('WORKSPACE_ID'))   // { month, cost, credits }
"
```

### Usage replicates about a minute late

A run's usage lands in the query layer roughly **60 seconds** after it finishes. Ask for cost
immediately after running something and you get an empty result — which reads exactly like "it was
free".

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const startedAt = new Date().toISOString()
// ... run the suite or the batch ...
const { events, complete } = await tela.waitForUsageEvents({ canvasId: 'CANVAS_ID', start: startedAt }, { minEvents: 31 })

if (!complete)
  console.log('usage has not finished replicating — do not quote a cost yet')
else
  console.log(tela.formatModelUsage(tela.groupUsageByModel(events)))
"
```

**An empty result right after a run is not a zero cost.** Say "not replicated yet" and wait, or come
back to it. Reporting no cost because the query returned nothing is the same class of mistake as
reporting a suite as passing because it executed.

Three properties of this data shape what you can ask:

- **Usage is attributed to the canvas, not to a version.** There is no `promptVersionId` on a usage
  event. To compare what two versions cost, separate them by **time window**, or by `model` when that
  is what changed — `getUsageCostByModel` exists for exactly that.
- **Always bound the window.** The store is time-partitioned; an unbounded query is slow and prone to
  failing.

Each event carries `cost` (what the provider charged) and `effectiveCost` (what the workspace is
billed, after the multiplier). Quote the one the user asked about, and say which it is.

Completion runs remain the right place to see **what was sent to the model** — just not what it cost.

## Per kind

| | Canvas | Workflow | Agent |
|---|---|---|---|
| Version scoreboard | `getVersionStats` | `getVersionStats` | `getAgentTestCaseStats(agentId, commitHash)` |
| Compare | `compareVersions` | `compareVersions` | `listAgentTestCaseRuns` across commits |
| Named metrics over attributes | — | — | `listAgentTestMetrics` |
| What was sent to the model | tasks only — **not test cases** | `getWorkflowRun` (step outputs) | `getAgentSessionThread` |
| Cost | `getUsageCost` / `getUsageCostByModel` | same | same (usage events carry `agentSessionId`) |

Agent versions are commits, so the agent side compares by commit hash rather than version id. The
shape of the question is the same; the key is not.

## Reporting honestly

- If `pending > 0` on either side, the comparison is inconclusive by construction. Say how many
  attributes are unjudged and on which version, rather than quoting the partial score as a result.
- If `verdict` is `inconclusive`, say so rather than picking the version with more `good`.
- If the two versions ran different suites, the comparison is void. Re-run both.
- **Spot-check the automatic verdicts.** An LLM-judged `good`/`bad` counts as a review, and the judge
  is wrong sometimes — it has been seen marking a semantically correct answer as bad. When a score
  decides whether something ships, read the failing attributes yourself before quoting it.
- If asked why a test case answered the way it did, say the resolved prompt is not retrievable for
  test case runs and offer the reconstruction, rather than implying you read the real payload.


## Close with the link

A comparison is more persuasive next to the thing it describes. End with the version you are
recommending, and when a specific step or run is at fault, link straight to it:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
console.log(tela.formatLinks(tela.getWorkflowLinks('PROMPT_ID', {
  promptVersionId: 'VERSION_ID',
  executionId: 'RUN_ID',
  nodeId: 'FAILING_NODE_ID',
})))
"
```

This matters most exactly where the skill is blind: you cannot retrieve the resolved prompt behind a
canvas test case, but the user can see it in the app. Send them there instead of guessing.
