---
name: test-case-review
description: Use when the user wants to evaluate, review, give feedback on, comment on, or set acceptance criteria for test case results — for a canvas, workflow, or agent — or asks whether a test suite actually passed.
---

# Test cases and reviewing their results

## Executed is not reviewed

This is the single most important thing in this file. A test case run that finishes reports
`completed`. So does a fully reviewed one. So does one that had **nothing to review**. The product
has no separate `reviewed` status, so "the run finished" and "a human checked it" collapse onto the
same word.

Never tell a user their suite passed because the runs completed. Either review the attributes, or say
plainly that the cases executed and were not reviewed.

Use `getTestCaseReview` to tell the three apart:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const testCase = await tela.getTestCase('TEST_CASE_ID')
const generation = tela.getLatestGeneration(testCase)
const version = await tela.getCanvasVersion(generation.promptVersionId)

const review = tela.getTestCaseReview(testCase, version)
console.log(review.status)         // pending_review | partially_reviewed | completed | failed | …
console.log(review.pending)        // attributes still awaiting a verdict
console.log(review.unreviewable)   // true => nothing to review; 'completed' means nothing
"
```

`unreviewable: true` means the prompt version declares no reviewable attributes. That is the state to
call out to the user — it is what makes a suite look green when nobody has looked at anything.

**`testCase.generations` is not in chronological order.** Re-running leaves the superseded generation
in the array, sometimes last, so `generations.at(-1)` can hand you an `aborted` run. Always use
`getLatestGeneration` — writing feedback onto the wrong generation records a review nobody will see.

### What makes an attribute reviewed

An attribute counts as reviewed when it has either:

- **manual** feedback — `attributeFeedback[key]` is `1` or `0`, or
- an **inferred** verdict — `validations[key].type` is `good` or `bad`.

`warning` and `missing` do not count. Manual feedback overrides the inferred verdict.

### Where the reviewable attributes come from

The prompt version's `configuration.structuredOutput.schema.properties`, flattened to dot-paths.

- Nested objects flatten: `{ address: { city } }` → `address.city`. Only leaves are keys.
- An array is a **single** key (`ingredients`), not one key per element.
- **A canvas with no structured output has no reviewable attributes at all.** Its test cases can be
  executed and never reviewed. If a user wants reviewable results, the canvas needs a
  `structuredOutput` first — see `create-canvas.md`.
- A **workflow** whose schema is empty falls back to the fields present in the generation content, so
  its attribute keys are derived per-run and can shift between runs.

## Reviewing: thumbs per attribute

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
await tela.updateAttributeFeedback('GENERATION_ID', 'vendor', 1)   // 1 up, 0 down, null clears
"
```

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
await tela.updateAttributeFeedbackBulk('GENERATION_ID', [
  { attributeKey: 'vendor', feedback: 1 },
  { attributeKey: 'total', feedback: 0 },
])
"
```

This one call does three things: records the verdict, logs who reviewed it, and files the generation's
value into the test case's **answer bank** (`answers[key].good` / `.bad`), which auto-classifies
matching values on future runs. Do not write `attributeFeedback` through `PATCH /generation/:id` —
that path skips all three.

Only rate the attributes the user actually spoke about. Leaving the rest pending is correct and
visible; guessing a verdict is not.

Two constraints worth knowing:

- The endpoint parses the generation's content unguarded, so it **500s on a generation whose content
  is empty or not JSON**. Only review finished runs.
- Recording review needs a user-bearing token. Under a workspace or API-key token the feedback still
  lands but the reviewer is anonymous.

## Acceptance criteria: validation rules

A rule is a human-written criterion attached to one attribute — and it is **not** just documentation.
Every subsequent run sends it to an LLM judge, which writes a verdict into `validations[key]`. That
verdict counts as a review.

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
await tela.addValidationRule('TEST_CASE_ID', 'total', 'The total must be formatted as Brazilian currency.')
await tela.addValidationRule('TEST_CASE_ID', 'vendor', 'The vendor must be the legal entity name, not the trading name.')
"
```

Write them as checkable statements. A rule that restates the field name teaches the judge nothing.

Note the side effect: **adding a rule suppresses the automatic exact-match "bad" verdict** for that
attribute, handing judgment to the LLM instead. Remove with
`deleteValidationRule(testCaseId, attributeKey, ruleId)`.

`expectedOutput` on a test case is inert — nothing reads it. Encode expectations as validation rules.

## Comments

Free text on a result, for the explanation a thumbs-down cannot carry. Comments take no part in the
review status.

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
await tela.createGenerationComment({
  content: 'Total is reading the subtotal line, not the total line.',
  generationId: 'GENERATION_ID',
  promptVersionId: 'VERSION_ID',
})
console.log(await tela.listGenerationComments('TEST_CASE_ID'))
"
```

`listGenerationComments` reads the comments off the test case, because the dedicated list endpoint
is broken: it requires an author filter and returns only *deleted* comments. Do not call
`/generation-feedback` directly to read.

## Tags

A tag belongs to exactly one canvas or workflow — there are no workspace-wide tags. `color` is
required and unvalidated; names are not unique.

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const tag = await tela.createTestCaseTag({ name: 'regression', color: '#4F46E5', promptId: 'PROMPT_ID' })
await tela.addTestCaseTags('TEST_CASE_ID', [tag.id])
await tela.bulkAddTestCaseTags(['TC_1', 'TC_2'], [tag.id])
console.log(await tela.getTestCaseTags('TEST_CASE_ID'))
"
```

The assignment calls return no useful body — re-read `getTestCaseTags` to confirm.

## Per kind

| | Canvas | Workflow | Agent |
|---|---|---|---|
| Entity | `/test-case` | **the same** `/test-case` | separate `/agent/:id/tests` |
| Reviewable attributes | `structuredOutput` schema | schema, falling back to run content | `output-format.json` |
| Feedback call | `updateAttributeFeedback(generationId, attributeKey, feedback)` | same | `updateAgentTestCaseAttributeFeedback(agentId, testId, commitHash, [{ path, feedback }])` |
| Run identity | generation | generation | `(testCaseId, commitHash)` — one run per agent version |
| Validation rules | yes | yes | no |
| Comments | yes | yes | no |
| Abort a run | yes | yes | **no** |
| Metrics over attributes | no | no | yes |
| Tags | yes | yes | yes (same tables, agent-scoped) |

The agent payload differs in three ways at once: the wrapper key is `attributes`, the field is `path`
rather than `attributeKey`, and index variants collapse (`items[0].name` → `items.name`). Do not
carry the canvas shape over.

## Listing test cases

Always pass `promptVersionIds` when you intend to reason about review state:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
const { data } = await tela.listTestCases('PROMPT_ID', { promptVersionIds: ['VERSION_ID'] })
console.log(data.length)
"
```

Without it, every row's `status` silently degrades to the raw generation status and the
reviewed/unreviewed distinction disappears entirely.

`limit` and `offset` are accepted and ignored — `meta` reports them but `data` is not paginated.


## Close with the link

Reviewing is something a human does better in the app, where the diff and the attachments render.
Whenever you report on test cases, end with the test tab:

```bash
bun --preload ~/.claude/skills/tela-studio/preload.ts -e "
console.log('Review them here: ' + tela.getCanvasTabUrl('PROMPT_ID', 'test'))
"
```

For a workflow, use `getWorkflowTabUrl(promptId, { tab: 'test', promptVersionId })` so the page opens
on the version you reviewed. There is no per-test-case URL — name the case in the text.
