# Eval Framework Schemas

JSON schemas for the skill-builder eval framework. These define the data structures used throughout the eval → grade → benchmark pipeline.

## evals.json

Defines the test cases for a skill. Lives at `{skill-dir}/evals/evals.json`.

```json
{
  "evals": [
    {
      "id": "string — short descriptive identifier, e.g. 'basic-usage'",
      "prompt": "string — the exact user message to test with",
      "files": {
        "optional — map of filename → content for files the eval needs present",
        "src/input.txt": "Example file content"
      },
      "assertions": [
        "string — verifiable expectation about the output",
        "Each assertion should be specific enough to grade as PASS/FAIL",
        "Good: 'Creates a file matching *.test.ts'",
        "Bad: 'Output is good quality'"
      ]
    }
  ]
}
```

### Assertion Writing Guidelines

Good assertions are:
- **Specific**: "Creates a file ending in .tsx" not "Creates a component"
- **Verifiable**: Can be checked by reading outputs, not by subjective judgment
- **Independent**: Each assertion tests one thing
- **Non-trivial**: Tests behavior unique to the skill, not generic capabilities

Avoid:
- Subjective assertions ("well-structured", "clean code", "good quality")
- Compound assertions ("Creates a file AND it has proper imports AND it compiles")
- Tautological assertions ("Produces output" — of course it does)

## timing.json

Captures resource usage for a single eval run. Created by the test runner and saved alongside outputs.

```json
{
  "eval_id": "basic-usage",
  "config": "with_skill | without_skill",
  "total_tokens": 4523,
  "duration_ms": 12450,
  "timestamp": "2025-01-15T10:30:00Z"
}
```

Fields:
- `total_tokens`: Combined input + output tokens for the run
- `duration_ms`: Wall clock time from prompt submission to final output
- `timestamp`: ISO 8601 timestamp of when the run started

## grading.json

Output from the grader agent. One per eval per configuration.

```json
{
  "eval_id": "basic-usage",
  "config": "with_skill | without_skill",
  "expectations": [
    {
      "assertion": "The original assertion text",
      "verdict": "PASS | FAIL",
      "evidence": "Specific quote or file reference supporting the verdict",
      "confidence": "HIGH | MEDIUM | LOW"
    }
  ],
  "pass_count": 3,
  "fail_count": 1,
  "pass_rate": 0.75,
  "claims": [
    {
      "claim": "An implicit claim extracted from the output",
      "verified": true,
      "evidence": "How the claim was verified"
    }
  ],
  "eval_feedback": {
    "weak_assertions": [
      "Assertions that are too vague to meaningfully grade"
    ],
    "missing_assertions": [
      "Important behaviors not tested"
    ],
    "suggested_additions": [
      "New assertions that would improve coverage"
    ]
  },
  "summary": "Brief human-readable summary of grading results"
}
```

## benchmark.json

Aggregated results across all evals for an iteration or multi-run benchmark.

```json
{
  "iteration": 1,
  "timestamp": "2025-01-15T10:30:00Z",
  "configs": {
    "with_skill": {
      "overall_pass_rate": 0.85,
      "total_tokens_mean": 4500,
      "total_tokens_stddev": 300,
      "duration_ms_mean": 12000,
      "duration_ms_stddev": 1500,
      "evals": [
        {
          "eval_id": "basic-usage",
          "pass_rate": 1.0,
          "pass_count": 3,
          "fail_count": 0,
          "tokens": 4200,
          "duration_ms": 11000
        }
      ]
    },
    "without_skill": {
      "overall_pass_rate": 0.60,
      "total_tokens_mean": 3200,
      "total_tokens_stddev": 250,
      "duration_ms_mean": 8000,
      "duration_ms_stddev": 1000,
      "evals": [
        {
          "eval_id": "basic-usage",
          "pass_rate": 0.67,
          "pass_count": 2,
          "fail_count": 1,
          "tokens": 3100,
          "duration_ms": 7500
        }
      ]
    }
  },
  "comparison": {
    "pass_rate_delta": 0.25,
    "token_overhead_percent": 40.6,
    "time_overhead_percent": 50.0,
    "non_discriminating_assertions": [
      "Assertions that pass in both configs"
    ]
  }
}
```

## feedback.json

User feedback collected from the HTML review viewer. Saved to the workspace root.

```json
{
  "timestamp": "2025-01-15T11:00:00Z",
  "iteration": 1,
  "eval_feedback": [
    {
      "eval_id": "basic-usage",
      "rating": "good | needs-work | bad",
      "comment": "Free-form user feedback on this eval's results"
    }
  ],
  "general_feedback": "Overall notes on the skill's performance",
  "action": "iterate | publish | stop"
}
```

## Directory Layout Reference

```
{skill-dir}/
├── SKILL.md
├── evals/
│   └── evals.json
├── agents/                    # Only if skill uses subagents
│   └── *.md
├── references/                # Supporting documentation
│   └── *.md
├── scripts/                   # Bundled deterministic scripts
│   └── *.py / *.js
└── assets/                    # Templates, HTML, configs
    └── *

{slug}-workspace/              # Created during testing, not committed
├── iteration-1/
│   ├── eval-{id}/
│   │   ├── with_skill/
│   │   │   ├── outputs/
│   │   │   ├── timing.json
│   │   │   └── grading.json
│   │   └── without_skill/
│   │       ├── outputs/
│   │       ├── timing.json
│   │       └── grading.json
│   ├── eval_metadata.json
│   └── benchmark.json
├── iteration-2/
│   └── ...
├── benchmark.json             # Multi-run statistical results
└── feedback.json              # User feedback
```
