# Comparison Report

## Comparison Gate

- Scenario:
- Identical full prompt confirmed for every mode: yes / no
- Same clean baseline and reset method confirmed: yes / no
- Required metadata present (host, model, Heli version, prompt, baseline, reset method, applicable metrics): yes / no

Do not score report completeness across modes unless the identical-prompt gate is yes. If either comparison gate is no, mark the comparison INVALID RUN.

## Runs Compared

| Mode | Run ID | Prompt match | Baseline/reset match | Applicable category average | Verdict |
|---|---|---|---|---|---|
| A | | | | | |
| B | | | | | |
| C | | | | | |
| D | | | | | |

## Evidence Comparison

| Metric | Mode A | Mode B | Mode C | Mode D | Notes |
|---|---|---|---|---|---|
| | | | | | |

| Category | Mode A | Mode B | Mode C | Mode D | Applicable metrics only |
|---|---|---|---|---|---|
| | | | | | |

## Observed Evidence

- What occurred in each mode:
- Differences supported by run logs and scorecards:
- Unavailable hooks, probes, or metrics (Not applicable):

## Limitations and Recommendation

- Limitations:
- Recommendation supported by observed evidence only:
