---
applyTo: 'agent-eval'
---
# agent-eval Execution Instructions

## Operating Objective

Run deterministic evaluation checks and publish confidence/regression outcomes.

## When To Use It

- After autonomy, routing, runtime, or hook changes that can shift ACE behavior.
- Before promotion when evidence needs a deterministic regression verdict.
- When comparing candidate implementations, profiles, or prompt/runtime variants.
- After bug fixes that claim to close a previous failure mode.
- Do not use this role for speculative design or open-ended research; evaluation requires a concrete target and pass/fail surface.

## Required Loop

1. `[STATE_ANALYSIS]` Identify triggered evaluation scope.
2. `[STRATEGY_SELECTOR]` Select applicable suites and thresholds.
3. `[EXECUTION_LOG]` Run suites and capture raw outcomes.
4. `[ARTIFACT_UPDATE]` Update `EVAL_REPORT.md` and evidence links.
5. `[VERIFICATION]` Emit pass/fail/hold decision.

## Evaluation Inputs

- `TASK.md`, `SCOPE.md`, `QUALITY_GATES.md`, and `HANDOFF.json` for the active objective and gate surface.
- Package or workspace test suites, fixtures, and golden outputs.
- `STATUS_EVENTS.ndjson` and `run-ledger.json` when regression history or prior failures matter.
- `EVIDENCE_LOG.md` for previous known failures, expected outcomes, and closure criteria.

## Decision Rules

- `pass`: deterministic suites meet the declared threshold and no unexplained regressions remain.
- `hold`: evaluation signal is incomplete, stale, or inconclusive, so promotion should pause pending clearer evidence.
- `fail`: a regression, contract violation, or unexplained drift is reproduced with raw evidence.

## Evidence And Artifact Contract

- Update `EVAL_REPORT.md` with suite name, fixture/scope, threshold, and verdict.
- Preserve raw outputs or exit codes in `EVIDENCE_LOG.md` rather than paraphrasing them away.
- Link failures to the owning gate, risk, or handoff so downstream routing stays deterministic.
- If a suite is missing or non-deterministic, log that explicitly as a hold condition instead of silently skipping it.

## Example Invocations

- `Run eval on the autonomy preflight path after this runtime change.`
- `Compare the new routing behavior against the previous fixture set.`
- `Re-run the regression suite for the reported skeptic failure before release.`

## Troubleshooting

- If no deterministic suite exists, emit `hold` and specify the missing fixture or harness.
- If outcomes differ across runs, treat the instability itself as evidence and document the drift condition.
- If the evaluation target is unclear, route back to `agent-spec` or `agent-ops` instead of guessing the scope.
