---
name: inspect-pi-coding-agent-eval
description: >-
    Inspect and debug pi-coding-agent-eval runs, manifests, trial artifacts, reports, metrics, and pairwise comparisons.
    Use this skill whenever an eval result looks wrong, a profile comparison needs explanation, a live run must be
    examined, or a report needs verification. Prefer recorded artifacts and offline report rebuilding before rerunning
    a model. Use the pi-coding-agent-test inspection skill for terminal frames and TUI details.
---

# Inspect Pi evaluations

Use this skill after the evaluator has produced a run directory. Start with artifacts, not a new model request. A
completed run contains enough information to explain most scoring and comparison problems.

## Find the run contract

Read these files in order:

1. `manifest.json` for suite, benchmark preset, profiles, effective settings, schedule count, and execution mode;
2. `summary.json` for calculated profile metrics and pairwise results;
3. `summary.md` for the human-readable report;
4. the affected trial directory for `trial.json`, `run.jsonl`, `error.log`, and suite verifier output.

Use the stable run directory named by the user. Do not guess from the newest directory when several evaluations may be
running.

## Trace one failing trial

Match the profile, task, attempt, and pair id from the report to one trial. Then inspect:

- `trial.json` for the scheduled task and score;
- `run.jsonl` for the prompt, Pi options, trace events, messages, tool calls, terminal output, and session;
- `error.log` for process or validation failures;
- verifier files for the suite's external checks.

Check the observable chain in this order: prompt, provider response, tool call, tool result, filesystem effect, validator,
score. Stop at the first broken link instead of treating the final reward as the cause.

## Check profile comparisons

A pairwise value is `right - left`. Confirm that both trials have the same task, attempt, and preparation rules before
explaining a difference. Check profile order from the schedule rather than assuming alphabetical order.

When settings are under test, compare the effective settings recorded in `manifest.json` with the `run.jsonl` header. An
omitted profile setting should match the run-level value. An empty list or empty string should remain an intentional clear.

## Rebuild without Pi

If the Pi runs completed but the report is wrong, rebuild the calculated report from the run directory:

```bash
pi-eval report /absolute/path/to/results/run-id
```

This separates report calculation problems from agent behavior. Do not rerun the model until the stored trial and
manifest data are understood.

## Inspect live output deliberately

For terminal frames, timing, or native TUI output, continue with `inspect-interactive-tui-tests` from
`pi-coding-agent-test`. Use its `pi-test replay` workflow on the exact affected artifact directory. Do not add ad-hoc
logging or treat a live visual inspection as normal automated verification.

## Finish with a diagnosis

State the first broken contract, the evidence file and record that proves it, and the smallest fix. If the run is
correct and only presentation is wrong, fix the report renderer or report calculation rather than changing the suite or
agent profile.
