---
name: harness-refine
description: >
  Turn PMOS's own run evidence into concrete, evidence-cited edits to the harness. Read the eval
  defects, recurring friction, run transcripts, and the calibration log; cluster them into specific
  proposed edits — each mapped to ONE artifact (an MCP tool description, a skill body, an OKF concept,
  or an AGENTS.md heuristic) and citing the run/friction/eval that motivates it — including prunes of
  scaffolding the model now handles natively. Emit a proposals report and open a PR for the PM. Use on
  demand: when the PM invokes it, at the D23 quarterly harness review, or when v_friction_recurrence /
  recurring eval defects show a harness gap. Never merges, never self-approves, never auto-applies.
maturity_mode: active
---

# harness-refine — the self-refining harness agent

**Situating context:** PMOS's leverage is the harness, and the harness accumulates evidence about its
own weak spots — `evals` defects, `v_friction_recurrence`, run transcripts, `eval-calibration.md`. The
PM's weekly review *spots* recurring friction; nothing *drafts the fix*. This skill closes that gap the
same way [okf-author](/build-skills/okf-author.skill) closed the doc-upkeep gap (the D36→D42 pattern):
a deterministic signal *detects*, an on-demand agent *drafts*, the **PM gates the PR**. It realizes the
field's "agents improving their own tools" loop while staying at the **L3 ceiling** — see
[self-refining-harness](/okf/core/concepts/self-refining-harness.md) and
[D47](/okf/products/pmos/adr/d47-self-refining-harness.md).

## When to use vs. not

- **Use** on demand: a PM invokes it, the **D23 quarterly harness-assumption review** runs, or a
  recurring signal (a `v_friction_recurrence` theme with no resolving decision, a repeated `evals`
  defect class) says the harness has a gap worth fixing.
- **Do not** run it on a schedule or unprompted (D47 — that drifts toward L4).
- **Do not** use it to change **application code, database schema, migrations, or MCP-tool
  *implementations***. This skill edits the harness *text* — instructions, skills, concepts, tool
  *descriptions* — not the backend ([control-plane](/okf/core/concepts/control-plane.md), D01). A defect
  that needs a code/schema fix is **logged as a discovery and surfaced**, not fixed here.
- **Do not** replace the PM weekly transcript review or the `okf-drift`/`okf-author` loop — it
  complements them.

## Procedure

1. **Read** `AGENTS.md` and this skill. Confirm the trigger (PM-invoked / D23 review / a named recurring
   signal) — if there is no trigger, stop.
2. **Gather the evidence** (read-only). Pull from the harness's own record:
   - `v_friction_recurrence` — themes seen ≥ 2× (the AGENTS.md 2× tripwire), and `discoveries.md`
     `#friction` for the raw notes.
   - Recurring `evals` defect classes — `get_agent_runs` + the `evals` table: what does the evaluator
     keep catching across runs?
   - `eval-calibration.md` — where PM and evaluator disagreed (the leniency/bias signal).
   - The context/retrieval signals — `v_context_budget`, `v_knowledge_reuse`.
   Prefer signals that **recur** or that **cost repeated corrections**; a one-off is rarely a harness
   defect.
3. **Cluster into candidate harness defects.** Group the evidence into a *small* set (aim for the
   highest-leverage 3–6, not an exhaustive list). For each, name the **one target artifact** it maps to:
   an **MCP tool description** (ACI clarity, D19), a **skill body**, an **OKF concept**, or an
   **`AGENTS.md` heuristic**. If a defect maps to more than one artifact, split it.
4. **Draft the specific edit** for each, at **optimal altitude** (concrete heuristic, not a brittle rule
   or a vague mandate — Principle 7). **Every proposal must cite the evidence** that motivates it (the
   run_id / friction note / eval row — provenance, Principle 9). Apply the **prune lens** (D23 /
   Principle 12): also propose *removing* scaffolding the model now handles natively, or that a newer
   decision superseded — **at least one prune** where the evidence supports it.
5. **Emit a proposals report** → `planning/harness-refine/batch-<nnn>.md`: one entry per proposal with
   `target artifact`, `evidence`, `proposed edit` (add / sharpen / **prune**), and `rationale`. The
   report **proposes; it does not apply.**
6. **Open ONE PR — never merge, never self-approve.** For the proposals the PM greenlights (or, when
   invoked to draft directly, for all of them), draft the concrete edits and open a single coherent PR
   for the **PM Acceptance Gate**. Keep `okf-drift` green (touch index/log/reference docs as any edited
   concept requires). Hand the PR to the PM; do not merge.
7. **Log** the run (`log_agent_run`) and any friction (`log_friction` / `discoveries.md`).

## Guardrails

- **Never auto-apply, never merge, never self-approve.** The PM Acceptance Gate is final (D12); this is
  L3, not L4 (D47).
- **On-demand only** — never scheduled/unprompted (D47, mirroring D42).
- **Control altitude only** — edit the harness (instructions, skills, concepts, tool *descriptions*),
  never app code / schema / migrations / tool *implementations* (D01). Surface backend gaps as
  discoveries; don't fix them here.
- **Every proposal cites its evidence.** A proposal with no run/friction/eval behind it is an opinion,
  not a harness-refinement — drop it. (Principle 9.)
- **Prefer recurrence over one-offs**, and **prune as readily as you add** (D23). Undifferentiated
  growth dilutes the harness; a good batch often removes as much as it adds.
- **Read before you write** — the actual evidence and the actual artifact you propose to edit, not your
  memory of them.

## Level 3 — references

- The pattern it realizes: [self-refining-harness](/okf/core/concepts/self-refining-harness.md); the
  bounding decision [D47](/okf/products/pmos/adr/d47-self-refining-harness.md).
- The sibling detect→draft→PR loop: [okf-author](/build-skills/okf-author.skill),
  [D42](/okf/products/pmos/adr/d42-okf-author-on-demand.md).
- The evidence surfaces: recurring friction themes and self-reported context pressure, read from
  file-mode state ([discoveries.md](/state/discoveries.md) and `runs.jsonl`; the hosted
  `v_friction_recurrence` / `v_context_budget` views went with the backend in D66) plus
  [eval-calibration.md](/planning/evals/eval-calibration.md).
- The cadence it serves: [D23](/okf/products/pmos/adr/d23-quarterly-cadence.md) (harness-assumption
  review); the altitude standard: [agent-harness](/okf/core/concepts/agent-harness.md).
