---
type: Concept
title: The Agent Harness
description: Everything wrapping the model that lets it finish work — and the Summary Gate, the input gate that compresses prior-session context into essential state.
tags: [harness, agent, context, summary-gate, architecture]
timestamp: 2026-06-28
---

# The Agent Harness

**Situating context:** PMOS's deliverables are not application code — they are harness. This concept
names what a harness is and introduces the **Summary Gate**, the pattern that explains why PMOS's
structured artifacts beat raw context compaction. It informs AGENTS.md, the OKF bundle's role, and
the [control plane](/okf/core/concepts/control-plane.md) identity.

## Agent = Model + Harness

A model on its own predicts tokens. An **agent** finishes work. The difference is the **harness**:
everything wrapping the model that lets a run complete.

The harness includes:

- **Instructions / rule files** — CLAUDE.md, AGENTS.md, DECISIONS.md, PRINCIPLES.md.
- **Knowledge** — the [OKF bundle](/okf/core/concepts/okf-governance.md): concepts, templates,
  playbooks.
- **Skills** — procedural competence loaded on demand
  ([SKILL-FORMAT](/skills/SKILL-FORMAT.md)).
- **Tools / MCP** — the [ACI surface](/build-skills/MCP-SPEC.md) the agent acts through.
- **Sandboxes** — where work runs safely.
- **Orchestration** — sequencing, concurrency, routing.
- **Guardrails** — the [three-tier eval gate](/okf/core/concepts/output-eval.md), kill conditions.
- **Observability** — `agent_runs`, [discoveries.md](/state/discoveries.md), transcript review.

**PMOS builds the harness, not the app.** Improving any harness element improves every future run —
that is the leverage the [control plane](/okf/core/concepts/control-plane.md) exists to capture.

## The Summary Gate

> **Summary Gate:** PMOS's OKF bundle + SKILL.md + AGENTS.md constitute a *Summary Gate* — an input
> gate that compresses prior-session context into essential state before a new agent session begins.

A model's context window is finite and a project's history is not. Something has to decide what
crosses from "everything that happened" into "what the next session actually needs." That decider is
the Summary Gate.

### What it does

- **Compresses** prior-session context into essential state: the decisions, the knowledge, the
  contracts — not the transcript.
- **Extracts key decisions and discards intermediate reasoning.** The *conclusion* ("we decided X
  because Y") crosses the gate; the deliberation that produced it does not.
- **Prevents context pollution** — stops the next session from inheriting stale, contradictory, or
  irrelevant intermediate state that would degrade its judgment.

### Why structured handoff beats raw compaction

Raw context compaction (summarizing a transcript) preserves *narrative* — it keeps a lossy trace of
what was said. A Summary Gate preserves *state* — it keeps the durable decisions and knowledge in
purpose-built artifacts (OKF concepts, skills, AGENTS.md) that are authored to be read fresh.

A new session reading a hand-authored decision record starts from clean, intentional state. A new
session reading a compacted transcript starts from a blurry recollection of a conversation. **This
is why PMOS invests in structured handoff artifacts instead of relying on context compaction** — the
artifacts *are* the gate, and they are designed to survive the session boundary intact.

## Optimal altitude

Harness instructions must be written at **optimal altitude**: concrete heuristics that guide
behavior, not a brittle if-else rules engine, and not a vague mandate that assumes shared context.
Too low and the harness is rigid and breaks on the first unforeseen case; too high and it gives no
real guidance. AGENTS.md and every skill body are held to this standard.

## Relationship to autonomy and drift

The harness is what makes higher [autonomy](/okf/core/concepts/output-eval.md) safe: the better the
guardrails, observability, and Summary Gate, the longer an agent can run unattended without producing
[silent drift](/okf/core/concepts/watermelon-flag.md). Harness quality and safe autonomy rise
together — which is why the quarterly harness assumption review exists.

The harness also **improves from its own runs**: evals, friction, and calibration data are evidence for
the next harness edit. See [self-refining-harness](/okf/core/concepts/self-refining-harness.md) — the
detect → draft → PM-gate loop that turns run evidence into harness edits, bounded at L3 so it never
becomes L4 self-mutation.
