---
type: Reference
title: "SKILL-FORMAT.md — Required Format for PMOS Skills"
timestamp: 2026-07-23
---

# SKILL-FORMAT.md — Required Format for PMOS Skills

This document defines the **required** structure for every `.skill` file in PMOS. It is the authority that all seed skills (`prd`, `eval-rubric`, `okr`, `strategic-brief`, `source-to-concept`) and every future skill must conform to.

## Why a fixed format

Skills are how PMOS gives an agent procedural competence on demand without flooding its context. The format exists to make that economical: cheap to keep many skills available, expensive to load only the one in use. The mechanism is **three-level progressive disclosure**.

## The three levels

### Level 1 — Frontmatter (always loaded)

YAML frontmatter containing **only `name` and `description`**. This is pre-loaded for **all** skills at agent startup, so it must stay small — it is the *trigger signal* the agent uses to decide whether the skill is relevant.

```yaml
---
name: prd
description: Write a Product Requirements Document for an initiative. Use when an initiative needs a spec before build work begins.
---
```

Rules:

- `name` — lowercase, matches the filename stem (`prd` → `prd.skill`). **Required.**
- `description` — one or two sentences. State **what** the skill does and **when** to use it. This is the only text the agent sees for an unloaded skill, so the "when" is what makes it fire. **Required.**
- `maturity_mode` — **optional**. Present **only** on skills that are not yet active (e.g. `bootstrap`); omit it entirely on active skills. It lives at L1 because "is this skill executable yet?" is an execution-gating signal a loader/CI must be able to read without parsing the body. Values: `bootstrap` (specified, non-executable) → graduates to active (key removed) by a PM decision in `DECISIONS.md`.
- **Optional Phase-B registry metadata (D49)** — a skill may also carry `version`, `owner`, `risk`, `category`, and `scope`. These are **optional**, producer-defined registry fields that the workspace's skill library reads to group/version/scope skills. They used to be sourced into a hosted `skills` table by a resync generator; [D66](/DECISIONS.md) removed that mirror along with the backend, so the frontmatter on disk is now the only registry there is. A skill with only `name` + `description` is still fully conformant (e.g. `build-skills/evaluator.skill`); carry the registry keys on durable `skills/core/` entries for consistency with their siblings.
- `lifecycle` — **optional**, in the same D49 registry slot. Values: `active` (the default when the key is absent) | `deprecated`. A deprecated skill **stays on disk** — a reader gets it in full and is *told* not to start new work on it. Retirement is a state you can read; deletion is a separate act (see the invariant below). Distinct from `maturity_mode`, which is about a skill not yet executable — `lifecycle` is about one on its way out.
- **No other L1 keys.** Beyond the required `name` + `description`, the optional `maturity_mode`, and the optional Phase-B registry keys above, procedural detail does not belong in L1 — it lives in the body.

### The disk ⇄ registry invariant — retired with the mirror (D66)

There used to be a second copy of every skill: a hosted `skills` table that `get_skill` served to agents in other people's repos. A `REGISTRY` check in the drift gate held the two in sync, because a mirror that has quietly drifted is a wrong answer delivered confidently — and it was not theoretical. A generator whose INSERT omitted the Phase-B columns silently nulled `version/owner/risk/category/scope` for **all 21 skills for about three weeks** in 2026-07, caught only by an incidental version pin in an RLS test.

[D66](/DECISIONS.md) deleted the mirror, the generator and that check together with the hosted backend. **The file on disk is now the only copy**, which is why the class of bug above cannot recur: there is nothing left to disagree with. The cost is that skill metadata is no longer validated by CI at all — if a registry key is wrong in the frontmatter, it is simply wrong. Read the frontmatter as authoritative and keep it accurate by hand.

### Level 2 — SKILL.md body (loaded when relevant)

The body below the frontmatter holds the **full procedural detail**: steps, heuristics, required fields, gate criteria, examples. The agent reads this only after Level 1 has signalled relevance.

This is where the actual "how to do the task" lives. Write it at optimal altitude — concrete heuristics that guide judgment, not a brittle decision tree and not a vague mandate.

### Level 3 — Linked sub-files (loaded on demand)

Rare sub-cases, long reference tables, or edge-case procedures live in **separate files**, referenced **by filename** from the Level 2 body. The agent reads them only when the body points it there for the specific case in hand.

Example reference from a body: *"For multi-product PRDs, follow `prd-multiproduct.md`."*

Use Level 3 to keep Level 2 focused. If a section applies to <20% of runs, it is a Level 3 candidate.

## Quality — what makes a skill good (not just well-formed)

The three levels are *structure*. This section is the *bar*. A skill can be perfectly structured and still be junk; these are the tests that keep it worth its context.

### Predictability is the root virtue

A skill exists to **wring determinism out of a stochastic system**: it makes the agent take the same *process* every run — not produce the same *output* (an evaluator should predictably *judge*; its words vary, its behavior doesn't). Every rule below serves predictability; cost and readability are symptoms of it, not rival goals. When unsure whether a line belongs, ask: *does it make the agent's process more repeatable?* If not, cut it.

### Per-step completion criteria

Every procedural step ends on a **checkable done-condition** — the agent must be able to tell done from not-done. Where it matters, make the criterion **exhaustive**, not vague: *"every living mirror of the rule accounted for"*, not *"produce a change list"*. A fuzzy criterion invites **premature completion** — the agent stops when it feels done, not when it is. The [semantic-pass rule](/skills/evaluator.skill) ("enumerate the live-doc set, then read each for meaning — a grep is evidence, never done") is the exemplar this generalizes: it exists because a vague "sweep it" criterion let real stale mirrors survive five "clean" passes.

### Pruning — hunt these named failure modes

A skill without a pruning discipline only ever grows. Hunt, sentence by sentence:

- **No-op** — a line the model already obeys by default, so you pay context to say nothing. Test: does it change behavior vs. the default? *"Be thorough"* to an already-thorough agent is a no-op; the fix is a stronger word (a leading word), not more words.
- **Sediment** — stale layers that settle because adding feels safe and removing feels risky. The default fate of any skill without this discipline.
- **Sprawl** — simply too long, even when every line is live. Cure: disclose reference to Level 3; split by branch.
- **Duplication** — the same meaning in more than one place. Costs tokens and maintenance, and inflates a meaning's apparent importance. Keep one source of truth.
- **Negation** — steering by prohibition backfires: *"don't think of an elephant"* names the elephant. State the **positive** target behavior; keep a prohibition only as a hard guardrail you can't phrase positively, and pair it with what to do instead.

**This is the checklist the quarterly harness-assumption review runs** (DECISIONS.md cadence): the review's question — *what scaffolding can be removed because the model now handles it natively?* — is the no-op and sediment hunt applied across every skill.

### Leading words

A **leading word** is a compact concept already in the model's pretraining that anchors a whole region of behavior in one token — *tight* (loop), *watermelon*, *default-to-refute*, *silent drift*, *Summary Gate*. Repeated across a skill, it accumulates a distributed definition and recruits priors the model already holds. A triad spelled out at three sites (*"fast, deterministic, low-overhead"*) is begging to **collapse** into one word (*tight*). You win twice: fewer tokens, and a sharper hook the agent hangs its behavior on. Hunt every skill for restatements a leading word retires.

### Invocation (advisory)

The Agent-Skills world splits skills by *who can invoke them* — **model-invoked** (description always loaded, agent can auto-fire it, pays permanent context load) vs **user-invoked** (invisible to the agent, human-only, zero context load). PMOS's model differs — skills are **retrieved on demand** (see the retrieval rationale under *Ecosystem compatibility* below), so Level-1 frontmatter is the always-available trigger and there is no auto-fire to gate. Treat the axis as an authoring lens — *is this skill's description earning its keep as an always-available trigger?* — **not** a required field. (A formal `invocation:` key is a deferred registry question, not part of this format.)

## Ecosystem compatibility — the Agent Skills standard

The open **Agent Skills standard** (Claude Code, Codex, and other harnesses) packages a skill as a **directory** `<skill-name>/SKILL.md` with YAML frontmatter (`name`, `description`, optionally `license`, `allowed-tools`) plus bundled sub-files. It **validates PMOS's three-level disclosure** — same frontmatter-as-trigger, body-as-procedure, sub-files-on-demand shape.

PMOS **diverges** deliberately: skills are flat `<name>.skill` files, **not** `<name>/SKILL.md` directories. The rationale is PMOS's retrieval model — skills are fetched on demand via `get_skill` or referenced from AGENTS.md, **not** auto-discovered by a harness scanning for `SKILL.md`, so the directory ceremony buys nothing here. The **mapping is lossless**, which is what keeps the divergence cheap:

- a `<name>.skill` file **≅** a standard `SKILL.md` body;
- PMOS's required `name` + `description` frontmatter is **already** the standard's required pair;
- Level-3 sub-files **≅** the standard's bundled files.

(One nuance for the follow-on below: PMOS's *optional* D49 registry keys — `version`/`owner`/`risk`/`category`/`scope`/`lifecycle` — are top-level custom keys; the standard would expect producer-defined extensions under a `metadata:` map. The **core** shape is lossless; relocating the registry keys belongs to the native-format packaging follow-on, not this flat form.)

### Kit native-format packaging (E6 — shipped)

The **kit** now ships every skill in native Agent-Skills directory form *as well as* the flat file — a lossless portability layer, generated **from** the flat `.skill` at build time (never hand-authored, so it cannot drift from its source). For each shipped skill, `kit/build-kit.sh` emits `skills-native/<name>/SKILL.md` where:

- the directory name **equals** the frontmatter `name` (the standard's hard rule), and required `name` + `description` are present;
- PMOS's D49 registry keys (`version`/`owner`/`risk`/`category`/`scope`/`lifecycle`) are relocated **under a `metadata:` map with a `pmos-` prefix** — the standard's producer-extension slot — never as top-level custom keys, so a conformant runtime ignores them cleanly;
- a **fail-closed conformance self-check** in `build-kit.sh` aborts the build on any violation (name≠dirname, missing required field, or a top-level custom key), the same posture as the containment check.

The flat `.skill` files remain the file-mode source of truth (the AGENTS.md state adapter references `pmos/skills/<name>.skill`); the native tree is the ecosystem-portable mirror. **Still a named follow-on (not built here):** wiring `create-pmos` to install the native tree into a runtime's auto-discovery locations (`.claude/skills/`, `.agents/skills/`). E6 delivers the conformant *packaging*; the *install location* is a separate change, revisited when native-discovery demand in kit repos is proven.

### Per-skill eval harness (E8 — convention shipped, benchmark deferred)

A skill may carry an **opt-in** eval file at `skills/evals/<name>.evals.json` (mapping to the standard's `<name>/evals/evals.json` in native packaging), adopting the skill-creator harness shape:

- **`description_triggers`** — `should_fire` / `should_not_fire` prompt lists that test whether the L1 description fires on the right requests and stays silent on the wrong ones (trigger-accuracy).
- **`cases`** — each a `{ name, prompt, expect, with_without }`, where `expect` is an **outcome property** a grader checks (never a script) and `with_without` toggles whether the case is meant for with-vs-without-the-skill benchmarking.

`scripts/check_skill_evals.py` (run in the `okf-drift` job) validates the **shape** of any present eval file and that it names a real skill. It **never runs a benchmark** — actually scoring with-vs-without needs labeled outcomes + compute and is **deferred** (there is no negative class yet); running models in CI would make the gate nondeterministic. The convention and a worked example (`skills/evals/prd.evals.json`) ship now; the verdict "this skill measurably helps" does not.

## File naming

- Skill files use the `.skill` extension: `prd.skill`, `eval-rubric.skill`.
- Level 3 sub-files are plain `.md` and live alongside the skill or in a named subfolder.

## Skill namespaces

- `/skills/core/` — seed skills, canonical across PMOS.
- `/skills/domains/` — domain-specific variants (P1).
- `/skills/pm/<pm_slug>/` — PM-specific variants (e.g. `/skills/pm/wawan/`).
- `/build-skills/` — skills for the coding agents that build PMOS itself (distinct from PM skills).

## Conformance checklist

A skill is well-formed when:

1. Level 1 frontmatter has `name` + `description` (both required), plus optionally `maturity_mode` (on non-active skills) and the Phase-B registry keys `version`/`owner`/`risk`/`category`/`scope`/ `lifecycle` (D49); no other L1 keys.
2. `name` matches the filename stem.
3. `description` states both what the skill does and when to use it.
4. The body (Level 2) carries the full procedure and is self-sufficient for the common case.
5. Edge cases are pushed to Level 3 sub-files and referenced by filename.

**Quality bar** (the "good", not just "well-formed" — see Quality above):

6. **Predictability served** — every line makes the agent's *process* more repeatable; cut what doesn't.
7. **Per-step completion criteria** — each procedural step ends on a checkable (and where it matters exhaustive) done-condition, not a vague "do X".
8. **Pruned** — no no-ops, no sediment, no duplication; prohibitions phrased as positive targets; sprawl disclosed to Level 3.
9. **Leading words** — restated qualities collapsed into pretrained tokens where one earns its keep.

All seed skills are authored against this checklist; the quarterly harness-assumption review re-runs items 6–9 across every skill.
