# Prompting

House prompt style — how production prompts in this repo are written. The library format is the `prompt-library` skill; quality measurement is the `evals` rule; this file is the craft.

## Structure

- **Be explicit, and say why.** State the task, the audience, the constraints, and the motivation — models generalize better from a stated *why* than from a pile of dos and don'ts.
- **Structure beats prose.** Separate instructions, context, examples, and output format with clear delimiters (XML-style tags are the house default: `<instructions>`, `<context>`, `<examples>`, `<output_format>`). Never interleave data with instructions.
- **Examples are load-bearing.** Two or three diverse, realistic examples beat ten repetitive ones; every example must be one the golden set would score as passing.
- **Permit uncertainty.** Tell the model what to do when it can't comply — a stated fallback ("say 'insufficient information'") beats a hallucinated answer, and rubrics score the fallback as correct.
- **Untrusted input stays untrusted.** Anything user- or document-supplied renders inside its own delimited block; instructions never come from inside it.

## Economy

- **Don't ask the model for what code can do.** Formatting, counting, filtering, deterministic transforms — do them in code around the prompt, not in the prompt.
- **Every token in a production prompt is a per-call cost.** Cut preamble, restated context, and defensive repetition; if a clause never changes an output (the evals will tell you), it goes.
- **One prompt, one job.** A prompt accreting second and third jobs gets split into family siblings — chaining certified prompts beats one uncertifiable mega-prompt.

## Iteration discipline

- The first draft rarely survives; iterate against the golden set (`prompt-eval`), not against one hand-picked example.
- Read failing outputs closely before editing — the fix for a reasoning failure (add motivation, add an example) differs from the fix for a format failure (tighten the output block).
- Record rejected variants and their eval deltas in the changelog entry; the next session inherits why the current wording won.
