# Autonomy guardrails

`workflow.autonomy_level` in `.sdd-agentic-flow/config.yml` is a **new axis orthogonal to** `workflow.execution_mode`
(`plan`/`guided`/`apply`/`review`/`full`).
`execution_mode` answers "what is a skill authorized to do"; `autonomy_level` answers "does a
skill need a human between it and the next one." Neither replaces the other, and neither changes
the safety boundary. Absent configuration resolves to `apply` + `supervised`, as defined in
[effective defaults](effective-defaults.md). Autonomous continuation requires explicit delegation.

## The three levels

- **`manual`**: every skill returns control completely. Nothing advances
  automatically, even when a skill reports success. Transition policy: `stop`.
- **`supervised`** (default): a skill executes, reports its evidence, and offers an explicit
  "continue to `<next skill>`?" recommendation; the human decides. Transition policy: `confirm`.
- **`autonomous`**: a skill executes and continues toward a verified outcome on its own. A
  positive result advances on the normal path; a recoverable negative result may take an
  authorized repair transition. Only an exceptional authority, safety, budget, or no-progress
  failure returns control to the human. Transition policy: `continue`, bounded by recovery.

## `execution_mode` × `autonomy_level` compatibility

| execution_mode | `manual` | `supervised` | `autonomous` |
| --- | --- | --- | --- |
| `plan` | valid | valid | **invalid** |
| `guided` | valid | valid | **invalid** |
| `apply` | valid | valid (effective default) | valid |
| `review` | valid | valid | valid |
| `full` | valid | valid | valid |

`plan` and `guided` never combine with `autonomous`. A plan-only workflow has nothing to
auto-advance into, and step-by-step confirmation is the entire point of `guided`. Pairing it with
unattended advance contradicts `guided`. `doctor --autonomy` flags either combination as `FAIL`.

## The 7 guardrails

The invoking host/agent evaluates these seven obligations from current evidence; the CLI
does not enforce their runtime evaluation. Scope and evidence adequacy require judgment.
An agent operating in
`autonomy_level: autonomous` re-checks them before treating a Skill result as permission for a
normal advance or an authorized repair transition. The checks govern transition admissibility;
they do not turn a recoverable negative result into a claim of completed work.

1. **Outcome classification** — the Skill's native status is classified as satisfied, recoverable,
   exceptional, or exhausted. `needs changes` and attributable `not ready` outcomes may authorize
   repair; a status label alone does not force human escalation.
2. **Evidence validation** — every artifact the skill's `autonomy_profile.evidence_required` lists
   actually exists and is non-empty.
3. **Verification integrity** — the skill's own required checks (tests, linter, spec consistency,
   and blocking findings) are executed or explicitly accounted for. Normal forward progression
   requires applicable checks to pass; an observed, attributable failure may instead authorize a
   repair transition. A skill never reports a positive completion status while a required check
   failed.
4. **Scope boundary** — the delegated semantic outcome remains bounded. Required implementation
   touchpoints may expand when repository evidence requires more files or modules; unrelated
   behavior and product scope may not be added silently.
5. **Skill transition validity** — the proposed next Skill is an authorized normal or repair path
   (see the main SDD flow diagram in `README.md`). Repair transitions such as check → implement
   → check are allowed when their owner and evidence conditions are explicit.
6. **Resource sufficiency** — the configured budget (`workflow.autonomy_budget` in
   `.sdd-agentic-flow/config.yml`: `max_iterations`, `max_tokens`, `max_runtime_hours`) is not exhausted, and
   `pause_on_warning` triggers a stop, not just a warning, once remaining budget drops below 20%.
7. **Human override gate** — no `pause=true` or `stop=true` is recorded in
   `.sdd-agentic-flow/autonomy/loop-state.md`. This is the one guardrail that is not evaluated automatically by
   construction — it exists specifically so a human can halt an in-flight autonomous run by
   editing state, without needing to kill a process.

If a guardrail identifies an exceptional blocker or exhausted execution, the agent stops, records
the reason in `.sdd-agentic-flow/autonomy/loop-state.md`, and waits for a human to resolve it. A
recoverable result records its native status, the authorized repair `Next`, and `Guardrails: PASS`,
then the invoking host/agent continues without a new human confirmation. SAF never starts the
next host turn itself.

## `autonomy_profile` sidecar

Each skill declares, in its `saf-contract.yml` sidecar, which levels it supports and what a `PASS`
means for it:

```yaml
autonomy_profile:
  supported_levels: [manual, supervised, autonomous]
  auto_continue_condition: 'spec.md and tasks.md present; design.md when required by profile or decision; no unresolved requirements'
  blocking_conditions: [missing_spec, inconsistent_design, unspecified_requirements]
  evidence_required: [spec.md, tasks.md, 'design.md when required by profile or decision']
```

- `supported_levels` — which of the three levels this skill can run under. A skill whose output is
  always a recommendation or explanation for a human to act on (never itself a link in the
  auto-advancing chain) omits `autonomous`.
- `auto_continue_condition` — one-line, human-readable statement of what "safe to continue
  automatically" means for this skill. Informational; the actual gate is the outcome
  classification and guardrails above.
- `blocking_conditions` — the specific failure modes that stop this skill from reporting `PASS`.
- `evidence_required` — the artifact(s) guardrail 2 checks for.

Repository maintenance validation checks that every installed Skill declares `autonomy_profile`
and that `supported_levels` is a subset of `{manual, supervised, autonomous}`.

## `.sdd-agentic-flow/autonomy/loop-state.md`

The execution-state file an agent maintains while running a workflow under `supervised` or
`autonomous`. It is the "memory of the loop": an agent (or a human) can read it to resume from the
last completed skill without replaying earlier ones, and `sdd-agentic-flow context autonomy-state`
/ `sdd-agentic-flow autonomous-resume` both read and update it. Minimal shape:

```markdown
# Loop State

Execution mode: full
Autonomy level: autonomous

## Current State

- Skill: saf-check-task (completed)
- Status: PASS
- Next: saf-validate
- Guardrails: PASS
- Human override: pause=false, stop=false

## Blocker History

None.
```

An agent appends a new "Current State" block after each skill completes; it never rewrites
history, only adds to it. A human halts an autonomous run by setting `pause=true` or
`stop=true` under Human override — guardrail 7 picks that up on the next check.

The `Skill:` value is not itself the SDD flow phase. The canonical workflow sequence
(`saf-brainstorm` through `saf-validate`) supplies the Plan/Prompt/Implement/Check/PR/Review/Fix/
Validate mapping. No new field is needed to make `loop-state.md` phase-inspectable; the data
already exists, this is just where to read it.

## Scope: what autonomy governs, and what it does not

`autonomy_level` governs **skill-to-skill transitions only**. It does not grant a skill any
authority `execution_mode` does not already grant — `autonomous` never implies "skip
`no_commit_by_default`" or "ignore an explicit scope boundary." A skill running in
`autonomy_level: manual` may still call any tool, including an available MCP integration, exactly
as it always could; autonomy only changes whether the agent asks before invoking the *next skill*.
There is no orchestration engine in this CLI that executes skills on a loop — `autonomy_level` is a
contract the skills and the invoking agent honor, validated statically by `doctor --autonomy`, not
a runtime this package hosts.
