---
description: Incident management—severity levels, response process, incident commander, communication. Detect, respond, mitigate, resolve; blameless postmortem.
alwaysApply: false
---

# Incident Management

Guidelines for detecting, responding to, and learning from incidents.

## Core Principles

1. **Detect Fast** - Monitoring and alerting before user reports.
2. **Communicate Clearly** - Stakeholders informed; status page and ETA.
3. **Mitigate First** - Stop the bleeding (rollback, scale, disable) before root cause.
4. **Learn Always** - Blameless postmortem; action items tracked.

## Severity Levels

- **SEV1 (Critical)**: Full outage or data loss; immediate response; war room; exec notification; status page.
- **SEV2 (Major)**: Significant degradation; many users; respond in 15–30 min; status page; escalate if long.
- **SEV3 (Minor)**: Limited impact; workaround exists; respond within SLA (e.g. 1–4 h).
- **SEV4 (Low)**: Minimal/cosmetic; next business day.

Define response time (ack, engage), communication (Slack, status page, exec), and escalation per level.

## Response Process

1. **Detect**: Alert or report; acknowledge quickly.
2. **Declare**: Create incident channel; assign incident commander (IC); set severity.
3. **Mitigate**: Rollback, scale, failover, or feature toggle; verify impact decreasing.
4. **Resolve**: Root cause fixed or documented; incident closed; postmortem scheduled.
5. **Postmortem**: Blameless; timeline, root cause, action items; share learnings.

## Incident Commander

- Single point of coordination; assign roles (comms, tech lead, scribe); make decisions when needed; escalate when stuck.
- Keep status updates regular (e.g. every 15–30 min); update status page and stakeholders.
- Ensure postmortem happens and action items are tracked.

## Communication

- Incident channel (e.g. inc-YYYYMMDD-shortname); pinned summary and runbook links.
- Status page: impact, ETA, next update. Honest and concise.
- Handoff: document current state and next steps when swapping IC or on-call.

## Definition of Done (Incident)

- [ ] Severity set; mitigation and resolution documented.
- [ ] Stakeholders and status page updated.
- [ ] Postmortem within 5 business days; action items owned and tracked.

## Common Pitfalls

- **Blame focus** - Focus on systems and process, not individuals.
- **Skipping postmortem** - Every SEV1/2 should have a postmortem and follow-up.
- **Vague communication** - “We’re looking into it” → “We’ve rolled back; investigating DB connection pool.”
