---
description: DevOps/SRE Overview
alwaysApply: false
---

# DevOps/SRE Overview

Guidelines for Site Reliability Engineering and operational excellence.

## Scope

- SRE practices, production operations, and 24/7 reliability
- Incident management, response, and blameless postmortems
- Monitoring, alerting, and observability (metrics, logs, traces)
- SLO/SLI definition and error budget management
- Capacity planning, disaster recovery, and chaos engineering
- Toil reduction, automation, and safe change management

## Core Principles

### Reliability is a Feature

- Treat reliability work as product work — users can't tell "slow" from "broken"
- Measure from the user's perspective; invest proportional to business impact

### Error Budgets Over Perfection

- 100% reliability is the wrong target — it means zero innovation
- Define SLOs, use error budgets to balance reliability and velocity
- Budget healthy → move fast; budget low → prioritize reliability

### Automate Toil Away

- If you're doing it manually more than twice, automate it
- Track toil as a metric; target < 50% of SRE time on toil

### Observability First

- Instrument everything from day one — logs, metrics, traces are not optional
- Design systems to be debuggable; correlate signals across the stack

### Blameless Culture

- Focus on systems, not individuals — ask "how did the system allow this?"
- Share postmortems widely; celebrate learning from failures

## Key Metrics

### Reliability (Four Golden Signals)

- **Availability**: % successful requests
- **Latency**: p50, p95, p99 response times
- **Error Rate**: % failed requests
- **Throughput**: requests per second

### Operational

- **MTTD/MTTR**: Mean time to detect / resolve incidents
- **MTBF**: Mean time between failures
- **Change Failure Rate**: % of changes causing incidents

### On-Call Health

- Pages per shift < 10, per night < 2
- False positive rate < 10%

### Toil

- Toil % of engineering time; manual interventions per deploy
- Automation coverage of runbooks

## Anti-Patterns

- **Alert on everything**: Every alert must be actionable and urgent — no "just in case" pages
- **Hero culture**: Don't rely on specific engineers; document everything, share knowledge
- **Postmortem graveyard**: Track action items to completion; measure recurring incidents
