---
description: Chaos engineering—hypothesis-driven experiments, minimal blast radius, production-like env. Pod/network/resource failure; validate resilience and alerting.
alwaysApply: false
---

# Chaos Engineering

Guidelines for controlled failure experiments.

## Core Principles

1. **Hypothesis** - Start with “If X fails, system will Y”; then test.
2. **Production-Like** - Staging or production (with safeguards); staging can miss real behavior.
3. **Minimal Blast Radius** - One pod, one AZ, small % of traffic; expand only after validation.
4. **Automate** - Repeatable experiments (e.g. Chaos Mesh, Litmus); not ad-hoc “kill a box.”

## What Chaos Is (and Isn’t)

- **Is**: Experiments to build confidence in resilience; discover weaknesses; validate monitoring and runbooks.
- **Isn’t**: Random breakage; testing in prod without controls; blame or “game day” without learning.

## Experiment Flow

1. **Steady state**: Define normal (e.g. latency, error rate, throughput).
2. **Hypothesis**: “If we kill one replica, latency stays under X and we get an alert.”
3. **Inject**: Run experiment (pod kill, network delay, CPU stress) in a bounded scope.
4. **Observe**: Did system behave as expected? Did alerts fire? Did runbooks work?
5. **Abort**: Have a way to stop immediately if impact is worse than expected.
6. **Improve**: Fix gaps (alerting, redundancy, runbooks); re-run to validate.

## Common Experiments

- **Pod failure**: One or few pods; verify restart and traffic shift.
- **Network**: Latency or partition; verify timeouts and fallbacks.
- **Resource**: CPU/memory stress; verify limits and scaling.
- **Dependency**: Downstream unavailable; verify circuit breakers and degradation.

Run in staging first; production only with small scope and rollback.

## Definition of Done (Chaos)

- [ ] Hypothesis and success criteria written; blast radius defined.
- [ ] Abort and rollback steps clear; on-call aware.
- [ ] Results documented; action items for gaps.

## Common Pitfalls

- **No hypothesis** - “Let’s break something” doesn’t yield learning; define expected behavior first.
- **Too big too soon** - Start small; avoid taking down whole service.
- **No follow-up** - Fix what you find; re-run to confirm.
