What is Chaos Engineering?
Chaos engineering is the practice of deliberately introducing failures into a system to discover weaknesses before they cause unplanned outages. The name comes from Netflix, which coined the term when they built Chaos Monkey in 2011 -- a service that randomly terminated EC2 instances during business hours to force their engineering teams to build resilience into every service. If Chaos Monkey could break you and nobody noticed, you were resilient. If it caused an incident, you had found a real problem worth fixing.
The core insight is that complex distributed systems will fail. The question is whether they fail predictably and recover gracefully, or fail catastrophically and unexpectedly. Chaos engineering shifts the discovery of failure modes from production incidents -- which affect users and happen on bad days -- to controlled experiments that you can observe, stop, and learn from.
Chaos engineering sits at the intersection of microservices architecture, reliability engineering, and DevOps culture. Systems that are architected for independent deployment (see blue-green deployment) tend to fare better under chaos experiments.
The Principles
Chaos engineering is not chaos. The Netflix team formalised its principles, which are worth understanding before reaching for a tool:
1. Define steady state. Before injecting a failure, establish a quantitative baseline: what does normal look like? P99 latency under 200ms, error rate below 0.1%, throughput above 1,000 req/s. This is what you are trying to preserve.
2. Form a hypothesis. Predict what will happen when the failure is injected: "If we terminate one of the three payment service pods, the load balancer will route traffic to the remaining two within 15 seconds and the error rate will stay below 1%."
3. Vary real-world events. Inject failures that mirror real incident patterns: instance termination, network partition, CPU saturation, disk exhaustion, dependency failure, clock skew.
4. Run experiments in production. Staging environments miss failures that only manifest at production traffic levels and with production data. Build confidence in staging first, then graduate to production with safeguards.
5. Automate to run continuously. A one-time experiment discovers weaknesses at a point in time. Running experiments continuously -- as Netflix does -- catches regressions before they ship.
A LitmusChaos Experiment on Kubernetes
LitmusChaos is an open-source chaos engineering framework for Kubernetes. It defines experiments as Kubernetes custom resources. Here is an experiment that kills a pod in the default namespace:
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
name: pod-delete-chaos
namespace: default
spec:
appinfo:
appns: default
applabel: "app=myapp"
appkind: deployment
chaosServiceAccount: litmus-admin
experiments:
- name: pod-delete
spec:
components:
env:
- name: TOTAL_CHAOS_DURATION
value: "30" # run for 30 seconds
- name: CHAOS_INTERVAL
value: "10" # delete a pod every 10 seconds
- name: FORCE
value: "false" # graceful termination
- name: PODS_AFFECTED_PERC
value: "50" # affect 50% of matching pods
Before running this experiment, you would have your monitoring dashboards open to observe whether the steady-state metrics (latency, error rate) remain within bounds. The experiment result -- ChaosResult -- records whether the steady-state was maintained.
GameDays
A GameDay is a scheduled, team-wide chaos experiment -- a planned event where engineers deliberately break things in a controlled way and practice incident response. Originally popularised by Amazon, GameDays have become a standard reliability engineering practice.
A well-run GameDay has: a defined scope and blast radius limit; monitoring dashboards visible to everyone; an incident commander who can halt the experiment; a predefined runbook the on-call team practices following; and a blameless post-mortem afterwards. The goal is not to surprise the team -- it is to practice the coordination and recovery process before a real incident forces you to improvise under pressure.
Where to Start
Most teams should not start chaos engineering by running Chaos Monkey in production on a Monday morning. A reasonable progression: instrument your system with metrics and alerting first (you need steady-state baselines); run table-top exercises (talk through what would happen if X failed); run chaos experiments in staging; identify and fix the first round of weaknesses; then graduate to limited production experiments with safeguards and a quick abort path.
Chaos engineering is covered as part of reliability and observability practices in the DevOps tools guide.
