What is Site Reliability Engineering (SRE)?
Site Reliability Engineering originated at Google in the early 2000s when Ben Treynor Sloss was asked to run a production team. His answer was to hire software engineers and give them an operations problem to solve. The result was a discipline that treats reliability as a software problem, not a staffing one.
The Core Premise
Traditional operations work was largely manual, reactive, and hard to scale. More servers meant more operators. SRE rejects that model. The fundamental bet is that if you write software to manage your systems, reliability scales with software -- which scales much better than headcount.
This shows up in a few key principles:
Toil has a cap. Toil is manual, repetitive, automatable work that does not improve the system. Google's SRE model limits toil to 50% of an SRE's working time. The rest must be engineering work: building automation, improving reliability tooling, or working with product teams on system design. If toil exceeds 50%, something is structurally broken.
Reliability is measured, not assumed. SREs define concrete reliability targets and measure against them continuously. You cannot improve what you cannot measure.
Risk is a product decision. Reliability costs money and slows feature development. SRE frames this as a business trade-off that product and engineering own together, not as a pure technical concern.
SLIs, SLOs, and SLAs
These three terms are the vocabulary of SRE reliability work:
Service Level Indicator (SLI) -- A specific metric that measures an aspect of service reliability. Common SLIs: availability (percentage of successful requests), latency (fraction of requests served below a threshold), error rate.
Service Level Objective (SLO) -- An internal target for an SLI over a time window. Example: "99.9% of API requests complete successfully over any rolling 30-day window." SLOs are operational contracts between the SRE team and the product team.
Service Level Agreement (SLA) -- A contractual commitment to customers, usually backed by financial penalties. SLAs are typically looser than SLOs. You want headroom between your internal target and your customer commitment so you have room to operate.
Defining good SLOs is harder than it sounds. The right SLI is the one that maps to user experience, not infrastructure health. An SLI measuring CPU utilisation is less useful than one measuring the fraction of homepage loads that completed in under 500ms.
Error Budgets in Practice
An error budget converts an SLO into something operational. If your SLO is 99.9% availability, you have 0.1% of requests (or 43.8 minutes per month) available as failures before you breach your objective.
Monthly error budget = (1 - SLO) x window
= (1 - 0.999) x 43,200 minutes
= 43.2 minutes
When the error budget is healthy, the team has permission to deploy frequently, run experiments, and accept the risk of minor incidents. When the budget is close to exhausted, the team freezes risky deployments and focuses on reliability improvements. This turns the reliability conversation from "ops says no" into "the data says slow down" -- a much more productive dynamic.
What SREs Do Day to Day
The SRE role spans three main activities:
Incident response and post-mortems. SREs are on-call for the services they own. When things break, they respond, mitigate, and then run blameless post-mortems to extract systemic improvements. The goal is never to punish individuals -- it is to make the system more resilient.
Toil reduction and automation. Every manual task that happens more than twice is a candidate for automation. SREs write tools, scripts, and services that reduce the human work required to keep systems healthy. This is the engineering half of the role.
Reliability consulting with product teams. The best time to fix a reliability problem is before it ships. SREs embed with product teams during design and implementation to raise reliability concerns early -- reviewing architecture for single points of failure, setting sensible SLOs, and ensuring new services have the observability needed to operate them.
SRE vs DevOps
People often ask whether SRE and DevOps are the same thing. The honest answer is: they are closely related but distinct.
DevOps is a cultural movement and a set of practices. It is broad and deliberately non-prescriptive. The goal is to break down walls between development and operations teams so software ships faster and more reliably.
SRE is a specific job function with a defined methodology. It emerged independently of the DevOps movement but aligns with its goals. Google's SRE book frames it as "what happens when a software engineer designs an operations function."
You can do DevOps without having dedicated SRE teams. You can also have SRE teams that operate in ways that conflict with DevOps principles. The concepts overlap but are not identical.
For a broader view of where SRE fits in the modern engineering organisation, see the CI/CD guide and infrastructure as code.
