Skip to content
all frameworks
07 · Author synthesis

SRE Maturity

From heroics to SLOs and adaptive learning
L1Reactive
What we measure:uptime
How we respond:heroics
How we learn:firefighting
L2Managed
What we measure:service SLIs
How we respond:runbooks
How we learn:postmortems
L3Proactive
What we measure:SLO + error budget
How we respond:automation
How we learn:blameless reviews
L4Resilient
What we measure:user-journey SLO
How we respond:self-healing
How we learn:game-days
L5Adaptive
What we measure:business-aligned SLI
How we respond:auto-rollback
How we learn:chaos → changes
SRE maturity is a reliability systemREACTIVE → ADAPTIVEL1 · ReactiveM · uptimeR · heroicsL · firefightingL2 · ManagedM · service SLIsR · runbooksL · postmortemsL3 · ProactiveM · SLO + error budgetR · automationL · blameless reviewsL4 · ResilientM · user-journey SLOR · self-healingL · game-daysL5 · AdaptiveM · business-aligned SLIR · auto-rollbackL · chaos → changesmeasure → respond → learn

This five-step ladder is an author synthesis of reliability practices rather than a canonical Google SRE model: Reactive → Managed → Proactive → Resilient → Adaptive.

Three lines evolve at every level. Measurement: uptime → basic service SLIs → SLO + error budget → user-journey SLO → business-aligned SLIs. Response: heroics → runbooks → automation → self-healing → auto-rollback. Learning: firefighting → postmortems → blameless reviews → game-days → chaos experiments that change policy and architecture.

An SLA is not a measurement: an SLI is an indicator, an SLO is a target, and an SLA is a business agreement with consequences. Use the ladder to name the current mode and adopt two or three practices from the next level without cargo culting Google.

How to use this model
01

Name the current maturity level without trying to look better than reality.

02

Compare what the team measures, how it responds, and how it learns after incidents.

03

Choose 2-3 practices that move the team to the next level.

Sources and related materials
Author materials