SRE Maturity
This five-step ladder is an author synthesis of reliability practices rather than a canonical Google SRE model: Reactive → Managed → Proactive → Resilient → Adaptive.
Three lines evolve at every level. Measurement: uptime → basic service SLIs → SLO + error budget → user-journey SLO → business-aligned SLIs. Response: heroics → runbooks → automation → self-healing → auto-rollback. Learning: firefighting → postmortems → blameless reviews → game-days → chaos experiments that change policy and architecture.
An SLA is not a measurement: an SLI is an indicator, an SLO is a target, and an SLA is a business agreement with consequences. Use the ladder to name the current mode and adopt two or three practices from the next level without cargo culting Google.
Name the current maturity level without trying to look better than reality.
Compare what the team measures, how it responds, and how it learns after incidents.
Choose 2-3 practices that move the team to the next level.