Taxi Reliability Tools — Alexander Fisher, Yandex Taxi
An interesting talk by Alexander Fisher about processes and tools for improving reliability. It is short and has six parts.
1. Taxi services and reliability An introduction to service criticality, load patterns, architecture and technologies, followed by reliability metrics:
- Standard measures: SLOs expressed in nines and MTTR, mean time to recovery.
- Less familiar ones: MTTRC, mean time to root cause, measuring how quickly the underlying cause is found; and ROCOF, rate of occurrences of failures, measuring failure frequency.
The rest of the talk presents tools for coping with complexity.
2. Chaos The team uses fault injection. Its data plane consists of Lua scripts in nginx traffic balancers, while a control plane manages the configuration of injected failures. The tool can:
- Slow a service down while remaining within its propagation deadline.
- Simulate a metastable state in which a service collapses under load and cannot recover without someone helping it :)
- Return HTTP 500 errors for a specified percentage of requests.
3. Load shedding: degradation, retries and limits The team has:
- A graceful-degradation mode that disables secondary functionality, adding 20% capacity, with 50% planned.
- Retry controls to address request amplification during failures.
- Service tiers based on how critical each service is to the core function: getting a taxi ride.
4. Virtual orders Virtual orders simulate load to test the entire system holistically rather than a single API endpoint. It is a useful way to examine nonlinear dependencies and the unpredictable behavior of a genuinely complex system under different load patterns.
5. Observability and an eventboard The speaker shows observability dashboards and an interesting board of production events, plus a button to “roll back everything from the last n minutes.” This helps reduce MTTR and speed up recovery.
6. Incident simulation This section covers incident coordinators, simulated failures and weekly incident-response drills.
At the end, the speaker describes experimental tools:
- Autorecovery, a bot that automatically resolves incidents. It is still being evaluated in dry-run mode.
- SRE GPT, a tool to help identify root causes.
I liked the talk: an engaging account, interesting ideas and practical recommendations that could be used in your own reliability-improvement processes.
#SRE #DistributedSystems #Reliability #Architecture #Software #Processes #Management