
Designing Reliable Systems
From choosing risk to an engineering platform and culture
Slide contents
1. Designing Reliable Systems
From choosing risk to an engineering platform and culture
2. Alexander Polomodov
Technical Director · Tinkoff
Responsible for architecture
Responsible for delivery management
Develops engineering practices and platform capabilities
3. Reliability starts with risk and ends with culture
Why reliability gets forgotten
How to choose and control risk
How to design reliable applications
How to implement, operate, and build culture
4. Availability is time; reliability is likelihood
5. 01. Why is reliability often forgotten?
Invisibility · assessment · evolution
6. Failure reveals reliability
7. 02. How do we choose and control risk?
Application class · SLI/SLO/SLA · monitoring
8. Business impact sets the service level
9. An SLO becomes an SLA when breach has consequences
10. White-box and black-box look from opposite sides
11. RED maps four signals to three
12. 03. How do we design reliable applications?
Domains · data · recovery · archetypes
13. Failure must stop at a domain boundary
14. Data survives by design
15. Recovery objectives choose the deployment model
16. Six patterns contain failure
17. Availability costs money
18. Reliable design prepares for change and recovery
19. 04. How do we implement and operate it?
Measures · continuous delivery · Spirit platform
20. Delivery speed need not trade away stability
21. Continuous delivery is a team effort
22. A platform brings Dev, Sec, Data, and Ops into one flow
23. Working code is only part of good design
A Philosophy of Software Design
Do not finish the current task by adding needless complexity
Design the system so it keeps working
Treat design as the primary goal, not a by-product
24. The platform spans the path from code to incident
25. 05. How do we create a culture of reliability?
Aristotle · Westrum · postmortems
26. Psychological safety is the first team factor
27. Generative culture investigates anomalies instead of hiding them
28. An incident review must change the system
29. Reliability pays when failure costs more
Choose the acceptable risk first
Encode it through SLI, SLO, and SLA
Bound failures through architecture and platform
Turn incidents into organizational learning
Yes — when failure costs more than reliability
30. Sources cited in the original
Site Reliability Engineering · Google
Deployment Archetypes for Cloud Applications
Building Secure and Reliable Systems
Accelerate · A Philosophy of Software Design
31. Thank you!
polomodov.tech
Slides, notes, and other talks are available on the site
Alexander Polomodov, Technical Director, Tinkoff
@book_cube