
Designing Reliable Systems — Is It Worth the Effort?
From choosing risk to an engineering platform and culture
Slide contents
1. Designing Reliable Systems — Is It Worth the Effort?
From choosing risk to an engineering platform and culture
2. Alexander Polomodov
Technical Director · Tinkoff
Responsible for architecture
Responsible for delivery management
Develops engineering practices and platform capabilities
3. Reliability starts with risk and ends with culture
Why reliability gets forgotten
How to choose and control risk
How to design reliable applications
How to implement, operate, and build culture
4. Availability is time; reliability is likelihood
5. 01. Why is reliability often forgotten?
Invisibility · assessment · evolution
6. Failure reveals reliability
7. 02. How do we choose and control risk?
Application class · SLI/SLO/SLA · monitoring
8. Business impact sets the service level
9. An SLO becomes an SLA when breach has consequences
10. White-box and black-box look from opposite sides
11. RED maps four signals to three
12. 03. How do we design reliable applications?
Domains · data · recovery · archetypes
13. Failure must stop at a domain boundary
14. Data survives by design
15. Recovery objectives choose the deployment model
16. Availability costs money
17. Reliable design prepares for change and recovery
18. 04. How do we implement and operate it?
Measures · continuous delivery · Spirit platform
19. Delivery speed need not trade away stability
20. Continuous delivery is a team effort
21. A platform brings Dev, Sec, Data, and Ops into one flow
22. Working code is only part of good design
A Philosophy of Software Design
Do not finish the current task by adding needless complexity
Design the system so it keeps working
Treat design as the primary goal, not a by-product
23. The platform spans the path from code to incident
24. 05. How do we create a culture of reliability?
Aristotle · Westrum · postmortems
25. Psychological safety is the first team factor
26. Generative culture investigates anomalies instead of hiding them
27. An incident review must change the system
28. Reliability pays when failure costs more
Choose the acceptable risk first
Encode it through SLI, SLO, and SLA
Bound failures through architecture and platform
Turn incidents into organizational learning
Yes — when failure costs more than reliability
29. Sources cited in the original
Site Reliability Engineering · Google
Deployment Archetypes for Cloud Applications
Building Secure and Reliable Systems
Accelerate · A Philosophy of Software Design
30. Thank you!
polomodov.tech
Slides, notes, and other talks are available on the site
Alexander Polomodov, Technical Director, Tinkoff
@book_cube