
Designing Reliable Systems — Is It Worth the Effort?
From choosing risk to an engineering platform and culture

From choosing risk to an engineering platform and culture
From choosing risk to an engineering platform and culture
Technical Director · Tinkoff
Responsible for architecture
Responsible for delivery management
Develops engineering practices and platform capabilities
Why reliability gets forgotten
How to choose and control risk
How to design reliable applications
How to implement, operate, and build culture
Invisibility · assessment · evolution
Application class · SLI/SLO/SLA · monitoring
Domains · data · recovery · archetypes
Measures · continuous delivery · Spirit platform
A Philosophy of Software Design
Do not finish the current task by adding needless complexity
Design the system so it keeps working
Treat design as the primary goal, not a by-product
Aristotle · Westrum · postmortems
Choose the acceptable risk first
Encode it through SLI, SLO, and SLA
Bound failures through architecture and platform
Turn incidents into organizational learning
Yes — when failure costs more than reliability
Site Reliability Engineering · Google
Deployment Archetypes for Cloud Applications
Building Secure and Reliable Systems
Accelerate · A Philosophy of Software Design
polomodov.tech
Slides, notes, and other talks are available on the site
Alexander Polomodov, Technical Director, Tinkoff
@book_cube