Skip to content
A stack of stones as a metaphor for balancing a reliable systemStachka 2023 conference logo
Stachka · Nizhny Novgorod · September 15, 2023

Designing Reliable Systems — Is It Worth the Effort?

From choosing risk to an engineering platform and culture

/ Designing Reliable Systems · Stachka 2023

Slide contents

  1. 1. Designing Reliable Systems — Is It Worth the Effort?

    From choosing risk to an engineering platform and culture

  2. 2. Alexander Polomodov

    Technical Director · Tinkoff

    Responsible for architecture

    Responsible for delivery management

    Develops engineering practices and platform capabilities

  3. 3. Reliability starts with risk and ends with culture

    Why reliability gets forgotten

    How to choose and control risk

    How to design reliable applications

    How to implement, operate, and build culture

  4. 4. Availability is time; reliability is likelihood

  5. 5. 01. Why is reliability often forgotten?

    Invisibility · assessment · evolution

  6. 6. Failure reveals reliability

  7. 7. 02. How do we choose and control risk?

    Application class · SLI/SLO/SLA · monitoring

  8. 8. Business impact sets the service level

  9. 9. An SLO becomes an SLA when breach has consequences

  10. 10. White-box and black-box look from opposite sides

  11. 11. RED maps four signals to three

  12. 12. 03. How do we design reliable applications?

    Domains · data · recovery · archetypes

  13. 13. Failure must stop at a domain boundary

  14. 14. Data survives by design

  15. 15. Recovery objectives choose the deployment model

  16. 16. Availability costs money

  17. 17. Reliable design prepares for change and recovery

  18. 18. 04. How do we implement and operate it?

    Measures · continuous delivery · Spirit platform

  19. 19. Delivery speed need not trade away stability

  20. 20. Continuous delivery is a team effort

  21. 21. A platform brings Dev, Sec, Data, and Ops into one flow

  22. 22. Working code is only part of good design

    A Philosophy of Software Design

    Do not finish the current task by adding needless complexity

    Design the system so it keeps working

    Treat design as the primary goal, not a by-product

  23. 23. The platform spans the path from code to incident

  24. 24. 05. How do we create a culture of reliability?

    Aristotle · Westrum · postmortems

  25. 25. Psychological safety is the first team factor

  26. 26. Generative culture investigates anomalies instead of hiding them

  27. 27. An incident review must change the system

  28. 28. Reliability pays when failure costs more

    Choose the acceptable risk first

    Encode it through SLI, SLO, and SLA

    Bound failures through architecture and platform

    Turn incidents into organizational learning

    Yes — when failure costs more than reliability

  29. 29. Sources cited in the original

    Site Reliability Engineering · Google

    Deployment Archetypes for Cloud Applications

    Building Secure and Reliable Systems

    Accelerate · A Philosophy of Software Design

  30. 30. Thank you!

    polomodov.tech

    Slides, notes, and other talks are available on the site

    Alexander Polomodov, Technical Director, Tinkoff

    @book_cube