Skip to content
SmartDev 2023 visual identity
SmartDev · September 21, 2023

Designing Reliable Systems

From choosing risk to an engineering platform and culture

/ Designing Reliable Systems · SmartDev 2023

Slide contents

  1. 1. Designing Reliable Systems

    From choosing risk to an engineering platform and culture

  2. 2. Alexander Polomodov

    Technical Director · Tinkoff

    Responsible for architecture

    Responsible for delivery management

    Develops engineering practices and platform capabilities

  3. 3. Reliability starts with risk and ends with culture

    Why reliability gets forgotten

    How to choose and control risk

    How to design reliable applications

    How to implement, operate, and build culture

  4. 4. Availability is time; reliability is likelihood

  5. 5. 01. Why is reliability often forgotten?

    Invisibility · assessment · evolution

  6. 6. Failure reveals reliability

  7. 7. 02. How do we choose and control risk?

    Application class · SLI/SLO/SLA · monitoring

  8. 8. Business impact sets the service level

  9. 9. An SLO becomes an SLA when breach has consequences

  10. 10. White-box and black-box look from opposite sides

  11. 11. RED maps four signals to three

  12. 12. 03. How do we design reliable applications?

    Domains · data · recovery · archetypes

  13. 13. Failure must stop at a domain boundary

  14. 14. Data survives by design

  15. 15. Recovery objectives choose the deployment model

  16. 16. Six patterns contain failure

  17. 17. Availability costs money

  18. 18. Reliable design prepares for change and recovery

  19. 19. 04. How do we implement and operate it?

    Measures · continuous delivery · Spirit platform

  20. 20. Delivery speed need not trade away stability

  21. 21. Continuous delivery is a team effort

  22. 22. A platform brings Dev, Sec, Data, and Ops into one flow

  23. 23. Working code is only part of good design

    A Philosophy of Software Design

    Do not finish the current task by adding needless complexity

    Design the system so it keeps working

    Treat design as the primary goal, not a by-product

  24. 24. The platform spans the path from code to incident

  25. 25. 05. How do we create a culture of reliability?

    Aristotle · Westrum · postmortems

  26. 26. Psychological safety is the first team factor

  27. 27. Generative culture investigates anomalies instead of hiding them

  28. 28. An incident review must change the system

  29. 29. Reliability pays when failure costs more

    Choose the acceptable risk first

    Encode it through SLI, SLO, and SLA

    Bound failures through architecture and platform

    Turn incidents into organizational learning

    Yes — when failure costs more than reliability

  30. 30. Sources cited in the original

    Site Reliability Engineering · Google

    Deployment Archetypes for Cloud Applications

    Building Secure and Reliable Systems

    Accelerate · A Philosophy of Software Design

  31. 31. Thank you!

    polomodov.tech

    Slides, notes, and other talks are available on the site

    Alexander Polomodov, Technical Director, Tinkoff

    @book_cube