Skip to content
PHDays Fest · May 22, 2025

Reliability and Security: Foundation or Optional?

Building systems where reliability and security aren't an afterthought

/ Reliability & Security · PHDays Fest 2025

Slide contents

  1. 1. Reliability and Security: Foundation or Optional?

    Building systems where reliability and security aren't an afterthought

  2. 2. What We'll Cover

    What's at stake: downtime, breaches, trust

    Emergent properties of systems and processes

    ATAM, risk, and threats as architecture process

    Google/Netflix/AWS cases, Spirit, takeaways

  3. 3. 01. Framing the Problem

    What's at stake when we talk reliability and security

  4. 4. Downtime and breaches cost real money

    Both hit revenue, reputation, and compliance

  5. 5. Complexity creates a limitless attack surface

    Hyper-connected: cloud, microservices, IoT

    Emergent behavior: a bug cascades

    Rapid change: hundreds of deploys a day

    Threats: more frequent and sophisticated

  6. 6. 02. Emergent Properties

    Reliability and security are whole-system properties, not features

  7. 7. Quality is a multidimensional system property

    From process and code to system and product

  8. 8. You can't buy a 'reliability module'

    Not plug-and-play — not a product

    Whole-system: code, infra, process, people

    Cross-cutting: spans every layer

    Foundational, not 'bolt it on later'

  9. 9. Engineer for failure

    Redundancy and graceful degradation

    Circuit breaker, bulkhead, backpressure

    SLOs for uptime/latency from the start

    Chaos and load tests before release

  10. 10. Fixing at the design stage is cheapest

    Shift left: the earlier you find it, the cheaper

  11. 11. 03. Architecture Processes

    Formalizing analysis: ATAM, risk, threats

  12. 12. ATAM turns decisions into deliberate trade-offs

    Quality attributes → analysis → risks and sensitivity points

  13. 13. Prioritize by likelihood × impact

    Risk Matrix and FMEA turn worries into data

  14. 14. Model threats systematically, not by gut

    STRIDE, PASTA, DREAD, MITRE ATT&CK, and more

  15. 15. 04. Big Tech Cases

    Google, Netflix, AWS — how the leaders do it

  16. 16. "Hope is not a strategy"

    Benjamin Treynor Sloss, Google

    SRE and SLOs/error budgets by default

    Design reviews with reliability/security

    Production Readiness Review before launch

    Blameless postmortems, learning culture

  17. 17. Resilience through chaos

    Chaos Monkey kills instances in prod

    Designs assuming things will break

    Team autonomy + responsibility

    Redundancy, fallback, load shedding

  18. 18. Security is Job Zero

    Werner Vogels, Amazon CTO

    Security before every other priority

    Encryption and strict IAM by default

    No launch with a known security issue

    Well-Architected: security + reliability

  19. 19. 5 principles from big tech

    Security is everyone's responsibility

    Reliability is a feature (SLOs, budgets)

    Design for failure: assume it breaks

    Leadership and culture set the tone

    continuous improvement is a journey, not a one-time project

  20. 20. 05. T-Technologies Case

    The Spirit platform: PaaS, observability, incidents, DevEx

  21. 21. The platform owns the infrastructure

    One portal, the whole service lifecycle underneath

  22. 22. Logs, metrics, and traces in one loop

    A unified search engine, alerting, and UI over the signals

  23. 23. An incident is a managed lifecycle

    From detection to postmortem and releases

  24. 24. An AI copilot with security built in

    Nestor — AI assistant in IDE

    Safeliner flags vulnerabilities in code

    Explanation + suggested fix

    Security review inside the dev flow

  25. 25. DevEx = flow minus cognitive load

    Feedback loops, cognitive load, flow state

  26. 26. Measure productivity across 5 axes

    Satisfaction, Performance, Activity, Communication, Efficiency

  27. 27. 06. Recommendations

    Strategy, culture, architecture, DevSecOps

  28. 28. Build — and avoid

    Build

    Reliability/security in OKRs

    Blameless security champions

    ATAM, threat modeling, DevSecOps

    Avoid

    Security later

    No incidents means fine

    InfoSec owns it

  29. 29. Where it's all heading

    AI ops helps, adds risk

    Zero trust as a security foundation

    Serverless/containers reshape resilience

    Regulators mandate 'secure by design'

  30. 30. 07. Key Takeaways & Next Steps

    What to remember and where to start on Monday

  31. 31. Foundation, not optional

    Reliability and security are built-in

    Mix architecture + risk + culture

    Learn from leaders, fit your context

    It's a journey: threats and complexity grow

    building securely and reliably lets the business move faster

  32. 32. Where to start on Monday

    Assess — evaluate your current state

    Quick wins — threat model + chaos

    Set targets — SLOs for critical services

    Plan — 12–18 month roadmap

  33. 33. Thank you!

    polomodov.tech

    All slides and links — in the Telegram channel

    Alexander Polomodov, Technical Director & Fellow, T-Technologies

    @book_cube