Reliability and Security: Foundation or Optional?
Building systems where reliability and security aren't an afterthought
Slide contents
1. Reliability and Security: Foundation or Optional?
Building systems where reliability and security aren't an afterthought
2. What We'll Cover
What's at stake: downtime, breaches, trust
Emergent properties of systems and processes
ATAM, risk, and threats as architecture process
Google/Netflix/AWS cases, Spirit, takeaways
3. 01. Framing the Problem
What's at stake when we talk reliability and security
4. Downtime and breaches cost real money
Both hit revenue, reputation, and compliance
5. Complexity creates a limitless attack surface
Hyper-connected: cloud, microservices, IoT
Emergent behavior: a bug cascades
Rapid change: hundreds of deploys a day
Threats: more frequent and sophisticated
6. 02. Emergent Properties
Reliability and security are whole-system properties, not features
7. Quality is a multidimensional system property
From process and code to system and product
8. You can't buy a 'reliability module'
Not plug-and-play — not a product
Whole-system: code, infra, process, people
Cross-cutting: spans every layer
Foundational, not 'bolt it on later'
9. Engineer for failure
Redundancy and graceful degradation
Circuit breaker, bulkhead, backpressure
SLOs for uptime/latency from the start
Chaos and load tests before release
10. Fixing at the design stage is cheapest
Shift left: the earlier you find it, the cheaper
11. 03. Architecture Processes
Formalizing analysis: ATAM, risk, threats
12. ATAM turns decisions into deliberate trade-offs
Quality attributes → analysis → risks and sensitivity points
13. Prioritize by likelihood × impact
Risk Matrix and FMEA turn worries into data
14. Model threats systematically, not by gut
STRIDE, PASTA, DREAD, MITRE ATT&CK, and more
15. 04. Big Tech Cases
Google, Netflix, AWS — how the leaders do it
16. "Hope is not a strategy"
Benjamin Treynor Sloss, Google
SRE and SLOs/error budgets by default
Design reviews with reliability/security
Production Readiness Review before launch
Blameless postmortems, learning culture
17. Resilience through chaos
Chaos Monkey kills instances in prod
Designs assuming things will break
Team autonomy + responsibility
Redundancy, fallback, load shedding
18. Security is Job Zero
Werner Vogels, Amazon CTO
Security before every other priority
Encryption and strict IAM by default
No launch with a known security issue
Well-Architected: security + reliability
19. 5 principles from big tech
Security is everyone's responsibility
Reliability is a feature (SLOs, budgets)
Design for failure: assume it breaks
Leadership and culture set the tone
continuous improvement is a journey, not a one-time project
20. 05. T-Technologies Case
The Spirit platform: PaaS, observability, incidents, DevEx
21. The platform owns the infrastructure
One portal, the whole service lifecycle underneath
22. Logs, metrics, and traces in one loop
A unified search engine, alerting, and UI over the signals
23. An incident is a managed lifecycle
From detection to postmortem and releases
24. An AI copilot with security built in
Nestor — AI assistant in IDE
Safeliner flags vulnerabilities in code
Explanation + suggested fix
Security review inside the dev flow
25. DevEx = flow minus cognitive load
Feedback loops, cognitive load, flow state
26. Measure productivity across 5 axes
Satisfaction, Performance, Activity, Communication, Efficiency
27. 06. Recommendations
Strategy, culture, architecture, DevSecOps
28. Build — and avoid
Build
Reliability/security in OKRs
Blameless security champions
ATAM, threat modeling, DevSecOps
Avoid
Security later
No incidents means fine
InfoSec owns it
29. Where it's all heading
AI ops helps, adds risk
Zero trust as a security foundation
Serverless/containers reshape resilience
Regulators mandate 'secure by design'
30. 07. Key Takeaways & Next Steps
What to remember and where to start on Monday
31. Foundation, not optional
Reliability and security are built-in
Mix architecture + risk + culture
Learn from leaders, fit your context
It's a journey: threats and complexity grow
building securely and reliably lets the business move faster
32. Where to start on Monday
Assess — evaluate your current state
Quick wins — threat model + chaos
Set targets — SLOs for critical services
Plan — 12–18 month roadmap
33. Thank you!
polomodov.tech
All slides and links — in the Telegram channel
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@book_cube
