Reliability and Security: Foundation or Optional?
Building systems where reliability and security aren't an afterthought
Building systems where reliability and security aren't an afterthought
Building systems where reliability and security aren't an afterthought
What's at stake: downtime, breaches, trust
Emergent properties of systems and processes
ATAM, risk, and threats as architecture process
Google/Netflix/AWS cases, Spirit, takeaways
What's at stake when we talk reliability and security
Both hit revenue, reputation, and compliance
Hyper-connected: cloud, microservices, IoT
Emergent behavior: a bug cascades
Rapid change: hundreds of deploys a day
Threats: more frequent and sophisticated
Reliability and security are whole-system properties, not features
From process and code to system and product
Not plug-and-play — not a product
Whole-system: code, infra, process, people
Cross-cutting: spans every layer
Foundational, not 'bolt it on later'
Redundancy and graceful degradation
Circuit breaker, bulkhead, backpressure
SLOs for uptime/latency from the start
Chaos and load tests before release
Shift left: the earlier you find it, the cheaper
Formalizing analysis: ATAM, risk, threats
Quality attributes → analysis → risks and sensitivity points
Risk Matrix and FMEA turn worries into data
STRIDE, PASTA, DREAD, MITRE ATT&CK, and more
Google, Netflix, AWS — how the leaders do it
Benjamin Treynor Sloss, Google
SRE and SLOs/error budgets by default
Design reviews with reliability/security
Production Readiness Review before launch
Blameless postmortems, learning culture
Chaos Monkey kills instances in prod
Designs assuming things will break
Team autonomy + responsibility
Redundancy, fallback, load shedding
Werner Vogels, Amazon CTO
Security before every other priority
Encryption and strict IAM by default
No launch with a known security issue
Well-Architected: security + reliability
Security is everyone's responsibility
Reliability is a feature (SLOs, budgets)
Design for failure: assume it breaks
Leadership and culture set the tone
continuous improvement is a journey, not a one-time project
The Spirit platform: PaaS, observability, incidents, DevEx
One portal, the whole service lifecycle underneath
A unified search engine, alerting, and UI over the signals
From detection to postmortem and releases
Nestor — AI assistant in IDE
Safeliner flags vulnerabilities in code
Explanation + suggested fix
Security review inside the dev flow
Feedback loops, cognitive load, flow state
Satisfaction, Performance, Activity, Communication, Efficiency
Strategy, culture, architecture, DevSecOps
Build
Reliability/security in OKRs
Blameless security champions
ATAM, threat modeling, DevSecOps
Avoid
Security later
No incidents means fine
InfoSec owns it
AI ops helps, adds risk
Zero trust as a security foundation
Serverless/containers reshape resilience
Regulators mandate 'secure by design'
What to remember and where to start on Monday
Reliability and security are built-in
Mix architecture + risk + culture
Learn from leaders, fit your context
It's a journey: threats and complexity grow
building securely and reliably lets the business move faster
Assess — evaluate your current state
Quick wins — threat model + chaos
Set targets — SLOs for critical services
Plan — 12–18 month roadmap
polomodov.tech
All slides and links — in the Telegram channel
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@book_cube