Skip to content
back to the archive page
#Architecture

[2/3] Reliability and Security: Optional Extras or the Foundation of Modern IT Systems? (Category Architecture)

Continuing the discussion of my PHDays talk, here are the remaining points and materials. The previous post ended with examples from Western big tech. Next, I wanted to explain how this works at T and offer some advice on making reliability and security fundamental to development rather than optional extras :)

Our company has a developer ecosystem centered on Spirit, a PaaS platform with many components. Let us start with the principles the team follows when building it:

— A repeatable, transparent code lifecycle within delivery processes. Code passes through the same stages: writing, building, different kinds of testing, documentation updates, deployment to environments, post-deployment quality checks, support, and retirement. Any developer should understand what is happening and which stages they need to go through. — Managing the codebase and assessing its quality at every stage. Codebase quality is a core characteristic of a technical product. Keeping it consistently high indicates stable, mature development processes. — Moving toward Inner Source within the company. Reusing colleagues’ code and work reduces development costs and speeds up delivery of business features. The platform should make sharing code quick and painless. — Focusing on outcomes. Giving developers all the tools they need to manage a product’s lifecycle saves a great deal of time. Creating or fixing features, running tests, and other work become easier. We want developers to spend 80% of their working time building business features, without distractions from routine tasks and infrastructure complexity. Users ultimately benefit.

The components are shown in this diagram. Some particularly relevant to reliability and security are:

  • Sage, an observability platform for centralized collection and analysis of telemetry from all company services. It makes business applications and IT infrastructure visible and helps keep services running.
  • FineDog, an incident-management platform that helps T-Bank detect service failures promptly, reduce recovery time, and prevent recurring incidents.
  • Nestor, a copilot for code suggestions and chat inside the IDE, along with other tools related to code and beyond.
  • Safeliner, an AI security assistant for development teams. It can run within CI/CD stages or integrate into the IDE as a plugin.

The Spirit team also actively works on developer experience. They discussed this in “Why DevEx Matters When Building an IDP and How to Measure It,” a talk I wrote about. This lets them collect engineering-experience metrics and identify what needs improving in the platform.

A mature platform and established reliability and security processes are great, but what if your company does not have them?

The changes need to happen at several levels:

  • Company strategy: make these issues priorities, define clear goals and include them in OKRs, and agree on resources.
  • Company culture: make security and reliability a shared responsibility of teams.
  • Architecture and design: account for these architectural characteristics and apply good practices.
  • Processes: integrate them into pipelines through DevSecOps and shift-left practices.

Manage all of this as a major change initiative, using change-management approaches.

#SRE #SystemDesign #Software #Architecture #Metrics #SoftwareArchitecture #Engineering