[6/7] Meta’s Hyperscale Infrastructure: Designing Scalable Systems (Category Infrastructure)
[6/7] Meta’s Hyperscale Infrastructure: Overview and Insights — Designing Scalable Systems (Category #Infrastructure)
We continue this excellent article from Meta, a company banned in Russia, with a look at how its engineers design scalable applications. Earlier parts: 1, 2, 3, 4 and 5.
Centralization versus decentralization Planet-scale infrastructure has historically been associated with decentralized architectures such as BGP and BitTorrent. They scale well without a SPOF, or single point of failure. Meta’s experience, however, suggests that inside a data center—where resources are relatively reliable and managed by one organization—centralized controllers often simplify systems while providing sufficient scalability. They can also make better decisions globally than many local agents. Meta has therefore deliberately moved many initially distributed designs toward centralized management. For example:
- The internal data-center network, or fabric, still uses BGP for interoperability. A central controller manages routing, reoptimizing traffic paths during congestion or link failures instead of waiting for BGP’s slow convergence.
- Meta’s global backbone, or WAN, initially used the decentralized bandwidth-reservation protocol RSVP-TE. It later moved to a central controller that calculates optimal paths for flows between data centers and provisions backup paths for common failures in advance. This improved bandwidth utilization substantially and simplified network management.
The article summarizes the approach in this insight:
Insight 9 : In a datacenter environment, we prefer centralized controllers over decentralized ones due to their simplicity and ability to make higher-quality decisions. In many cases, a hybrid approach - a centralized control plane combined with a decentralized data plane-provides the best of both worlds.
One detailed example is the hybrid service mesh ServiceRouter, an attempt to get the best of both worlds. It handles billions of calls per second between microservices distributed across millions of L7 software routers. Traditional service meshes, such as Istio, give each application a local proxy through which all incoming and outgoing calls pass. Meta abandoned that arrangement in ServiceRouter: as mentioned earlier, ~99% of requests run without a sidecar proxy. Instead:
- The control plane is centralized. It aggregates service information and global network metrics, computes optimal routing rules and stores them in a RIB (Routing Information Base). This is built on the distributed database Delos, which uses the Paxos protocol and is therefore distributed and fault-tolerant. The central ServiceRouter controllers compute global decisions, while the data plane performs the actual routing.
- The data plane consists of decentralized L7 routers. They automatically retrieve the information they need from the RIB, caching a small relevant subset, and operate autonomously without constant involvement from a central coordinator.
This design provides:
- Simple management: the whole picture is visible centrally.
- Scalability: there is no bottleneck through which all traffic must pass.
The result is full service-mesh functionality—load balancing, retries, discovery and monitoring—with minimal resource use and the ability to optimize load distribution globally.
The final post in this series will discuss future directions for Meta’s infrastructure and architecture, one of the most interesting sections.
#Infrastructure #PlatformEngineering #Architecture #DistributedSystems #SystemDesign #Engineering #Software #DevEx #DevOps