[2/2] The History and Evolution of Google TPUs (Category Engineering)
I’ll continue the story of Google’s TPUs from 2021, starting with TPU v4.
4. TPU v4 (2021): Optical Circuit Switching TPU v4 introduced optical circuit switching to accelerate communication between chips, increasingly important for more complex AI models. It delivered 275 TFLOPS per chip with improved optical interconnects.
5. TPU v5 and v5e (2023): Cost optimization TPU v5e and v5p focused on cost-effective training at scale, with improved energy efficiency, dynamic scaling and support for sparsity.
6. TPU v6 Trillium (2024): Performance optimization Trillium, the sixth generation, offers a 4.7-fold increase in compute performance per chip over TPU v5e. Its other characteristics include:
- Twice the High Bandwidth Memory (HBM) capacity and bandwidth.
- Twice the interchip interconnect bandwidth.
- 67% greater energy efficiency than TPU v5e.
- Scaling to 256 TPUs in a low-latency pod.
7. TPU v7 Ironwood (2025): Back to inference Ironwood, introduced in April 2025, is again a TPU specifically designed for inference, like TPU v1. Its headline specifications are:
- Scaling to 9,216 liquid-cooled chips.
- 42.5 exaflops of compute, 24 times the capacity of the most powerful supercomputer, El Capitan.
- 4,614 TFLOPS per chip and 192 GB of HBM, 6 times Trillium’s memory.
- Twice Trillium’s energy efficiency.
Google’s processors have come a long way. But how do they compare with NVIDIA? Here are the figures:
- NVIDIA H100: 3,958 TFLOPS (FP8), 80 GB HBM3, 3.35 TB/s memory bandwidth.
- NVIDIA H200: 3,958 TFLOPS (FP8), 141 GB HBM3e, 4.8 TB/s memory bandwidth.
- TPU v6 Trillium: around 2 PFLOPS FP16 for tensor operations.
- TPU v7 Ironwood: 4,614 TFLOPS per chip, 192 GB HBM, 7.37 TB/s bandwidth.
The FLOPS figures look healthy. On efficiency, the independent comparison cited here reports a 50-70% lower cost per billion training tokens for TPU v5e than for NVIDIA H100 clusters. TPU v5e also uses substantially less energy for a comparable workload: an H100 can consume around 5 times as much as a loaded TPU v5e chip. The reported practical comparisons are roughly:
- For training GPT-scale models, TPUs are 4-10 times as cost-effective as GPUs.
- For inference, TPU v5e provides 3 times the throughput per dollar.
- TPU v4 runs 1.2-1.7 times as fast as NVIDIA A100 while using 1.3-1.9 times less energy.
TPUs therefore have both advantages and drawbacks: (+) Specialization in tensor operations and deep learning. (+) High energy and cost efficiency. (+) Google Cloud integration and optimization for TensorFlow/JAX. (+) Scalability within the Google Cloud ecosystem. (-) Availability only through Google Cloud. (-) Less flexibility than GPUs across different types of computation. (-) A smaller development ecosystem than CUDA’s. (-) Less memory per chip than the newest GPUs, until recently.
Google seems well positioned with its processors for generative AI and ML, and likely to keep building infrastructure that gives it a substantial competitive advantage during this AI-app boom. For everyone else, the chips either mean vendor lock-in if used directly, or provide a benchmark for what to aim for in the future.
#AI #ML #Software #Engineering #Architecture #Infrastructure #Data