Skip to content
#Engineering

[1/2] The History of Google TPU and its Evolution (Category Engineering)

#Engineering #AI #ML #Software #Architecture #Infrastructure #Data

Literally yesterday I told About the report "CodeFest Russia: Where is the iron for neurons rolling?", and today I decided to talk about how Google got its own processors for multiplying matrices. In fact, it all began in the early 2000s, when Google was actively implementing ML models in its products. (search, translator, photo). They did it so successfully that with the advent of complex neural networks. (We remember the extravaganza with CNN Networks and ImageNet. 2012) They wanted to incorporate them into their products, but the computing power of both learning and inferencing is exponential. In 2013 Google has learned that if nothing is changed, will have to double Number of data centers on existing equipment (then existing CPU and GPU). As a result, the guys thought and came up with a project to create a TPU with such goals. Application-Specific Integrated Circuit (ASIC)which will provide 10Multiple cost/performance advantage in inferencing compared to GPU

  • Build a solution quickly. (ASAP or on short notice) Achieve high performance at scale with new workloads out of the box while remaining cost-effective

The project was run by Norman Jupi. (Norman "Norm" Jouppi)He is a computer architect and Google Fellow. Norman previously distinguished himself in the design of MIPS processors. Prior to Google, he worked at HP Labs, where he led a lab of advanced architectures. According to Jonathan Ross, one of the first TPU engineers. (later founder of the company Groq)Three separate groups at Google were developing AI accelerators, but it was the design of the TPU that was ultimately chosen for implementation.

As far as the results are concerned, they are good, especially considering that the seventh version of the TPU is already available. And this is what they looked like in dynamics. (I was guided by the article.TPU transformation: A look back at 10 years of our AI-specialized chipsfrom Google Cloud)

1. TPU v1 (2015) - Infernance. It was developed at a record speed - just for the 15 months from the start of the project to the deployment in Google data centers at the beginning 2015 years. This speed was achieved through the use of "obsolete" 28Nanometer process technology and relatively low clock frequency 700 MHz, which made it relatively easy to meet deadlines. The energy consumption was 40 Watts, and productivity 92 TOPS for 8- bit integers. This processor was designed only for inferencing. The chip showed performance in 15-30 higher than the current CPU and GPU, 30-80Multiple advantages in energy efficiency. 2. TPU v2 (2017) - Inferencing + Training By the end. 2014 When the TPU v1 was in production, Google realized that learning was becoming a limiting factor. TPU v2, presented in 2017 It was a revolutionary step – it was no longer just a chip, but a full-fledged supercomputer system. Key innovations of TPU v2: Support for both learning and inferencing

  • TPU Pod - network of 256 TPU v2 chips with high-bandwidth interconnection network
  • Productivity: 180 TFLOPS
  • Memory: 64 GBM 3. TPU v3 (2018) - Liquid cooling. TPU v3 introduced liquid cooling for efficient heat management, allowing for higher performance levels. Productivity has increased to 420 TFLOPS, improved interconnection network and memory bandwidth.

Continuation of history in next post.

#AI #ML #Software #Engineering #Architecture #Infrastructure #Data