Retrospective on "In-Datacenter Performance Analysis of a Tensor Processing Unit" (Рубрика AI)
Интересная заметка на две страницы про то, как и почему появился TPU в Google, продолжая тему прошлого поста про железо для ML/AI. TPU оказался отличным решением и поддерживал продуктовизацию deep learning инициатив внутри Google уже 10 лет подряд, начиная с начала 2015 года, когда он появился в проде. Завтра будет заметка побольше про всю историю эволюции TPU, а для завтравки рекомендую прочитать этот мини whitepaper, где есть такое объяснение старту проекта
A key signal soon afterward was that matrix multiplication exceeded 1% of CPU fleet cycles in Google Wide Profiling. Another signal was the analysis by Jeff Dean (a Google Fellow, now the Chief Scientist) that processing a few minutes of speech or video by 100M users would require doubling or tripling the size of the CPU fleet. Other options were clearly required.
#AI #Infrastructure #Engineering #Architecture