Skip to content
back to the archive page
#AI

CodeFest Russia: Where Is AI Hardware Heading? (Category AI)

An interesting talk by Valentin Mamedov from Sber, analyzing the AI hardware market. He begins with the market situation at the time:

  • ChatGPT is the world's 5th most popular website;
  • NVIDIA controls 92% of the data-center GPU market;
  • An H100 costs $30,000, while a system with 8 of these cards costs $250,000.

Comparing consumer and server GPUs:

  • An RTX 4090 costs about $2k and an H100 about $30k—a 15-fold difference despite similar peak performance;
  • The RTX 4090 has 1 TB/s of memory bandwidth, compared with 3 TB/s for the H100;
  • The H100 is optimized for BF16, a numerical format used in neural networks.

The main challenge, however, is the development ecosystem:

  • NVIDIA has spent 20 years developing it, with mature frameworks, bug fixes, and all the supporting work;
  • Flash Attention arrived on AMD a year after NVIDIA;
  • Switching to alternatives involves risk and lost development time.

On lower-cost alternatives:

  • An Apple M4 Pro laptop is comparable to an RTX 4090 for many ML tasks;
  • A $20,000 Apple M cluster can run DeepSeek v3 at 20 tokens per second;
  • Consumer GPUs cannot be efficiently combined into a training cluster; the talk describes this capability as available only on NVIDIA's industrial cards.

The talk's main conclusions:

  • NVIDIA's monopoly rests on software rather than hardware: the CUDA ecosystem is crucial;
  • Competitors are growing, attracted by NVIDIA's estimated 57% margin on industrial cards. ||As in a gold rush, the shovel seller wins.||
  • Competitors include Cerebras ($2.5M per chip), AWS Trainium, and Google TPU v7. ||The speaker estimates that Google's TPU chips cost it half as much as using NVIDIA cards.||
  • An RTX 4090 is enough for personal inference or even a small business;
  • An Apple M cluster is an option for inference at a larger scale.

#AI #Engineering #Software #ML #Hardware #Future #DevEx

Open video on YouTube