CodeFest Russia: Where Is AI Hardware Heading? (Category AI)
An interesting talk by Valentin Mamedov from Sber, analyzing the AI hardware market. He begins with the market situation at the time:
- ChatGPT is the world's 5th most popular website;
- NVIDIA controls 92% of the data-center GPU market;
- An H100 costs $30,000, while a system with 8 of these cards costs $250,000.
Comparing consumer and server GPUs:
- An RTX 4090 costs about $2k and an H100 about $30k—a 15-fold difference despite similar peak performance;
- The RTX 4090 has 1 TB/s of memory bandwidth, compared with 3 TB/s for the H100;
- The H100 is optimized for BF16, a numerical format used in neural networks.
The main challenge, however, is the development ecosystem:
- NVIDIA has spent 20 years developing it, with mature frameworks, bug fixes, and all the supporting work;
- Flash Attention arrived on AMD a year after NVIDIA;
- Switching to alternatives involves risk and lost development time.
On lower-cost alternatives:
- An Apple M4 Pro laptop is comparable to an RTX 4090 for many ML tasks;
- A $20,000 Apple M cluster can run DeepSeek v3 at 20 tokens per second;
- Consumer GPUs cannot be efficiently combined into a training cluster; the talk describes this capability as available only on NVIDIA's industrial cards.
The talk's main conclusions:
- NVIDIA's monopoly rests on software rather than hardware: the CUDA ecosystem is crucial;
- Competitors are growing, attracted by NVIDIA's estimated 57% margin on industrial cards. ||As in a gold rush, the shovel seller wins.||
- Competitors include Cerebras ($2.5M per chip), AWS Trainium, and Google TPU v7. ||The speaker estimates that Google's TPU chips cost it half as much as using NVIDIA cards.||
- An RTX 4090 is enough for personal inference or even a small business;
- An Apple M cluster is an option for inference at a larger scale.
#AI #Engineering #Software #ML #Hardware #Future #DevEx