Trends Across the AI Frontier (AI column)
I watched a short one recently. speech George Cameron, Co-Founder and CPO Company Artificial AnalysisAn independent AI benchmarking company founded in October 2023 years. They specialize in independent testing. 150 A variety of AI models across a wide range of metrics, including intelligence, performance, cost, and speed, posting the results on the site. artificialanalysis.ai. The report itself was short but rich in analytics and forecasts, with key insights presented below.
1. Multiple boundaries in AI George emphasizes the central idea of his talk: there are not one, but several boundaries in AI. (frontiers)Developers do not always have to use the most intelligent model. The report explores four key boundaries:
- Reasoning models (Reasoning Models)
- Open weights. (Open Weights)
- Cost. (Cost)
- Speed. (Speed)
2. The Boundary of Reasoning Models: Productivity Tradeoffs Georg shows analytics with incisions on an order of magnitude between conventional and reasoning models:
- Verbality of models -- GPT-4.1: 7 Millions of tokens to run an intelligence index -- O4 Mini High: 72 million (into 10 more) -- Gemini 2.5 Pro: 130 million
- Delayed response. -- GPT-4.1: 4,7 seconds for a full answer
- O4 Mini High: more 40 seconds This is critical for agent systems where 30 successive requests may be 5 minutes 30 seconds.
3. The Open Limit: Bridging the Gap George It shows a dramatic reduction in the gap between open and proprietary models. A special role is played by Chinese AI laboratories: DeepSeek leads in both reasoning and non-reasoning Alibaba with the Qwen series 3 Takes second place in reasoning Meta and Nvidia also compete with Llama-based models DeepSeek R1, released in January 2025 year, shows performance comparable to the leading proprietary models, achieving 79.8% on AIME 2024 against 79.2O1 has percent.
4. The Cost Boundary: Dramatic Decline Ordinal changes in value:
- O3: $2,000 to run their intelligence index benchmark
- 4.1: in30 cheaper than O1 The cost of accessing a new level of intelligence is halved in a few months (The cost of access to GPT intelligence4 fell over 100 midway 2023 year). According to research, the cost of training frontier models is growing in 2,4 yearly 2016 years.
5. Speed Boundary: Significant productivity growth Increasing the speed of withdrawal of tokens: GPT-4 into 2023 year: ~40 Tokens per second. The same level of intelligence now: more 300 Tokens per second. Such speed changes are tied to technological improvements.
- Mixture of Experts (MoE) Models activate only part of the parameters at inferencing Distillation – creating smarter smaller models Optimization of software - flash attention, speculative decoding Improvement of equipment - B200 achieves more 1000 tokens per second
Main conclusions
- Despite improved efficiency and lower cost, Cameron predicts continued growth in computing demand: Larger models - DeepSeek has more 600 billion-dollar Insatiable demand for intelligence “Speaky” reasoning models require more inferencing calculations
- Agents with 20-30-100+ successive requests
- Strategic recommendations Measure tradeoffs in your apps rather than relying only on the price per token Build with future cost reductions in mind – what is not possible today can become available through 6 months Consider the cost structure of the application when choosing models
In general, Georg shows interesting results of benchmarks, which show interesting trends. And I personally after the speech with interest poked other studies from the site. Artificial Analysis.
#AI #Software #Engineering #Management #Metrics #ML #Leadership