
The Economics of AI Development
Why tokens get cheaper while budgets grow—and why accepted work is the unit that matters

Why tokens get cheaper while budgets grow—and why accepted work is the unit that matters
Why tokens get cheaper while budgets grow—and why accepted work is the unit that matters
Original forecast synthesized from market, FinOps, and agentic-workload evidence
Three independent signals with different sample boundaries
A synthesis of historical trends, tariffs, and demand expansion
Spend expands through tasks, calls, context, and infrastructure
Money, steps, time, and authority are bounded across the workflow
Seat, exploration, production, platform, and risk have different owners
Expose cost and outcome before hard allocation
Tokens are a billing unit but a weak value unit
Only routes that pass one acceptance threshold are comparable
Observed tension
METR: 19% slower
Participants expected +20%
Narrow 2025 sample
The useful conclusion
Not a universal penalty
Tools already changed
Build your own baseline
Prompts, tools, evals, and habits move more slowly than HTTP clients
State, data, evals, and acceptance belong to the company
SLA, capacity, data terms, and criticality matter before discounts
Auto, Fast, Deep, and Sensitive hide market volatility
Policy, observability, and evals form independent control loops
A learned router is the last step, not the starting point
Align the quality gate before comparing fully loaded cost
Local handles stable flow; frontier handles the difficult tail
Visibility, guardrails, and portfolio management arrive in sequence
The operating measurement model for every task class
Count accepted tasks
Put budgets on traces
Budget at natural ownership levels
Own state, evals, and contracts
Start routing with telemetry
Cheap intelligence expands demand faster than it shrinks the bill