all longreads
Longread#AI4SDLC

The Economics of AI Development: A Complete Analysis from Tokens to Accepted Work

The price of reaching a fixed intelligence level keeps falling, yet a company that finds productive AI use cases is likely to spend more in total. The reason is not one expensive token. It is a larger task portfolio, agentic call chains, infrastructure, verification, and the cost of failure. Development economics must therefore be measured in accepted work at a target quality and risk level—not in model calls.

July 21, 2026≈ 25 min

The full version of a talk delivered at Podlodka AI Club on July 21, 2026. Prices and market data were reviewed on July 16, 2026; tariffs, models, and commercial terms change faster than conventional enterprise software. The forecast ranges, reserves, and negotiation thresholds below are author heuristics, not industry standards.

01

Tokens get cheaper. The budget does not

Discussions about AI budgets often mix three different prices. The first is the cost of reaching a known quality level. It falls quickly as newer models, hardware, and inference optimization make yesterday's outcome cheaper. The second is the price of the new frontier, long context, high speed, and guaranteed capacity. That layer remains premium. The third is the company's total budget, driven by task volume, automation depth, and the number of production use cases.

diagram 01 · fixed-quality price and total budget diverge
Unit intelligence gets cheaper while total AI spend growsFIXED QUALITY3–10× ↓cost by 2029ACCEPTEDWORKis the unitTOTAL AI BUDGET2–5× ↑scaled adoptersWorking forecast for 2026 → 2029 · ranges, not promises

Market signals point to demand expansion. Menlo Ventures estimated U.S. enterprise GenAI spend at $37 billion in 2025, up from $11.5 billion a year earlier. This is a venture firm's model based on roughly five hundred enterprise decision-makers—not a census of every invoice. The State of FinOps 2026 says 98% of respondents now manage AI spend, versus 63% a year earlier and 31% two years earlier. McKinsey likewise reports broad adoption but a much narrower group scaling AI across the business.

diagram 02 · money and adoption outpace outcome management
Enterprise demand is expanding faster than governance$11.5B → $37BUS enterprise GenAI2024 → 2025$4B for development31% → 63% → 98%FinOps manages AI2024 → 2026cloud-heavy sample88% ≠ ⅓adoption vs scaleEBIT impact is rareoutcome gapMenlo Ventures 2025 · State of FinOps 2026 · McKinsey 2025

These observations cannot yield a precise 2029 budget. They can support a management scenario. In this scenario, today's fixed quality becomes 3–10 times cheaper, a typical accepted task becomes 2–5 times cheaper, the new frontier stays around the same order of list price or falls by up to half, while the total budget of a successful AI portfolio grows 2–5 times. These are author ranges with different confidence levels, not a market forecast to the last percentage point.

diagram 03 · working forecast through 2029
The 2029 forecast moves in opposite directions20262029TODAY'S QUALITY3–10× ↓FRONTIER LIST PRICE1–2× ↓ACCEPTED DEV TASK2–5× ↓SCALED COMPANY BUDGET2–5× ↑Medium confidence except frontier price and total budget · source synthesis
02

Spend is a product, not a price sheet

Optimizing the price per million tokens sees only one multiplier. The actual bill depends on task volume, model calls per task, context per call, call price, seats, platform layers, and reserved capacity. A cheaper call makes new use cases viable, while an agent turns one user action into a branching workflow.

AI spend = tasks × calls per task × tokens per call × token price + seats + platform + fixed capacity

The FinOps Foundation gives a useful range: one agent interaction may generate 5, 10, or 50 model calls. Every call can resend context, append tool results, retry, or open another branch. An aggregate invoice cannot distinguish a valuable complex trace from a looping agent. Cost must be attributable to use_case, task_class, andtrace_id.

diagram 04 · the multipliers behind AI spend
Cheaper tokens multiply through a larger systemTASKSmore use cases×CALLS / TASK5 · 10 · 50×TOKENS / CALLcontext grows×UNIT PRICEfalls+ seats + platform + fixed capacityvolume can outrun the unit-price declineFinOps Foundation: one agent interaction may create 5, 10, or 50 model calls

Lower unit prices therefore do not guarantee lower total cost. They create room for more demand. The management question is not whether consumption rose, but what outcome each additional call created and where the workflow developed a heavy tail.

03

Bound the task, not the person

A global per-user token cap mixes autocomplete, architectural analysis, and a production agent with authority to change a system. It punishes a useful power user yet offers weak protection from one runaway workflow. For a copilot, the model call often costs less than a few minutes of an engineer's time. For an autonomous agent, the danger lies in branches, retries, growing context, and repeated tool calls.

diagram 05 · guardrails cover the whole trace
Guard the task trace, not the person's token countTRACEaccepted or handed off$ / TRACESTEPSWALL CLOCKTOOL CALLSSAFE FALLBACK / HUMANAlso cap retries, repeated calls, context size, and concurrency

A minimum production control plane bounds several resources at once:

  • max_cost_usd, steps, retries, and repeated tool calls;
  • tool-output size, model timeout, workflow timeout, and concurrency;
  • tool authority, permitted environments, and human approval requirements;
  • loop detection, kill switches, and safe fallback or handoff;
  • mandatory team, product, environment, use-case, and trace tags.

Raising a limit is itself a management decision. It needs an owner, a reason, and an observable outcome. Otherwise, “temporarily give the agent more budget” quietly becomes a permanent mode without feedback.

04

Five budget baskets need different owners

The argument over budgeting by person, team, or project is a false choice. Seats naturally belong to a person or function, exploration to a team, production inference to a product or workflow, and shared gateways, evals, and security to a platform. Putting all of them under one cost center hides both value and the reason spend grows.

diagram 06 · five AI budget baskets
Each cost belongs at its natural levelSEATSpersonEXPLORATIONteamPRODUCTIONproduct / use caseSHARED PLATFORMcentral functionRISK RESERVEportfolioone budget hierarchy would distort ownershipForecast production with P50 and P90; keep 15–25% reserve while history is thin
BasketOwnerManagement unit
SeatsPerson or functionActive licenses, adoption, and cost of a useful hour
ExplorationTeamA safe sandbox budget and validated hypotheses
Production inferenceProduct or workflowCost per accepted task and accepted-work volume
Shared platformCentral platformGateway, observability, evals, security, and reserved capacity
Risk reservePortfolioSpend tails, pilots, emergency capacity, and uncertainty

With limited history, planning P50 and P90 plus an uncertainty reserve is more honest than one point estimate. A 15–25% reserve is a practical starting heuristic, not a FinOps standard. It should shrink as stable workload and quality profiles emerge.

Showback before chargeback

Early hard chargeback encourages teams to hide experiments and debate false precision in shared-platform allocation. For the first one or two quarters, expose cost and outcome by product and team; then add budgets and soft thresholds. Chargeback becomes useful for stable production workloads with a clear owner and unit economics.

diagram 07 · the path from visibility to chargeback
Showback should precede chargeback1VISIBILITY2ALLOCATION3FORECAST4GOVERNANCE5OPTIMIZATIONSHOWBACK · first 1–2 quartersCHARGEBACK · stable workloadsMake cost and outcome visible before creating hard allocation incentives
05

The accepted task is the economic denominator

Tokens are a convenient billing unit but a weak value unit. A generated patch is not yet an outcome: it may fail tests, require lengthy review, or introduce a regression. The denominator of unit economics should contain only work that passes a common acceptance criterion.

cost per accepted task = (model + tools + retrieval + gateway + compute + human review + rework + expected failure loss) / accepted tasks

diagram 08 · the full cost of an accepted task
The accepted task is the economic denominatorMODELTOOLSRETRIEVALREVIEWREWORKFAILURE LOSSACCEPTED TASKSPATCH✓ tests✓ review✓ no regressionCheap failure can still be expensive after retries and human review

“Accepted” must be defined for each task class. For a bug fix, it may be a hidden regression test and a green full suite; for review, a useful finding without noise; for a migration, a correct artifact with no regression; for a production agent, a completed workflow with permitted effects. Price and latency become comparable only after candidates pass one quality threshold.

diagram 09 · quality gate before price optimization
Quality gates come before cost optimizationREAL TASKSbugs · testsreview · migrationROUTE A$ · pass@1ROUTE B$ · pass@1QUALITY GATEtestsblind reviewno regressionsPARETOqualitycostlatencyRepeat stochastic runs and count the whole trace: tools, retries, fallback, and human time

Run every candidate several times on the same real tasks and count the entire trace: tools, retries, fallback, and human time. A cheaper model at 70% success can beat a more expensive one at 95% only while the review and rework difference remains small. One extra minute of engineer time—or one rare expensive failure—can reverse the choice.

Self-estimates are not outcomes

METR's 2025 RCT across 16 experienced open-source developers and 246 tasks found a 19% slowdown with early-2025 tools, even though participants believed afterward that they had accelerated by roughly 20%. The authors now mark that result as out of date. Their 2026 update shifted toward likely acceleration but encountered strong selection effects and led to a redesign of the next experiment.

The useful conclusion is neither “AI always slows developers down” nor “new models solved productivity.” A narrow sample cannot represent all software work, and tools change quickly. The study demonstrates why perception, activity, and benchmarks cannot replace an internal baseline. DORA offers the more durable framing: AI amplifies the strengths and weaknesses of the delivery system around it.

06

The deepest lock-in sits above the API

An HTTP client or OpenAI-compatible endpoint is usually the easiest layer to replace. Migration fails higher up: prompts were tuned to one model's style, tools expect particular behavior, recovery assumes familiar failures, and the eval set misses silent degradation. Deeper still are conversation state, data, commercial commitment, and team skills.

diagram 10 · six layers of vendor lock-in
Behavior is harder to move than an APIORGANIZATIONAL SKILLSCOMMERCIAL COMMITMENTSTATE + DATA + EVALSBEHAVIORMODEL FEATURESHTTP / APIHIGHLOWThe hidden dependency lives in prompts, tools, evals, and operating habits

The useful abstraction is not a lowest-common-denominator API but the business task contract: input, expected artifact, acceptance criteria, and permitted effects. Business code calls review_patch orsummarize_incident; configuration selects the model. State, documents, audit trails, and evals remain under organizational control.

diagram 11 · a portable task contract
Own the task contract, not the lowest common denominatorTASK CONTRACTinputartifactacceptanceOWN STATEdocs · traces · auditOWN EVALSgolden tasksALIASESfast · balanced · deepPROVIDER AadapterPROVIDER BadapterLOCAL ROUTEadapterKeep critical 80% portable; isolate valuable provider-specific 20% behind adapters

A lowest common denominator also has a cost because it blocks valuable provider capabilities. A practical author heuristic is to make the critical 80% portable and isolate the valuable provider-specific 20% behind adapters, regression evals, and an exit plan. A critical workflow needs a certified secondary route, not merely a compatible request schema.

07

An enterprise contract starts with risk, not discounts

Account-team engagement should not wait for a large invoice. A critical process, sensitive data, SLA requirements, regional constraints, or guaranteed capacity can justify it earlier. Sustained spend of $10–25k per month with one provider, or a credible $100–250k annual forecast, is only an economic heuristic for procurement and legal overhead—not an official threshold published by any provider.

diagram 12 · enterprise-contract triggers
A contract starts with risk, not a spend thresholdENTERPRISECONTRACTCRITICAL FLOWSLA / CAPACITYSENSITIVE DATARATE LIMITSPROCUREMENTCOMMITMENT VALUEEconomic conversation often starts around $10–25k/month; security and SLA can start sooner
AreaWhat to establish
CriticalitySLA, support, incident path, accountability, and escalation rights
CapacityProvisioned throughput, burst, rate limits, regions, and scarcity behavior
DataRetention, training use, ZDR/DPA, audit evidence, and usage export
ChangeVersion pinning, deprecation windows, canaries, and regression evals
EconomicsPortfolio discounts, P50–P60 commitment, PAYG peaks, and exit rights

Commit around the base load—P50–P60, for example—and keep peaks on PAYG when provider economics allow it. A contract without usage export, model versions, deprecation windows, and exit rights can offer an attractive discount while making future optimization more expensive.

08

People need modes; the platform needs outcome-driven routing

Users should not have to track price sheets, regional availability, and the latest leaderboard. They need to recognize intent and error cost. A practical interface exposes modes: Auto for most requests, Fast for simple transformations, Deep for architecture and difficult defects, and Sensitive for an approved local or contracted route.

diagram 13 · modes instead of model names
People choose intent; the router chooses modelsPERSONtask + riskPOLICY+ROUTERAUTOpolicy decidesFASTsmall modelDEEPfrontier + budgetSENSITIVElocal / ZDR / regionTeach task classes and error cost—not a catalog of twenty changing model names

Behind that interface sit three independent loops. The gateway handles authentication, secrets, budgets, retries, logging, and enforcement. The router selects an allowed route by risk, capability, cost, health, and latency. Evals independently verify quality so the router cannot certify its own decisions.

diagram 14 · routing closes on the accepted outcome
The router needs an outcome feedback loopIDECISERVICESAGENTSTASKCONTRACT+ SDKGATEWAYauth · budgetstrace · retryPOLICY+ ROUTERrisk · healthCLOUD ACLOUD BLOCALBATCHevals + FinOps + accepted/rejectedGateway enforces; router selects; evals independently verify

A learned router is a poor starting point. Begin with complete telemetry and task tags, then aliases and static rules, certified fallback and canaries, followed by cascades that escalate on failed tests, invalid schemas, or insufficient retrieval. Learning becomes appropriate only after real outcome labels accumulate.

diagram 15 · the routing maturity ladder
Routing maturity begins with telemetry1ONE PROVIDERtelemetry + tags2ALIASESstatic rules3FALLBACKcertified + canary4CASCADEsignal-based escalation5LEARNED ROUTERonly with mature evalsStart with deterministic policy; optimize only after routes are measurable

RouteLLM and FrugalGPT demonstrate the potential of cascades, but their savings percentages do not transfer to another company's task distribution and evals. Routing is not about finding the cheapest model; it is about finding the cheapest route that still passes acceptance.

09

Local inference wins only after the quality gate

A local model is not free. The API invoice becomes GPUs or cloud commitment, idle capacity, HA headroom, storage, inference engineering, SRE, security, and model updates. Variable cost may be lower, but fixed cost and underutilization make small or unpredictable workloads expensive.

Q* = fully loaded local fixed cost / (cloud variable cost per accepted task − local variable cost per accepted task)

diagram 16 · local-inference break-even
Local inference wins only after quality and utilizationQ*accepted tasks / yearCLOUD · variableLOCAL · fixed + idleQUALITY GATEmust pass firstreview time countsno frontier needQ* = fixed local cost / (cloud cost per accepted task − local variable cost)

The formula is meaningful only after a common quality gate. If the local candidate needs more review, rework, or escalation, its variable cost per accepted task may exceed cloud. Local is strongest for data residency, air gaps, stable high-volume classification, extraction, embeddings, reranking, and narrow tasks after distillation.

Hybrid beats infrastructure ideology

In a mature system, local is one route. It handles repeatable flow, preprocessing, and tasks with a strict data path; the complex or uncertain tail escalates to a frontier API under the same acceptance gate. Privacy, latency, and utilization no longer require giving up frontier quality where it genuinely pays off.

diagram 17 · local as a route in a hybrid system
Local is strongest as one route in a hybrid systemINCOMINGtasksdata classLOCALclassifyredactrepeatable workFRONTIER APIcomplex / uncertaindeep budgetLOCAL OUTPUTfast / privatequality-passedACCEPTtestshumantraceUse local for privacy or stable volume—not as an ideological default
10

Ninety days are enough to expose the economics

The first quarter should not begin with a large model marketplace or a GPU purchase. Make volume, quality, cost, and failure consequences visible across a few high-volume use cases; then add guardrails; only afterward should routing and the commercial layer become more sophisticated.

diagram 18 · practical plan from day 0 to day 90
Ninety days are enough to expose the economics0–30SEEinventory · tagsaccepted outcome31–60BOUNDtrace budgetsshowback · evals61–90MANAGErouter · canarycontract · local pilotVisibility → guardrails → portfolio management

Days 0–30: inventory and baseline

Find seats, keys, gateways, endpoints, and owners. Select 3–5 high-volume use cases, define the accepted outcome, and capture volume, cost, pass rate, review time, and P50/P90/P99. Without this, negotiation and optimization remain anchored to the invoice rather than the work.

Days 31–60: guardrails and showback

Add trace budgets, loop detection, kill switches, the sandbox/production boundary, product showback, and golden evals. Validate caching and batch against the real workload profile; a low hit rate may not justify another layer of complexity.

Days 61–90: routes and resilience

Certify a secondary provider, add canaries and static modes, run a failover drill, prepare a negotiation package, and launch one local pilot with fully loaded TCO. By the end, every decision should have an owner, a baseline, and a stop criterion.

11

One scorecard connects quality, cost, and risk

The operating system cannot collapse into one average cost or composite score. Each task class needs acceptance, cost, distribution tails, human time, technical failures, and failure consequences together. P99 matters as much as P50 because loops, overload, and rare expensive incidents live in the tail.

diagram 19 · operating scorecard by task class
One scorecard connects quality, cost, and riskACCEPTEDTASKby classACCEPTANCEpass@1UNIT COST$ / acceptedTAILP50 · P90 · P99HUMAN TIMEreview + reworkFAILUREschema · tools · lossTokens remain a billing measure; outcomes become the management measure
LayerWhat to measure
OutcomeAcceptance/pass@1 and the share of genuinely completed tasks
CostCost per accepted task and P50/P90/P99 cost
TimeTrace latency, human review, and rework
ReliabilityRetries, fallback, schema/tool errors, and escalations
RiskExpected failure loss, privacy incidents, and unsafe actions

This scorecard gives engineering, platform, FinOps, risk, and procurement one language. The router receives outcome feedback, teams see review cost, the platform sees heavy traces, and procurement sees the volume worth committing. The objective is no longer “spend fewer tokens.” It is “maximize accepted work at the required quality and risk level.”

Takeaways

What to carry into practice

  1. 01Tokens remain the billing unit, but the management unit should be an accepted task including review, rework, and expected failure loss.
  2. 02Cheaper fixed-quality intelligence expands demand: more use cases and longer agent traces can increase the total budget even as calls get cheaper.
  3. 03Guardrails belong on the whole workflow—money, steps, retries, time, and authority—with a safe fallback or human handoff.
  4. 04The deepest vendor lock-in lives in behavior, state, evals, and organizational skills; portability begins with an owned task contract and independent acceptance.
  5. 05Routing, local models, and commercial commitments make sense only after telemetry and a quality gate on the real task distribution.
Sources

Data, research, and documentation

Share