AIOps / LLMOps for AI Products
Executive decisions, economics and FinOps for the production AI loop
Executive decisions, economics and FinOps for the production AI loop
Executive decisions, economics and FinOps for the production AI loop
Technical Director & Fellow, T-Technologies
Architecture and engineering practices
AI adoption at development scale
Focus: managed production AI
Executives choose behavior, launch path and control
Vocabulary is known — Gateway, context, evals and guardrails.
Focus on decisions — SLA, autonomy and quality cost.
Money anchor — Support bot tests TCO/FinOps.
AIOps / LLMOps starts after the platform appears
Skip the platform map
Choose service by scenario
Count solved-task cost
Focus: AI product manageability
Do not repeat the AI platform map; translate it into executive decisions
Owners, budget and control loop appear
Quality: scenarios and thresholds
Risk: rights, audit, fallback
Cost: not only tokens
Ownership: product, platform, risk
One model can be assistant, service or agent
What should happen
Which scenario improves
Which baseline it beats
What must not break
Error cost and data risk
Cost ceiling per operation
Before: scenario and ceiling
During: autonomy and fallback
After: economics and routing
The operating model matters more than any single model
Scenario, autonomy, error cost, SLA, quality budget and ownership
Otherwise AI grows faster than control
Value — Value and scenario — Value, baseline and SLA.
Risk — Risk and control — Allowed autonomy and human-in-the-loop.
Cost — Economics and ownership — Quality cost; product, data, risk and bill owners.
Higher consequence means stricter control
Low error cost
Drafts, hints, search
Fast models and cache
Metric: speed and deflection
High error cost
Money, contracts, prod
Policy gate, approval, audit
Metric: incidents and escalations
Users need a safe result
Service health
Latency, timeouts, fallback
Peaks, queues, rate limits
SLO includes fallback path
Quality health
Sources, completeness, policy
No context means escalation
Measure quality by scenario
The same stack can be right or wrong
Value: process and baseline
Risk: error and approval
Economics: unit cost
Operability: release and rollback
Explain the decision, not only the stack
API-first, managed platform, rented GPU, owned self-host and hybrid
Choice follows data, control, horizon
Buy more ready-made
API-first: fast launch
Managed: control and lock-in
SaaS: demand validation
Take more control
Rented GPU: OPEX over CAPEX
Owned self-host: control and ops
Hybrid: routes by scenario
One model, different risk and governance
Read-only assistant
Value: search and explanation
Risk: error and leakage
Control: ACL, citations, fallback
Action agent
Value: executes steps
Risk: money and prod
Path follows scenario, risk and cost
Low risk uses cheap routes.
Expensive errors need premium/review.
Failure means partial/read-only/handoff.
Cost growth changes routing.
API-first validates the launch
Rented GPU gives transitional control
Owned self-host needs stable load
Hybrid fits different scenarios
Runtime cannot be chosen without scenarios and financial horizon
Support bot: load, tokens, GPUs, team and the management decision
200,000 monthly dialogs · 10 messages · 5-second SLA
What we calculate — Load, tokens, peak, GPU, CAPEX/OPEX, team.
What we compare — API-first, rented GPU, own self-host, hybrid.
What we decide — Year-one path and review metrics.
This is no longer a toy demo, but still not a load that automatically justifies owned hardware
1,000,000 LLM answers per month create a peak of about 9-10 concurrent generations for 5 sec SLA
With 1,400 input and 180 output tokens per answer, the LLM part costs about $4,970 per month
The risk is silent unit-economics growth
Why launch is convenient
Near-zero CAPEX
Faster production launch
Thinner surrounding platform
What must be controlled
Output, cache, retries
Provider limits and degradation
For a peak near 10 concurrent answers, the estimate gives 6-8 A100 GPUs for 14B and about 12 GPUs for 32B
In the example, 6-8 A100 GPUs cost 1.17-1.56M RUB/month before network, storage and team
Tokens become an infrastructure product
Rented GPU
Tests the self-host hypothesis
OPEX depends on utilization
Needs MLOps/SRE
Owned hardware
Maximum control
CAPEX and delivery lead time
Makes sense with stable load
For the given load, API-first is usually rational unless there is a strict on-prem requirement
Launching via API does not remove the need to recalculate under traffic, compliance or quality-cost growth
Launch API-first unless owned perimeter is mandatory.
Track unit economics immediately.
Set review triggers.
Self-host must pass calculation, SLA and operations readiness.
The first decision is speed with control, not eternal architecture
Dashboard, cost allocation, unit economics review and a 120-day plan
Product-linked spend grows slower
Each scenario has a limit.
Cost is visible by route.
Shared buckets hide choices.
Price explains the trade-off.
The Habr/TCO example shows the order: one configuration around 202M RUB, prod/dev/test/reserve around 606M RUB
Power, cooling, licenses, support and team turn hardware into a service
GPU hours, storage, kWh, traffic, containers, models and users connect costs to product behavior
Separate token charts do not create a decision
Quality and risk
Quality by scenario
Fallback and policy violations
Incidents with evidence pack
Economics and control
Cost per outcome
Cache, retries, route mix
Start with a management loop
0-30: scenarios, baseline, owners
30-60: observability, fallback, dashboard
60-90: limited release, unit economics
90-120: FinOps and scaling
Scale managed scenarios
Without evals, quality looks higher.
Without fallback, risk grows.
Without unit economics, economics break.
Without ownership, disputes begin.
Managed AI starts with manager decisions
AIOps, risk and economics
OpenAI pricing and docs.
OWASP, NIST, Langfuse.
Habr TCO, FinOps Foundation.
Feedback form for day two