Skip to content
LectureHSE · May 23, 2026

AIOps / LLMOps for AI Products

Executive decisions, economics and FinOps for the production AI loop

/ AIOps / LLMOps for AI Products · HSE 2026

Slide contents

  1. 1. AIOps / LLMOps for AI Products

    Executive decisions, economics and FinOps for the production AI loop

  2. 2. Alexander Polomodov

    Technical Director & Fellow, T-Technologies

    Architecture and engineering practices

    AI adoption at development scale

    Focus: managed production AI

  3. 3. The platform exists. Now manage it

    Executives choose behavior, launch path and control

    Vocabulary is known — Gateway, context, evals and guardrails.

    Focus on decisions — SLA, autonomy and quality cost.

    Money anchor — Support bot tests TCO/FinOps.

  4. 4. Basic components are assumed

    AIOps / LLMOps starts after the platform appears

    Skip the platform map

    Choose service by scenario

    Count solved-task cost

    Focus: AI product manageability

  5. 5. 01. From basics to governance

    Do not repeat the AI platform map; translate it into executive decisions

  6. 6. After launch, AI becomes a service

  7. 7. The manager chooses the frame

  8. 8. AIOps is a decision loop

  9. 9. 02. Manager decisions

    Scenario, autonomy, error cost, SLA, quality budget and ownership

  10. 10. Scaling needs decisions first

    Otherwise AI grows faster than control

    Value — Value and scenario — Value, baseline and SLA.

    Risk — Risk and control — Allowed autonomy and human-in-the-loop.

    Cost — Economics and ownership — Quality cost; product, data, risk and bill owners.

  11. 11. Error cost defines autonomy

  12. 12. AI SLA is broader than API uptime

  13. 13. Evaluation starts from product

    The same stack can be right or wrong

    Value: process and baseline

    Risk: error and approval

    Economics: unit cost

    Operability: release and rollback

    Explain the decision, not only the stack

  14. 14. 03. Options and trade-offs

    API-first, managed platform, rented GPU, owned self-host and hybrid

  15. 15. AI runtime is a portfolio

  16. 16. Assistant and agent are different products

  17. 17. You need routing rules

  18. 18. Choice is a portfolio decision

    API-first validates the launch

    Rented GPU gives transitional control

    Owned self-host needs stable load

    Hybrid fits different scenarios

    Runtime cannot be chosen without scenarios and financial horizon

  19. 19. 04. Money case

    Support bot: load, tokens, GPUs, team and the management decision

  20. 20. Calculate the support bot

    200,000 monthly dialogs · 10 messages · 5-second SLA

    What we calculate — Load, tokens, peak, GPU, CAPEX/OPEX, team.

    What we compare — API-first, rented GPU, own self-host, hybrid.

    What we decide — Year-one path and review metrics.

  21. 21. Case parameters set the decision scale

  22. 22. Translate parameters into RPS and parallelism

  23. 23. GPT-5.2 API gives a clear starting OPEX

  24. 24. API-first buys speed

    The risk is silent unit-economics growth

    Why launch is convenient

    Near-zero CAPEX

    Faster production launch

    Thinner surrounding platform

    What must be controlled

    Output, cache, retries

    Provider limits and degradation

  25. 25. Size self-host from SLA and peaks

  26. 26. GPU rental trades CAPEX for OPEX

  27. 27. Self-host changes the problem

  28. 28. Strategy comparison shows the year-one decision

  29. 29. The decision must include a review point

  30. 30. The case decision

    Launch API-first unless owned perimeter is mandatory.

    Track unit economics immediately.

    Set review triggers.

    Self-host must pass calculation, SLA and operations readiness.

    The first decision is speed with control, not eternal architecture

  31. 31. 05. Manageability and FinOps

    Dashboard, cost allocation, unit economics review and a 120-day plan

  32. 32. FinOps starts with bill ownership

  33. 33. Owned hardware makes the bill platform-wide

  34. 34. After CAPEX, operations begin

  35. 35. Shared platforms need clear allocation drivers

  36. 36. The dashboard connects quality, risk and money

  37. 37. 120 days: scenario to scale

  38. 38. No measurement, no management

    Without evals, quality looks higher.

    Without fallback, risk grows.

    Without unit economics, economics break.

    Without ownership, disputes begin.

    Managed AI starts with manager decisions

  39. 39. References

    AIOps, risk and economics

    OpenAI pricing and docs.

    OWASP, NIST, Langfuse.

    Habr TCO, FinOps Foundation.

  40. 40. Feedback form for day two

    Feedback form for day two