Skip to content
back to the episode
Episode summary2026CTO

PRD == evals: How AI Brings Product and ML Closer

Alexander Polomodov hosts Albina Munirova, a colleague from T-Bank who moved from ML engineer and team lead into product roles and helped create the ML product manager specialisation. Nearly an hour and a half on why requirements for a probabilistic product are now written as evals, who owns an assistant's quality, and how its team changes.

Code of Leadership · episode #689 min read

The summary is written from the transcript of the recording. Linked below: the recording.

The main thread of the material
01

A requirement that became a test

Munirova opens with a caveat: standard PRDs never reached ML teams anyway, because the product is non-deterministic, so evals are not new to her, only sharper now. Eight years ago the business brought a task and the ML engineer broke it into pieces, picked the metric, assembled the dataset. What changed is that model APIs and coding agents arrived: a product manager can build a prototype alone, while the engineer works less on the model and more on fitting it into a business process. The roles collapsed together, and where the boundary runs, she says, is no longer clear. How the product behaves and what value it brings is sewn into the evals, so that is the control point, and whoever owns them carries most of the responsibility; before, it was split between product and analytics with no single owner of quality. Polomodov adds that wrappers around the black box have partly become a commodity while taming the box has not, and an autonomous product has no human in the loop, the way internal automations do.

Munirova joined eight years ago as an intern, among the company's first ten ML engineers, when nobody knew what the role was. Her first task was forecasting cash in ATMs: money had to be present across the whole network but not in excess: funding it is expensive. The main lesson was not about the model: solving it well required understanding the business process, and she made plenty of mistakes there. The engineer was also analyst and manager, deployed and built datasets alone, so iterations were few. Out of that grew an AutoML platform, Etna, and an open-sourced time-series forecasting framework: stars on GitHub, a community, mentions in other companies' articles. Polomodov recalls a translated forecasting book whose cover, as he remembers it, carried both Etna and a well-known foreign library, which Munirova adds was then called Facebook Prophet. Around then Munirova realised that building convenient things for others interested her more than staying an ML engineer.

02

An investment assistant and the cost of checks

Polomodov asks for a concrete case: a feature helping a beginner investor understand an individual investment account without giving personal investment advice. Munirova unfolds it into a working loop. You take fresh production logs — request-and-answer pairs from the RAG assistant — and the product manager reads the traffic with their own eyes, taking notes: here we sent no link to the source, here we inserted a stale number, here we lied, here we answered a shopping question that should have been routed elsewhere. The team and invited experts annotate too, but she stresses feeling the traffic yourself, or clustering the failures gets hard. The notes go to a model that clusters them, and the largest cluster is taken first: too little data in the index means adding data, a freshness problem means plugging in search, a broken tone of voice means no longer answering beginners like seasoned trading investors. Fixing each error separately grows a cascade of rules that generalises to nothing.

A good eval, in her account, is fresh and carries a regression part: what worked must keep working, what was repaired must not break again. A perfect score on a frozen dataset means overfitting, not a finished product, so new logs keep flowing in. In annotation the team does not converge on categories, and binary classes beat a scale. Their own 2023 evals were built empirically: models were far dumber, the familiar metrics did not exist yet, and they made so many mistakes collecting the suite that quality suffered. Checks are split by price — an ordinary test covers a deterministic answer, reasoning calls for LLM-as-a-judge against pre-written criteria, everyday appropriateness goes to annotation platforms, and the most expensive tier is investment experts and their support desk, who know the new regulations. All of that costs money, so runs are scheduled. Online, guardrails hold the line: the disclaimer about not being investment advice went everywhere, since they chose not to risk anything, and the children's assistant runs filtering models and stop-word dictionaries after generation. She recalls Gemini streaming an answer and then apologising that it cannot answer: more stubs is better, but latency has to be watched.

03

Domains, the bitter lesson and the builder

Why cut assistants by domain? Divide and rule, she says: in production it is simply easier. A general assistant needs a strong technological base from day one, and good in-house models are expensive to build and maintain. The company had a general assistant before the LLM era: its first chatbot was trained on answers from mail.ru, and asked who owned the bank it named a banker from elsewhere. The first versions appeared, she recalls, between 2017 and 2019, and it never matched what users expected. In 2023 the team launched a universe of six specialised assistants, among them shopping, travel, children, finance and investments. Domains were chosen without a scoring table: a broad audience, an explicit pain, complementarity to the bank; an assistant for construction made no sense. The honest drawback: users must work out which tab their question belongs to. In exchange, an investor and a child get different guardrails and tone of voice — that domain judgement is what makes a product good.

Polomodov raises the bitter lesson: does domain expertise lose once a bigger model arrives? Munirova answers from experience: since 2023 they invested heavily in prompt engineering, and smarter models need fewer wrappers — you state the task and the model reasons out the steps. The same happened in machine translation, where decades of architecture work were devalued when LLMs appeared. Her conclusion runs both ways: think ahead about whether your work survives the next class of models, and still build only the wrappers you need. The AI Product Builder is a concept not yet settled across their teams: half ML engineer, half manager, moving from a user need to a prototype and a solution, setting up the improvement loop and answering for the result. A one-person team does not work, but five or six beat ten to fifteen, where communication eats everything; in Cagan's Empowered, she recalls, the point is not gathering top performers but blurring role boundaries. Product managers should build prototypes and speed up discovery; ML developers should move towards the user and the business. Polomodov adds that a builder needs gates — a design system, tests, shift-left from security, a GitOps pipeline — so an agent can check the same rules.

Takeaways

What to take away

  1. 01Evals are the requirement: scenarios, good-answer criteria, prohibitions and regressions describe a probabilistic product more precisely than PRD prose, and give quality one owner.
  2. 02Fresh production logs matter more than a flattering number: a perfect score on a frozen dataset means an overfitted suite, not a ready product.
  3. 03Split checks by price and risk: an ordinary test for deterministic answers, LLM-as-a-judge for reasoning, annotation platforms for everyday appropriateness, experts for regulated content.
  4. 04Grow into the adjacent craft: product managers need prototypes and evals, ML engineers need the user and the business process, and wrappers should be built as temporary.

Sources