Skip to content
back to the episode
Episode summary2026CTO

AI Adoption: It's the Protocol, Not the Model

The third and final AMA closes a marathon of nearly five hours: Alexander Polomodov and Aleksey Litvinov — a principal engineer and author of a programme on AI-assisted engineering — work through six listener questions, from adoption that never adds up to a system to the requirement ledger, ROI, OpenSpec against Spec Kit and greenfield evals.

Code of Leadership · episode #699 min read

The summary is written from the transcript of the recording. Linked below: the recording.

The main thread of the material
01

Task protocol, gurus and Shadow AI

The first question comes from a company with an AI policy, a golden standard of tools, a platform over LLM providers and purchased licences — and still no system: the same task is solved in eight thousand lines or in forty thousand, teams drag in Shadow AI and local models, and the arguments are about tokens and access. Forty thousand lines astonish Aleksey by their sheer volume, and he recalls the meme about an agent that changed three hundred files, wrote a thousand tests, and nothing ran. The cause, in his experience, is rarely the model and usually the protocol around the task: how it is framed on the way in and what counts as each stage's output. Hence standardisation — a mandatory task format in the tracker, filters and background agents that hold back a vaguely worded task — and segmentation: who autocompletes and who delegates. It grows through the local gurus: until they buy the practice themselves, the policy stays a page in the wiki.

Shadow AI, Aleksey argues, is a diagnostic signal: work out what people drag in and why, then legalise it where possible. His example is a company that bought Claude while the people doing heavy fact-checking preferred Codex; research suggests Codex follows instructions better, Claude invents more, and by Aleksey's subjective impression that still holds on Opus 5. Actual usage is best watched through pull requests — up to an agent that runs hourly, gathers analytics and drops non-contradictory suggestions into a backlog, needing only filtered data and read/write access. Alexander adds the view from a large company: centrally you enable budgets, token limits, access and the model gateway, but you cannot walk the maturity levels for a team. One has reached spec-driven development, another launches a swarm of agents, and that experience does not transfer — what works is demo days and watching the neighbours.

02

The requirement ledger and ROI metrics

The second question goes to Alexander: is a requirement ledger — an accounting book of requirements — worth keeping, and should you measure how many accepted constraints an agent lost or reopened? Alexander admits he had not thought about it: strong models with a good harness already hold to the markdown structure, check themselves, and at the end offer options where nothing was settled at the start. In a closed loop with weaker models requirements are visibly lost, and there a ledger of original, clarified, implemented and verified items could pay off — though a generation or two of models later it gets thrown out as over-engineering. Aleksey doubts the value: such a register mixes functional requirements, non-functional ones and ADRs, and it is unclear how to measure it. He recalls event sourcing an agent's actions, an idea heard a year ago and never seen implemented, and reduces the task to fitness functions and an independent reviewer — a background agent checking consistency per pull request or every N merged changes.

The third question is how to calculate and defend the effect of AI in front of a CIO or CEO. Aleksey thinks it arrives late: the top first demands a tenfold speed-up, and only when the bill for credits lands does it emerge that nothing got faster and bugs multiplied. In audits he often sees roughly forty per cent more returns from QA and longer fixes: the developer does not know how the feature was built. Measuring a whole organisation is near impossible; a team or a product can be measured — lead time, time to market and the five DORA metrics including rework rate, the recently added fifth. Alexander recalls a sausage-factory metaphor and an Avito platform team lead's talk at IT-Piknik: they could compute throughput, so they sold the business an inverted 'effective cycle time' — a task moving some fifteen to twenty per cent faster. But the average hides the distribution, and saved time is no saving until it turns into other useful work or a lower cost of running the same service — until then it never reaches the P&L.

03

Methodologies, evals and the limits of autonomy

The fourth question, from a staff engineer, is how to sell an AI-native operating model at CEO level. Aleksey's answer depends on the company: a non-digital business — a Walmart, say — is served by supportive assistants and chatbots; a vertical GenAI product needs speed or it gets eaten by the providers or by competitors; a comfortable market leader says 'of course we are AI-native', points at a hundred and twenty-five working groups and drowns the idea in sub-committees — no urgency, just as Kodak had no need of a digital camera. Alexander recalls rolling out legal AI in a hundred-and-twenty-year-old conglomerate that gives the topic half an hour a week. The fifth question, from a chat of technical directors: how do OpenSpec and GitHub Spec Kit differ? Both are worth trying; migrating current projects, Aleksey would not. In Spec Kit the unit of work is a feature run through a micro-waterfall from constitution and specify to tasks and implementation; it copes badly with deltas and tight coupling. OpenSpec targets an existing codebase, runs an explore, propose and apply cycle, and is easily adjusted through files in the project.

The choice between tools Aleksey reduces to a garden: if all you hold is a spade, the lawn will not improve — the ten methodologies in his book are a map for choosing. Alexander adds that running a methodology on a pet project is cheaper than knowing it by hearsay. A separate trouble, for Aleksey, is that an LLM never shows what you do not know while making everything feel like it works. The last question is the order of evals in a greenfield project: Alexander would build the system alongside the first business feature — architecture skeleton, CI/CD, contour tests, linters, fitness functions — and put evals on top of the deterministic metrics, reading the reports himself rather than starting an autonomous loop. The price of unbounded autonomy is Aleksey's own background agent, which 'fixed' a broken mailing on his site by flipping its feature flag from one to zero; he noticed the missing letters two weeks later. Alexander himself lets an agent into production only through GitOps: configuration in a repository, Argo CD or Flux with checks. Three sessions closed nineteen questions; tomorrow brings the first episode of a new three-way podcast with Zhenya Sergeev.

Takeaways

What to take away

  1. 01A spread from eight to forty thousand lines on one task is fixed by the protocol, not the model: the framing format, the exit criteria of each stage, and knowing who autocompletes and who delegates.
  2. 02Shadow AI is a request for a missing model, limit or permission: people doing heavy fact-checking wanted Codex while the company issued Claude.
  3. 03Freed time becomes a saving only when it turns into other useful work or into a lower cost of running the same product.
  4. 04Autonomy rests on deterministic checks and narrow permissions: the background agent that silenced a mailing through a feature flag went unnoticed for two weeks.

Sources