Skip to content
back to the episode
Episode summary2026Fellow

AI Writes More Code. Why Doesn't Delivery Speed Up?

In the second 3 AImigo episode, Evgeny Sergeev, Alexander Polomodov, and Aleksey Litvinov examine a paradox: AI makes an individual engineer faster, while end-to-end cycle time barely improves. Instead of looking for a stronger tool, the hosts trace constraints across the delivery system, feedback inside the agent loop, organizational incentives, and metrics that distinguish activity from value.

3 AImigo · season 1, episode 26 min read

A summary based on the complete YouTube recording and Russian captions through 1:41:06. The text is condensed; not every contribution can be confidently attributed to a speaker.

The main thread of the material
01

More code does not mean more throughput

A local optimization does not remove a system constraint; it moves the queue downstream. Once coding accelerates, alignment, verification, integration, or a person's ability to supervise several agent tasks may become the new limit. The hosts start with an accepted user outcome and work backward: map the complete path, locate today's bottleneck, and predict where it will move after automation. One stage's speed matters only against the cycle time of the whole delivery system. Otherwise a team receives more pull requests than it can responsibly accept, and visible activity masks a growing inventory of unfinished work. The practical constraint question is which step now governs value delivery and what will happen if that step accelerates next.

Aleksey describes an anonymized company of roughly ten people that rebuilt its workflow around agents. It created a task tracker, memory, and incident tools so context and actions were programmatically available. High autonomy then exposed another constraint: how much concurrent conversation and decision context one person could retain. Orchestration and cognitive load became more important than generation speed. This is not a universal instruction to rewrite every tool; it demonstrates systematic constraint discovery. Custom tools were justified by their fit with a new operating model: agents needed structured context and the ability to act without manual copying. Yet every internal system adds maintenance, so the gain must be measured across its lifecycle rather than at generation time alone.

02

Move feedback inside the agent loop

Late review makes deviations expensive: work has crossed several handoffs before a lawyer, security specialist, editor, or API owner sees it. The alternative is to make rules explicit inside the agent's short loop. A human reviewer steps in only for exceptions, while each issue returns as a policy, prompt update, or evaluation case. Human-by-exception does not remove the expert; it shifts the work from repairing outputs to improving the mechanism that prevents recurrence. The key move is turning implicit reviewer knowledge into an executable signal close to the action. Some judgment remains human, but repeatable defects no longer wait for the same manual check on every cycle.

An end-to-end agent workflow requires more than adding an assistant to every product. Wardley Mapping informs what to buy, adapt, or build; bounded contexts and Team Topologies define boundaries and interactions. Existing platform interfaces may be hostile to agents: an API or MCP tool returning a huge JSON payload for a human UI burns tokens and breaks automation. Platform owners need compact machine contracts and interoperable integrations, or fast local agents never become one fast delivery flow. A team boundary also becomes a boundary for context, authority, and accountability. When an agent crosses it through an unstable or excessive interface, local speed is paid for with integration errors and costly state recovery.

03

The constraint lives in organization design and metrics

Buying licenses is not process change. A pilot needs internal expertise, room to bypass some old rituals temporarily, and a sponsor able to carry the change into normal work. Resistance may be rational when acceleration threatens a familiar role, influence, or stability, so AI transformation needs change management alongside technical depth. A consultant can identify a systemic cause, but without an internal owner the organization returns to its previous incentives and habits. Sponsorship is operational rather than ceremonial: it removes cross-team blocks and changes rules that the pilot proves obsolete. The internal owner preserves context after a temporary expert group leaves and turns learning into routine practice.

Exploration and exploitation should not share one scorecard. Early work benefits from bounded hypotheses, clear direction, and criteria for moving to scale instead of demanding an immediate return from every experiment. Once practice stabilizes, token use, license counts, and AI-generated-code share remain activity measures—modern equivalents of Lines of Code. Value appears in time to an accepted outcome, quality, rework, and business effect. If those measures do not move, extra code is simply inventory in the system. The metric must include consequences after release: incidents, maintenance, and repeated work. Otherwise the cost of acceleration quietly moves from the developer to reviewers, operators, or users.

Takeaways

What to take away

  1. 01Judge faster coding by the complete path to an accepted user outcome, not by an individual engineer's output.
  2. 02Move policies, checks, and expert feedback into the short agent loop, keeping people for exceptions and system improvement.
  3. 03End-to-end gains require new domain boundaries, platform interfaces, team interactions, and organizational incentives.
  4. 04Tokens, licenses, and generated-code share measure activity; cycle time, quality, and business outcomes measure value.

Sources