Skip to content
all episodes
Research Insights Made Simple · episode 11

Measuring AI Code Assistants and Agents

33:36
Conversation

What we discussed on the recording

Alexander Polomodov and Evgeny Sergeev review DX’s framework for measuring AI code assistants and agents. Development suits such experiments because the work is digitized and surrounded by automation and checks. Evgeny adds experience from Flo Health and a roundtable with the authors.

The model is a sequence, not a promise of immediate return. Start with adoption and utilization: who uses the tool, for which jobs, and how often. Then assess impact on time, quality, and DX Core 4 before economics. Models perform better on deliberately boring stacks because examples are abundant.

An incident postmortem illustrates a measurable use case: an agent assembles the timeline and actions, while the team checks time saved and output quality. If utilization cannot be defined, stop the experiment or choose another job. Simple metrics create a fast feedback loop; impact requires richer telemetry, clear jobs, and carefully designed surveys.

Metrics must not become individual targets; psychological safety is essential for honest data. Organization-level measures guide process improvement. Benchmarks and percentiles add context about gaps, but comparisons help only when methods, use cases, and data limits are understood.

AI in SDLCDeveloper productivityResearch methodology