Skip to content
back to the archive page
#AI

Does AI Actually Boost Developer Productivity? (100k Devs Study) (Category AI)

I watched an interesting 20-minute talk by Stanford University’s Yegor Denisov-Blanch about measuring engineering productivity. He presented a three-year Stanford study of 100k developers across 600 companies. Its reported finding is an average AI productivity gain of only 15–20%, varying sharply with task context, code maturity, language popularity, and repository size. Here are the main ideas from the talk and the research page.

1. Why are existing productivity estimates unreliable?

  • Most published studies are funded by tool vendors such as GitHub Copilot or Sourcegraph Cody, affecting the sample and metrics.
  • Counting commits or PRs without accounting for task size and quality creates misleading output growth. There may be more code, but how useful is it?
  • Greenfield and toy projects overstate AI’s benefits: LLMs readily generate boilerplate, while real enterprise code rarely starts from scratch.
  • Self-reported productivity is rarely accurate, although surveys can measure other factors, such as satisfaction or well-being, well.

2. What was Stanford’s methodology?

  • Connect Git repositories, mostly private ones, to capture the team’s context.
  • Have 10–15 architects assess code quality, maintainability, and complexity. Their ratings correlated well. In effect, this created labels for supervised learning.
  • Train a model to reproduce expert assessments, with high correlation on the expert-reviewed project sample, so evaluation could scale without costly reviews.
  • Classify changes as adding functionality, deletion, refactoring, or rework of recently written code.
  • Apply the model retrospectively from 2019 to 2025 to examine the effects of COVID, LLM adoption, and other changes.

3. The main numerical findings

  • Raw code volume grew by 30–40% after AI adoption, including both useful work and rework.
  • Adjusted productivity grew by 15–20%, accounting for bug fixes.
  • For complex tasks, gains were only slightly above 0%, with substantial variation and sometimes a decline.
  • Put this into a 2×2 matrix, as consultants like to do, and the axes are task complexity and project maturity:

Complexity / project type: Greenfield | Brownfield Low complexity: 30–40% gains | 15–20% gains High complexity: 10–15% gains | 0–10% gains, sometimes losses

  • Language matters: LLMs work better with popular languages and worse with esoteric ones.
  • Codebase size also matters: benefits fall logarithmically as it grows. The proposed explanation is context-window limits, more noise, and greater coupling between parts of the code.

The practical advice seems to be:

  • Assess your project’s typical task complexity and maturity before adopting LLM assistants widely.
  • Favor popular technologies, for which the best models have more training data and offer better suggestions.
  • For legacy monoliths, pilot on small modules. Off-the-shelf assistants, such as basic Cursor, may not improve productivity.
  • Monitor the rework share: rapid code growth can create an illusion of productivity.
  • Combine quantitative Git analysis with qualitative review, rather than relying only on a team-satisfaction survey.

P.S. Yegor Denisov-Blanch also authored the research on “ghost” developers that caused a stir last year. I could not find the study itself, though there are many posts about it.

P.P.S. Compare this with the design of the METR experiment :)

#Engineering #AI #Metrics #Software #DevEx #Productivity

Open video on YouTube