Skip to content
#AI

[1/2] Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (AI column)

#AI #ML #Management #Leadership #Software #SoftwareDevelopment #Architecture #Metrics #Devops #Processes

I got to know this pretty quickly. whitepaperIt shows developers slowing down when using AI. But initially I did not want to write about it - the methodology seemed interesting to me, but the sample size is as much as possible. 16 I thought engineers were too small to make loud statements about slowing down development. But I changed my mind after the headlines came out.Scientist rapes journalistAnd then the headline-grabbing people started sending in this study with the argument, "Here's the evidence: mustache, paws and tail."

After that, I went and read not only the brief. squeeze from METR (Model Evaluation & Threat Research)but all 50 pages of the main whitepaper, where the authors described the entire methodology of the experiment, their thoughts about the results and their causes (Of course, without any significance.) And even instructions for participants in experiments. I will share my thoughts on this experiment.

To begin with, why this study has received so much attention, and there are several reasons. 1. Contradiction to conventional wisdom The findings directly contradict the widely held belief that AI tools inevitably increase developer productivity. The study showed that AI tools actually slow down experienced developers. 19This is a shocking discovery even for the authors themselves. 2. Methodological accuracy The authors conducted a study using randomized controlled trials (In fact, the a/b test) To assess the real impact of AI on the productivity of developers in their normal work environment. The study differs from previous work by using fixed tasks and experienced developers working in familiar codebases. However, here the participants knew about the experiment itself, so it could manifest itself. hawthorne 3. Relevance for the industry The results of research on the impact of AI have huge implications for assessing the multibillion-dollar investment in AI development tools. Previous studies have taken a different approach to evaluating the effect and all have shown increased productivity. And this study raises questions about the real return on investment in AI development tools.

Let’s talk about why we should trust METR researchers. METR is a non-profit research organization specializing in assessing the capabilities of frontline AI systems and their potential risks. The organization works with leading AI companies, including OpenAI and Anthropic, and collaborates with government AI security institutes. The team of researchers included technicians Joel Becker, Nate Rush, David Rein, as well as Beth Barnes - founder and principal investigator of METR, former OpenAI and DeepMind researcher, leading expert on AI security and evaluation of AI systems.

And what did the methodology of randomized controlled trials look like?

  • Participated. 16 Experienced developers of open-source projects with average experience with repositories 5 years 1500 commissions on average.
  • They decided. 246 Real problems from large open-source repositories (average 23,000 stars, 1,100,000 line). The design of the experiment assumed that each task was randomly assigned to one of two groups. AI-permitted group: developers could use any AI tools (Cursor Pro with Claude 3.5/3.7 Sonnet) AI-prohibited group: the use of generative AI is prohibited The time of the task was measured, which the developers tracked themselves. Additionally, screen recordings, interviews and detailed analytics of AI usage were collected. In addition, all tasks underwent a standard code review process and had to meet high repository quality standards.

About the results and why its results should be taken with caution, I will tell you in continuation.

#AI #ML #Management #Leadership #Software #SoftwareDevelopment #Architecture #Metrics #Devops #Processes