Measuring AI Code Assistants and Agents (Category AI)
I read a fresh—||yesterday’s||—research report on measuring AI’s effect on software development by Abi Noda, co-founder and CEO of the developer-productivity platform DX. It is interesting because the DX team helps shape thinking about developer productivity:
- They created the DevEx model, which I covered in “DevEx: What Actually Drives Productivity” and “DevEx in Action.”
- Their team includes Nicole Forsgren, who drove the development of DORA metrics, co-authored Accelerate and contributed to the SPACE framework.
With that background, let’s look at the report’s main ideas.
1. AI’s impact on development The report argues that AI is fundamentally changing engineering in technology companies: the constraint increasingly lies in how much AI augments engineers’ capabilities rather than simply how many engineers a company has. It cites several company results:
- Booking.com introduced AI tools for more than 3500 engineers and achieved a 16% increase in throughput within a few months.
- Intercom nearly doubled the use of coding assistants and reported a 41% increase in developers’ time savings.
AI also broadens the meaning of “developer.” Product managers, designers and business analysts can now build working software with AI, blurring the boundary between technical and nontechnical roles. Remember vibe coding? :)
2. Key metrics The authors extend their DX framework with a DX AI component focused on three things:
- Utilization: track adoption and use of AI tools. Research suggests that even leading organizations reach only around 60% active use.
- Impact: measure the actual effect on productivity. They recommend combining direct measures, such as developer time savings, with indirect measures from DX Core 4, including PR throughput, perceived delivery speed and the developer experience index.
- Cost: track spending and net benefit: developer time saved minus costs.
Interestingly, DX provides its customers with benchmarks for these measures.
3. Balancing speed and quality Organizations need to balance efficiency measures with quality indicators to avoid undermining development speed over the long term. AI-generated code can be less intuitive for people to understand, creating potential problems.
4. Measuring AI agents The report recommends treating autonomous agents as extensions of development teams rather than independent contributors. Each developer will increasingly act as the manager of a team of AI agents, with their results evaluated through that team’s outcomes.
5. Introducing a measurement system It is important to communicate why AI-related metrics are being introduced. The authors recommend emphasizing three points:
- The metrics will not be used to assess individual employee performance. ||They won’t, will they?||
- The purpose is to understand how AI-assisted work affects developer experience and software quality, rather than to micromanage.
- The data supports investment decisions by identifying tools and workflows that provide real value.
6. Avoiding pitfalls The report strongly warns against top-down mandates and using these metrics for individual performance assessment. Measures such as the volume of generated code are especially easy to game. Proactive communication is essential; otherwise, speculation and fear can fill the information vacuum.
Overall, the report shows why AI-specific recommendations need to be combined with measures of overall developer productivity to understand AI’s effect on organizational effectiveness.
#AI #ML #PlatformEngineering #Software #Architecture #Processes #DevEx #Devops