Skip to content
#AI4SDLC

[2/2] Claude Code and Expertise: Why Session Success Is Not a Sign of Productivity (Category AI4SDLC)

#AI4SDLC #AI #Research #Engineering #Agents #Metrics #DevEx

Keep going. analysis research by Anthropic"Agentic coding and persistent returns to expertise" about the use of Claude Code. In the first part, we talked about the methodology and results of the study, and now we would like to understand how to interpret them and how to assess the weight of the evidence given in the study. I would put them on a scale like this.

1High confidence in the description of the selected traffic Claude Code What modes of operation were met, how many internal actions were triggered by the user's request and how planning and execution were divided within the sessions. 2Average confidence in the relationship between assessed expertise and outcome The difference persists after statistical controls and in Anthropic's internal validation, where code from sessions with higher expertise is more likely to fall into the main. But this is an observation, not an experiment. 3Low confidence in causality conclusions It's about team performance, quality or the job market. The study simply did not measure these results.

The main methodological problem of the study is the widespread bias of methods. (CMB, common-method bias). The bottom line is that both human expertise and session success are extracted from a single transcript model. Moreover, the definition of expertise includes the exact statement of the problem, targeted verification and the ability to correct Claude. And success includes past tests, commits, and user confirmation. It turns out a partial intersection of structures. A person who knows how to request a good test and clearly communicate the result looks both more competent and more successful to the classifier.

Nana 198 SWE-chat sessions were compared to Mythos Preview. For ordinal metrics, the exact match was 53–68%, and with a tolerance of the neighboring level - 78–99%; for categorical - 78–98%. The man only checked. 15 discrepancies that the model judge considered significant. It's acceptance of a strong model. (strong-model agreement), not reference data marked by a person (human ground truth).

The second interesting thing is the verified success. (verified success)which requires explicit confirmation of the result. But the researchers don't see whether the code was accepted by the team, deployed, secure, and useful to the user. That is why the success of a session cannot be quietly renamed productivity.

The third interesting point is economic value. Anthropic reports that the model cost of the average task has increased by 27%. But the price is built by turning the session into a conditional freelance job and matching with ads. Appendix to the appraiser log R² ≈ 0,38: small tasks he overestimates in approximately 2,4 The biggest ones are underestimating about 7 once. The correct conclusion is more modest: the estimated composition of tasks changed, not the measured economic value.

To understand the implications for AI’s impact on development processes, one source is not enough. Next, it is useful to put a separate randomized experiment Anthropic.How AI Impacts Skill Formation" (from 28 January 2026 year). There. 52 The developer studied an unfamiliar Python library Trio with or without GPT-4o, and then passed the same test without an assistant. The AI group finished about two minutes faster, but the difference was not statistically significant. In the test, she scored on average. 50percentage 67% of the group without AI; the biggest gap was in debugging. In qualitative analysis, explanation and self-attempt were associated with better learning than full delegation.

These studies cannot be compared to the forehead, because The first one observes the experience shown in real Claude Code sessions. The second experimentally tests the short-term formation of a new skill on a small sample. Together, they create an organizational paradox: higher expertise in tasks is associated with greater agent activity and success of sessions, and full delegation in a narrow controlled experiment is associated with weaker skill acquisition.

#AI #AI4SDLC #Research #Engineering #Agents #Metrics #DevEx