Skip to content
back to the archive page
#AI4SDLC

[1/2] Claude Code and Expertise: Findings from 398 Thousand Sessions (Category AI4SDLC)

I was adding Anthropic’s new study, Agentic coding and persistent returns to expertise, dated 16 June 2026, to our AI4SDLC 2026 meta-study. Its headline conclusion is that coding agents do not make expertise obsolete. What interests me more is how the authors measure the division of labour between people and Claude Code in real work sessions. This is a large observational source, not an experiment: a random sample of 398 198 interactive sessions from 234 751 users, spanning October 2025 to April 2026. It includes Claude Code in the CLI, Claude.ai and the desktop app. Internal Anthropic employee sessions, third-party IDEs, SDKs and headless runs were excluded. It is substantial observational evidence from one product, not an independent study of all software development.

The method is interesting. Model classifiers analysed transcripts with privacy safeguards and labelled:

  • the kind of work being done;
  • who made planning and execution decisions;
  • the user’s demonstrated domain competence;
  • the session’s outcome.

The authors checked part of the classification against independent telemetry: model and tool calls, output volume, and lines added and deleted. A separate transcript-based success classifier looked for commits, tests and user confirmation. Researchers did not read individual private transcripts and published only aggregates.

Expertise meant competence in the current task, not seniority, years of experience or job title. A five-level classifier assessed how precisely users framed the task, understood terminology, asked for specific risks to be checked and substantively corrected Claude. A senior developer using Rust for the first time could be a beginner at a Rust task. An accountant with no Python experience who specifies reconciliation rules precisely and catches a month-end closing error could be an expert at that accounting task. This is not a comparison of junior and senior engineers.

The findings are interesting:

1️⃣ A division of labour is already visible In a typical session, the person made around 70% of decisions about what to do, but only 20% about how to do it. The user largely retained the goal and criteria, while Claude chose files, commands and implementation details. This fits the sequence I described in State of AI4SDLC: intent → context → plan → tasks → implementation → verification. Anthropic’s analysis does not prove that the whole SDLC has changed, but it shows a similar pattern within interactive agent work.

2️⃣ Different competence levels were associated with different levels of delegation In novice-labelled sessions, one request prompted roughly 5 Claude actions and 600 words of output; in expert sessions, around 12 actions and 3 200 words. The association remained after accounting for work mode, task value, month, occupation and model family. But more actions and text do not necessarily mean more useful work. A long chain may represent autonomous execution or unnecessary iterations and rework. This measures delegation, not productivity.

3️⃣ Higher expertise ratings were associated with more frequent verified success After statistical adjustment, the rate rose from 14.5% for novices to 20.9% for beginners, 28.3% for intermediate users, 29.5% for advanced users and 32.9% for experts. The largest jump was between novice and intermediate; the increase from intermediate to expert was much smaller. Here, Verified success means that the classifier considered the goal achieved and found at least one signal: a passing test or command, an appropriate commit or PR, or explicit user confirmation. This is stricter than a model verdict alone, but does not mean a change reached production and delivered value.

I have already discussed Anthropic’s Agentic Coding Trends and the shift from writing code to directing agents. This study adds an observed pattern: the agent handles most execution, while task-specific expertise is associated with more frequent session success.

In the next part, I will examine the awkward questions: how reliably expertise and success were measured, why correlation is not causation, and what a separate randomised experiment says about skill development.

#AI #AI4SDLC #Research #Engineering #Agents #Metrics #DevEx