Skip to content
back to the archive page
#AI4SDLC

Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance (Category AI4SDLC)

I read this paper, which shows why agent rankings that ignore task type are not very useful. The right answer to “Which coding agent is best?” is “It depends” :). The authors used the AIDev dataset of GitHub pull requests, which I discussed recently. They analysed merge rates while accounting for the fact that one agent might get more documentation tasks, another more bug fixes and features. You cannot compare them as though their workloads were identical.

The work is by Giovanni Pinna, Jingzhi Gong, David Williams and Federica Sarro, from the University of Trieste, King’s College London and University College London. It was accepted to the MSR ’26 Mining Challenge Track. It is observational research rather than a randomised experiment, so it reveals associations, not causal effects.

Methodology 1️⃣ The authors used a subset of AIDev-POP: AI-generated pull requests from GitHub repositories with 100+ stars. They retained 7 156 of 33 596 PRs: closed PRs in MIT/Apache-2.0 repositories, with at least one review or comment from someone other than the author before closure. Success was defined by acceptance rate: whether the PR was merged. 2️⃣ Rather than putting OpenAI Codex, GitHub Copilot, Devin, Cursor and Claude Code in one comparison table, they split PRs into task types: docs, feat, fix, test, refactor, chore, build, ci, perf and so on. They then examined weekly trends, task distributions and pairwise comparisons within each task type.

Results 1️⃣ Overall rates look tempting: OpenAI Codex 77.9%, Cursor 74.5%, Claude Code 71.9%, GitHub Copilot 68.0% and Devin 61.6%. But context matters. Claude Code has only 139 PRs in the sample, versus 2 252 for Devin and 2 194 for Copilot. Their workloads differ too: Copilot has more fixes, Claude Code more features, and Cursor a shorter, concentrated observation period. 2️⃣ For me, the main result is the size of the confounding factor, not the winner. Task type produces larger gaps than many differences between agents. Chore tasks were accepted in 84.0% of cases, versus 55.4% for perf, a gap of roughly 29 percentage points. Among common tasks, docs reached 82.1% and features 66.1%. An agent assigned more documentation may look “smarter” simply because its workload is easier. 3️⃣ Devin stands out over time: the authors report the only sustained positive trend, +0.77% acceptance rate per week across 32 active weeks, rising roughly from 60% to 80%. They caution against a causal interpretation: model improvements, users learning, changing task mixes, concentration in different repositories or a combination could explain it. 4️⃣ Task-stratified comparisons make a neat leaderboard harder. Codex performs consistently across categories, from 59.6% to 88.6%, especially fix and refactor. Claude Code leads in docs and features, but some samples are small, making generalisation risky. Cursor does well on fix/test tasks. After Bonferroni correction for multiple comparisons, significant differences appear mainly in fix and feat, rather than across all categories.

Limitations deserve separate attention Acceptance rate is not code quality. A merged PR can introduce bugs, technical debt or security issues. Public GitHub is not enterprise development. Repositories have different review cultures and acceptance rules. The chi-square tests do not explicitly model repository-level clustering. Task labels come from the AIDev pipeline, not manual annotation of every PR. Agent samples differ in size and time coverage. These limitations are precisely why the paper is useful: it encourages sound adoption metrics rather than choosing an agent by one attractive headline number.

For a team adopting coding agents, I would take one practical rule: stratify internal analysis by task. Compare documentation with documentation, bug fixes with bug fixes, tests with tests and features with features. Separately examine review load, comment counts, diff size, CI, rollbacks, static analysis, discovered defects and maintenance after merging.

#AI #AI4SDLC #Engineering #Research #Agents #Software #Evals