Skip to content
back to the archive page
#AI4SDLC

SWE-Together: Evaluating Coding Agents Through Dialogue (Category AI4SDLC)

SWE-Together: Evaluating Coding Agents Through Dialogue Rather Than in One Turn (Category #AI4SDLC)

I read the new paper “SWE-Together: Evaluating Coding Agents in Interactive User Sessions” from a team at Meta, a company banned in Russia (arXiv, 29 June 2026). The authors target a weakness in almost all coding-agent benchmarks: they are static. An agent receives the complete task specification at once and is judged on its final code. Real work with an agent is different: it is a dialogue in which the user gradually reveals their intention, clarifies requirements and corrects errors along the way. Broadly, the authors wanted to test whether a stronger model needs fewer user interventions while working. ||And that really was confirmed.||

Benchmark problems are familiar: some become saturated and stop distinguishing top models—see my discussion of the SWE-rebench talk. They also do not always measure how engineers actually interact with agents. SWE-Together (leaderboard, GitHub) tries to move from a single-turn specification to multi-turn dialogue, closer to real work with agents.

The dataset collection process is interesting:

  • The input was 11,260 real user sessions with coding agents from four public Hugging Face datasets: DataClaw, Pi-staging, Hyperswitch and SWE-chat.
  • The output was 109 reproducible repository-level tasks, a conversion rate of 0.97%.
  • Strict filters sat between them: a recoverable repository state, a clear objective, a verifiable outcome and a final change written predominantly by the agent rather than a human.
  • Each resulting task is a sandbox with a pinned commit, environment and verifiers.

The most interesting engineering component is the user simulator. You cannot replay an original session word for word: a new agent will take a different path, making the original user’s messages irrelevant. The LLM simulator is therefore anchored in the original session’s intentions but responds to the evaluated agent’s live trajectory. After each turn, it makes one decision: stay silent, ask a question, redirect the agent, add a requirement or request a check of an external artifact. In a blind test, annotators could not distinguish the simulator from real users: it was judged human 46% of the time, against a chance level of 50%. The difference was not statistically significant—a kind of modified Turing test passed.

The evaluation also addresses a familiar problem: fixed tests can distort results in either direction. Narrow tests fixate on incidental implementation details, while broad ones demand behavior nobody requested. Here, an agent judge evaluates the final repository state against a frozen rubric. Weighted behavioral goals are created once, offline, before and independently of any candidate, then applied equally to all models. The measure is behavioral completeness, not similarity to the original patch.

There is also a User Correction metric: how often the user explicitly had to correct the agent on the way to the result, with softer nudges weighted at 0.2. Two agents with identical final scores may require very different levels of effort from their users.

Results for seven models using the same opencode harness:

  • Claude Opus 4.8 leads with pass@1 of 63%, an average judge score of 0.801 and the fewest corrections, 1.38 per session. It also uses the most output and reasoning tokens: 74k per task.
  • GPT-5.5 has the second-highest average judge score, 0.763, and is the most economical: 29.9k tokens and 10.7 minutes per task.
  • Capability and correction count have an almost linear relationship, with a Pearson correlation of −0.92.

The hypothesis that stronger models need fewer user interventions thus received quantitative support within this benchmark. An amusing detail: Meta’s own models are absent from the table.

The authors have managed to measure a tax on verifying agent output through User Correction. AI has made code production cheaper, but not trust in changes, and “how often a person had to steer the agent” is a reasonable proxy for that trust. Two practical implications follow: 1️⃣ Evaluate agents with a pair of metrics—success and the cost of steering—rather than pass@1 alone. 2️⃣ Your own agent sessions are ready-made input for such evals. The paper describes the session-to-task pipeline in detail, and the code is open.

P.S. Scale AI’s related SWE-Interact whitepaper, published on the same day as SWE-Together, is covered in the next post.

#AI #AI4SDLC #Engineering #Agents #Evals #Research