SWE-Interact: The Cost of Interactivity for Coding Agents (Category AI4SDLC)
I decided to cover a second paper today on evaluating coding agents through dialogue: Scale AI’s “SWE-Interact: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions.” An interesting coincidence: it appeared on arXiv on 29 June 2026, the same day as SWE-Together from Meta (banned in Russia), which I discussed earlier today. Both teams arrived at almost the same idea at almost the same time: single-turn benchmarks no longer reflect how people actually work with agents; we need to measure dialogue. Multi-turn evaluation seems to be becoming a field of its own. Or perhaps Meta and Scale AI are somehow connected… ||Oh right, Meta bought 49% of Scale AI for $14 billion and also hired its founder.||
Scale approaches the benchmark from the opposite direction. Meta works bottom-up: it reconstructs 109 tasks from 11,260 real sessions and anchors the simulator in real users’ intentions. Scale works top-down: it takes 75 ready-made tasks from three existing benchmarks—SWE-bench Pro, SWE Atlas (Refactoring) and DeepSWE, the first two of which are its own—hides the full specification inside a user simulator and makes it reveal requirements gradually. The original task verifier still checks the solution. It is an elegant move: the same task and verifier, with only the delivery of requirements changing. The difference in resolve rate becomes an almost pure measure of the cost of interactivity.
That cost turns out to be high. The strongest models solve around 50% of tasks in a single turn, but only about a quarter of those same tasks through dialogue: GPT-5.5 falls from 48.0% to 24.7%, and Opus 4.8 from 50.7% to 26.7%. They typically take 3–4 times as many steps and up to 4.5 times as many tokens, while cost per task is 2–3.5 times higher: $11.80 versus $5.09 for Opus 4.8. Working with a user adds a separate dimension of difficulty.
The user simulator here is more interesting than Meta’s.
1️⃣ The authors built a persona from statistics They analyzed real sessions from the SWE-chat dataset and modeled the most common type: a meticulous expert in vibecoding mode. This is a senior developer who writes brief, casual messages and does not touch the code personally—in this mode, SWE-chat shows the agent writing more than 99% of it—but closely checks exact API signatures and raises comments one at a time.
2️⃣ The simulator itself is an agent It has shell access to the coding agent’s working copy, including git, grep, sed and find, and inspects diffs and commits before responding. Meta’s simulator sees only text: trajectory summaries and tool outputs. The SWE-Together authors list this as a limitation themselves. Here, the experiment shows how demanding that meticulous persona is: trajectory lengths grow by 30–50% for most models, the number of user messages almost doubles, and resolve rate falls for 4 of the 5 models compared with a neutral simulator that simply provides requirements on request.
The paper also analyzes failures, auditing 287 unsuccessful trajectories.
- Some models ask good questions and include more than 80% of the hidden goals in their plans even from a vague description. By the end of the dialogue, GPT-5.5, Opus 4.8 and Sonnet 4.6 cover more than 90% on average.
- The failure labels break down as follows:
- Technical implementation errors: ~34% of labels.
- Forgotten requirements: ~34%, where a requirement from an early message never reaches the final code.
- In ~12%, the simulator itself never disclosed a necessary requirement. This is probably a false positive from the benchmark, as the authors acknowledge.
The study has limitations too.
- It contains only 75 tasks, compared with Meta’s 109.
- It uses just one persona; modeling different users would be interesting.
- The single-turn baseline was run once.
- Results depend noticeably on the simulator model: with GPT-5.5 as the user, Opus 4.8 solves 30.7%; with Opus 4.7, it solves 22.7%.
- Each agent uses its native harness—Claude Code, Codex CLI or Kimi CLI—while Gemini uses the third-party OpenCode. Model comparisons are therefore mixed with harness comparisons, unlike SWE-Together’s shared opencode harness.
- Scale AI sells data and model evaluation and promotes its own benchmarks. Two of the four benchmarks from which tasks were drawn were Scale AI’s.
Taken together, the two papers give us signals from two directions:
- Meta measured how many user corrections an agent needs to reach a result; stronger agents need fewer.
- Scale measured how much capability is lost when requirements arrive in pieces; even strong agents lose half.
In practice: 1️⃣ When choosing a model to work with real people, look at the multi-turn gap as well as the autonomous resolve rate. 2️⃣ Of the two main failure modes, technical errors are a model problem, while forgotten requirements can be addressed through process: a plan file, an explicit requirements list, and a commit and review for each iteration. That is a matter for your harness, not just the benchmark.
#AI #AI4SDLC #Engineering #Agents #Evals #Research