
SWE-Together × SWE-INTERACT
Two ways to measure how coding agents work with people

Two ways to measure how coding agents work with people
Two ways to measure how coding agents work with people
Single-turn
Full specification upfront
Final code score
Steering stays invisible
Multi-turn
Requirements arrive gradually
Review and corrections
Dialogue cost becomes measurable
SWE-Together · Meta*
Bottom-up: real sessions
How much steering to succeed?
109 tasks · frozen rubric
SWE-INTERACT · Scale AI
Top-down: benchmark tasks
What is lost in dialogue?
75 tasks · original verifier
SWE-Together diagram showing three stages from a real session to a reproducible benchmark task
Raghavendra et al. · SWE-Together · Figure 2 · PDF p. 3 · arXiv:2606.29957v1 · CC BY 4.0 · cropped
Three contexts
Session intent anchors
Agent's latest turn
Simulator memory
One action
Stay silent · ask
Redirect · add requirement
Check an external artifact
SWE-Together user-simulator context diagram with anchors, current turn, memory, and a structured action
Raghavendra et al. · SWE-Together · Figure 4 · PDF p. 5 · arXiv:2606.29957v1 · CC BY 4.0 · cropped
SWE-Together plots showing an inverse relationship between model capability and User Correction
Raghavendra et al. · SWE-Together · Figure 6 · PDF p. 9 · arXiv:2606.29957v1 · CC BY 4.0 · cropped
109 of 11,260 · 0.97%
LLM simulator + LLM judge
Two runs per task
Text only, no interruptions
SWE-INTERACT diagram turning autonomous tasks into interactive workflows with progressive requirement disclosure
Wu et al. · SWE-INTERACT · Figure 1 · PDF p. 2 · arXiv:2606.30573v1 · CC BY 4.0 · cropped
SWE-INTERACT architecture with separate agent and user-simulator containers and shell access to the workspace
Wu et al. · SWE-INTERACT · Figure 2 · PDF p. 4 · arXiv:2606.30573v1 · CC BY 4.0 · cropped
SWE-INTERACT table comparing single-turn and multi-turn resolve rate, steps, tokens, and cost
Wu et al. · SWE-INTERACT · Table 1 · PDF p. 5 · arXiv:2606.30573v1 · CC BY 4.0 · cropped
SWE-INTERACT chart showing five failure modes and their definitions
Wu et al. · SWE-INTERACT · Figure 5 · PDF p. 8 · arXiv:2606.30573v1 · CC BY 4.0 · cropped
SWE-Together
109 real sessions · opencode
Text replay · frozen rubric
U-Corr · natural, selective
SWE-INTERACT
75 benchmark tasks · native harness
Shell user · original verifier
Gap · controlled, persona-bound
Success / verifier reward
User Correction
Single-turn → multi-turn gap
Requirement ledger
Commit + review each iteration
A recommendation, not an internal experiment report
Open-ended ambiguity
Other user personas
Production toolchains
Pure model-only comparison
Book Cube
Paper links and new engineering research breakdowns are in the channel
Alexander Polomodov, Technical Director & Fellow, T-Technologies
@Book_Cube