Skip to content
Research Insights Made Simple logo
Recording · July 17, 2026#AI4SDLC · Multi-turn evals

SWE-Together × SWE-INTERACT

Two ways to measure how coding agents work with people

/ Research Insights Made Simple #21 · Multi-turn evals

Slide contents

  1. 1. SWE-Together × SWE-INTERACT

    Two ways to measure how coding agents work with people

  2. 2. One-shot scores hide the dialogue cost

    Single-turn

    Full specification upfront

    Final code score

    Steering stays invisible

    Multi-turn

    Requirements arrive gradually

    Review and corrections

    Dialogue cost becomes measurable

  3. 3. Same day, two research questions

    SWE-Together · Meta*

    Bottom-up: real sessions

    How much steering to succeed?

    109 tasks · frozen rubric

    SWE-INTERACT · Scale AI

    Top-down: benchmark tasks

    What is lost in dialogue?

    75 tasks · original verifier

  4. 4. Only 109 tasks survive strict filtering

  5. 5. A session becomes a reproducible task

    SWE-Together diagram showing three stages from a real session to a reproducible benchmark task

    Raghavendra et al. · SWE-Together · Figure 2 · PDF p. 3 · arXiv:2606.29957v1 · CC BY 4.0 · cropped

  6. 6. The simulator reacts to the live trajectory

    Three contexts

    Session intent anchors

    Agent's latest turn

    Simulator memory

    One action

    Stay silent · ask

    Redirect · add requirement

    Check an external artifact

  7. 7. Every reply combines three contexts

    SWE-Together user-simulator context diagram with anchors, current turn, memory, and a structured action

    Raghavendra et al. · SWE-Together · Figure 4 · PDF p. 5 · arXiv:2606.29957v1 · CC BY 4.0 · cropped

  8. 8. Success and steering are separate axes

  9. 9. Stronger models need fewer corrections

    SWE-Together plots showing an inverse relationship between model capability and User Correction

    Raghavendra et al. · SWE-Together · Figure 6 · PDF p. 9 · arXiv:2606.29957v1 · CC BY 4.0 · cropped

  10. 10. Realism comes through severe selection

    109 of 11,260 · 0.97%

    LLM simulator + LLM judge

    Two runs per task

    Text only, no interruptions

  11. 11. Seventy-five tasks become dialogues

  12. 12. Delivery changes; the verifier stays fixed

    SWE-INTERACT diagram turning autonomous tasks into interactive workflows with progressive requirement disclosure

    Wu et al. · SWE-INTERACT · Figure 1 · PDF p. 2 · arXiv:2606.30573v1 · CC BY 4.0 · cropped

  13. 13. Expert Nitpicker can inspect before replying

  14. 14. Shell grounds feedback in the workspace

    SWE-INTERACT architecture with separate agent and user-simulator containers and shell access to the workspace

    Wu et al. · SWE-INTERACT · Figure 2 · PDF p. 4 · arXiv:2606.30573v1 · CC BY 4.0 · cropped

  15. 15. Dialogue nearly halves resolve rate

  16. 16. Interaction costs more across every resource

    SWE-INTERACT table comparing single-turn and multi-turn resolve rate, steps, tokens, and cost

    Wu et al. · SWE-INTERACT · Table 1 · PDF p. 5 · arXiv:2606.30573v1 · CC BY 4.0 · cropped

  17. 17. Bugs and forgotten requirements make up 68%

  18. 18. Revealed requirements can disappear before submission

    SWE-INTERACT chart showing five failure modes and their definitions

    Wu et al. · SWE-INTERACT · Figure 5 · PDF p. 8 · arXiv:2606.30573v1 · CC BY 4.0 · cropped

  19. 19. Two methods reveal the same signal

    SWE-Together

    109 real sessions · opencode

    Text replay · frozen rubric

    U-Corr · natural, selective

    SWE-INTERACT

    75 benchmark tasks · native harness

    Shell user · original verifier

    Gap · controlled, persona-bound

  20. 20. Evals must price outcomes and steering

    Success / verifier reward

    User Correction

    Single-turn → multi-turn gap

    Requirement ledger

    Commit + review each iteration

    A recommendation, not an internal experiment report

  21. 21. These papers do not cover production yet

    Open-ended ambiguity

    Other user personas

    Production toolchains

    Pure model-only comparison

  22. 22. Multi-turn is a separate capability axis

    Book Cube

    Paper links and new engineering research breakdowns are in the channel

    Alexander Polomodov, Technical Director & Fellow, T-Technologies

    @Book_Cube