Skip to content
#AI4SDLC

Research Insights Made Simple 21How to Measure Coding Agents in Dialogue (Category AI4SDLC)

#AI4SDLC #AI #Agents #Evals #Engineering #Research

What changes if the requirements for coding agents do not come in perfect promptom, but are clarified along the way? 17 July 16:00 MSC discuss Alexei Litvinov on the podcast "Research Insights Made Simple" #21" We’ll talk about the new SWE-Together and SWE-INTERACT benches released on the same day. I’ve been talking to both of you on this channel. (SWE-Together and SWE-INTERACT). Together, they show one important thing: the ability of an agent to autonomously code a complete, comprehensive TK. (Like a classic SWE-bench.) This does not mean that it will be useful in real collaboration. In practice, requirements rarely come as a perfect prompt - they are refined on the fly, and here the ability to adapt to a changing context comes to the fore.

Alexey Litvinov – Principal Engineer, author of the book andEducational Program in AI-Assisted Engineering. He explores and implements ways to turn working with AI agents from a set of successful experiments into a managed, reproducible and verifiable engineering process on real codebases. Losha also has a Telegram channel - @tip\ podcast.

If we go back to the benches themselves,

  • SWE-Together (Researchers from the banned company Meta in Russia) Goes from real user sessions and enters the User Correction metric - how many times a person had to return an agent to the course
  • SWE-INTERACT (from researchers at Scale.ai) He does the opposite experiment: he takes the same tasks and the same tests, but he reveals the requirements gradually. In strong models, the proportion of solved problems in this mode drops by about half, and the work becomes longer and more expensive. The interest of these studies is not in the next ranking of models. They allow you to see the price of the dialogue: corrective steering, forgotten requirements, repeated checks and the dependence of the result on the binding of the agent. Multi-turn work turns out to be a separate engineering ability.

We will talk about which of the two techniques is closer to real development, whether you can trust the user simulator, why the agent forgets the requirements from early replicas, and how teams build their own multi-turn evals.

17 July, 16:00 MSK. Come. figure out And argue with us.

#AI #AI4SDLC #Agents #Evals #Engineering #Research