SWE-agent: how an interface helps a model work with code (Series #AI4SDLC)
In Research Insights Made Simple #33, I explored SWE-agent, the Princeton team’s NeurIPS 2024 paper. It is interesting to read it alongside the next chapter of the story: the authors first built a dedicated interface for the model, then returned to a minimal bash-based design in mini-swe-agent.
Fixing a bug in a real project means finding the right files, understanding their relationships, making changes, and checking the result. Humans have IDEs, search, and error highlighting to help them. The authors proposed designing a workspace for a language model too: an Agent-Computer Interface, or ACI. They did not change the model’s weights.
In SWE-agent, search helps narrow the scope of work gradually. The file viewer shows line numbers and the position within the document. The edit command replaces a range of lines and immediately returns the updated fragment. If an edit introduces certain linter errors, the system rejects it and shows what went wrong. Older tool responses are shortened so that previous file versions do not fill the context.
These choices were tested experimentally. On 300 SWE-bench Lite tasks, GPT-4 Turbo with the full SWE-agent solved 18% of tasks, while an agent with only a shell and a solution demonstration solved 11%. Removing the special editor reduced the result to 10.3%; keeping the editor without the linter gave 15%. These are results for specific 2024 configurations; the percentages cannot be carried over to today’s models.
While reading, I paid particular attention to the cost of constraints. A linter helps the agent avoid getting stuck in its own errors, but it may forbid an intermediate state needed for a large edit. Limiting search results saves context but forces extra queries. Every such decision changes the sequences of actions available to the agent.
Then models became better at using the shell. In mini-swe-agent, the same team kept bash, a linear message history, and independent command executions. The authors associate this simplification with the models’ growing capabilities. This is a continuation of the story, not a controlled experiment from the original paper.
I think SWE-agent’s four principles still make sense: simple actions, compact actions, meaningful concise feedback, and help recovering from errors. The implementation is worth rechecking, though. A 100-line window and the last five observations are settings for a particular system, not a universal recipe.
The authors’ working method is worth adopting too: read trajectories, find a recurring failure, change a component, and measure the result. After changing the model, check again what value each part of the harness provides. Yesterday’s improvement may well become today’s constraint.
P.S. I attached the whitepaper with my notes. Interestingly, it runs to 100+ pages, but the main part fits into 10 pages, with another 10 providing an expanded account of the ACI itself. After that come program listings, prompts, and so on.
#AI4SDLC #Research #Agents #Evals #Engineering
Files from the post
- Annotated_SWE_agent_Agent_Computer_Interfaces_Enable_Automated_Software.pdfDownload PDF