Skip to content
back to the episode
Episode summary2026Fellow

SWE-agent: Interfaces Change Outcomes

In episode 33, Alexander Polomodov examines SWE-agent through the question of a language model’s working environment. Why does the same model solve tasks differently when search, editing, and environment responses change? The story moves from a specialized interface in 2024 to a minimal agent using bash, showing why an agent’s design must be revisited as models improve.

Research Insights Made Simple #338 min read

A summary based on the episode audio transcript, supported by the slides. SWE-agent results refer to experiments from 2024; the subsequent research is discussed separately.

The main thread of the material
01

Real repository tasks require navigation as well as code generation

SWE-bench tests agents on real GitHub issues. A solution requires understanding a project, locating the fault, and producing a patch that passes tests. This differs from a self-contained algorithm exercise in which the relevant function is already supplied. The episode recalls how models at the start of this story solved only a small fraction of tasks. The benchmark authors therefore investigated ways to improve the model’s interaction with its computer. Human engineers have editors, search, navigation, and highlighting. A text-based agent has a different workplace: it reads textual tool responses, spends tokens on that information, and cannot use a graphical interface in the same way as a person. The agent–computer interface, or ACI, makes a useful action easy to express and its outcome understandable. Usability is evaluated through the ability to complete a working patch, covering the chain from the initial problem description to submission.

SWE-agent follows a thought–action–observation loop. The model proposes a step, sends a command, and reads the environment’s response. The system discussed requires one thought and one action; an invalid response triggers a format error. Execution takes place in a prepared environment with both an ordinary shell and purpose-built commands. Even the absence of command output is stated explicitly. Submitted patches are evaluated by tests. The main metric is the proportion of tasks solved in one attempt under a limited monetary budget. The full SWE-bench contains 2,294 tasks, while experiments on individual components use the 300-task Lite subset. The episode compares BM25 retrieval followed by immediate patch generation with a shell-only agent and the specialized interface. GPT-4 Turbo and Claude 3 Opus benefit differently, showing that interface value depends on the model. Comparisons must retain the distinction between full and Lite datasets and between settings with and without a demonstration.

02

An action should return the information needed for the next step

The first practical question is where to change the code. Chains of standard commands require exact arguments, sometimes calculations to select line ranges, and can flood the context with irrelevant output. SWE-agent separates search by scale: finding a file, finding matches within a directory, and locating content inside a file. The agent progressively narrows its working area. Opening a file returns its name, line numbers, the selected fragment, and information about the remaining content. The agent can jump to a line or inspect adjacent regions without calculating its position itself. The edit command replaces a specified line range and immediately returns the updated window. No separate read is needed to discover whether the change was applied. A single useful command therefore replaces several small operations. Window size remains an experimental parameter: excessive text can be as unhelpful as missing context. Line numbers provide stable references for the next action.

The next layer checks syntax after an edit. When the linter detects a problem, the change is rejected. The response contains the error, the proposed code, and the original fragment. Each part supports recovery: the error provides a diagnosis, the rejected edit identifies an action that should not simply be repeated, and the original code avoids another search. On SWE-bench Lite, the configuration without edit solved 10.3% of tasks, edit without linting solved 15%, and edit with linting solved 18%. These figures belong to a particular experiment. Validation also restricts the order of work: an intermediate state needed for a larger change may be rejected before the whole change is complete. The host illustrates this tradeoff with a sequence that modifies a definition and then its uses. A guard changes the trajectories available to the agent. Its value must therefore be assessed through both the errors it prevents and the useful work it obstructs.

03

Measure the system and remain willing to change its design

Context grows after every step and accumulates old versions of files. The authors retained the last five observations in full and replaced earlier outputs with a note that they had been omitted, while keeping thoughts and commands. This reduced both volume and stale information. A hundred-line viewing window and that history depth worked well in their configuration; the episode cautions against treating those settings as universal recipes. Successful action sequences included creating a script, inserting code to reproduce the issue, and executing it, or narrowing a search from directory to file to relevant location. Successful solutions often arrived quickly, while difficulties consumed more steps. Trajectory length can reflect task difficulty, however, so the correlation does not establish a stopping rule. Several independent runs solved more tasks, but these were fresh attempts rather than continuation of a stalled conversation. Wrong implementations, overly specific fixes, and unfinished work show limits that better tools alone cannot remove.

The interface-development process is central to the story. The authors ran tasks, inspected trajectories, formed a hypothesis, changed a configuration, and measured again. The same loop can evaluate a viewing window, history handling, or an individual tool. The resulting principles are accessible: simple and compact actions, concise useful feedback, and restrictions that support recovery. Their implementation nevertheless changes over time. The closing discussion connects SWE-agent to SWE-smith, which generates verifiable training tasks, and mini-SWE-agent, which returns to bash and a message history. Stronger models changed the value of the earlier specialized interface. Evaluation evolved too, as illustrated by SWE-bench Verified and its subsequent audit. The engineering lesson is to reassess the complete combination regularly. Measure each component’s contribution on tasks that matter to the team, and be willing to remove yesterday’s improvement when a new model makes it redundant. Historical percentages explain the experiment; they do not guarantee performance for a current system.

Takeaways

What to take away

  1. 01An agent’s interface includes actions and environment responses. Search, line numbers, and updated code after an edit reduce intermediate operations in an unfamiliar repository.
  2. 02Rejecting a bad edit works together with an explanation and both code versions. The restriction also changes the order of edits the agent can carry out.
  3. 03Ablations measure a component’s contribution under specific conditions. Keep full SWE-bench separate from Lite, and one attempt separate from several independent runs.
  4. 04As models improve, a specialized interface may become redundant. Transfer the method of evaluating actions, feedback, and restrictions, and revisit the selected tools regularly.

Sources