Skip to content
back to the episode
Episode summary2026Fellow

Long-running Agents: Work Handoffs

An agent can preserve its code and still lose the ability to resume confidently. In episode 34, Alexander Polomodov examines what must survive a context reset: requirements, an account of progress, a verified state, and the next step. Anthropic’s experience shows how small, verifiable changes support long-running work and why that workflow must evolve with the model.

Research Insights Made Simple #348 min read

A summary based on the audio transcript and episode slides. The reviewed post describes engineering experience; proposed experiments are distinguished from the authors’ results.

The main thread of the material
01

Code survives; an understanding of the work may not

Alexander Polomodov continues the discussion of coding agents after the SWE-agent episode. That review considered an agent solving an individual issue through a useful interface. This episode asks how an agent can build an application over several sessions. Its source is Anthropic’s November 2025 engineering post “Effective harnesses for long-running agents.” The post describes a team’s experience rather than a scientific study measuring each component’s contribution. The example is a web application reproducing the claude.ai interface.

When a session ends, code remains on disk, but the next worker may not know the intent, unfinished changes, or test results. The host compares this with an engineer arriving at work and checking assigned tasks: a verified starting point, current status, and next priority are needed. Conversation compaction keeps context from overflowing, yet a short account may omit instructions needed to resume. A statement that work was completed does not establish that the application works.

02

Two roles and four external artifacts

Anthropic makes the first session special. An initializer prepares the environment, and a coding agent handles subsequent sessions. Their harnesses are similar; the principal difference is the prompt. The initializer creates a feature list, a progress log, an init.sh startup script, and an initial commit. These artifacts live outside the model’s context and survive a reset. The design concerns sequential work over time: having two roles does not itself imply multiple agents working in parallel.

Each artifact answers a different question. The feature list records what must work and what remains unfinished. The log explains the previous session’s activity and stopping point. The startup script avoids rediscovering how to run the application. Git preserves changes and provides a way back to a verified state. File names can vary across approaches, but the responsibilities remain distinct: requirements, intent, environment setup, and code history cannot substitute for one another.

03

Completion means passing an observable scenario

Without this structure, an agent may tackle chat, themes, conversation loading, and error handling at once. Its context ends with several features unfinished. The next session must first restore a working application and infer what its predecessor intended. Another failure is premature completion: an agent sees a partially working interface and declares the whole project finished. When requirements exist only in the conversation, the agent can infer completion from whatever already exists.

Instead of “build chat,” a feature should describe a user journey: open the main screen, press New Chat, see a new conversation and greeting, and verify that the conversation appears in the sidebar. This serves as both a requirement and an acceptance scenario. The coding agent may change the passing flag after completing all steps, but may not remove requirements or rewrite their meaning. In the reviewed design, this restriction is a prompt instruction; a separate external check protecting requirements is not demonstrated.

The authors report better results with structured JSON than with a Markdown checklist, without quantifying the difference or explaining its cause. The host suggests that a separate status field makes a diff easier to inspect than a line mixing requirement text and a completion mark. That is a hypothesis, not an established property of JSON. An external change check could strengthen the rule, but it should not be credited to the original implementation.

04

One feature per step, checked as a user

A coding session starts by locating its working directory and reading the progress log, feature list, and Git history. The agent then starts the application and checks its foundation with a basic end-to-end scenario, such as creating a chat and receiving a reply. Only then does it select a feature, implement it, and get its check to pass. It commits the code and updates the log. Several iterations can fit within one session: the rule limits each step, not the number of tasks in a session.

For a web application, passing unit tests and a successful curl request to an API are insufficient. The agent needs to exercise the interface as a user. The article uses Puppeteer MCP, but native alert dialogs were poorly exposed by the tool in that setup, leaving more bugs in features that relied on them. This connects back to SWE-agent: results depend on the actions an agent can take and the consequences it can observe. Browser access alone does not establish coverage of all user behavior.

05

Give the next session verifiable reference points

Git history becomes a sequence of states: project setup, new chat, message sending, and theme switching. If a later change breaks the application, the agent can return to a previously verified state. The progress log adds an explanation of the work. A useful entry records accomplishments, completed tests, unresolved problems, next steps, and the number of passing scenarios. A commit reference ties the account to a specific version of the code.

The report helps another session resume, but remains an agent’s claim. The next session therefore checks the foundation, and feature status changes follow scenario execution. Preparation must be paired with repeated behavior: the initializer provides the starting structure; the coding agent reads it on entry and updates it on exit. A log that nobody reads, verifies, or maintains cannot solve state loss. Handing over work belongs inside each completed step.

06

Revisit the harness when the model changes

The engineering post offers a mechanism, an example application, a session trajectory, and quickstart code. Its improvement claims are qualitative. The host proposes a test: force context resets, change one mechanism at a time, and compare performance on the same tasks, model, and budget. This would help establish whether state handoffs contribute to the outcome. It is a proposed experiment from the episode, not one performed in the original article.

The later history changes the perspective. The host mentions the compiler project in which several agents coordinated through Git, then a follow-up on long-running application development. That follow-up clarified that the early harness had been built around Sonnet 4.5. With Opus 4.5, premature stopping became weak enough in the reported experience to replace explicit restarts with a continuous session using compaction. Separate browser-based acceptance also emerged. This does not eliminate the value of external state; it changes the necessary combination of mechanisms. When the model changes, earlier constraints should be retested while preserving a way to establish what actually works.

Takeaways

What to take away

  1. 01A handoff needs external requirements, progress notes, a startup procedure, and code history; a conversation summary alone is insufficient.
  2. 02Feature completion should follow a user-scenario check, and each new session should verify its starting state.
  3. 03Tool blind spots constrain autonomy: an inaccessible browser dialog can leave a bug unnoticed.
  4. 04A harness’s usefulness depends on the model; qualitative engineering observations need testing on your own tasks.

Sources