Skip to content
back to the archive page
#AI4SDLC

Anthropic’s agent harness: handing work over between coding sessions (Series #AI4SDLC)

For Research Insights Made Simple #34, I’m exploring “Effective harnesses for long-running agents,” an engineering post by Justin Young at Anthropic from almost a year ago, November 26, 2025. One question caught my attention: what will remain of an agent’s work when its context runs out and the next session has to continue the project?

In this article, Anthropic describes its experience building a claude.ai clone. The agent tried to do too much in one go, left a feature unfinished without a clear description, and the next instance spent time reconstructing what had happened. Or it saw code already written and cheerfully declared the entire project finished. Compacting the conversation history alone did not solve these problems.

It becomes development in shifts, with each new engineer arriving with no memory of the previous one. Leaving “I did some work here, it’s almost ready” is not much of a handover :)

The authors build a protocol around several things:

— An initializer prepares the environment, a startup script, and a feature list with verification criteria. All features start out marked as unfinished. — Subsequent sessions take one feature at a time, check the result, save changes in Git, and update a progress log. — At the start of a new session, the agent reads the log and commit history, checks that the app works, and chooses the next task. — For a web app, verification includes browser scenarios. Code can look convincing until the user clicks a button.

Incidentally, the “two agents” here are two phases with different initial prompts. A footnote clarifies that the system prompt, tools, and harness are the same. This provides no evidence for the benefits of a swarm. Another curious detail: the authors observed that the agent was less likely to make unwanted changes to a requirements list in JSON than in Markdown. There are no numbers or comparison conditions, so I would treat it as a hypothesis to test. JSON itself imposes no prohibition: if preserving requirements matters, changes need separate checks. Asking “don’t change the tests” does not create a technical constraint.

What interests me here is the difference between a saved account of the work and a verified project state. The log helps explain intent, Git shows changes, the feature list recalls commitments, and tests verify the result. Each artifact has its own role. A passes: true checkbox that the agent ticked itself remains its opinion without reliable verification.

I would not call this post a scientific breakthrough: it describes a vendor team’s web development experience, without a control group or quantitative comparisons. It does show well how much ordinary engineering discipline is needed for an agent to continue a large task.

The recipe changes with the model too. In a March 2026 follow-up, Anthropic described a separate evaluator and abandoning forced context resets when moving from Sonnet 4.5 to Opus 4.5. It is useful to examine which specific weakness each harness component compensates for, and whether it is needed in the next configuration.

Tomorrow, October 7, we’ll discuss this article in more detail on the Research Insights #34 livestream: how to hand work over between sessions, what to trust in an agent’s report, and when separating roles actually helps. Come along; there is plenty to discuss.

#AI4SDLC #AI #Agents #Engineering #Research

Files from the post

  • Annotated-Effective-harnesses-for-long-running-agents.pdfPDF · 2 MB
    Download PDF

Open video on YouTube