Skip to content
all episodes
Research Insights Made Simple · episode 34

Long-running Agents: Work Handoffs

41:46

Episode participants

A solo episode without invited guests.

Conversation

What we discussed on the recording

In episode 34, Alexander Polomodov reviews Anthropic’s engineering post “Effective harnesses for long-running agents.” The claude.ai clone example shows why code can survive a session change while intent, unfinished changes, and test results are lost. Context compaction alone does not guarantee that work can resume.

The first session prepares a feature list, progress log, init.sh, and an initial commit. Later sessions read the state, start the application, check its foundation, and implement one feature at a time. Completion is checked through a browser user scenario; Puppeteer MCP’s limitations around native dialogs show how tool blind spots leave bugs.

The review closes by separating the authors’ qualitative experience from a proposed context-reset experiment and examining the follow-up work. Harnesses evolve with models: some restrictions become unnecessary, while verified state and outcome acceptance remain important.

Answering a question about reading papers, the host describes saving a PDF, annotating it on a tablet, and then discussing the text and his objections with a language model. Questions about software factories and open-source projects are deferred to later reviews.

AI in SDLCDeveloper productivityPlatform engineeringResearch methodology