Long-running Agents: Work Handoffs
Episode participants
A solo episode without invited guests.
What we discussed on the recording
In episode 34, Alexander Polomodov reviews Anthropic’s engineering post “Effective harnesses for long-running agents.” The claude.ai clone example shows why code can survive a session change while intent, unfinished changes, and test results are lost. Context compaction alone does not guarantee that work can resume.
The first session prepares a feature list, progress log, init.sh, and an initial commit. Later sessions read the state, start the application, check its foundation, and implement one feature at a time. Completion is checked through a browser user scenario; Puppeteer MCP’s limitations around native dialogs show how tool blind spots leave bugs.
The review closes by separating the authors’ qualitative experience from a proposed context-reset experiment and examining the follow-up work. Harnesses evolve with models: some restrictions become unnecessary, while verified state and outcome acceptance remain important.
Answering a question about reading papers, the host describes saving a PDF, annotating it on a tablet, and then discussing the text and his objections with a language model. Questions about software factories and open-source projects are deferred to later reviews.