Skip to content
back to the episode
Episode summary2026Fellow

Long-running App Harnesses

An agent can spend hours writing code and confidently declare an application complete while it still fails. In episode 36, Alexander Polomodov reviews Anthropic’s experience separating creation from evaluation, agreeing on outcomes before implementation, and revisiting a harness when the model improves. The examples reveal the cost of more complete implementation and the limits of browser-based verification.

Research Insights Made Simple #368 min read

A summary based on the audio transcript and episode slides. The reported runs belong to the article’s models and conditions; the host’s conclusions are identified separately.

The main thread of the material
01

Two problems lead to a separate critic

The fourth harness review continues the story of handing work between sessions. Its source is Prithvi Rajasekaran’s March engineering post from Anthropic Labs, “Harness design for long-running application development.” The author combines two projects: improving frontend design and building complete applications from a single request. Designs often looked generic, while applications could remain incomplete. Additional prompting helped without resolving both limitations.

The author distinguishes context problems from self-evaluation. During lengthy tasks, models lose coherence; some start wrapping up because they expect to exhaust their context window. Meanwhile, a creator tends to praise its own output, particularly when judging subjective qualities. Inspiration from generative adversarial networks suggests separating a generator from a critic. This is an analogy for organizing work: pretrained language models exchange artifacts and feedback without training their weights. Tuning a separate evaluator for skeptical review proved easier than making the creator equally demanding of itself.

02

Context handoffs have their own cost

The earlier Sonnet 4.5 harness used an initializer to prepare a feature list, startup instructions, and a progress log. A coding agent implemented one feature per step and left external artifacts for the next session. Resets provided a fresh context but required orchestration, rereading state, and handoff time. Compaction works differently: earlier conversation is summarized, and work continues within the same session.

With Opus 4.5, the author could switch to continuous work with compaction through the Claude Agent SDK. A mechanism’s value therefore depends on the model: something that rescued a weaker worker can later add unnecessary cost. Self-evaluation remained a separate problem, addressed through a dedicated checking role with explicit criteria. Changing context management alone does not establish that the application works.

03

Make taste explicit enough to evaluate

For design, the author defines four criteria: overall design quality, originality, craft, and functionality. Colors, typography, composition, and imagery should communicate a coherent mood. Originality distinguishes deliberate choices from generic templates. Craft includes hierarchy, spacing, and contrast; functionality concerns clear actions. The first two receive greater emphasis because technical execution and interface behavior were already stronger. Several scored examples help the evaluator apply those principles consistently.

The generator produces HTML, CSS, and JavaScript, while the evaluator uses Playwright MCP to inspect the live page and return criticism. The creator can refine the current design or change direction entirely. A Dutch museum site moves from a clean dark landing page at iteration nine to a three-dimensional gallery at iteration ten. The author sometimes preferred an intermediate candidate to the final one, and complexity could grow. The host highlights the stopping decision: a higher score does not settle which candidate should be accepted.

04

Agree on completion before implementation

Full application development adds a planner. It expands a short request into an ambitious product specification and a broad technical design. The host compares this role to a product manager: describe the intended result without dictating every implementation detail. The generator and evaluator resemble a developer and tester who agree on scope and acceptance conditions.

In the Opus 4.5 version, they work in sprints. Before changing code, the generator proposes what to build and how to verify success. The evaluator checks whether that proposal fits the specification. They exchange files until they agree on a contract. After implementation, the evaluator exercises user journeys in a browser, finds discrepancies, and returns bug reports. Both sides know the criteria beforehand, so acceptance depends on an agreed outcome rather than a persuasive completion report.

05

A more complete result needs a larger budget

The retro game maker is the clearest comparison. A direct Claude Code run with Opus 4.5 took roughly twenty minutes and cost nine dollars, but core features did not work. The full harness ran for six hours and cost two hundred dollars; most features in its application worked. The host asks what the extra time and money bought. This compares functionality under different budgets, so the demonstration cannot establish a universal economic advantage.

Traces show the evaluator holding implementation to the sprint contract. It tested behavior through Playwright MCP and reported violations. That level of checking required tuning. According to the author, Claude could otherwise find a real problem, persuade itself that it was unimportant, and approve the work. Inspecting those decisions helps refine quality rules. A separate agent does not automatically become a reliable tester, and minor defects remained even after tuning.

06

Test one component at a time with a new model

After Opus 4.6 arrived, the author initially stripped out much of the harness and added new ideas. He could not reproduce the previous quality, and simultaneous changes made the cause unclear. He then returned to incremental testing: change or remove one component, inspect the outcome, and decide whether it remains useful. The host connects this approach with the ablations discussed in the SWE-agent review.

The updated workflow no longer needed sprints. The generator did most of its work continuously, and an evaluator stepped in after it reported completion. The planner remained. A stronger model needed less help on simple tasks, with evaluation contributing chiefly near its capability limits. This calls for measuring each role or rule again. Earlier usefulness does not guarantee that a mechanism justifies its cost after a model upgrade.

07

Evaluation is limited by available feedback

The next example is a digital audio workstation in a browser. Generation consumes most of the budget, and the evaluator still finds real gaps, although fewer repair cycles are needed. It can inspect the interface, APIs, and database state, but cannot hear playback. An application for making sound therefore has a verification blind spot in its central user outcome.

The host transfers the design lesson to audio: the worker needs access to the result and principles for judging it. Checking a playback button cannot replace listening, just as correct spacing cannot establish expressive design. A harness can support lengthy autonomous work while acceptance remains bounded by what the evaluator can observe. A new task requires checking both the criteria and whether the available feedback channels are sufficient.

08

Simplification does not remove action controls

The author describes a shift in engineering work: a stronger model removes old constraints and opens up new tasks. The host asks what human involvement might become if models could rebuild their own harnesses. This is a question about the future, not a finding from the reported runs. The history helps identify temporary mechanisms and test which limitations they address today.

The epilogue considers external checks on action risk and coordination through an event log. Answering a question about stronger models, Alexander distinguishes execution quality from permission controls: a classifier for dangerous actions limits risk and does not automatically become obsolete as the model improves. In Managed Agents, an event log sits outside the harness so a new process can resume work. Responsibility boundaries may last while the particular arrangement of sessions and roles changes.

Takeaways

What to take away

  1. 01A separate evaluator needs explicit criteria, tuning, and access to the user outcome; dividing roles alone does not ensure quality.
  2. 02Acceptance starts by agreeing on a verifiable outcome before implementation and ends by checking the application’s behavior.
  3. 03After a model upgrade, test each mechanism separately and compare quality alongside time and cost.
  4. 04Compensating for model weaknesses and externally controlling risk serve different purposes; simplifying the former does not remove the latter.

Sources