
Engineering article reviewResearch Insights Made Simple
Long-running development: who checks the agent?
Planner, generator, evaluator: what pays off, and what goes when the model improves
/ Research Insights Made Simple · Agent harnesses 4/7 · Planner and evaluator
Slide contents
1. Long-running development: who checks the agent?
Planner, generator, evaluator: what pays off, and what goes when the model improves
2. Every harness component encodes a model assumption
3. On long tasks, context and self-assessment fail
4. Resets trade continuity for a clean slate
5. Tuning a skeptic beats tuning self-critique
6. Four criteria turn taste into grades
7. The last iteration is not always best
8. Three agents split scope, build and acceptance
9. “Done” is agreed before any code
10. One criterion below threshold fails the sprint
11. Full harness: over 20× the solo cost
12. The solo build’s game did not work
13. Evaluator findings name a likely cause
14. Untuned, Claude is a poor QA agent
15. Simplify one component at a time
16. Each new model retired some scaffolding
17. The evaluator pays off beyond solo reliability
18. Building took 91% of the DAW budget
19. QA caught display-only features
20. What the evaluator cannot observe stays unchecked
21. The harness space moves rather than shrinks