Skip to content
Research Insights Made Simple
Engineering article reviewResearch Insights Made Simple

Long-running development: who checks the agent?

Planner, generator, evaluator: what pays off, and what goes when the model improves

/ Research Insights Made Simple · Agent harnesses 4/7 · Planner and evaluator

Slide contents

  1. 1. Long-running development: who checks the agent?

    Planner, generator, evaluator: what pays off, and what goes when the model improves

  2. 2. Every harness component encodes a model assumption

  3. 3. On long tasks, context and self-assessment fail

  4. 4. Resets trade continuity for a clean slate

  5. 5. Tuning a skeptic beats tuning self-critique

  6. 6. Four criteria turn taste into grades

  7. 7. The last iteration is not always best

  8. 8. Three agents split scope, build and acceptance

  9. 9. “Done” is agreed before any code

  10. 10. One criterion below threshold fails the sprint

  11. 11. Full harness: over 20× the solo cost

  12. 12. The solo build’s game did not work

  13. 13. Evaluator findings name a likely cause

  14. 14. Untuned, Claude is a poor QA agent

  15. 15. Simplify one component at a time

  16. 16. Each new model retired some scaffolding

  17. 17. The evaluator pays off beyond solo reliability

  18. 18. Building took 91% of the DAW budget

  19. 19. QA caught display-only features

  20. 20. What the evaluator cannot observe stays unchecked

  21. 21. The harness space moves rather than shrinks