
Engineering article reviewResearch Insights Made Simple
Agent evals: telling whether an agent succeeded
What counts as success, who checks it, and when a low score means a broken test
/ Research Insights Made Simple · Agent harnesses 7/7 · Agent evals
Slide contents
1. Agent evals: telling whether an agent succeeded
What counts as success, who checks it, and when a low score means a broken test
2. A “done” reply is not the outcome
3. We evaluate the model with its harness
4. Suite, task and trial are different denominators
5. The trajectory explains; the outcome confirms
6. Code, models and humans check different things
7. Scoring rules change what scores mean
8. Fix failing tests without breaking passing ones
9. Conversations are graded on three dimensions
10. A citation does not prove a claim
11. A confirmation page is not an order
12. One successful run can hide instability
13. Track growth and regressions in separate suites
14. Start with 20–50 tasks from real failures
15. Test both action and restraint
16. Every trial starts from a clean environment
17. One judge per dimension, checked against experts
18. Low scores can come from broken tests
19. Read trajectories before trusting the score
20. No single layer catches every failure
21. Evals turn “feels worse” into evidence