Skip to content
Research Insights Made Simple
Engineering article reviewResearch Insights Made Simple

Agent evals: telling whether an agent succeeded

What counts as success, who checks it, and when a low score means a broken test

/ Research Insights Made Simple · Agent harnesses 7/7 · Agent evals

Slide contents

  1. 1. Agent evals: telling whether an agent succeeded

    What counts as success, who checks it, and when a low score means a broken test

  2. 2. A “done” reply is not the outcome

  3. 3. We evaluate the model with its harness

  4. 4. Suite, task and trial are different denominators

  5. 5. The trajectory explains; the outcome confirms

  6. 6. Code, models and humans check different things

  7. 7. Scoring rules change what scores mean

  8. 8. Fix failing tests without breaking passing ones

  9. 9. Conversations are graded on three dimensions

  10. 10. A citation does not prove a claim

  11. 11. A confirmation page is not an order

  12. 12. One successful run can hide instability

  13. 13. Track growth and regressions in separate suites

  14. 14. Start with 20–50 tasks from real failures

  15. 15. Test both action and restraint

  16. 16. Every trial starts from a clean environment

  17. 17. One judge per dimension, checked against experts

  18. 18. Low scores can come from broken tests

  19. 19. Read trajectories before trusting the score

  20. 20. No single layer catches every failure

  21. 21. Evals turn “feels worse” into evidence