Skip to content
Research Insights Made Simple
Research reviewResearch Insights Made Simple

Agent harnesses: what actually helps

Planning, tools and context: a controlled comparison

/ Research Insights Made Simple · Agent harnesses 5/7 · Harness design study

Slide contents

  1. 1. Agent harnesses: what actually helps

    Planning, tools and context: a controlled comparison

  2. 2. Harnesses can be tested part by part

    Planning · actions · context around one loop

  3. 3. 176 settings leave some combinations untested

    4 models × 2 benchmarks × 22 settings

  4. 4. The two benchmarks exercise different work

    Repository repair differs from terminal task completion

  5. 5. The plan lives outside conversation history

    update_plan → stored plan → next model input

  6. 6. The whole action interface changes

    Predefined operations ↔ bash commands

  7. 7. Five policies combine three mechanisms

    Elision · recall · summarization

  8. 8. Cheap elision comes before summarization

    T4 preserves the task preamble and recent tail

  9. 9. Smaller windows make context management more valuable

    Success rate by window size, both benchmarks

  10. 10. Most gains come from avoiding overflow

    T0 overflow failures and trajectory length

  11. 11. T4 is the cheapest managed tier

    T4 triggers costly summarization less often

  12. 12. Models rarely recall output behind placeholders

    T2 vs T1: recall added, all else equal

  13. 13. Planning helps the weaker model persist

    Nemotron-3 30B · SWE-bench · T4/128k

  14. 14. Planning trims redundant checks for stronger models

    Nemotron-3 550B and Mistral Medium 3.5 · SWE-bench · T4/128k

  15. 15. Bash helps one model and hurts another

    T4/128k · planning enabled

  16. 16. One model changes preference across benchmarks

    Mistral Medium 3.5 · T4/128k

  17. 17. Each component leaves its own mark

    Trajectory length · stopping point · action size

  18. 18. Behavior annotations need validation too

    LLM judge + a human validation sample

  19. 19. Results depend on models, tasks and implementation

    Four models · two families · one attempt per task

  20. 20. Symptoms point to components

    Our synthesis from the study

  21. 21. A harness must earn its cost

    Context management removes window limits

    Elision before summarization cuts cost

    Planning changes when agents stop

    Interfaces depend on model and workload

    Recheck value after model changes

    Measure the failure, then change the mechanism