
Agent harnesses: what actually helps
Planning, tools and context: a controlled comparison
Slide contents
1. Agent harnesses: what actually helps
Planning, tools and context: a controlled comparison
2. Harnesses can be tested part by part
Planning · actions · context around one loop
3. 176 settings leave some combinations untested
4 models × 2 benchmarks × 22 settings
4. The two benchmarks exercise different work
Repository repair differs from terminal task completion
5. The plan lives outside conversation history
update_plan → stored plan → next model input
6. The whole action interface changes
Predefined operations ↔ bash commands
7. Five policies combine three mechanisms
Elision · recall · summarization
8. Cheap elision comes before summarization
T4 preserves the task preamble and recent tail
9. Smaller windows make context management more valuable
Success rate by window size, both benchmarks
10. Most gains come from avoiding overflow
T0 overflow failures and trajectory length
11. T4 is the cheapest managed tier
T4 triggers costly summarization less often
12. Models rarely recall output behind placeholders
T2 vs T1: recall added, all else equal
13. Planning helps the weaker model persist
Nemotron-3 30B · SWE-bench · T4/128k
14. Planning trims redundant checks for stronger models
Nemotron-3 550B and Mistral Medium 3.5 · SWE-bench · T4/128k
15. Bash helps one model and hurts another
T4/128k · planning enabled
16. One model changes preference across benchmarks
Mistral Medium 3.5 · T4/128k
17. Each component leaves its own mark
Trajectory length · stopping point · action size
18. Behavior annotations need validation too
LLM judge + a human validation sample
19. Results depend on models, tasks and implementation
Four models · two families · one attempt per task
20. Symptoms point to components
Our synthesis from the study
21. A harness must earn its cost
Context management removes window limits
Elision before summarization cuts cost
Planning changes when agents stop
Interfaces depend on model and workload
Recheck value after model changes
Measure the failure, then change the mechanism