
Agent anatomy: seven harness responsibilities
What eleven systems' source code reveals — and what it cannot establish
Slide contents
1. Agent anatomy: seven harness responsibilities
What eleven systems' source code reveals — and what it cannot establish
2. The harness turns model output into action
Agent = model + harness; evaluation wraps the agent
3. Eleven systems reveal different engineering choices
Four model-vendor systems, seven open-source, Omnigent above
4. Source code establishes structure, not effectiveness
Seven dimensions on pinned versions
5. Every system addresses all seven responsibilities
Table 1: the simplest and the fullest implementation of each subsystem
6. A small loop already does the work
Query → act → observe → repeat; the rest is operations
7. Completion can have its own check
Aider checks inside the loop, Hermes at the exit
8. Portability costs model-specific maintenance
Five strategies, from one provider to plugins
9. A tool defines an action contract
Exact replacement, patch, fuzzy cascade or shell
10. Beyond 15 tools, load descriptions on demand
From one tool to 109, and search instead of a catalog
11. Session history differs from model context
The log keeps everything; the model sees a projection
12. Working context is only one memory layer
Context compaction converged; memory writing did not
13. Permission and isolation solve different problems
May it run? · What can the process touch?
14. Delegation needs context, authority and verifiable results
Six patterns, from one agent to a cross-process protocol
15. Skills, MCP and hooks complement one another
How to work · what to connect · when to intervene
16. An absent technology is not proven useless
A corpus observation is not a causal experiment
17. A second snapshot can overturn earlier findings
Eight systems pinned in April and again in July 2026
18. The harness turns from tool into platform
SDKs, extensions, marketplaces, and a meta-harness on top
19. Claim strength depends on source and version
Source · version · measurement · interpretation
20. Start with 90 lines; grow on failure
The authors' scaffold → observed failure → new layer → measurement
21. An architecture map helps choose the test
Agent = model + a seven-subsystem harness
Most harness code serves operations, not task solving
Permission and isolation are different mechanisms
Inventories age in weeks; structure holds so far
Source observations are hypotheses for experiments
Copy the responsibility; validate the implementation on your tasks