Skip to content
all longreads
Longread#AI4SDLC#Architecture

The Agent Harness: Past, Present, and Future of the Layer Between the Model and the World — Seven Parts, Four Eras, Eight Bets

In May 2026 the authors of a Google report estimated that the experience of working with a coding agent is ten percent the model and ninety percent the harness around it. Two months later the creator of Claude Code said that for the new model his team had deleted more than eighty percent of the system prompt, and quality did not drop. Both statements are made with confidence, and both are true: they are about different parts of the same harness. Fewer instructions, more walls around them, and more freedom inside the walls. This article is the written version of my Giga Conf 2026 talk on the past, present and future of harnesses: what a harness is, which parts it is made of, how it has changed over four years and where it is going.

1 October 2026≈ 46 minprimary sources ↓

Product documentation, engineering blogs, research and talk recordings were checked against primary sources on 18 September 2026. Every claim is assigned to one of five classes — verifiable fact, product or market snapshot, author's forecast, management heuristic, author's conclusion — and the class is named in the text or in the section's evidence boundary. Companies' statements about their own products are presented as claims. OpenAI's 29 September conference is not part of this snapshot; its announcements will be added in a separate revision after the talk.

01

What a harness is and which seven parts it consists of

I already have a definition, and I am not going to rewrite it. In the September article on what remains engineering I called the agent harness the execution environment: the working loop, tools, state, constraints and feedback. The Claude Code documentation says almost the same thing more briefly: a harness is the tools, context management and execution environment that turn a language model into a working agent; Claude Code is the harness, Claude is the model inside. Birgitta Böckeler on Martin Fowler's site is shorter still: everything in the agent except the model. The three definitions do not contradict one another, but they do not help you design anything either. Design needs a decomposition.

First, the term's neighbors, because in 2026 they get confused even in glossaries. A framework is a library the harness is assembled from; it does not run the agent and does not own its boundaries. A runtime is where the harness runs at a vendor. Scaffold is a word from the 2024 era; in Anthropic's 2025 documents it meant the whole wrapper around the model, while in the Hugging Face glossary of May 2026 it means only the behavior layer, and the execution layer is what they call the harness. There is no single definition, and the glossary's authors say so honestly. Finally, the word has a second meaning that usually gets lost: the training environment. In July 2026 Prime Intellect split the training environment into exactly two parts — a task set and a harness, the program that solves those tasks and produces trajectories; one configuration serves both evaluation and training. It is the same object seen from the lab's side, and it will return later in the article.

Term
Framework
What it is
A library for assembling a harness: LangChain, ADK, Agents SDK, Deep Agents
What it is not
Does not run the agent itself and does not own its boundaries
Term
Scaffold
What it is
The 2024-era wrapper: chains, graphs, roles; in Anthropic's 2025 usage — the whole wrapper around the model
What it is not
In the 2026 Hugging Face glossary it is the opposite, only the behavior layer — there is no single definition
Term
Harness
What it is
The model's protocol for working with the world: loop, context, tools and environment, verification, orchestration, observability, boundaries
What it is not
Not the model, not the task, and not the framework it was assembled from
Term
Runtime
What it is
Where a harness runs at a vendor: Managed Agents, Foundry Hosted Agents, Managed Agents in the Gemini API
What it is not
Does not replace your task, context, and success criteria
Term
Agent
What it is
Model plus harness plus task — the thing you actually run
What it is not
Not a property of the model: one model in different harnesses yields different results
Term
Training environment
What it is
The same notion from the lab's side: a task set plus the harness the model is trained in
What it is not
Does not guarantee that the deployed harness matches the training one

Now the decomposition. I split the harness into seven parts, and each part answers a question the model asks the world. What can I see — context and memory. What can I do — tools and environment. How long, and when to stop — the loop. Who checks — verification. How to scale — orchestration. Who sees what happened — observability and evals. Who permits — trust and boundaries. Historically the harness answered the first three questions with words: instructions in the prompt, tool descriptions, skills. The last four it answered with walls: code that runs regardless of what the model decided. What walls are not matters just as much: they do not lead the model by the hand at every step, they fence the field — what it may touch, where it may go, who will see. This split into words and walls is the spine of the whole article, and the rest of it shows three movements at once: words shrink, walls grow, and there is more freedom inside the walls.

figure 02 · seven harness parts: words and walls around the model
Seven parts of the harness: words and walls around the modelSEVEN PARTS OF THE HARNESS: WORDS AND WALLSMODELweights and trained behaviorLOOP AND STOPPINGhow long, and when to stopCONTEXT AND MEMORYwhat I can seeTOOLS AND ENVIRONMENTwhat I can doOBSERVABILITYwho sees what happenedTRUST AND BOUNDARIESwho permitsORCHESTRATIONhow to scaleVERIFICATIONwho checks● words: instructions, shrinking● walls: boundaries and environment, growingSeven questions the harness answers for the model: three with words, four with walls, more freedom inside

The site's corpus already has two decompositions of the same entity, and a third one without a reconciliation would look like inconsistency. In the July longread on the co-evolving stack the harness is one of six layers, between the model and the tools, and it is responsible for context, planning, compaction, state, isolation and confirmations. The same article audited six months of history of three open harnesses and split them into ten mechanisms. The seven parts here are not a replacement but a different cut: not “what lives in the repository” but “which question of the model's does this answer”. The table's last column ties the cuts together: each part corresponds to one to three mechanisms from the July audit.

Part
Loop and stopping
The model's question
How long, and when to stop
What the model took over by 2026
Reasoning between tool calls, running tests unprompted, tracking the remaining context
What stays in the harness
Stopping criteria, budgets, checkpoints, a judge on another model
Audit mechanisms
planning · state
Part
Context and memory
The model's question
What I can see
What the model took over by 2026
A million-token window, the claimed “trained” compaction
What stays in the harness
The compaction policy, a repository map instead of instructions, memory across sessions
Audit mechanisms
context · compaction
Part
Tools and environment
The model's question
What I can do
What the model took over by 2026
Function calling, screen use, choosing the procedure without a script
What stays in the harness
Tool definitions, the sandbox, worktrees, the MCP protocol
Audit mechanisms
editing · discovery · sandbox
Part
Verification
The model's question
Who checks
What the model took over by 2026
Self-checking — claimed by vendors, not measured independently
What stays in the harness
Tests and linters as the feedback channel, hooks, a fresh model as judge, replayable episodes
Audit mechanisms
telemetry and evals
Part
Orchestration
The model's question
How to scale
What the model took over by 2026
Spawning subagents in swarm-trained models
What stays in the harness
Durable execution, agent teams, schedules, state ownership
Audit mechanisms
orchestration
Part
Observability and evals
The model's question
Who sees what happened
What the model took over by 2026
Nothing
What stays in the harness
Traces, the session log, cost accounting, evaluation on episodes
Audit mechanisms
telemetry and evals
Part
Trust and boundaries
The model's question
Who permits
What the model took over by 2026
Nothing — deliberately
What stays in the harness
Permissions and an action classifier, agent identity, audit
Audit mechanisms
policy · sandbox
02

Four eras: from a bash loop to a managed runtime

The word is older than agents. In testing, a harness is the set of stubs and drivers that let you exercise a component outside its real surroundings. It came into the world of language models by the same road: EleutherAI's lm-evaluation-harness is an evaluation runner, not a wrapper for actions. As late as August 2024, OpenAI's description of SWE-bench Verified called the wrapper around the model a scaffold and the harness the containerized test runner. In January 2025 Anthropic defined an agent as “a combination of a model and the software scaffolding around it”, and in the Claude 4 and Sonnet 4.5 announcements the SWE-bench results came with a note about “a simple scaffold with two tools — bash and file editing”. The shift from scaffold to harness happened when the wrapper became a product: in September 2025 Anthropic described the Agent SDK as “the harness that Claude Code runs on”, and by 2026 METR had replaced scaffold with agent harness in its reports. The discipline arrived in February 2026: Mitchell Hashimoto wrote that he did not know an accepted term and started calling his practice harness engineering, and six days later OpenAI's post of the same name came out.

figure 01 · where the word came from: four meanings on one line
Where the word harness came from: four senses on one timelineONE WORD, FOUR SENSESpre-LLM202420252026TEST HARNESSstubs and drivers for teststhe pre-LLM eraEVAL HARNESSlm-evaluation-harness, EleutherAIthe evaluation runnerSCAFFOLDSWE-agent ACI · May 2024Anthropic: model + scaffolding · Jan 2025HARNESS AS A PRODUCTClaude Code · Feb–May 2025Agent SDK · Sep 2025A DISCIPLINEHashimoto Feb 5 · OpenAI Feb 11, 2026LangChain · Anthropic · Böckeler · Mar–Apr 2026THE TRAINING SENSERL environment = taskset + harnessverifiers v1 · Jul 2026One word, four senses: test runner → scaffold around the model → product → discipline and training environment

Behind the word stand four eras, and in each the harness compensated for a specific weakness of the model. In 2022–2023 the model could not act: no tools, no memory. The think–call–read loop was hand-written around the API, and the first harness function moved into the API as early as June 2023, when OpenAI added function calling to gpt-4-0613. In 2024 the model could act but got lost in a repository and on a long task — this is the era of scaffolds: SWE-agent's agent–computer interface, the role graphs of LangGraph and CrewAI, the planner and the critic, Claude's screen control in October and the MCP context protocol in November. In 2025 the model was already writing files and whole features while the products around it offered autocomplete and chat: Boris Cherny calls this product overhang, and Claude Code grew out of the observation that Sonnet 3.5 could do more than it was allowed to. One general-purpose agent in the file space became a product — Claude Code, Codex, Gemini CLI — and within a year it was surrounded by a repository instruction file, a harness SDK, a skills standard and a standards foundation under the Linux Foundation.

2026 is the era of discipline and infrastructure. The harness is described as an engineering practice with its own posts, courses and exams, and at the same time it is rented out whole: in April Anthropic released Managed Agents and called it a “meta-harness”, in May Google showed Managed Agents in the Gemini API, in June Microsoft — Foundry Hosted Agents. In July the MCP protocol became stateless, in August GitHub released a plugin standard, and Antigravity replaced Gemini CLI on the free tier. The model weakness that the 2026 harness compensates for sounds paradoxical: the model does more than the product allows, and the harness itself goes stale faster than it is written. Konstantin Krestnikov's periodization from Data Fest — from bare language models through ReAct and chains to graphs and finally to one agent with files — matches mine up to the era boundaries, and his metaphor is exact: the model is the pulling power, the files are the field, the harness is the yoke.

Era
2022–2023 · text in, text out
Model weakness
The model cannot act: no tools, no memory
What the harness did
The think–call–read loop was hand-written around the API
Artifact
The ReAct loop, the first autonomous scripts, function calling in the API
Representatives
AutoGPT, BabyAGI, function calling (June 2023)
Era
2024 · scaffolds
Model weakness
The model acts but gets lost in a repository and on long tasks
What the harness did
An agent–computer interface, role graphs, a planner and a critic
Artifact
The ACI, LangGraph and CrewAI graphs, screen control, the context protocol
Representatives
SWE-agent, Devin, computer use (October 2024), MCP (November 2024)
Era
2025 · harness as a product
Model weakness
The model already writes files and features while products around it offer autocomplete
What the harness did
One general agent in the file space; loop, tools and file system in the box
Artifact
Terminal agents, the repository instruction file, a harness SDK, skills, a standards foundation
Representatives
Claude Code, Codex, Gemini CLI, AGENTS.md, Agent SDK, Skills, AAIF
Era
2026 · discipline and infrastructure
Model weakness
The model can do more than the product allows; the harness goes stale faster than it is written
What the harness did
The harness is described as an engineering discipline and rented out whole
Artifact
The “harness engineering” posts, managed runtimes, a stateless protocol, a plugin standard
Representatives
OpenAI (February), Anthropic (March–April), Google and Microsoft (May–June), MCP 2026-07-28
03

The mechanism: how the model and the harness evolve together

Anthropic stated the main law of this evolution in March 2026 in a single sentence: every harness component encodes an assumption about what the model cannot do on its own. Everything else follows from that. Once the assumption stops being true, the component becomes not useless but harmful: it constrains the model to what the model has already outgrown. In July I called the harness a temporary theory of the current model's weaknesses; now we can trace which of those theories have already been refuted and which the labs keep deliberately.

Here is the dated line of what moved into the model. Function calling — June 2023. Grammar-constrained structured output — August 2024: the JSON mode of a year earlier promised only valid syntax, not schema conformance, and in Simon Willison's experience even the syntax was not guaranteed. Screen use — October 2024, when Anthropic wrote outright that it was teaching the model general computer skills instead of building tools for each task. Reasoning between tool calls — in Claude 4 it was switched on with a beta header, in 4.6 it switches on by itself, in 4.7 the manual mode is rejected: a harness flag dissolved into the model's behavior. Tracking the remaining context — Sonnet 4.5 receives budget tags from the API, Opus 4.7 and newer no longer do. Compacting the history — in November 2025 OpenAI called GPT-5.1-Codex-Max “the first model trained to work across multiple context windows”; that is a claim, not a measurement, and Simon Willison promptly reminded everyone that Claude Code had been doing the same thing at the harness level. The most telling line is in the Claude Code changelog: the task-tracking tools are now offered only to models older than Opus 4.7. A component that was mandatory for a year has been removed for newer models because they keep the plan in their heads.

figure 03 · what the model took over and what the harness kept, 2023–2026
What the model absorbed and what the harness kept, 2023 to 2026ABSORBED BY THE MODEL2023202420252026function callingJun 2023structured outputsAug 2024computer useOct 2024interleaved thinkingMay 2025 → defaultcompaction: claimedNov 2025task tracker retired2026KEPT IN THE HARNESSpermissions enforcedby the harness, not the modelhooks33 eventscompaction policystays with the developerOS sandbox−84 % promptsidentityand auditThe July longread's figure 05, now dated: what can be trained moves into weights, what cannot be trusted stays a wall

Now what the labs keep in the harness deliberately, and here the wording matters more than the dates. Claude Code documentation: permission rules are enforced by Claude Code, not by the model; instructions in the prompt determine what the model tries to do but do not change what it is allowed to do. Also there: CLAUDE.md is not a hard enforcement layer, and to block an action regardless of the model's decision you need a hook before the tool call. An operating-system-level sandbox, by Anthropic's data, cut confirmation prompts by eighty-four percent and covers any subprocesses — not because the model is unreliable, but because the boundary has to hold even if it were. Codex describes two layers outside the model — the sandbox mode and the approval policy — and explains it directly: you trust not the agent's intentions but the fact that it acts inside enforced boundaries. The compaction policy also stays with the developer: Anthropic's server-side compaction has a default threshold, but the summary prompt is replaceable and “model-dependent”. The pattern is visible to the naked eye: what can be trained goes into the weights; what cannot be entrusted even to a well-trained model stays in the harness.

The training environment and the deployed harness converge

The word's second meaning from the first section becomes a mechanism here. In October 2025 Cursor wrote that in reinforcement learning the Composer model “can call any tool of the Cursor Agent harness”, and in the Composer 2 report in March 2026 — that training runs “in realistic Cursor sessions with the same tools and harness as the deployed model”; model checkpoints ship roughly every five hours, and the user has become part of the training environment. Cognition, in the SWE-1.5 announcement, talks about retraining the model on the updated Cascade harness. OpenAI trained GPT-5-Codex “on real engineering work” and recommends using it “only in Codex or Codex-like environments” — the training harness is not named, but the recommendation speaks for itself. Zhipu assembled more than ten thousand environments for GLM-5 in the Harbor format — that is, in the format of the Terminal-Bench benchmark. The minimal preset of DeepSeek's open harness is described in the repository as “RL-agent composition”: a short fixed prompt, a persistent bash and a line editor reproduce the training environment. Anthropic makes no direct statement and speaks only of “the same infrastructure as Claude Code” — and the absence of a statement is not the absence of a practice.

If the model learns in its own harness, a foreign harness at the frontier should lose — and should lose by a little, because frontier models are trained to transfer skill across wrappers. That is exactly what the measurements show. In June 2026 Harbor-Index pitted the models' native harnesses against the neutral Terminus-2 on the same tasks: the native one is “usually slightly ahead, but no comparison is statistically significant”. In February METR wrote that Claude Code and Codex are “not obviously better” than the lab's default scaffolds: Opus 4.5 in Claude Code beats a simple ReAct in only half of the bootstrap samples, GPT-5 in Codex beats Triframe in one of seven. And at the same time, below the frontier the spread is enormous. On Claw-SWE-Bench GLM 5.1 solves nineteen percent of tasks with a minimal adapter and seventy-three with the full one; averaged across models, the choice of harness shifts the result by twenty-seven points. The GigaChain team sees a twenty-to-thirty-point spread between harnesses on its own measurements with one and the same model.

Study
Claw-SWE-Bench · June 2026
Model and benchmark
GLM 5.1 · Claw-SWE-Bench
Effect
Minimal adapter 19.1% vs full 73.4%; average effect 27.4 pp
Who published it
Independent paper (arXiv)
Class
Verifiable fact
Study
GigaChain · August 2026
Model and benchmark
Sonnet · 104 contest tasks
Effect
Claude Code 70 tasks, DeepAgents 67; 20–30 pp spread across harnesses
Who published it
Sber, a Data Fest talk
Class
Vendor claim
Study
Model profile · August 2026
Model and benchmark
GigaChat 3.5 · 391 tasks
Effect
77.2 → 87.0% success, tokens −50%
Who published it
Sber on Habr
Class
Vendor claim
Study
AHE · April 2026
Model and benchmark
GPT-5.4 · Terminal-Bench 2.0
Effect
69.7% (seed) → 71.9% (Codex CLI) → 77.0% (evolved)
Who published it
Independent paper (arXiv)
Class
Verifiable fact
Study
icat-agent · June 2026
Model and benchmark
GPT-5.4-xhigh · SWE-bench Pro
Effect
59.1 → 67.4%
Who published it
Independent paper (arXiv)
Class
Verifiable fact
Study
Vercel d0 · December 2025
Model and benchmark
Model not named · 5 internal queries
Effect
Tools −80%: 4 of 5 → 5 of 5, 3.5× faster, tokens −37%
Who published it
Vercel
Class
Vendor claim
Study
Meta-Harness · March 2026
Model and benchmark
Opus 4.6 and Haiku 4.5 · Terminal-Bench 2.0
Effect
74.7 → 76.4% and 33.7 → 37.6% under a “up to 6×” headline
Who published it
Stanford (arXiv)
Class
Verifiable fact
Study
Google whitepaper · May 2026
Model and benchmark
Model not named · Terminal-Bench 2.0
Effect
From outside the top 30 into the top 5 by changing the harness; LangChain +13.7 points
Who published it
Google, the document's authors
Class
Snapshot
Study
Nvidia · August 2026
Model and benchmark
Opus 5 · ARC-AGI-3
Effect
30% without a harness → 100% with Nvidia's harness
Who published it
TechCrunch on Nvidia's work; ARC Prize verification not mentioned
Class
Snapshot
Study
Harbor-Index · June 2026
Model and benchmark
Several models · Terminal-Bench
Effect
Native harness vs neutral Terminus-2 — no statistically significant difference
Who published it
Terminal-Bench
Class
Verifiable fact
Study
METR · February 2026
Model and benchmark
Opus 4.5, GPT-5 · METR tasks
Effect
Claude Code beats ReAct in 50.7% of samples, Codex beats Triframe in 14.5% — “not obviously better”
Who published it
METR
Class
Verifiable fact

Harnesses that rewrite themselves

The logical continuation of co-evolution is a harness changed not by an engineer but by another agent. The idea is older than it looks: in August 2024 Hu, Lu and Clune described a meta-agent that iteratively programs new agents from an archive of earlier discoveries. Sakana's Darwin Gödel Machine in May 2025 raised its SWE-bench score from twenty to fifty percent by rewriting its own code — and in the same paper the authors report that the machine removed the markers the reward function used to catch hallucinations, against an explicit instruction not to. In 2026 the direction became academic routine: Meta-Harness from Stanford, Agentic Harness Engineering with GPT-5.4 going from seventy to seventy-seven percent on Terminal-Bench, DemoEvolve with the honest conclusion that self-evolution on a sparse reward is “brittle”. In June Anthropic released “a harness for every task” — Claude writes the orchestration script on the fly. All of this works, and all of it optimizes exactly the reward it was given. An ACL 2026 paper adds an unpleasant touch: experience accumulated on harmless tasks still undermines safety in risky scenarios, because it strengthens the tendency to act rather than to refuse. A self-modifying harness without a hidden judge and a ledger of assumptions is Goodhart's law handed a machine gun.

04

What shrinks: instructions turn into maps

Three 2026 figures describe one movement. Boris Cherny in a July interview: “we deleted eighty percent of the system prompt — Opus 5 just does it”; there is an environment variable that removes all prompts altogether, and in that bare mode the model turns out even slightly better on the team's measurements. Nick Nisi of WorkOS at AI Engineer Europe: they generated ten thousand seven hundred lines of skill instructions from documentation, after deleting ninety-five percent the result got better, and flow control moved out of the skill into a TypeScript state machine. Vercel in December 2025: they removed eighty percent of the tools from an internal agent, leaving a single bash — three and a half times faster, a third fewer tokens, and success on the internal set went from four of five queries to five of five. All three are self-reports, and Vercel's “one hundred percent” is five queries. But the direction at three independent teams is the same, and Vercel states it most sharply: the model makes better choices when we stop making choices for it.

The words do not disappear — they change genre. In OpenAI's post on harness engineering the AGENTS.md file is about a hundred lines and serves as “a map with pointers to documentation” in a structured catalog: the agent reads what it needs now, not everything at once, because, as the authors write, “anything the agent cannot reach in context does not exist for it”. The skills standard is built the same way: a hundred tokens of metadata are always in context, instructions of up to five thousand tokens load on demand, resources as needed; the standard's showcase lists forty-six clients. A study of one hundred twenty-four PRs from ten repositories links the presence of AGENTS.md to a nearly thirty percent reduction in median time and sixteen percent in output tokens at comparable task completion — a correlation, not an experiment, but it agrees with Vercel's internal evals, where an eight-kilobyte documentation index in AGENTS.md scored one hundred percent against seventy-nine for skills. The map beats the script because the script encodes an assumption about how the model should go, and the map only about where things are.

figure 04 · from words to walls, 2024 → 2026: an illustration, not a measurement
From words to walls, 2024 to 2026FROM WORDS TO WALLS, 2024 → 2026WORDSWALLSWORDSWALLS20242026−80 % of the system prompt(Claude Code, Opus 5)−95 % of skills (WorkOS)−80 % of tools (Vercel d0)33 hook eventsOS sandbox: −84 % promptsAn illustration, not a measurement: bar heights are notional, the numbers come from different products and measurements

The loop: from a bash one-liner to a judge on another model

The loop is the second harness part that historically lived in words, and its evolution over the year is instructive. In July 2025 Geoffrey Huntley described Ralph: “in its pure form it is a bash loop” that feeds the agent the same prompt over and over. A year later at Anthropic it is an official plugin, implemented through a session-stop hook with an iteration limit, and next to it are built-in commands: repetition on a schedule or self-paced with a one-week lifetime, and a goal whose completion is checked by “a fresh model, not the one doing the work”. Every turn creates a checkpoint, the last hundred are kept, and the documentation honestly warns that changes made by bash commands are not tracked. The Agent SDK sets hard limits on the number of turns and on the budget in dollars, and the budget covers subagents. Kyle Mistele of HumanLayer suggests looking at all of this as a textbook control loop: the setpoint is the desired property of the codebase, the sensor is a deterministic check, the controller picks the next small change, the actuator is the agent, and a model is far from necessary in every link. The direction is the same as for instructions: from “repeat until it works” to “stop when an independent check said it worked”.

Context: a million tokens, and compaction all the same

In 2026 context is a finite resource even with a seemingly infinite window. Anthropic's models from 4.6 onward got a million-token window by default and at the standard price, and Claude Code still compacts the history at ninety-seven percent of the window: first it clears old tool outputs, then it summarizes, then it re-reads up to five changed files. The reason is named in the documentation itself with the words context rot — Chroma showed on eighteen models that quality declines as input length grows, even on simple tasks. Back in September 2025 Anthropic described three techniques for tasks longer than the window: compaction, structured notes in external memory, subagents with condensed summaries — and within a year all three became API or product features. The economics of the prefix cache and the warning about memory that promotes any past answer to a rule I covered in the July and September articles; here one thing matters: even where the harness compresses instructions, context management remains its job, because the model cannot see its own window from the outside.

05

What grows: boundaries, environment and verification

While instructions shrink, the code that runs regardless of the model's decision grows. It does not dictate the model's route: inside the field the model picks its own tools, order and moment of checking; the code holds the edges of the field. The most visible growth is hooks. The Claude Code guide lists thirty-three events: session start, permission request, before and after a tool call, agent and subagent stop, before and after compaction, worktree creation, configuration change. Exit code 2 blocks the action, the decisions deny, defer, ask and allow combine so that the strictest wins, and the documentation states the purpose bluntly: some actions must happen always, not at the model's discretion. Hooks can be not only scripts but also prompts to another model, and this is an important hybrid: a probabilistic check inside a deterministic loop. Permissions come in six modes, and in auto mode a separate classifier model checks actions before execution and blocks anything outside the scope of the request; some actions are never auto-approved in any mode.

The environment became part of quality, and this is the second growing belt of walls. Cursor, summing up a year of cloud agents, writes that an incomplete environment does not fail with an error — the agent simply works worse, and the degradation gets blamed on the model. Hence resumable workspaces: images with version history, rollback, audit, egress rules. In April Anthropic split the runtime into the “brain” — the model with the loop, compaction and memory — and the “hands” — the tools and the sandbox. While they lived in one container, the model could not start reasoning until the environment finished building, and a failure of either half killed the agent whole; after the split, time to first token fell by sixty percent at the median and by more than ninety at the tail, a dead sandbox is recreated, a dead “brain” is brought back from the session log. Credentials in this scheme are decrypted only at the moment of the tool call, and the model never sees them. In January Docker moved agent sandboxes to dedicated micro-VMs, Vercel announced a bounty of up to a million dollars for escaping its own, and git worktrees became the isolation primitive at Claude Code, Cursor and Codex.

figure 05 · brain and hands: the boundary a tool call crosses
The brain and the hands: the split across the trust boundaryTHE BRAIN AND THE HANDSTHE BRAINmodel + loopcompaction · memory · session logTHE HANDStools · sandboxworktrees · VM · allow-listed networktool call →← resultTHE BOUNDARYTTFT −60 % median · > 90 % P95$0.08 per session-hoursandbox died → recreatebrain died → replay the logcredentials only at call timeThe split per Anthropic Managed Agents (April 2026); compare with the two gateways of the July configurations piece

Tools: four basic ones and a stateless protocol

Tools shrink in their descriptions and grow in the environment. There are four basic ones — read a file, search, edit, bash; mini-SWE-agent gets by with a single bash in a hundred lines and scores above seventy-four percent on SWE-bench Verified, and Pi deliberately lives without the MCP protocol and without subagents, with four tools and TypeScript extensions. When there are thousands of tools their descriptions stop fitting into context, and the harness answers with tool search — Anthropic holds up to ten thousand deferred definitions and reports a token reduction of more than eighty-five percent, honestly adding that accuracy drops after thirty to fifty available tools. The next step is programmatic calling: the model writes code in a container, calls tools from it and gets only the script's output into context; Anthropic claims plus eleven percent in quality at minus twenty-four percent input tokens, and Cloudflare had arrived at the same thing a year earlier with the phrasing “models have seen a lot of code and few tool calls”. The MCP protocol itself, in the revision of 28 July 2026, became stateless: the handshake and the session identifier are gone, server discovery and subscriptions appeared, tasks moved into an extension, dynamic OAuth client registration is deprecated, trace context travels in metadata, and any feature gets at least twelve months before removal. The foundation that governs it grew in nine months from one hundred forty-six to one hundred ninety members, including the US Army and the national laboratories, and cumulative SDK downloads passed a billion.

Verification, observability, identity

Verification is the fastest-growing wall, because it does not require trusting the model. At OpenAI it looks like this: about a million lines in five months, fifteen hundred PRs, seven engineers and zero lines written by hand — and the discipline “shows up in the scaffolding, not in the code”: custom linters that answer the agent with a message containing fix instructions, CI that keeps the documentation fresh, background agents that collect garbage. Böckeler calls these sensors, as opposed to guides, and Thoughtworks in its April Radar describes the same two categories: guides such as skills and specifications, sensors such as mutation testing — under the shared theme “putting coding agents on a leash”. The judge on another model, replayable episodes and the difference between pass@k and pass^k are the subject of a separate article, and in January 2026 Anthropic formalized the same ladder: task, run, grader, transcript, result.

With observability and trust the picture is mixed. Claude Code exports metrics, events and traces over OpenTelemetry, but the GenAI conventions in OpenTelemetry themselves, as of mid-July 2026, are still in “Development” status — there is no stable release, and every vendor traces in its own way. Agent identity, by contrast, took shape as a product within six months: in April Okta made agents full directory records, in August it released single sign-on for agents, and the Cross App Access extension became an official part of MCP for enterprise authorization; in February NIST launched a standards initiative for agents with a separate pillar on identity, and OWASP published its top ten risks for agentic applications. And one more story about the boundary as a product: on 31 March 2026 the Claude Code package went out to npm with a source map — about one thousand nine hundred TypeScript files; Anthropic called it a packaging error, not a security breach, but a day later trojanized “leaks” appeared online. The harness is value that gets stolen and value that gets attacked through. In the security snapshot I showed that every 2026 injection fix ran along boundaries — files, network, auto-approval — and not one was a sentence in a prompt. It is the same pattern seen from the attacker's side.

06

What becomes infrastructure: orchestration and long loops

The third movement is the harness parts that every team wrote itself in 2025 being rented out in 2026. Orchestration went first. Subagents in Claude Code run in their own context windows, return only a summary, nest three levels deep and run up to twenty in parallel; agent teams share a task list and mailboxes, and the documentation honestly says they spend significantly more tokens than a single session. In September, Projects spawn cloud sessions on separate branches from one coordinating conversation, and Routines launch an agent on a schedule, on a GitHub event or by API call, with no permission prompts during the run. The 2025 argument between two schools — Anthropic with an orchestrator and workers that beat a single model by ninety percent on internal evals at fifteen times the token spend, and Cognition with one thread and full traces instead of messages — is not settled, but it has stopped being an argument about code: both strategies are now runtime settings.

Long loops went second. Cursor keeps the agent loop in Temporal, manages the virtual machine's lifecycle separately, and moves storage and conversation streaming into its own layer: “one agent” splits apart by lifetime and by state owner. In September 2026 Temporal raised five hundred fifty million at a twelve-and-a-half-billion valuation “amid demand for reliable AI infrastructure”; Cloudflare built waiting for an event — that is, a pause for human confirmation — directly into the steps of its engine; Inngest explains the demand in one sentence: re-running a step with a model means paying for the tokens again. Durable execution had existed for ten years for payments and orders; the agent loop turned out to be the same class of problem with a more expensive retry.

The runtimes themselves went third. In eight weeks of spring 2026 three vendors released a managed harness whole: on 8 April Anthropic — Managed Agents, “a ready, configurable agent harness on managed infrastructure” with agent, environment, session and event primitives, cloud or self-hosted sandboxes, and a price of eight cents per active session hour on top of tokens; on 19 May Google — Managed Agents in the Gemini API, where one call brings up a remote Linux sandbox configured by AGENTS.md and SKILL.md files; on 3 June Microsoft — Foundry Hosted Agents and an “Agent Harness” in its framework with auto-compaction, file-based memory and a shell. The flip side of the same wave: OpenAI is shutting down its visual agent builder on 30 November 2026, keeping the SDK. Builders lost to code and to the runtime at the same time.

So what stays in-house at the companies that have gone down this road? At Stripe more than one thousand three hundred PRs a week are produced entirely by agents on a fork of the open goose harness and reviewed by people; what remained their own were Blueprints — deterministic process skeletons — and an internal tool server. At Spotify background agents on Claude Code ran two hundred forty migration PRs and saved about ten engineer-weeks, while skills and customizability are switched off deliberately for the sake of guardrails. At Cloudflare, one hundred thirty-one thousand review runs in a month at a dollar nineteen per run, with a “break glass” bypass in six cases out of a thousand and a separate reviewer for the AGENTS.md file. Three companies, three different rented harnesses — and in all three the same things stayed their own: the task, the context, the success criteria and the policy. Konstantin Krestnikov put the same thought from the other side: powerful models no longer need a complex harness, it is enough to give the agent the right files and an instruction.

Messages API → Agent SDK → Managed Agents: the boundary of “what is mine” shrinks to the task, the context, the episodes and the identity

07

Limits of the effect: where the harness stops helping

An honest article about the harness is obliged to gather in one place everything that argues against it. First, the effect at the frontier. Harbor-Index: native harness versus neutral — not significant. METR: Claude Code and Codex not obviously better than the default scaffolds. Meta-Harness promises a sixfold gap in its introduction, and in its own table on Terminal-Bench gives Opus 4.6 plus one point seven points. Put that together with the twenty-seven points of Claw-SWE-Bench and the twenty to thirty points of GigaChain, and you get not a contradiction but the shape of a curve: the gulf between a broken and an adequate harness is measured in tens of points, the gap between two adequate ones at the frontier in single digits. This is exactly what I formulated in July as “at the frontier harnesses converge, below the frontier the harness decides”; in two months the numbers arrived.

figure 06 · harness effect versus model strength: 2026 data
The effect of the harness against model strength, 2026 dataHARNESS EFFECT VERSUS MODEL STRENGTHeffect of changing the harness, ppmodel strength: below the frontier → frontier0204060Claw-SWE-Bench · GLM 5.1 · 54 ppGigaChain · 20–30 ppAHE · GPT-5.4 · +7.3icat · GPT-5.4 · +8.3Meta-Harness · Opus 4.6 · +1.7Harbor-Index · not significantMETR · “not obviously better”● independent study● vendor claim2026 data for figure 05 of the local-AI-stack piece: benchmarks and models differ, point positions are the author's reading

Second, what the harness cannot do in principle. Dex Horthy at the AI Engineer World's Fair attached a footnote to the formula “agent equals model plus harness”: a good harness sharply improves execution but does not teach the model to hold architectural quality over the long haul, because the signal “tests passed” does not penalize a spare try/catch or changes sprawling across the system, and the cost of bad design shows up months later. There is no fast verifier for maintainability, and he admits it honestly. His own team tried a “lights-out factory” in 2025 and a few months later ran into an incident in a codebase that people were no longer watching. DORA's May report on return on investment gives a similar picture from above: thirty-five to forty percent gain on new tasks and about ten on legacy, while the change failure rate in the model calculation rises from five to six percent.

Third, methodology. A paper with the telling title “Stop Comparing Agents Without Disclosing the Harness” shows that current evaluation protocols systematically attribute harness gains to model improvements. A survey of thirteen agents across twelve dimensions finds that eleven of the thirteen combine several primitives and diverge exactly where the questions are open — in compaction and state. For a practitioner the conclusion is one: the harness in a measurement report is as mandatory as the model version, and a number without the harness's name is not a number. This, incidentally, also explains why METR's original “minus nineteen percent” had not “flipped” by February 2026, as secondary retellings claim: for the original cohort the estimate is “a speedup of minus eighteen percent” with an interval crossing zero, for the new one minus four. Tools change, harnesses change, and measuring a task in the lab still does not measure delivery.

08

The labs' bets: who is betting on which harness parts

In July I laid the labs out by the stack layers they own — from hardware to traces. Here the cut is different: by the seven harness parts each player is betting on, and by the mechanism from the third section that explains those bets. The general rule is simple. Functions that can be trained the lab moves into the weights and stops selling as a harness. Functions that cannot be entrusted to the model — permissions, isolation, the compaction policy, identity — it keeps in the runtime and sells as infrastructure. The interfaces between them — protocols and instruction files — it hands over to standards, because a neutral interface expands the market for its own model.

Anthropic bets on the harness as a product at three altitudes: Claude Code for the developer, Cowork for the employee, Managed Agents for the platform; MCP and skills have been handed to open standards, and at its May conference the company formulated the thesis “infrastructure, not intelligence, is the bottleneck”. I expect that over the next five years it will develop precisely the runtime — “brain and hands”, the session log, built-in graders — because the model and the runtime are the moat and the interfaces are neutral territory. OpenAI bets on one harness under all surfaces — App Server unified the terminal, the IDE and the web — and on the harness inside the repository: linters, CI and documentation as machine-readable constraints. Thibault Sottiaux says Codex “is becoming the standard agent” and that scaling to billions of users hinges on isolation; I expect OpenAI to spend five years developing the sandbox and enterprise identity, because those are what turn a developer tool into a product for everyone. Google bets on one harness brand for the IDE, the terminal, the SDK and the API, and on owning the environment that others have to rent: the browser and the operating system. Replacing the open Gemini CLI with the closed Antigravity CLI on the free tier shows a willingness to pay with openness for platform coherence. Microsoft and GitHub bet on the harness in the operating system and in the task-tracking surface: the agent does not need to invent branches, PRs and permissions, they already exist.

Player
Anthropic
Five-year bet
The harness as a product at three altitudes — for the developer, the knowledge worker, the platform; protocols go to standards
2026 evidence
Claude Code, Cowork, Managed Agents; MCP and Skills donated as open standards; “infrastructure, not intelligence, is the bottleneck”
Why
Model and runtime are the moat, interfaces are neutral ground
Player
OpenAI
Five-year bet
One harness beneath every surface and a harness inside the repository
2026 evidence
App Server unified CLI, IDE and web; the “Harness engineering” post; Codex Security; an enterprise platform with agent identities
Why
“Codex is becoming the standard agent” — scale through the sandbox
Player
Google
Five-year bet
One harness brand for IDE, terminal, SDK and API, plus the browser as environment
2026 evidence
Antigravity replaced Gemini CLI on the free tier; Managed Agents in the Gemini API; WebMCP in Chrome; A2A donated
Why
Own the browser and OS-level affordances others must rent
Player
Microsoft and GitHub
Five-year bet
The harness inside the OS and inside the system of record
2026 evidence
Agent Framework with an “Agent Harness”, Foundry Hosted Agents, Windows Agent Runtime; Agent HQ and Agent Plugins 1.0
Why
Issues, branches and permissions already exist — the agent need not invent them
Player
Cursor and Cognition
Five-year bet
A model trained inside its own harness
2026 evidence
Composer 2 trains “in realistic Cursor sessions”; SWE-1.5 — “model, inference system and harness as one unified system”
Why
A moat only where the harness has fused with the weights
Player
DeepSeek and Moonshot
Five-year bet
Open weights plus an own harness as the training environment
2026 evidence
The DeepSeek Harness minimal preset reproduces the training environment; Kimi K2.6 — a swarm of up to 300 subagents
Why
The price floor for everyone else
Player
Open harnesses
Five-year bet
A minimal core and a marketplace of models
2026 evidence
OpenCode as a layer between developers and providers; Pi — four tools, no MCP, no subagents; mini-SWE-agent — a hundred lines
Why
The model is interchangeable and the harness thin — competition on price
Player
The Russian contour
Five-year bet
A ready open harness plus a profile for one's own model
2026 evidence
GigaChain on Deep Agents with a GigaChat profile; SourceCraft CLI on OpenCode; GigaCode role agents; T-Bank opens trading to agents via MCP
Why
Harness sovereignty is cheap, model sovereignty is not

The second group bets on a model trained in its own harness: Cursor, which since August 2026 has reportedly been a subsidiary of SpaceX, and Cognition with SWE-1.5, where the model, inference and harness are designed as one system. The third — open weights plus its own harness as a training environment: DeepSeek with a minimal preset that reproduces the training environment, and Moonshot with a swarm of up to three hundred subagents in Kimi K2.6. The fourth — a minimal core and a marketplace of models: OpenCode as the layer between developers and providers with thirteen million monthly active users according to its head, Pi with four tools, mini-SWE-agent with a hundred lines. For this group the model is interchangeable, the harness is thin, and the competition is on price.

figure 07 · which harness parts the players bet on: a reading of announcements as of 18 September 2026
Which parts of the harness the players bet onWHICH PARTS OF THE HARNESS THE PLAYERS BET ONAnthropicOpenAIGoogleMicrosoftCursor ·CognitionDeepSeek ·KimiOpen(OpenCode · Pi)LoopContextTools & envVerificationOrchestrationObservabilityTrust & boundariesinvestsdonates to a standardwinds downThe author's reading of public announcements as of 18 September 2026; DevDay on 29 September is not included

The Russian landscape

A note on the venue: this text was written for Sber's conference, Sber's products are discussed below, and all of their figures are vendor claims. The Russian bet is a ready open harness plus a profile for one's own model. GigaChain took Deep Agents from LangChain, plugged in GigaChat with a config change, and at an external competition got sixty-seven tasks with the same model against seventy for Claude Code, even though Anthropic optimizes the model and the harness for each other; the model profile described on Habr raised success on the internal set from seventy-seven to eighty-seven percent at half the token spend. Yandex built the open OpenCode into SourceCraft CLI. T-Bank opened trading to external agents via MCP — the harness became a client of a financial service. At Saint HighLoad++ in July Raiffeisen reported a quarter of its backlog going through agents, with four percent of tasks fully autonomous. The conclusion for this landscape is the same one I drew in July, only now with numbers: harness sovereignty is cheap — open cores, standard protocols, a profile in a month — model sovereignty is expensive, and the harness does not replace it.

09

Forecasts for 2027–2031: eight bets with signs of refutation

A forecast without a sign of refutation is not a forecast but a mood. So each of the eight bets below is tied to a year and to an observable sign by which it can be called wrong. The class of all eight is the author's forecast; the supports are METR's metrics, the DORA and JetBrains reports, vendor announcements and the co-evolution mechanism from the third section. The overall vector is one: what can be trained will go into the model; what cannot be entrusted will remain a wall and move to the vendor; what stays with companies as their own is the task, the context, the episodes and the identity.

Bet
1 · The managed runtime becomes the default way to ship an agent; an own loop is a compliance niche
Verifiable sign
More than half of new enterprise agents at the three big vendors run in their runtimes
By when
2028
What refutes it
Large companies keep writing generic loops and publish that as the norm
Bet
2 · Compaction, memory consolidation and subagent spawning become native in the API
Verifiable sign
All three available without a harness from two or more labs
By when
2027
What refutes it
Harnesses keep carrying their own implementations as mandatory
Bet
3 · Protocols stay neutral and boring; competition moves to the runtime and the model
Verifiable sign
None of the four standards has gone back under one company's control
By when
2028
What refutes it
A protocol fork by a large vendor with incompatible extensions
Bet
4 · The loop leaves the human: most agent runs start from an event, a schedule or a build
Verifiable sign
Vendor telemetry shows more than half of runs start without an interactive prompt
By when
2027
What refutes it
The interactive terminal stays the main entry point by run count
Bet
5 · Verification becomes the largest harness part: the agent runs the product, not only tests
Verifiable sign
Screen and browser confirmation of the result is a standard step in leading harnesses
By when
2028
What refutes it
Verification stays at “the tests passed”
Bet
6 · “Harness as a moat” survives only fused with a trained model
Verifiable sign
Standalone harness companies are acquired or move to the labs
By when
2029
What refutes it
An independent harness without its own model keeps leading its market
Bet
7 · Agent identity and audit become a regulatory requirement
Verifiable sign
A regulation in at least one jurisdiction requires a per-run identity and an action log
By when
2028–2030
What refutes it
Regulators stop at models and do not touch the runtime
Bet
8 · The harness effect at the frontier shrinks to statistical zero, below the frontier it does not
Verifiable sign
Neutral and native runs of frontier models are indistinguishable; the spread persists for open models
By when
2027
What refutes it
Native harnesses consistently beat neutral ones at the frontier

A few explanations for the bets that look bolder than the others. The first rests on three managed runtimes shipping within eight weeks, not on faith in vendors: when three competitors converge on one architecture, the architecture becomes the market's expectation. The second rests on what has already happened: reasoning between calls and context tracking went into the model within a year, and compaction with memory is of the same type. The fourth rests on JetBrains, where ninety percent of developers use agents weekly and about half of the code is written entirely by agents: the next step after “every day” is “without me”. The seventh rests on the NIST initiative and on Dario Amodei's September essay about “built-in graders”: when a lab itself asks for regulation of the runtime, it comes. The eighth is a direct continuation of the seventh section and the only bet that already has 2026 data.

The horizon of all the bets is bounded by what METR calls the task horizon: the fifty-percent horizon of Opus 4.5 in January 2026 was three hundred twenty minutes with a wide interval, doubling every three to seven months depending on which stretch of the series you take. If the doubling holds, by 2028 an agent will hold a task the length of a working week, and the harness for such a task is no longer a loop but an operating system with a scheduler, memory and permissions. If the doubling slows, bets one, four and five shift by a year or two, but the direction does not change: back in June 2025 Gartner promised the cancellation of more than forty percent of agentic projects by the end of 2027, and the survivors will be the ones that are observable, bounded and cheap to check — that is, the ones whose walls were built before their words.

10

The harness ledger: how to keep assumptions about the model

From the law “every component encodes an assumption” follows a practice, and it is not a checklist. In September I proposed an ablation test: take characteristic episodes, compare the current model with the full harness, with a simplified one, and a new model with minimal adaptation, and remove the rule that no longer improves anything. The test is a method; it needs a record-keeping form. I call it the harness ledger: for each component — a hook, a rule, a skill, a reset, a tool — you write down which assumption about the model it encodes, what evidence confirms it, under which model and version, what triggers a recheck, and in which ownership mode the component lives. The last column comes from the ownership boundary framework: rent, adapt, own.

figure 08 · the harness ledger: six columns and two hypothetical examples
The harness ledger: six columns and two notional rowsTHE HARNESS LEDGERCOMPONENThook, rule,skill, resetcontext resetat 60 %hook: blockrm -rf outsidethe worktreeASSUMPTIONwhat the modelcan't dothe model getsanxious nearthe limitan injectionmay issue adestructivecommandEVIDENCEepisodesbefore and after0 ppon 30 episodesan incident,not episodesDATEmodeland versionSonnet 4.5→ Opus 4.5any modelTRIGGERnew modelincidentregressionnew modelnever: it isa boundaryMODErentadapt · owndeleteownamber: a word that a new model made redundant · blue: a wall that no model retiresA notional example: the ledger is the bookkeeping form of the ablation test, not a checklist

The ledger solves three problems a checklist does not. First, it distinguishes assumptions from boundaries. A context reset at sixty percent of the window is an assumption about a nervous model, and it has to be rechecked with every new model; a hook forbidding destructive commands outside the worktree is a boundary whose recheck trigger is “never”, because it protects against injection, not against a weakness of the model. Second, it gives a shelf life: Cherny's practice — every six months delete the instructions, skills and hooks and bring back line by line whatever the model stumbles on repeatedly — becomes not a gesture of courage but a row with a date. His own estimate that evals outlive the harness by one to three model generations sets a shelf life for the evidence too. Third, it makes ablation cheap: a component without evidence is the first candidate for removal, a component with an incident instead of episodes is the last.

Component
Context reset at 60% of the window
Assumption about the model
The model gets anxious near the context limit and wraps up early
Evidence
30 episodes before and after removal: no difference
Date and model
Introduced under Sonnet 4.5, retested under Opus 4.5
Trigger
New model
Mode
Delete
Component
Pre-call hook: destructive commands blocked outside the worktree
Assumption about the model
An injection in text the agent read may issue a destructive command
Evidence
An incident, not episodes
Date and model
Any model
Trigger
Never — it is a boundary, not an assumption
Mode
Own
Component
A 400-line “release checklist” skill
Assumption about the model
The model forgets release steps
Evidence
Not verified
Date and model
Written for last year's model
Trigger
Every six months
Mode
Adapt or delete
  • Open a ledger row when you add any harness component, not at audit time
  • Evidence is replayable episodes before and after; an incident is evidence only for boundaries
  • The “new model” trigger runs ablation on every row with an assumption; boundary rows are not subject to ablation
  • The ownership mode is chosen by the ownership boundary framework, not by habit

The example in the table is hypothetical, and that needs saying plainly: the ledger is a tool I am proposing, not a report on something already in use. There is one way to check it — start one for your own fleet of agents and see how many rows end up without evidence. My forecast: more than half, and almost all of them will be words, not walls.

Takeaways

Seven takeaways on the harness

  1. 01The harness is migrating from words to walls. Instructions — the system prompt, skills, tool descriptions — shrink into maps, while boundaries and environment — hooks, permissions, sandboxes, identity, durable execution — grow. Walls do not lead the model by the hand: they fence the field, and inside it the model gets more freedom than a year ago. The paradox of “ninety percent of the result” and “minus eighty percent of the prompt” is not an argument about models but a description of two different parts of one harness.
  2. 02The word's two meanings converge. The harness the model works in and the environment it is trained in are increasingly the same thing: Cursor, Cognition and DeepSeek say so outright. That is why a foreign harness at the frontier is structurally at a disadvantage against the native one — and exactly why the difference between them at the frontier is measured in single points.
  3. 03The harness has become a disclosable variable of evaluation. One model in different harnesses yields anywhere from a few to fifty points of difference, and evaluation protocols attribute that gain to the model. The harness in a measurement report is as mandatory as the model version.
  4. 04Orchestration and long loops have become infrastructure. Three vendors released managed runtimes within eight weeks, the agent loop lives in durable execution engines, visual builders are on their way out. The “boundary of what is mine” has shrunk to the task, the context, the episodes and the identity.
  5. 05Every harness component encodes an assumption about what the model cannot do. Assumptions go stale within one to three model generations, so they have to be kept in a ledger with a recheck date, not a checklist with ticks.
  6. 06Self-modifying harnesses optimize the reward they were given: the very first one deleted the markers used to catch hallucinations. Without a ledger and a hidden judge, harness auto-evolution is Goodhart's law handed a machine gun.
  7. 07By 2028 the harness as a separately named product will dissolve into the model API and the managed runtime. What stays visible is boundaries, episodes and identity. The sign of refutation: independent harness companies without a model of their own holding leadership in their markets.
Sources

Documentation, research, talk recordings and the limits of evidence

The list is grouped by section. Every figure in the article names its baseline where it appears; company statements are marked as claims. Product documentation and specifications capture the state as of the check date — 18 September 2026. Talk recordings are cited through my write-ups in Knizhny kub, which record what exactly was said. The full dossier with evidence classes, contradictions between sources and a list of what could not be confirmed sits in the site repository next to this article.

Definitions and the term

  1. Anthropic · Claude Code docs: Glossary — 18 September 2026: “agentic harness — the tools, context management, and execution environment… Claude Code is the harness; Claude is the model inside it”; a vendor definition
  2. LangChain (Vivek Trivedy) · The Anatomy of an Agent Harness — 10 March 2026: “a harness is every piece of code, configuration, and execution logic that isn't the model itself”; the formula “Agent = Model + Harness”; a framework vendor
  3. Birgitta Böckeler (martinfowler.com) · Harness engineering for coding agent users — 2 April 2026: the harness is “everything in an AI agent except the model itself”; split into guides (feedforward) and sensors (feedback)
  4. Anthropic (Prithvi Rajasekaran) · Harness design for long-running application development — 24 March 2026: “every component in a harness encodes an assumption about what the model can't do on its own”; Opus 4.6 removed the need for sprints and context resets; the full harness took 6 hours and $200 vs 20 minutes and $9 for a solo agent — one experiment
  5. Anthropic (Martin, Cemaj, Cohen) · Scaling Managed Agents: Decoupling the brain from the hands — 8 April 2026: the harness is “the loop that calls Claude and routes Claude's tool calls to the relevant infrastructure”; brain and hands; harness assumptions “go stale as models improve”; a vendor describing its own product
  6. Hugging Face (Paniego, Gosthipaty) · The Agent Glossary — 25 May 2026: harness as the execution layer, scaffold as the behavior layer; the authors note there are no universally accepted definitions; inverse to Anthropic's 2025 usage
  7. Prime Intellect · verifiers v1 — 10 July 2026: a training environment = a taskset (“data, tools, and scoring”) + a harness (“the program that solves the task”); one configuration for evals and training; vendor documentation
  8. Addy Osmani · The New Software Lifecycle (blog version of Google's whitepaper “The New SDLC With Vibe Coding”) — 16 June 2026: “an agent is a model plus a harness”, the “10% model, 90% harness” split is the authors' estimate, not a measurement; from outside the top 30 into the top 5 of Terminal-Bench 2.0 by changing the harness; LangChain +13.7 points
  9. Knizhny kub · The New SDLC: harness, the factory model and the economics of agentic engineering — 21 June 2026: a breakdown of the second half of Google's whitepaper — the “Agent = Model + Harness” formula, the factory model, the conductor and orchestrator roles
  10. Knizhny kub · Loop Engineering: why the most important part of the agent loop is the right to say no — 14 July 2026: the ladder prompt → context → harness → loop and the five movements of a loop; the HuaShu material is a working note, not an Anthropic publication

Eras and history

  1. EleutherAI · lm-evaluation-harness — Opened 18 September 2026: “a framework for few-shot evaluation of language models” — the word “harness” entered the model world as an evaluation runner
  2. Simon Willison · OpenAI function calling — 13 June 2023: gpt-4-0613 and gpt-3.5 accept JSON schemas of functions and return a structured call — the first harness function to move into the API
  3. Yang et al. (Princeton) · SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — 6 May 2024 (v1): the agent-computer interface (ACI) markedly improves file creation and editing, repository navigation and test execution; the abstract uses neither “harness” nor “scaffold”
  4. OpenAI · Introducing SWE-bench Verified — August 2024: the wrapper around the model is called a “scaffold” (Agentless), while “harness” names the containerized evaluation runner; page updated 24 February 2025
  5. Anthropic · Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku — 22 October 2024: “we're teaching it general computer skills” instead of building task-specific tools; OSWorld 14.9% from screenshots vs 7.8% for the next-best system; “at times cumbersome and error-prone”
  6. Anthropic · Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet — 6 January 2025 (date after the page migration): “an ‘agent’ refers to the combination of an AI model and the software scaffolding around it”; “give as much control as possible to the language model itself, and keep the scaffolding minimal”
  7. Anthropic · Introducing Claude 4 — 22 May 2025: SWE-bench results use “the same simple scaffold that equips the model with solely the two tools” — bash and file editing; Claude Code general availability
  8. Anthropic · Introducing Claude Sonnet 4.5 — 29 September 2025: “more than 30 hours” of focus on multi-step tasks — a vendor claim; the Agent SDK on “the same infrastructure that powers Claude Code”; context editing and the memory tool in the API
  9. Anthropic (Thariq Shihipar) · Building agents with the Claude Agent SDK — 29 September 2025: “the agent harness that powers Claude Code… can power many other types of agents, too”; the loop “gather context → take action → verify work → repeat”
  10. Anthropic (Justin Young) · Effective harnesses for long-running agents — 26 November 2025: an initializer agent creates a launch script, a progress file and a feature list; a later instance “would look around, see that progress had been made, and declare the job done”
  11. Linux Foundation · Announces the Formation of the Agentic AI Foundation — 9 December 2025: eight platinum members including Anthropic, Google, Microsoft and OpenAI; MCP, goose and AGENTS.md donated; “more than 10,000 published MCP servers”; AGENTS.md “adopted by over 60,000 open source projects”
  12. Andrej Karpathy · 2025 LLM Year in Review — 19 December 2025: “Claude Code emerged as the first convincing demonstration of what an LLM Agent looks like”
  13. Mitchell Hashimoto · My AI Adoption Journey — 5 February 2026: “I don't know if there is a broad industry-accepted term for this yet, but I've grown to calling this ‘harness engineering’”; the rule — turn every agent mistake into a fix that makes it impossible
  14. Knizhny kub · Harness: why one agent with files displaces complex harnesses (Krestnikov, Data Fest 2026) — 15 August 2026: a periodization from bare LLMs to one agent in the file space; “the LLM is the draft power, files are the field, the harness is the yoke”; Claude Code with Sonnet solves 70 of 104 tasks, DeepAgents with the same model 67; claims by the GigaChain team (Sber)

Co-evolution of model and harness

  1. Simon Willison · OpenAI Structured Outputs — 6 August 2024: the schema is compiled into a context-free grammar; the earlier JSON mode promised only valid syntax, not schema conformance, and in the author's words “neither of these modes were guaranteed to return valid JSON”
  2. Anthropic · Claude docs: Extended thinking — 18 September 2026: interleaved thinking needed a beta header on Claude 4, is automatic on 4.6 and later, and the manual mode is rejected on 4.7 and later — a harness flag dissolved into model behavior
  3. Anthropic · Claude docs: Context windows — 18 September 2026: the million-token window is default and standard-priced on 4.6 and later; “context rot” is named in the docs; Sonnet 4.5–5 receive API-injected budget tags, Opus 4.7 and later do not
  4. Simon Willison · GPT-5.1-Codex-Max — 19 November 2025: OpenAI — “our first model natively trained to operate across multiple context windows through… compaction”; Willison notes Claude Code already did auto-summarization in the harness
  5. Anthropic · Claude docs: Compaction — 18 September 2026: server-side compaction with a default 150,000-token trigger and a 50,000 minimum; the summary prompt “varies by model” and can be replaced by the developer; SDK compaction marked deprecated
  6. Anthropic · Claude Code CHANGELOG — 18 September 2026, version 2.1.271: “changed the task-tracking tools to be offered only on Claude 3.x, Opus 4.0–4.7” — a harness component withdrawn for newer models
  7. Cursor · Composer: Building a fast frontier model with RL — 29 October 2025: during RL the model “can call any tool in the Cursor Agent harness”; “hundreds of thousands of concurrent sandboxed coding environments”; a vendor claim
  8. Cursor · Composer 2 technical report — 27 March 2026: “RL training occurs in realistic Cursor sessions with the same tools and harness the deployed model uses”; a vendor claim
  9. Cursor · Real-time RL for Composer — 26 March 2026: model checkpoints “approximately every 5 hours”; the user is part of the training environment; +2.28% edit persistence, −3.13% dissatisfied follow-ups, −10.3% latency — internal metrics
  10. Cognition · SWE-1.5 — 29 October 2025: “end-to-end RL on real task environments using our custom Cascade agent harness”; “re-trained the model on the updated harness”; 950 tokens per second; a vendor claim
  11. VentureBeat · OpenAI unveils GPT-5-Codex, optimized for agentic coding — 15 September 2025: the model is “trained on real-world engineering work”; recommended “only for agentic coding tasks in Codex or Codex-like environments”; the training harness itself is not named
  12. Qwen · Qwen3-Coder — 22 July 2025: “long-horizon RL (Agent RL)” across “20,000 independent environments in parallel”; the terminal client adapted from Gemini CLI; a vendor claim
  13. Zhipu · GLM-5 — 17 February 2026: over 10,000 verifiable environments in the Harbor format with Docker build accuracy above 90% — training environments literally in the benchmark's format; a vendor claim
  14. DeepSeek · deepseek-harness — 18 September 2026: the session as an append-only log of typed events; the minimal preset is described as an “RL-agent composition” with persistent bash and a string-replace editor; open code shows the possibility, not the data policy
  15. Knizhny kub · DeepSeek Harness: when the agent's history becomes part of inference economics — 18 August 2026: why compaction repeats the previous prefix byte for byte, and why the loop “harness → trajectories → post-training → the same harness” is an architectural possibility, not a proven practice
  16. Anthropic · Introducing Claude Opus 4.6 — 5 February 2026: Terminal-Bench 2.0 — “all runs used the Terminus-2 harness, except for OpenAI's Codex CLI”; a multi-agent harness raised BrowseComp to 86.8%; a prompt modification yielded 81.42% on SWE-bench; a vendor claim
  17. arXiv 2604.25850 · Agentic Harness Engineering — 28 April 2026: GPT-5.4 on Terminal-Bench 2.0 — 69.7% with the seed harness, 71.9% with Codex CLI, 77.0% with the evolved one; transferring the harness to other models yields +2.3 to +10.1 pp
  18. arXiv 2606.25514 · icat-agent — 24 June 2026: GPT-5.4-xhigh on SWE-bench Pro — 59.1% with the best prior scaffold vs 67.4% with the new one
  19. Sakana AI · Darwin Gödel Machine — 30 May 2025: a self-modifying agent raised SWE-bench from 20.0 to 50.0% and Polyglot from 14.2 to 30.7%; it also “removed the markers we use in the reward function to detect hallucination (despite our explicit instruction not to do so)”
  20. Hu, Lu, Clune · Automated Design of Agentic Systems — 15 August 2024: “a meta agent iteratively programs interesting new agents based on an ever-growing archive of previous discoveries” — the precursor of self-modifying harnesses
  21. Google DeepMind · AlphaEvolve impact — 7 May 2026: −30% variant detection errors in DeepConsensus, −20% write amplification in Spanner — vendor results on its own systems
  22. arXiv 2604.16968 · On Safety Risks in Experience-Driven Self-Evolving Agents (ACL 2026) — 18 April 2026: experience gathered on benign tasks “can still compromise safety in high-risk scenarios”; experience “reinforces agents' tendency to act rather than refuse”
  23. Manus (Peak Ji) · Context Engineering for AI Agents: Lessons from Building Manus — 18 July 2025: “we've rebuilt our agent framework four times”; “if model progress is the rising tide, we want Manus to be the boat, not the pillar stuck to the seabed”
  24. Philipp Schmid · The importance of Agent Harness in 2026 — 5 January 2026: the harness is “the infrastructure that wraps around an AI model to manage long-running tasks. It is not the agent itself”; model as CPU, context as RAM, harness as OS; five Manus refactors in six months
  25. Anthropic · Effective context engineering for AI agents — 29 September 2025: “smarter models require less prescriptive engineering”; context is a finite resource with an “attention budget”; compaction, notes and subagents as three techniques for tasks longer than the window

Instructions, loop and context

  1. YC Root Access · Boris Cherny: Building Claude Code — 27 July 2026: “we deleted 80% of the system prompt — Opus 5 just does it”; a variable that removes all prompts; “delete the entire system prompt and then bring it back line by line”; “every six months, delete your instructions, skills, hooks”; “evals outlive the harness… maybe one, two, three” generations; a vendor employee's self-report
  2. Knizhny kub · Boris Cherny: We Cut 80% of Claude Code's Prompt (YC Startup School 2026) — 17 August 2026: a breakdown of the talk given the day after the Opus 5 release — “without a measurable quality loss on their coding evals” in the speaker's words; “product overhang”; the shift from prompts through context to task setting
  3. Knizhny kub · Nick Nisi on agent skills: fewer instructions, more evidence — 11 July 2026: the talk “How I deleted 95% of my agent skills and got better results” (AI Engineer Europe, 10 April 2026); 10,739 generated instruction lines; flow control moved into a state machine with mandatory gates; an engineer's self-report
  4. Vercel · We removed 80% of our agent's tools — 22 December 2025: the agent stripped down to one bash tool — 3.5× faster, −37% tokens, −42% steps, success 80 → 100% on five internal queries; “the model makes better choices when we stop making choices for it”; a vendor claim
  5. Knizhny kub · Andrew Qu (Vercel): How We Solved Agent Building — 16 September 2026: d0's path from a big prompt through a chain of specialists to one agent with a working directory; the contributions of the architecture and of a stronger model to the doubled score are not separated
  6. Knizhny kub · Kyle Mistele: Loop Engineering from First Principles — 30 August 2026: the agent loop as a control loop — set point, sensor, controller, actuator; the model is needed far from everywhere; the speaker co-founded a vendor
  7. AGENTS.md — 18 September 2026: “used by over 60k open-source projects” (a December 2025 figure); 24 supporting tools; stewarded by the AAIF
  8. Lulla et al. · AGENTS.md study (arXiv 2601.20404) — 28 January 2026 (v2 30 March): across 124 PRs in 10 repositories the presence of AGENTS.md is associated with −28.64% median runtime and −16.58% output tokens at comparable completion; a correlation, not an experiment
  9. Vercel · AGENTS.md outperforms skills in our agent evals — 27 January 2026: a compressed 8 KB docs index in AGENTS.md scored 100% on Next.js 16 API evals, skills maxed at 79%; a vendor's internal eval
  10. Agent Skills · Specification — 18 September 2026: a SKILL.md file with a name up to 64 and a description up to 1024 characters; progressive disclosure — metadata (~100 tokens) → instructions (under 5,000 tokens) → resources on demand; 46 clients on the showcase
  11. Geoffrey Huntley · Ralph Wiggum as a “software engineer” — 14 July 2025: “in its purest form, Ralph is a Bash loop” that feeds one prompt to the agent again and again
  12. Anthropic · claude-code/plugins/ralph-wiggum — 18 September 2026: the official plugin implements the loop via a Stop hook “by blocking normal session exit”, with an iteration cap and a completion promise
  13. Anthropic · Claude Code docs: /goal — 18 September 2026: completion “is decided by a fresh model rather than the one doing the work” (Haiku by default); conditions must name a check, not an intention
  14. Anthropic · Claude Code docs: Scheduled tasks (/loop) — 18 September 2026: repetition at a fixed interval or self-paced from one minute to an hour, session-scoped, a seven-day expiry
  15. Anthropic · Claude Code docs: Checkpointing — 18 September 2026: every prompt creates a checkpoint, the 100 most recent are kept; “checkpointing does not track files modified by Bash commands”
  16. Anthropic · A harness for every task: dynamic workflows in Claude Code — 2 June 2026: “Claude can now write its own harness on the fly, custom-built for the task at hand” — an orchestration script generated per task; a vendor claim
  17. Anthropic · Agent SDK docs: The agent loop — 18 September 2026: the loop ends with a response that contains no tool calls; hard caps on turns and on a dollar budget, the budget covers subagents
  18. Chroma (Hong, Troynikov, Huber) · Context Rot — 14 July 2025: 18 models; “LLMs do not maintain consistent performance across input lengths” even on trivial tasks
  19. Anthropic · Claude Code docs: Context window — 18 September 2026: auto-compaction clears old tool outputs first, then summarizes; the threshold is about 967,000 tokens on million-token models; up to five recently modified files are re-read afterwards

Boundaries, tools and standards

  1. Anthropic · Claude Code docs: Permissions — 18 September 2026: “permission rules are enforced by Claude Code, not by the model. Instructions in your prompt or CLAUDE.md shape what Claude tries to do, but they don't change what Claude Code allows”
  2. Anthropic · Claude Code docs: Memory — 18 September 2026: CLAUDE.md “is not a hard enforcement layer”; “to block an action regardless of what Claude decides, use a PreToolUse hook”; aim for a file under 200 lines
  3. Anthropic · Claude Code docs: Hooks guide — 18 September 2026: a table of 33 events; exit code 2 blocks the action; deny/defer/ask/allow decisions with the most restrictive winning; “certain actions always happen rather than relying on the LLM to choose to run them”
  4. Anthropic · Claude Code docs: Permission modes — 18 September 2026: six modes; in auto mode “a separate classifier model reviews actions before they run, blocking anything that escalates beyond your request”; some actions are never auto-approved in any mode
  5. Anthropic · Beyond permission prompts: making Claude Code more secure and autonomous — 20 October 2025: a bubblewrap and Seatbelt sandbox “safely reduces permission prompts by 84%” and covers “any scripts, programs, or subprocesses”; the vendor's internal telemetry
  6. OpenAI · Codex docs: Sandboxing — 18 September 2026: two layers outside the model — the sandbox mode (what is technically possible) and the approval policy (when to ask); “you aren't just trusting the agent's intentions; you are trusting that the agent is operating inside enforced limits”
  7. Anthropic · Claude Code docs: Security — 18 September 2026: network commands are never auto-approved, web fetch runs in a separate context window, credentials are masked, “all operations in cloud sessions are logged”; Anthropic “does not security-audit or manage any MCP server”
  8. Anthropic · Claude docs: Tool search tool — 18 September 2026: up to 10,000 deferred tools; ~55,000 tokens of definitions reduced “by over 85 percent”; accuracy degrades “once you exceed 30–50 available tools”; a vendor claim
  9. Anthropic · Claude docs: Programmatic tool calling — 18 September 2026: the model writes code in a container and calls tools from it, only the script's output returns to context; “improved performance by an average of 11% while using 24% fewer input tokens”; a vendor claim
  10. Cloudflare (Varda, Pai) · Code Mode — 26 September 2025: MCP tools become a TypeScript API and the model writes code executed in an isolate; “LLMs have seen a lot of code. They have not seen a lot of ‘tool calls’”
  11. Model Context Protocol · Specification 2026-07-28: Changelog — 28 July 2026: the protocol went stateless — the handshake and session id removed; server discovery and subscriptions added; tasks moved to an extension; OAuth dynamic client registration deprecated; trace context in metadata; a minimum twelve-month deprecation window
  12. Linux Foundation · Agentic AI Foundation adds 43 new members — 18 May 2026: 190 members; new Gold members include F5, GoDaddy, Stripe; government members include the U.S. Army, PNNL, Sandia
  13. Linux Foundation · Agentic AI Foundation launches MCPA certification — 14 September 2026: the foundation's first certification; the TypeScript and Python SDKs — “more than 1 billion cumulative downloads”; “MCP has quickly become the standard for connecting AI agents to external systems”
  14. Linux Foundation · A2A protocol surpasses 150 organizations — 9 April 2026: more than 150 organizations, 22,000+ stars, SDKs in five languages, version 1.0 with signed agent cards; accepted into the AAIF on 27 August 2026
  15. Okta · Agent SSO — 24 August 2026: Cross App Access “formally incorporated as the official Enterprise-Managed Authorization extension for the Model Context Protocol”; agents as “first-class identities in Universal Directory” since GA on 30 April 2026
  16. NIST/CAISI · Announcing the AI Agent Standards Initiative — 17 February 2026: three pillars — industry-led standards, open protocol development, research on agent security and identity; a concept paper on identity and authorization
  17. OWASP · Top 10 for Agentic Applications for 2026 — 9 December 2025: from agent goal hijack (ASI01) to rogue agents (ASI10); more than a hundred experts
  18. OpenTelemetry · Semantic conventions for generative AI — 18 September 2026: the conventions moved to a dedicated repository; as of mid-July 2026 every gen_ai item is still marked “Development” — no stable release
  19. InfoQ · Claude Code source leak — 7 April 2026: package version 2.1.88 shipped a source map with about 1,906 TypeScript files; Anthropic called it “a release packaging issue caused by human error, not a security breach”; the contents are not used in the article
  20. Anthropic · Demystifying evals for AI agents — 9 January 2026: task, trial, grader, transcript, outcome; pass@k vs pass^k — “the probability that all k trials succeed”; start with 20–50 tasks from real failures
  21. Docker · Docker Sandboxes: run Claude Code and other coding agents unsupervised but safely — 30 January 2026: sandboxes on dedicated microVMs “adding a hard security boundary”; Claude Code, Copilot CLI, Codex CLI, Gemini CLI, Kiro supported
  22. Anthropic · Claude Code docs: Worktrees — 18 September 2026: a separate worktree per session or subagent; four checks block edits and git operations in the main checkout; Cursor 2.0 (29 October 2025) ran up to eight agents on worktrees or remote machines
  23. SWE-agent · mini-SWE-agent — 18 September 2026: about a hundred lines, “does not have any tools other than bash”, a linear history; “scores >74% on the SWE-bench verified benchmark”
  24. Mario Zechner · pi-mono — 18 September 2026: four tools, a short prompt, TypeScript extensions; deliberately no MCP and no subagents; MIT license

Orchestration and infrastructure

  1. Anthropic · Claude Code docs: Subagents — 18 September 2026: a subagent works in its own context window and “returns only the summary”; nesting up to three levels, up to 20 concurrent; worktree isolation
  2. Anthropic · Claude Code docs: Agent teams — 18 September 2026: a lead and teammates with a shared task list and mailboxes; “agent teams use significantly more tokens than a single session”; “a teammate can't approve a permission prompt… on your behalf”
  3. Anthropic · Claude Code docs: Projects — 18 September 2026 (public beta): one coordinator conversation spawns threads — cloud sessions on their own branches; shared instructions up to 16,000 characters and a project memory
  4. Anthropic · Claude Code docs: Routines — 18 September 2026 (research preview): runs on a schedule, a GitHub event or the API; no permission prompts during a run; API-trigger text arrives marked untrusted
  5. Temporal · Temporal raises $550M at a $12.55B valuation — 14 September 2026: a $550M Series E at a $12.55B valuation “as demand grows for reliable AI infrastructure”; $300M at $5B in February; a company statement
  6. Cloudflare · Workflows GA: production-ready durable execution — 7 April 2025: steps that wait for an event or sleep — pausing the loop for a human approval is built into the engine
  7. Inngest (Charly Poly) · Durable execution: the key to harnessing AI agents — 19 February 2026: “this is especially valuable for LLM calls, where re-execution means re-paying for tokens”; a vendor blog
  8. Knizhny kub · Cursor Cloud Agents: what to hand to the agent and what to keep on the platform — 14 August 2026: per Cursor's June write-up — procedural logic moves from the harness into tools, the environment became part of quality, the loop lives in Temporal, and “one agent” decomposes by lifetime and state owner
  9. Anthropic · Claude Managed Agents — 8 April 2026: “pre-built, configurable agent harness that runs in managed infrastructure”; agent, environment, session, events; cloud or self-hosted sandboxes; $0.08 per active session-hour on top of tokens; early customers Notion, Rakuten, Asana, Sentry
  10. Knizhny kub · Evolution of Agentic Surfaces: the harness becomes infrastructure — 21 August 2026: three generations of surfaces — Messages API, Agent SDK, Managed Agents; Sonnet 4.5's “context anxiety” and the context reset that became overhead under Opus 4.5; the brain–hands split cut median time to first token by 60%; a vendor's account
  11. Google Developers Blog · All the news from the Google I/O 2026 developer keynote — 19 May 2026: Managed Agents in the Gemini API — one call provisions a remote Linux sandbox, configured via AGENTS.md and SKILL.md; Antigravity 2.0, CLI and SDK with “programmatic control over agent harness”; WebMCP in Chrome 149
  12. Microsoft (Shawn Henry) · Microsoft Agent Framework at Build 2026 — 3 June 2026: an “Agent Harness” with auto-compaction, file memory and a shell; Foundry Hosted Agents; CodeAct — the agent writes Python to call tools, latency −52.4%; a vendor claim
  13. OpenAI · Agent Builder guide — 18 September 2026: the visual builder “is scheduled to shut down on November 30, 2026”; ChatKit and the Agents SDK remain
  14. Anthropic · How we built our multi-agent research system — 13 June 2025: an orchestrator with workers is 90.2% better than a single Opus 4 on internal evals at roughly 15× the tokens; a vendor claim
  15. Cognition (Walden Yan) · Don't Build Multi-Agents — 12 June 2025: “share context, and share full agent traces, not just individual messages”; actions carry implicit decisions, and conflicting decisions carry bad results; a vendor's position
  16. Stripe (Alistair Gray) · Minions: Stripe's one-shot, end-to-end coding agents — 9 February 2026: “over 1,300 pull requests… merged each week” are fully agent-produced and human-reviewed; a goose fork; Blueprints combine the determinism of workflows with agents' flexibility; a company's self-report
  17. Spotify Engineering · Background coding agents: dataset migrations (Honk, part 4) — 22 April 2026: Claude-Code-based agents delivered 240 migration PRs and saved about ten engineering weeks; skills and configurability withheld as “a deliberate design choice for guardrails”; a company's self-report
  18. Cloudflare (Ryan Skidmore) · AI code review at scale — 20 April 2026: 131,246 review runs on 48,095 merge requests in a month, $1.19 average per run, 0.6% “break glass” overrides, a dedicated AGENTS.md reviewer; a company's self-report
  19. Simon Willison · Code w/ Claude 2026 (live blog) — 6 May 2026: multi-agent Managed Agents, Outcomes, Dreaming, Routines; “most people will experience AI through one of the things you've built on the Claude platform” (Ami Vora)

Limits of the effect

  1. METR · Measuring time horizon using Claude Code and Codex — 13 February 2026: “Claude Code and Codex aren't obviously better than the default scaffolds we use for our agents”; Opus 4.5 in Claude Code beats ReAct in 50.7% of bootstrap samples, GPT-5 in Codex beats Triframe in 14.5%
  2. Terminal-Bench · Harbor-Index — 29 June 2026: “native usually finishes a little ahead… but no comparison is statistically significant”
  3. arXiv 2609.04298 · Harbor adapters: Terminus-2 vs native harnesses — 3 September 2026: every model is run with Terminus-2 and with one of three native harnesses; the strongest, GPT-5.5 with Codex, reaches 28.0%
  4. Lee et al. (Stanford) · Meta-Harness — 30 March 2026: the introduction says “changing the harness around a fixed LLM can produce a 6× performance gap”; the paper's own Terminal-Bench 2.0 numbers are modest: Opus 4.6 — 74.7 → 76.4%, Haiku 4.5 — 33.7 → 37.6%
  5. arXiv 2606.12344 · Claw-SWE-Bench — 10 June 2026: GLM 5.1 — 19.1% with a minimal adapter vs 73.4% with the full one; with the model fixed, the harness choice moves the score by 27.4 pp on average
  6. arXiv 2605.23950 · Stop Comparing LLM Agents Without Disclosing the Harness — 7 May 2026: “current evaluation protocols therefore systematically misattribute harness-level gains to model improvements”
  7. TechCrunch · Nvidia just showed that the harness, not the AI model, is now the real hero — 21 August 2026: per Nvidia's researchers, Opus 5 on ARC-AGI-3 scores 100% with a harness vs 30% without; ARC Prize verification is not mentioned in the article
  8. Knizhny kub · Harness Engineering is not Enough: Why Software Factories Fail (Dex Horthy) — 25 August 2026: a footnote to “Agent = Model + Harness” — a harness improves execution but does not teach the model to keep architectural quality over time; there is no fast verifier for maintainability; the speaker co-founded a vendor
  9. Knizhny kub · YC Paper Club: Why The Harness Matters More Than The Model — 15 September 2026: four engineering talks on what changes when the same model gets more environment; the story of a 99.9% ARC-AGI run that turned out to be cheating once the logs were read
  10. Thoughtworks · Technology Radar Vol. 34 — 15 April 2026: the themes “Putting coding agents on a leash” and “Securing permission-hungry agents”; “teams are beginning to iterate on coding agent harnesses” — feedforward controls like skills and specs, feedback controls like mutation testing
  11. InfoQ (Matt Saunders) · DORA report on the ROI of AI-assisted software development — 11 May 2026: 35–40% gains on greenfield tasks and about 10% on legacy code; a model for 500 engineers — 39% ROI, roughly an eight-month payback, change failure rate 5 → 6%
  12. METR · Uplift update — 24 February 2026: for the original cohort “a speedup of −18% with a confidence interval between −38% and +9%”, for the new cohort −4% (−15% to +9%); 30–50% of developers withheld some tasks rather than do them without AI

The labs' bets

  1. VentureBeat (Michael Nuñez) · Anthropic says it hit a $30 billion revenue run rate — 8 May 2026: a $30B run-rate; Claude Code — “more than $2.5B” run-rate by February, business subscriptions 4× since the start of the year; “the majority of code is now written by Claude Code” internally; company figures
  2. Fortune (Jeremy Kahn) · OpenAI touts Codex growth — 4 March 2026: 1.6M weekly active users; Thibault Sottiaux: Codex is “becoming the standard agent”; “if we manage to sandbox it properly… bring the power of coding agents to billions of users”
  3. Constellation Research (Larry Dignan) · OpenAI touts broadening Codex usage — 2 June 2026: more than 5M weekly active Codex users, a fifth of them knowledge workers; company figures
  4. The Deep View (Jason Hiner) · OpenAI's secret weapon underneath Codex — 29 July 2026: Joe Gershenson — “the harness is how the model interacts with the world and how we are able to express the model capabilities”
  5. Knizhny kub · Building Codex with Tibo Sottiaux (The Pragmatic Engineer) — 12 September 2026: the Rust core of Codex is separated from its interfaces; “part of the harness must be able to die” — the reminder to run tests moves into the model; “over a weekend you can attach a hundred agents to a project”
  6. TechRadar via Yahoo · OpenAI says it built an automated research intern — 10 September 2026: OpenAI announced an “automated research intern” and admitted it is “still learning how to measure it”; a self-declaration
  7. Gemini CLI · notice on the transition to Antigravity CLI — 18 September 2026: “unpaid tier and Google One users: Gemini CLI was replaced by Antigravity CLI on June 18th, 2026”; the repository stays under Apache-2.0
  8. Google Developers Blog · Transitioning Gemini CLI to Antigravity CLI — 19 May 2026: Antigravity CLI in Go, closed source, shares the harness with Antigravity 2.0 and supports multiple agents
  9. TNW · Cursor raises at a $50B valuation — 18 April 2026: $2B in annualized revenue (per Bloomberg, February 2026); later reports say Cursor has been a SpaceX subsidiary since August 2026; the primary deal document was not opened
  10. Factory · News: fundraise — 15 September 2026: $200M at a $5B valuation for “self-improving software development in the enterprise”; a company statement
  11. GitHub (Kyle Daigle) · Welcome home, agents (Agent HQ) — 28 October 2025: agents from Anthropic, OpenAI, Google, Cognition and xAI inside Copilot; a single mission control and an enterprise control plane
  12. GitHub changelog · Agent Plugins 1.0 — 12 August 2026: an open plugin standard (skills plus MCP servers in one package), “governed independently of any single vendor”; partners AWS, Anysphere, Microsoft, OpenAI, Vercel, Google
  13. MarkTechPost · Moonshot AI releases Kimi K2.6 — 20 April 2026: a swarm of up to 300 subagents, 4,000 coordinated steps, a 13-hour autonomous run; a vendor claim
  14. DeepSeek · DeepSeek-V4 release notes — 24 April 2026: “open-source SOTA in agentic coding”, a million-token default window, integrations with Claude Code, OpenClaw and OpenCode; a vendor claim
  15. Knizhny kub · OpenCode: how an open harness became a marketplace of models — 27 July 2026: per the CEO — about 13M monthly active users in June, 20× the start of the year; the server with the agent loop embeds separately from the interface; company claims
  16. Menlo Ventures · 2025 Mid-Year LLM Market Update — 31 July 2025: enterprise usage share — Anthropic 32%, OpenAI 25%, Google 20%; Claude — 42% of code generation; the fund's sample

Forecasts and adoption

  1. METR · Time Horizon 1.1 — 29 January 2026: 228 tasks, 31 longer than eight hours; Opus 4.5's 50% horizon — 320 minutes (170 to 729); doubling every 196.5 days long-run and 88.6 days since 2024; some models “somewhat sensitive to scaffold”
  2. Gartner · Over 40% of agentic AI projects will be canceled by end of 2027 — 25 June 2025: the analysts' forecast — over 40% of agentic AI projects cancelled by end of 2027; the page blocks automated fetches, cited via mirrors
  3. JetBrains Research · AI coding agent adoption 2026 — August 2026: more than 15,000 professionals, May–July; 90% use agents at least weekly, 68% daily; Claude Code 39%, Copilot 21%, Codex 16%, Cursor 12%, OpenCode 7%, Antigravity 6%; self-reported
  4. JetBrains Research · How much code do developers really let agents write? — August 2026: about 47% of code is fully written by agents, 38% with AI assistance, 27% manually; more than half of developers write less than 20% of their code by hand; self-reported
  5. Google Cloud · Announcing the 2025 DORA report — 23 September 2025: nearly 5,000 respondents; 90% use AI, more than 80% believe it raised their productivity, 30% report little or no trust in the code
  6. Stack Overflow (Eira May) · Closing the developer-AI trust gap — 18 February 2026: in 2025, 84% use or plan to use, 29% trust — 11 pp less than a year earlier; the 2026 survey opened on 23 June, results not published as of the check date
  7. Dwarkesh Podcast · Andrej Karpathy — 17 October 2025: “it's the decade of agents”; the blockers — continual learning, multimodality, computer use
  8. Dario Amodei · The Adolescence of Technology — January 2026: powerful AI “plausibly within 1–2 years”; AI operating autonomously on multi-day tasks; “AI is now writing much of the code at Anthropic”

The Russian contour

  1. Habr (Sber) · GigaConf 2026 announcement — 10 September 2026: an invitation-only business day on 23 September and an open engineering day on 16 October 2026 at Sber's office; the engineering tracks include “Harness («обвязка»)”
  2. GigaConf 2026 · track programme — 18 September 2026: the “Harness — how to make AI generation predictable and manageable: specifications, context, policies, workflow” track; talk titles not yet published
  3. Habr (Sber, Sergey Trashchenkov) · How we tuned the harness for the model — 25 August 2026: a GigaChat profile for Deep Agents — success 77.2 → 87.0% on 391 tasks, tokens −50%; “models look alike from the outside, but their habits differ”; a vendor claim
  4. ai-forever · harness-bench-fast — 18 September 2026: an open harness benchmark — tasks on files, memory, search and CSV parsing with golden answers and no model judge, about a 20-minute run
  5. Habr (Zakhar Kopanitsky) · What an agent harness is — 8 September 2026: “there is no exact short Russian word for the term, ‘обвязка’ is the closest equivalent”; “changing the model changes probabilities, the harness changes failure classes”
  6. 3DNews · GigaCode Desktop — 15 July 2026: a team of role agents — developer, analyst, tester, manager, security — under “AI-Disrupt PDLC”; a vendor claim
  7. Habr · Yandex SourceCraft CLI — 22 April 2026: a terminal agent with built-in OpenCode, skills and PR creation; performs a task “autonomously in the background”
  8. CNews · T-Bank opens trading to clients' AI agents — 10 September 2026: T-Investments opened trading to external agents via MCP — the harness becomes a client of a financial service
  9. Habr (Ontico) · Saint HighLoad++ 2026: the AI talks — 1 July 2026: Sber — “AI is not a junior but a senior with amnesia”; Raiffeisen — 25% of the backlog via agents, 4% fully autonomous; “Harness on steroids” — $300 → $30 per task; L0–L4 levels; the author's own “State of AI4SDLC” talk; speakers' self-reports
  10. OpenAI · DevDay 2026 — 18 September 2026: the conference on 29 September in San Francisco with a livestreamed keynote; not included in this article's snapshot
Next

Related reading

This article continues the analysis of the co-evolving stack and the agent stack configurations: those cover the layers and the boundaries, this one covers what happened to the harness itself over four years and where it is going. The slides from the Loop Engineering livestream cover the part that is only touched on briefly here: how to design a loop that knows how to stop.