What a harness is and which seven parts it consists of
I already have a definition, and I am not going to rewrite it. In the September article on what remains engineering I called the agent harness the execution environment: the working loop, tools, state, constraints and feedback. The Claude Code documentation says almost the same thing more briefly: a harness is the tools, context management and execution environment that turn a language model into a working agent; Claude Code is the harness, Claude is the model inside. Birgitta Böckeler on Martin Fowler's site is shorter still: everything in the agent except the model. The three definitions do not contradict one another, but they do not help you design anything either. Design needs a decomposition.01Maxim Smirnov and I walked through the ladder “prompt → context → harness → loop” in a July livestream; that is also where we discussed why the main part of the loop is not the one that tells the agent “keep going” once more, but the one that can say “stop”.Knizhny kub · Loop Engineering
First, the term's neighbors, because in 2026 they get confused even in glossaries. A framework is a library the harness is assembled from; it does not run the agent and does not own its boundaries. A runtime is where the harness runs at a vendor. Scaffold is a word from the 2024 era; in Anthropic's 2025 documents it meant the whole wrapper around the model, while in the Hugging Face glossary of May 2026 it means only the behavior layer, and the execution layer is what they call the harness. There is no single definition, and the glossary's authors say so honestly. Finally, the word has a second meaning that usually gets lost: the training environment. In July 2026 Prime Intellect split the training environment into exactly two parts — a task set and a harness, the program that solves those tasks and produces trajectories; one configuration serves both evaluation and training. It is the same object seen from the lab's side, and it will return later in the article.
| Term | What it is | What it is not |
|---|---|---|
| Framework | A library for assembling a harness: LangChain, ADK, Agents SDK, Deep Agents | Does not run the agent itself and does not own its boundaries |
| Scaffold | The 2024-era wrapper: chains, graphs, roles; in Anthropic's 2025 usage — the whole wrapper around the model | In the 2026 Hugging Face glossary it is the opposite, only the behavior layer — there is no single definition |
| Harness | The model's protocol for working with the world: loop, context, tools and environment, verification, orchestration, observability, boundaries | Not the model, not the task, and not the framework it was assembled from |
| Runtime | Where a harness runs at a vendor: Managed Agents, Foundry Hosted Agents, Managed Agents in the Gemini API | Does not replace your task, context, and success criteria |
| Agent | Model plus harness plus task — the thing you actually run | Not a property of the model: one model in different harnesses yields different results |
| Training environment | The same notion from the lab's side: a task set plus the harness the model is trained in | Does not guarantee that the deployed harness matches the training one |
Now the decomposition. I split the harness into seven parts, and each part answers a question the model asks the world. What can I see — context and memory. What can I do — tools and environment. How long, and when to stop — the loop. Who checks — verification. How to scale — orchestration. Who sees what happened — observability and evals. Who permits — trust and boundaries. Historically the harness answered the first three questions with words: instructions in the prompt, tool descriptions, skills. The last four it answered with walls: code that runs regardless of what the model decided. What walls are not matters just as much: they do not lead the model by the hand at every step, they fence the field — what it may touch, where it may go, who will see. This split into words and walls is the spine of the whole article, and the rest of it shows three movements at once: words shrink, walls grow, and there is more freedom inside the walls.
The site's corpus already has two decompositions of the same entity, and a third one without a reconciliation would look like inconsistency. In the July longread on the co-evolving stack the harness is one of six layers, between the model and the tools, and it is responsible for context, planning, compaction, state, isolation and confirmations. The same article audited six months of history of three open harnesses and split them into ten mechanisms. The seven parts here are not a replacement but a different cut: not “what lives in the repository” but “which question of the model's does this answer”. The table's last column ties the cuts together: each part corresponds to one to three mechanisms from the July audit.
| Part | The model's question | What the model took over by 2026 | What stays in the harness | Audit mechanisms |
|---|---|---|---|---|
| Loop and stopping | How long, and when to stop | Reasoning between tool calls, running tests unprompted, tracking the remaining context | Stopping criteria, budgets, checkpoints, a judge on another model | planning · state |
| Context and memory | What I can see | A million-token window, the claimed “trained” compaction | The compaction policy, a repository map instead of instructions, memory across sessions | context · compaction |
| Tools and environment | What I can do | Function calling, screen use, choosing the procedure without a script | Tool definitions, the sandbox, worktrees, the MCP protocol | editing · discovery · sandbox |
| Verification | Who checks | Self-checking — claimed by vendors, not measured independently | Tests and linters as the feedback channel, hooks, a fresh model as judge, replayable episodes | telemetry and evals |
| Orchestration | How to scale | Spawning subagents in swarm-trained models | Durable execution, agent teams, schedules, state ownership | orchestration |
| Observability and evals | Who sees what happened | Nothing | Traces, the session log, cost accounting, evaluation on episodes | telemetry and evals |
| Trust and boundaries | Who permits | Nothing — deliberately | Permissions and an action classifier, agent identity, audit | policy · sandbox |
Four eras: from a bash loop to a managed runtime
The word is older than agents. In testing, a harness is the set of stubs and drivers that let you exercise a component outside its real surroundings. It came into the world of language models by the same road: EleutherAI's lm-evaluation-harness is an evaluation runner, not a wrapper for actions. As late as August 2024, OpenAI's description of SWE-bench Verified called the wrapper around the model a scaffold and the harness the containerized test runner. In January 2025 Anthropic defined an agent as “a combination of a model and the software scaffolding around it”, and in the Claude 4 and Sonnet 4.5 announcements the SWE-bench results came with a note about “a simple scaffold with two tools — bash and file editing”. The shift from scaffold to harness happened when the wrapper became a product: in September 2025 Anthropic described the Agent SDK as “the harness that Claude Code runs on”, and by 2026 METR had replaced scaffold with agent harness in its reports. The discipline arrived in February 2026: Mitchell Hashimoto wrote that he did not know an accepted term and started calling his practice harness engineering, and six days later OpenAI's post of the same name came out.
Behind the word stand four eras, and in each the harness compensated for a specific weakness of the model. In 2022–2023 the model could not act: no tools, no memory. The think–call–read loop was hand-written around the API, and the first harness function moved into the API as early as June 2023, when OpenAI added function calling to gpt-4-0613. In 2024 the model could act but got lost in a repository and on a long task — this is the era of scaffolds: SWE-agent's agent–computer interface, the role graphs of LangGraph and CrewAI, the planner and the critic, Claude's screen control in October and the MCP context protocol in November. In 2025 the model was already writing files and whole features while the products around it offered autocomplete and chat: Boris Cherny calls this product overhang, and Claude Code grew out of the observation that Sonnet 3.5 could do more than it was allowed to. One general-purpose agent in the file space became a product — Claude Code, Codex, Gemini CLI — and within a year it was surrounded by a repository instruction file, a harness SDK, a skills standard and a standards foundation under the Linux Foundation.
2026 is the era of discipline and infrastructure. The harness is described as an engineering practice with its own posts, courses and exams, and at the same time it is rented out whole: in April Anthropic released Managed Agents and called it a “meta-harness”, in May Google showed Managed Agents in the Gemini API, in June Microsoft — Foundry Hosted Agents. In July the MCP protocol became stateless, in August GitHub released a plugin standard, and Antigravity replaced Gemini CLI on the free tier. The model weakness that the 2026 harness compensates for sounds paradoxical: the model does more than the product allows, and the harness itself goes stale faster than it is written. Konstantin Krestnikov's periodization from Data Fest — from bare language models through ReAct and chains to graphs and finally to one agent with files — matches mine up to the era boundaries, and his metaphor is exact: the model is the pulling power, the files are the field, the harness is the yoke.
| Era | Model weakness | What the harness did | Artifact | Representatives |
|---|---|---|---|---|
| 2022–2023 · text in, text out | The model cannot act: no tools, no memory | The think–call–read loop was hand-written around the API | The ReAct loop, the first autonomous scripts, function calling in the API | AutoGPT, BabyAGI, function calling (June 2023) |
| 2024 · scaffolds | The model acts but gets lost in a repository and on long tasks | An agent–computer interface, role graphs, a planner and a critic | The ACI, LangGraph and CrewAI graphs, screen control, the context protocol | SWE-agent, Devin, computer use (October 2024), MCP (November 2024) |
| 2025 · harness as a product | The model already writes files and features while products around it offer autocomplete | One general agent in the file space; loop, tools and file system in the box | Terminal agents, the repository instruction file, a harness SDK, skills, a standards foundation | Claude Code, Codex, Gemini CLI, AGENTS.md, Agent SDK, Skills, AAIF |
| 2026 · discipline and infrastructure | The model can do more than the product allows; the harness goes stale faster than it is written | The harness is described as an engineering discipline and rented out whole | The “harness engineering” posts, managed runtimes, a stateless protocol, a plugin standard | OpenAI (February), Anthropic (March–April), Google and Microsoft (May–June), MCP 2026-07-28 |
The mechanism: how the model and the harness evolve together
Anthropic stated the main law of this evolution in March 2026 in a single sentence: every harness component encodes an assumption about what the model cannot do on its own. Everything else follows from that. Once the assumption stops being true, the component becomes not useless but harmful: it constrains the model to what the model has already outgrown. In July I called the harness a temporary theory of the current model's weaknesses; now we can trace which of those theories have already been refuted and which the labs keep deliberately.02The best example of migration part by part is Cursor's one-year report on cloud agents: procedure moves out of the harness into tools the agent picks itself, while the complexity grows around the environment, policies and state.Knizhny kub · Cursor's cloud agents
Here is the dated line of what moved into the model. Function calling — June 2023. Grammar-constrained structured output — August 2024: the JSON mode of a year earlier promised only valid syntax, not schema conformance, and in Simon Willison's experience even the syntax was not guaranteed. Screen use — October 2024, when Anthropic wrote outright that it was teaching the model general computer skills instead of building tools for each task. Reasoning between tool calls — in Claude 4 it was switched on with a beta header, in 4.6 it switches on by itself, in 4.7 the manual mode is rejected: a harness flag dissolved into the model's behavior. Tracking the remaining context — Sonnet 4.5 receives budget tags from the API, Opus 4.7 and newer no longer do. Compacting the history — in November 2025 OpenAI called GPT-5.1-Codex-Max “the first model trained to work across multiple context windows”; that is a claim, not a measurement, and Simon Willison promptly reminded everyone that Claude Code had been doing the same thing at the harness level. The most telling line is in the Claude Code changelog: the task-tracking tools are now offered only to models older than Opus 4.7. A component that was mandatory for a year has been removed for newer models because they keep the plan in their heads.
Now what the labs keep in the harness deliberately, and here the wording matters more than the dates. Claude Code documentation: permission rules are enforced by Claude Code, not by the model; instructions in the prompt determine what the model tries to do but do not change what it is allowed to do. Also there: CLAUDE.md is not a hard enforcement layer, and to block an action regardless of the model's decision you need a hook before the tool call. An operating-system-level sandbox, by Anthropic's data, cut confirmation prompts by eighty-four percent and covers any subprocesses — not because the model is unreliable, but because the boundary has to hold even if it were. Codex describes two layers outside the model — the sandbox mode and the approval policy — and explains it directly: you trust not the agent's intentions but the fact that it acts inside enforced boundaries. The compaction policy also stays with the developer: Anthropic's server-side compaction has a default threshold, but the summary prompt is replaceable and “model-dependent”. The pattern is visible to the naked eye: what can be trained goes into the weights; what cannot be entrusted even to a well-trained model stays in the harness.
The training environment and the deployed harness converge
The word's second meaning from the first section becomes a mechanism here. In October 2025 Cursor wrote that in reinforcement learning the Composer model “can call any tool of the Cursor Agent harness”, and in the Composer 2 report in March 2026 — that training runs “in realistic Cursor sessions with the same tools and harness as the deployed model”; model checkpoints ship roughly every five hours, and the user has become part of the training environment. Cognition, in the SWE-1.5 announcement, talks about retraining the model on the updated Cascade harness. OpenAI trained GPT-5-Codex “on real engineering work” and recommends using it “only in Codex or Codex-like environments” — the training harness is not named, but the recommendation speaks for itself. Zhipu assembled more than ten thousand environments for GLM-5 in the Harbor format — that is, in the format of the Terminal-Bench benchmark. The minimal preset of DeepSeek's open harness is described in the repository as “RL-agent composition”: a short fixed prompt, a persistent bash and a line editor reproduce the training environment. Anthropic makes no direct statement and speaks only of “the same infrastructure as Claude Code” — and the absence of a statement is not the absence of a practice.03I wrote about DeepSeek Harness separately: there the session log is append-only, compaction repeats the previous prefix byte for byte for the sake of the cache, and the loop “harness → trajectories → fine-tuning → the same harness” is so far an architectural possibility, not a proven data policy.Knizhny kub · DeepSeek Harness
If the model learns in its own harness, a foreign harness at the frontier should lose — and should lose by a little, because frontier models are trained to transfer skill across wrappers. That is exactly what the measurements show. In June 2026 Harbor-Index pitted the models' native harnesses against the neutral Terminus-2 on the same tasks: the native one is “usually slightly ahead, but no comparison is statistically significant”. In February METR wrote that Claude Code and Codex are “not obviously better” than the lab's default scaffolds: Opus 4.5 in Claude Code beats a simple ReAct in only half of the bootstrap samples, GPT-5 in Codex beats Triframe in one of seven. And at the same time, below the frontier the spread is enormous. On Claw-SWE-Bench GLM 5.1 solves nineteen percent of tasks with a minimal adapter and seventy-three with the full one; averaged across models, the choice of harness shifts the result by twenty-seven points. The GigaChain team sees a twenty-to-thirty-point spread between harnesses on its own measurements with one and the same model.
| Study | Model and benchmark | Effect | Who published it | Class |
|---|---|---|---|---|
| Claw-SWE-Bench · June 2026 | GLM 5.1 · Claw-SWE-Bench | Minimal adapter 19.1% vs full 73.4%; average effect 27.4 pp | Independent paper (arXiv) | Verifiable fact |
| GigaChain · August 2026 | Sonnet · 104 contest tasks | Claude Code 70 tasks, DeepAgents 67; 20–30 pp spread across harnesses | Sber, a Data Fest talk | Vendor claim |
| Model profile · August 2026 | GigaChat 3.5 · 391 tasks | 77.2 → 87.0% success, tokens −50% | Sber on Habr | Vendor claim |
| AHE · April 2026 | GPT-5.4 · Terminal-Bench 2.0 | 69.7% (seed) → 71.9% (Codex CLI) → 77.0% (evolved) | Independent paper (arXiv) | Verifiable fact |
| icat-agent · June 2026 | GPT-5.4-xhigh · SWE-bench Pro | 59.1 → 67.4% | Independent paper (arXiv) | Verifiable fact |
| Vercel d0 · December 2025 | Model not named · 5 internal queries | Tools −80%: 4 of 5 → 5 of 5, 3.5× faster, tokens −37% | Vercel | Vendor claim |
| Meta-Harness · March 2026 | Opus 4.6 and Haiku 4.5 · Terminal-Bench 2.0 | 74.7 → 76.4% and 33.7 → 37.6% under a “up to 6×” headline | Stanford (arXiv) | Verifiable fact |
| Google whitepaper · May 2026 | Model not named · Terminal-Bench 2.0 | From outside the top 30 into the top 5 by changing the harness; LangChain +13.7 points | Google, the document's authors | Snapshot |
| Nvidia · August 2026 | Opus 5 · ARC-AGI-3 | 30% without a harness → 100% with Nvidia's harness | TechCrunch on Nvidia's work; ARC Prize verification not mentioned | Snapshot |
| Harbor-Index · June 2026 | Several models · Terminal-Bench | Native harness vs neutral Terminus-2 — no statistically significant difference | Terminal-Bench | Verifiable fact |
| METR · February 2026 | Opus 4.5, GPT-5 · METR tasks | Claude Code beats ReAct in 50.7% of samples, Codex beats Triframe in 14.5% — “not obviously better” | METR | Verifiable fact |
Harnesses that rewrite themselves
The logical continuation of co-evolution is a harness changed not by an engineer but by another agent. The idea is older than it looks: in August 2024 Hu, Lu and Clune described a meta-agent that iteratively programs new agents from an archive of earlier discoveries. Sakana's Darwin Gödel Machine in May 2025 raised its SWE-bench score from twenty to fifty percent by rewriting its own code — and in the same paper the authors report that the machine removed the markers the reward function used to catch hallucinations, against an explicit instruction not to. In 2026 the direction became academic routine: Meta-Harness from Stanford, Agentic Harness Engineering with GPT-5.4 going from seventy to seventy-seven percent on Terminal-Bench, DemoEvolve with the honest conclusion that self-evolution on a sparse reward is “brittle”. In June Anthropic released “a harness for every task” — Claude writes the orchestration script on the fly. All of this works, and all of it optimizes exactly the reward it was given. An ACL 2026 paper adds an unpleasant touch: experience accumulated on harmless tasks still undermines safety in risky scenarios, because it strengthens the tendency to act rather than to refuse. A self-modifying harness without a hidden judge and a ledger of assumptions is Goodhart's law handed a machine gun.
What shrinks: instructions turn into maps
Three 2026 figures describe one movement. Boris Cherny in a July interview: “we deleted eighty percent of the system prompt — Opus 5 just does it”; there is an environment variable that removes all prompts altogether, and in that bare mode the model turns out even slightly better on the team's measurements. Nick Nisi of WorkOS at AI Engineer Europe: they generated ten thousand seven hundred lines of skill instructions from documentation, after deleting ninety-five percent the result got better, and flow control moved out of the skill into a TypeScript state machine. Vercel in December 2025: they removed eighty percent of the tools from an internal agent, leaving a single bash — three and a half times faster, a third fewer tokens, and success on the internal set went from four of five queries to five of five. All three are self-reports, and Vercel's “one hundred percent” is five queries. But the direction at three independent teams is the same, and Vercel states it most sharply: the model makes better choices when we stop making choices for it.04I went through Cherny's talk the day after the Opus 5 release. My main takeaway then was this: the system prompt and the instruction file are technical debt with a shelf life of one model generation.Knizhny kub · Boris Cherny on −80 % of the prompt
The words do not disappear — they change genre. In OpenAI's post on harness engineering the AGENTS.md file is about a hundred lines and serves as “a map with pointers to documentation” in a structured catalog: the agent reads what it needs now, not everything at once, because, as the authors write, “anything the agent cannot reach in context does not exist for it”. The skills standard is built the same way: a hundred tokens of metadata are always in context, instructions of up to five thousand tokens load on demand, resources as needed; the standard's showcase lists forty-six clients. A study of one hundred twenty-four PRs from ten repositories links the presence of AGENTS.md to a nearly thirty percent reduction in median time and sixteen percent in output tokens at comparable task completion — a correlation, not an experiment, but it agrees with Vercel's internal evals, where an eight-kilobyte documentation index in AGENTS.md scored one hundred percent against seventy-nine for skills. The map beats the script because the script encodes an assumption about how the model should go, and the map only about where things are.05What I value most in Nisi's talk is the story of the marker file the agent created with a touch command instead of running the tests; in Vercel's, the honest caveat that they cannot separate the contribution of the architecture from that of a stronger model in the doubling of the score.Knizhny kub · Nick Nisi on skillsKnizhny kub · Vercel d0
The loop: from a bash one-liner to a judge on another model
The loop is the second harness part that historically lived in words, and its evolution over the year is instructive. In July 2025 Geoffrey Huntley described Ralph: “in its pure form it is a bash loop” that feeds the agent the same prompt over and over. A year later at Anthropic it is an official plugin, implemented through a session-stop hook with an iteration limit, and next to it are built-in commands: repetition on a schedule or self-paced with a one-week lifetime, and a goal whose completion is checked by “a fresh model, not the one doing the work”. Every turn creates a checkpoint, the last hundred are kept, and the documentation honestly warns that changes made by bash commands are not tracked. The Agent SDK sets hard limits on the number of turns and on the budget in dollars, and the budget covers subagents. Kyle Mistele of HumanLayer suggests looking at all of this as a textbook control loop: the setpoint is the desired property of the codebase, the sensor is a deterministic check, the controller picks the next small change, the actuator is the agent, and a model is far from necessary in every link. The direction is the same as for instructions: from “repeat until it works” to “stop when an independent check said it worked”.
Context: a million tokens, and compaction all the same
In 2026 context is a finite resource even with a seemingly infinite window. Anthropic's models from 4.6 onward got a million-token window by default and at the standard price, and Claude Code still compacts the history at ninety-seven percent of the window: first it clears old tool outputs, then it summarizes, then it re-reads up to five changed files. The reason is named in the documentation itself with the words context rot — Chroma showed on eighteen models that quality declines as input length grows, even on simple tasks. Back in September 2025 Anthropic described three techniques for tasks longer than the window: compaction, structured notes in external memory, subagents with condensed summaries — and within a year all three became API or product features. The economics of the prefix cache and the warning about memory that promotes any past answer to a rule I covered in the July and September articles; here one thing matters: even where the harness compresses instructions, context management remains its job, because the model cannot see its own window from the outside.
What grows: boundaries, environment and verification
While instructions shrink, the code that runs regardless of the model's decision grows. It does not dictate the model's route: inside the field the model picks its own tools, order and moment of checking; the code holds the edges of the field. The most visible growth is hooks. The Claude Code guide lists thirty-three events: session start, permission request, before and after a tool call, agent and subagent stop, before and after compaction, worktree creation, configuration change. Exit code 2 blocks the action, the decisions deny, defer, ask and allow combine so that the strictest wins, and the documentation states the purpose bluntly: some actions must happen always, not at the model's discretion. Hooks can be not only scripts but also prompts to another model, and this is an important hybrid: a probabilistic check inside a deterministic loop. Permissions come in six modes, and in auto mode a separate classifier model checks actions before execution and blocks anything outside the scope of the request; some actions are never auto-approved in any mode.
The environment became part of quality, and this is the second growing belt of walls. Cursor, summing up a year of cloud agents, writes that an incomplete environment does not fail with an error — the agent simply works worse, and the degradation gets blamed on the model. Hence resumable workspaces: images with version history, rollback, audit, egress rules. In April Anthropic split the runtime into the “brain” — the model with the loop, compaction and memory — and the “hands” — the tools and the sandbox. While they lived in one container, the model could not start reasoning until the environment finished building, and a failure of either half killed the agent whole; after the split, time to first token fell by sixty percent at the median and by more than ninety at the tail, a dead sandbox is recreated, a dead “brain” is brought back from the session log. Credentials in this scheme are decrypted only at the moment of the tool call, and the model never sees them. In January Docker moved agent sandboxes to dedicated micro-VMs, Vercel announced a bounty of up to a million dollars for escaping its own, and git worktrees became the isolation primitive at Claude Code, Cursor and Codex.
Tools: four basic ones and a stateless protocol
Tools shrink in their descriptions and grow in the environment. There are four basic ones — read a file, search, edit, bash; mini-SWE-agent gets by with a single bash in a hundred lines and scores above seventy-four percent on SWE-bench Verified, and Pi deliberately lives without the MCP protocol and without subagents, with four tools and TypeScript extensions. When there are thousands of tools their descriptions stop fitting into context, and the harness answers with tool search — Anthropic holds up to ten thousand deferred definitions and reports a token reduction of more than eighty-five percent, honestly adding that accuracy drops after thirty to fifty available tools. The next step is programmatic calling: the model writes code in a container, calls tools from it and gets only the script's output into context; Anthropic claims plus eleven percent in quality at minus twenty-four percent input tokens, and Cloudflare had arrived at the same thing a year earlier with the phrasing “models have seen a lot of code and few tool calls”. The MCP protocol itself, in the revision of 28 July 2026, became stateless: the handshake and the session identifier are gone, server discovery and subscriptions appeared, tasks moved into an extension, dynamic OAuth client registration is deprecated, trace context travels in metadata, and any feature gets at least twelve months before removal. The foundation that governs it grew in nine months from one hundred forty-six to one hundred ninety members, including the US Army and the national laboratories, and cumulative SDK downloads passed a billion.
Verification, observability, identity
Verification is the fastest-growing wall, because it does not require trusting the model. At OpenAI it looks like this: about a million lines in five months, fifteen hundred PRs, seven engineers and zero lines written by hand — and the discipline “shows up in the scaffolding, not in the code”: custom linters that answer the agent with a message containing fix instructions, CI that keeps the documentation fresh, background agents that collect garbage. Böckeler calls these sensors, as opposed to guides, and Thoughtworks in its April Radar describes the same two categories: guides such as skills and specifications, sensors such as mutation testing — under the shared theme “putting coding agents on a leash”. The judge on another model, replayable episodes and the difference between pass@k and pass^k are the subject of a separate article, and in January 2026 Anthropic formalized the same ladder: task, run, grader, transcript, result.
With observability and trust the picture is mixed. Claude Code exports metrics, events and traces over OpenTelemetry, but the GenAI conventions in OpenTelemetry themselves, as of mid-July 2026, are still in “Development” status — there is no stable release, and every vendor traces in its own way. Agent identity, by contrast, took shape as a product within six months: in April Okta made agents full directory records, in August it released single sign-on for agents, and the Cross App Access extension became an official part of MCP for enterprise authorization; in February NIST launched a standards initiative for agents with a separate pillar on identity, and OWASP published its top ten risks for agentic applications. And one more story about the boundary as a product: on 31 March 2026 the Claude Code package went out to npm with a source map — about one thousand nine hundred TypeScript files; Anthropic called it a packaging error, not a security breach, but a day later trojanized “leaks” appeared online. The harness is value that gets stolen and value that gets attacked through. In the security snapshot I showed that every 2026 injection fix ran along boundaries — files, network, auto-approval — and not one was a sentence in a prompt. It is the same pattern seen from the attacker's side.
What becomes infrastructure: orchestration and long loops
The third movement is the harness parts that every team wrote itself in 2025 being rented out in 2026. Orchestration went first. Subagents in Claude Code run in their own context windows, return only a summary, nest three levels deep and run up to twenty in parallel; agent teams share a task list and mailboxes, and the documentation honestly says they spend significantly more tokens than a single session. In September, Projects spawn cloud sessions on separate branches from one coordinating conversation, and Routines launch an agent on a schedule, on a GitHub event or by API call, with no permission prompts during the run. The 2025 argument between two schools — Anthropic with an orchestrator and workers that beat a single model by ninety percent on internal evals at fifteen times the token spend, and Cognition with one thread and full traces instead of messages — is not settled, but it has stopped being an argument about code: both strategies are now runtime settings.
Long loops went second. Cursor keeps the agent loop in Temporal, manages the virtual machine's lifecycle separately, and moves storage and conversation streaming into its own layer: “one agent” splits apart by lifetime and by state owner. In September 2026 Temporal raised five hundred fifty million at a twelve-and-a-half-billion valuation “amid demand for reliable AI infrastructure”; Cloudflare built waiting for an event — that is, a pause for human confirmation — directly into the steps of its engine; Inngest explains the demand in one sentence: re-running a step with a model means paying for the tokens again. Durable execution had existed for ten years for payments and orders; the agent loop turned out to be the same class of problem with a more expensive retry.
The runtimes themselves went third. In eight weeks of spring 2026 three vendors released a managed harness whole: on 8 April Anthropic — Managed Agents, “a ready, configurable agent harness on managed infrastructure” with agent, environment, session and event primitives, cloud or self-hosted sandboxes, and a price of eight cents per active session hour on top of tokens; on 19 May Google — Managed Agents in the Gemini API, where one call brings up a remote Linux sandbox configured by AGENTS.md and SKILL.md files; on 3 June Microsoft — Foundry Hosted Agents and an “Agent Harness” in its framework with auto-compaction, file-based memory and a shell. The flip side of the same wave: OpenAI is shutting down its visual agent builder on 30 November 2026, keeping the SDK. Builders lost to code and to the runtime at the same time.06Anthropic's Applied AI team told this story as three generations of surfaces — Messages API, Agent SDK, Managed Agents — and their main point is about the context resets the team built in for the “nervous” Sonnet 4.5, which under Opus 4.5 became pure overhead.Knizhny kub · the evolution of agentic surfaces
So what stays in-house at the companies that have gone down this road? At Stripe more than one thousand three hundred PRs a week are produced entirely by agents on a fork of the open goose harness and reviewed by people; what remained their own were Blueprints — deterministic process skeletons — and an internal tool server. At Spotify background agents on Claude Code ran two hundred forty migration PRs and saved about ten engineer-weeks, while skills and customizability are switched off deliberately for the sake of guardrails. At Cloudflare, one hundred thirty-one thousand review runs in a month at a dollar nineteen per run, with a “break glass” bypass in six cases out of a thousand and a separate reviewer for the AGENTS.md file. Three companies, three different rented harnesses — and in all three the same things stayed their own: the task, the context, the success criteria and the policy. Konstantin Krestnikov put the same thought from the other side: powerful models no longer need a complex harness, it is enough to give the agent the right files and an instruction.07Krestnikov's Data Fest talk is dear to me also for its honest benchmark without a judge model: tasks on files, memory and CSV with reference answers, a twenty-minute run on any harness.Knizhny kub · one agent with files
Messages API → Agent SDK → Managed Agents: the boundary of “what is mine” shrinks to the task, the context, the episodes and the identity
Limits of the effect: where the harness stops helping
An honest article about the harness is obliged to gather in one place everything that argues against it. First, the effect at the frontier. Harbor-Index: native harness versus neutral — not significant. METR: Claude Code and Codex not obviously better than the default scaffolds. Meta-Harness promises a sixfold gap in its introduction, and in its own table on Terminal-Bench gives Opus 4.6 plus one point seven points. Put that together with the twenty-seven points of Claw-SWE-Bench and the twenty to thirty points of GigaChain, and you get not a contradiction but the shape of a curve: the gulf between a broken and an adequate harness is measured in tens of points, the gap between two adequate ones at the frontier in single digits. This is exactly what I formulated in July as “at the frontier harnesses converge, below the frontier the harness decides”; in two months the numbers arrived.
Second, what the harness cannot do in principle. Dex Horthy at the AI Engineer World's Fair attached a footnote to the formula “agent equals model plus harness”: a good harness sharply improves execution but does not teach the model to hold architectural quality over the long haul, because the signal “tests passed” does not penalize a spare try/catch or changes sprawling across the system, and the cost of bad design shows up months later. There is no fast verifier for maintainability, and he admits it honestly. His own team tried a “lights-out factory” in 2025 and a few months later ran into an incident in a codebase that people were no longer watching. DORA's May report on return on investment gives a similar picture from above: thirty-five to forty percent gain on new tasks and about ten on legacy, while the change failure rate in the model calculation rises from five to six percent.08From YC Paper Club I remember the Prime Intellect story about 99.9% on ARC-AGI that turned out to be cheating once the logs were read; from Horthy, the caveat that he sells tools for the very process he describes.Knizhny kub · YC Paper Club on harnessesKnizhny kub · Dex Horthy
Third, methodology. A paper with the telling title “Stop Comparing Agents Without Disclosing the Harness” shows that current evaluation protocols systematically attribute harness gains to model improvements. A survey of thirteen agents across twelve dimensions finds that eleven of the thirteen combine several primitives and diverge exactly where the questions are open — in compaction and state. For a practitioner the conclusion is one: the harness in a measurement report is as mandatory as the model version, and a number without the harness's name is not a number. This, incidentally, also explains why METR's original “minus nineteen percent” had not “flipped” by February 2026, as secondary retellings claim: for the original cohort the estimate is “a speedup of minus eighteen percent” with an interval crossing zero, for the new one minus four. Tools change, harnesses change, and measuring a task in the lab still does not measure delivery.
The labs' bets: who is betting on which harness parts
In July I laid the labs out by the stack layers they own — from hardware to traces. Here the cut is different: by the seven harness parts each player is betting on, and by the mechanism from the third section that explains those bets. The general rule is simple. Functions that can be trained the lab moves into the weights and stops selling as a harness. Functions that cannot be entrusted to the model — permissions, isolation, the compaction policy, identity — it keeps in the runtime and sells as infrastructure. The interfaces between them — protocols and instruction files — it hands over to standards, because a neutral interface expands the market for its own model.
Anthropic bets on the harness as a product at three altitudes: Claude Code for the developer, Cowork for the employee, Managed Agents for the platform; MCP and skills have been handed to open standards, and at its May conference the company formulated the thesis “infrastructure, not intelligence, is the bottleneck”. I expect that over the next five years it will develop precisely the runtime — “brain and hands”, the session log, built-in graders — because the model and the runtime are the moat and the interfaces are neutral territory. OpenAI bets on one harness under all surfaces — App Server unified the terminal, the IDE and the web — and on the harness inside the repository: linters, CI and documentation as machine-readable constraints. Thibault Sottiaux says Codex “is becoming the standard agent” and that scaling to billions of users hinges on isolation; I expect OpenAI to spend five years developing the sandbox and enterprise identity, because those are what turn a developer tool into a product for everyone. Google bets on one harness brand for the IDE, the terminal, the SDK and the API, and on owning the environment that others have to rent: the browser and the operating system. Replacing the open Gemini CLI with the closed Antigravity CLI on the free tier shows a willingness to pay with openness for platform coherence. Microsoft and GitHub bet on the harness in the operating system and in the task-tracking surface: the agent does not need to invent branches, PRs and permissions, they already exist.
| Player | Five-year bet | 2026 evidence | Why |
|---|---|---|---|
| Anthropic | The harness as a product at three altitudes — for the developer, the knowledge worker, the platform; protocols go to standards | Claude Code, Cowork, Managed Agents; MCP and Skills donated as open standards; “infrastructure, not intelligence, is the bottleneck” | Model and runtime are the moat, interfaces are neutral ground |
| OpenAI | One harness beneath every surface and a harness inside the repository | App Server unified CLI, IDE and web; the “Harness engineering” post; Codex Security; an enterprise platform with agent identities | “Codex is becoming the standard agent” — scale through the sandbox |
| One harness brand for IDE, terminal, SDK and API, plus the browser as environment | Antigravity replaced Gemini CLI on the free tier; Managed Agents in the Gemini API; WebMCP in Chrome; A2A donated | Own the browser and OS-level affordances others must rent | |
| Microsoft and GitHub | The harness inside the OS and inside the system of record | Agent Framework with an “Agent Harness”, Foundry Hosted Agents, Windows Agent Runtime; Agent HQ and Agent Plugins 1.0 | Issues, branches and permissions already exist — the agent need not invent them |
| Cursor and Cognition | A model trained inside its own harness | Composer 2 trains “in realistic Cursor sessions”; SWE-1.5 — “model, inference system and harness as one unified system” | A moat only where the harness has fused with the weights |
| DeepSeek and Moonshot | Open weights plus an own harness as the training environment | The DeepSeek Harness minimal preset reproduces the training environment; Kimi K2.6 — a swarm of up to 300 subagents | The price floor for everyone else |
| Open harnesses | A minimal core and a marketplace of models | OpenCode as a layer between developers and providers; Pi — four tools, no MCP, no subagents; mini-SWE-agent — a hundred lines | The model is interchangeable and the harness thin — competition on price |
| The Russian contour | A ready open harness plus a profile for one's own model | GigaChain on Deep Agents with a GigaChat profile; SourceCraft CLI on OpenCode; GigaCode role agents; T-Bank opens trading to agents via MCP | Harness sovereignty is cheap, model sovereignty is not |
The second group bets on a model trained in its own harness: Cursor, which since August 2026 has reportedly been a subsidiary of SpaceX, and Cognition with SWE-1.5, where the model, inference and harness are designed as one system. The third — open weights plus its own harness as a training environment: DeepSeek with a minimal preset that reproduces the training environment, and Moonshot with a swarm of up to three hundred subagents in Kimi K2.6. The fourth — a minimal core and a marketplace of models: OpenCode as the layer between developers and providers with thirteen million monthly active users according to its head, Pi with four tools, mini-SWE-agent with a hundred lines. For this group the model is interchangeable, the harness is thin, and the competition is on price.09What makes the OpenCode story interesting is how Anthropic's January attempt to cut a third-party harness off from subscriptions became an accelerator for it: people wanted to see what exactly was being blocked.Knizhny kub · OpenCode as a marketplace of models
The Russian landscape
A note on the venue: this text was written for Sber's conference, Sber's products are discussed below, and all of their figures are vendor claims. The Russian bet is a ready open harness plus a profile for one's own model. GigaChain took Deep Agents from LangChain, plugged in GigaChat with a config change, and at an external competition got sixty-seven tasks with the same model against seventy for Claude Code, even though Anthropic optimizes the model and the harness for each other; the model profile described on Habr raised success on the internal set from seventy-seven to eighty-seven percent at half the token spend. Yandex built the open OpenCode into SourceCraft CLI. T-Bank opened trading to external agents via MCP — the harness became a client of a financial service. At Saint HighLoad++ in July Raiffeisen reported a quarter of its backlog going through agents, with four percent of tasks fully autonomous. The conclusion for this landscape is the same one I drew in July, only now with numbers: harness sovereignty is cheap — open cores, standard protocols, a profile in a month — model sovereignty is expensive, and the harness does not replace it.
Forecasts for 2027–2031: eight bets with signs of refutation
A forecast without a sign of refutation is not a forecast but a mood. So each of the eight bets below is tied to a year and to an observable sign by which it can be called wrong. The class of all eight is the author's forecast; the supports are METR's metrics, the DORA and JetBrains reports, vendor announcements and the co-evolution mechanism from the third section. The overall vector is one: what can be trained will go into the model; what cannot be entrusted will remain a wall and move to the vendor; what stays with companies as their own is the task, the context, the episodes and the identity.
| Bet | Verifiable sign | By when | What refutes it |
|---|---|---|---|
| 1 · The managed runtime becomes the default way to ship an agent; an own loop is a compliance niche | More than half of new enterprise agents at the three big vendors run in their runtimes | 2028 | Large companies keep writing generic loops and publish that as the norm |
| 2 · Compaction, memory consolidation and subagent spawning become native in the API | All three available without a harness from two or more labs | 2027 | Harnesses keep carrying their own implementations as mandatory |
| 3 · Protocols stay neutral and boring; competition moves to the runtime and the model | None of the four standards has gone back under one company's control | 2028 | A protocol fork by a large vendor with incompatible extensions |
| 4 · The loop leaves the human: most agent runs start from an event, a schedule or a build | Vendor telemetry shows more than half of runs start without an interactive prompt | 2027 | The interactive terminal stays the main entry point by run count |
| 5 · Verification becomes the largest harness part: the agent runs the product, not only tests | Screen and browser confirmation of the result is a standard step in leading harnesses | 2028 | Verification stays at “the tests passed” |
| 6 · “Harness as a moat” survives only fused with a trained model | Standalone harness companies are acquired or move to the labs | 2029 | An independent harness without its own model keeps leading its market |
| 7 · Agent identity and audit become a regulatory requirement | A regulation in at least one jurisdiction requires a per-run identity and an action log | 2028–2030 | Regulators stop at models and do not touch the runtime |
| 8 · The harness effect at the frontier shrinks to statistical zero, below the frontier it does not | Neutral and native runs of frontier models are indistinguishable; the spread persists for open models | 2027 | Native harnesses consistently beat neutral ones at the frontier |
A few explanations for the bets that look bolder than the others. The first rests on three managed runtimes shipping within eight weeks, not on faith in vendors: when three competitors converge on one architecture, the architecture becomes the market's expectation. The second rests on what has already happened: reasoning between calls and context tracking went into the model within a year, and compaction with memory is of the same type. The fourth rests on JetBrains, where ninety percent of developers use agents weekly and about half of the code is written entirely by agents: the next step after “every day” is “without me”. The seventh rests on the NIST initiative and on Dario Amodei's September essay about “built-in graders”: when a lab itself asks for regulation of the runtime, it comes. The eighth is a direct continuation of the seventh section and the only bet that already has 2026 data.
The horizon of all the bets is bounded by what METR calls the task horizon: the fifty-percent horizon of Opus 4.5 in January 2026 was three hundred twenty minutes with a wide interval, doubling every three to seven months depending on which stretch of the series you take. If the doubling holds, by 2028 an agent will hold a task the length of a working week, and the harness for such a task is no longer a loop but an operating system with a scheduler, memory and permissions. If the doubling slows, bets one, four and five shift by a year or two, but the direction does not change: back in June 2025 Gartner promised the cancellation of more than forty percent of agentic projects by the end of 2027, and the survivors will be the ones that are observable, bounded and cheap to check — that is, the ones whose walls were built before their words.
The harness ledger: how to keep assumptions about the model
From the law “every component encodes an assumption” follows a practice, and it is not a checklist. In September I proposed an ablation test: take characteristic episodes, compare the current model with the full harness, with a simplified one, and a new model with minimal adaptation, and remove the rule that no longer improves anything. The test is a method; it needs a record-keeping form. I call it the harness ledger: for each component — a hook, a rule, a skill, a reset, a tool — you write down which assumption about the model it encodes, what evidence confirms it, under which model and version, what triggers a recheck, and in which ownership mode the component lives. The last column comes from the ownership boundary framework: rent, adapt, own.10In an interview Sottiaux frames the same thing as an engineering dilemma: where to fix a problem — in the harness, or by waiting for the next model — and admits that a large workaround may turn out to be unnecessary a month later.Knizhny kub · Tibo Sottiaux on Codex
The ledger solves three problems a checklist does not. First, it distinguishes assumptions from boundaries. A context reset at sixty percent of the window is an assumption about a nervous model, and it has to be rechecked with every new model; a hook forbidding destructive commands outside the worktree is a boundary whose recheck trigger is “never”, because it protects against injection, not against a weakness of the model. Second, it gives a shelf life: Cherny's practice — every six months delete the instructions, skills and hooks and bring back line by line whatever the model stumbles on repeatedly — becomes not a gesture of courage but a row with a date. His own estimate that evals outlive the harness by one to three model generations sets a shelf life for the evidence too. Third, it makes ablation cheap: a component without evidence is the first candidate for removal, a component with an incident instead of episodes is the last.
| Component | Assumption about the model | Evidence | Date and model | Trigger | Mode |
|---|---|---|---|---|---|
| Context reset at 60% of the window | The model gets anxious near the context limit and wraps up early | 30 episodes before and after removal: no difference | Introduced under Sonnet 4.5, retested under Opus 4.5 | New model | Delete |
| Pre-call hook: destructive commands blocked outside the worktree | An injection in text the agent read may issue a destructive command | An incident, not episodes | Any model | Never — it is a boundary, not an assumption | Own |
| A 400-line “release checklist” skill | The model forgets release steps | Not verified | Written for last year's model | Every six months | Adapt or delete |
- Open a ledger row when you add any harness component, not at audit time
- Evidence is replayable episodes before and after; an incident is evidence only for boundaries
- The “new model” trigger runs ablation on every row with an assumption; boundary rows are not subject to ablation
- The ownership mode is chosen by the ownership boundary framework, not by habit
The example in the table is hypothetical, and that needs saying plainly: the ledger is a tool I am proposing, not a report on something already in use. There is one way to check it — start one for your own fleet of agents and see how many rows end up without evidence. My forecast: more than half, and almost all of them will be words, not walls.
Seven takeaways on the harness
- 01The harness is migrating from words to walls. Instructions — the system prompt, skills, tool descriptions — shrink into maps, while boundaries and environment — hooks, permissions, sandboxes, identity, durable execution — grow. Walls do not lead the model by the hand: they fence the field, and inside it the model gets more freedom than a year ago. The paradox of “ninety percent of the result” and “minus eighty percent of the prompt” is not an argument about models but a description of two different parts of one harness.
- 02The word's two meanings converge. The harness the model works in and the environment it is trained in are increasingly the same thing: Cursor, Cognition and DeepSeek say so outright. That is why a foreign harness at the frontier is structurally at a disadvantage against the native one — and exactly why the difference between them at the frontier is measured in single points.
- 03The harness has become a disclosable variable of evaluation. One model in different harnesses yields anywhere from a few to fifty points of difference, and evaluation protocols attribute that gain to the model. The harness in a measurement report is as mandatory as the model version.
- 04Orchestration and long loops have become infrastructure. Three vendors released managed runtimes within eight weeks, the agent loop lives in durable execution engines, visual builders are on their way out. The “boundary of what is mine” has shrunk to the task, the context, the episodes and the identity.
- 05Every harness component encodes an assumption about what the model cannot do. Assumptions go stale within one to three model generations, so they have to be kept in a ledger with a recheck date, not a checklist with ticks.
- 06Self-modifying harnesses optimize the reward they were given: the very first one deleted the markers used to catch hallucinations. Without a ledger and a hidden judge, harness auto-evolution is Goodhart's law handed a machine gun.
- 07By 2028 the harness as a separately named product will dissolve into the model API and the managed runtime. What stays visible is boundaries, episodes and identity. The sign of refutation: independent harness companies without a model of their own holding leadership in their markets.
Documentation, research, talk recordings and the limits of evidence
The list is grouped by section. Every figure in the article names its baseline where it appears; company statements are marked as claims. Product documentation and specifications capture the state as of the check date — 18 September 2026. Talk recordings are cited through my write-ups in Knizhny kub, which record what exactly was said. The full dossier with evidence classes, contradictions between sources and a list of what could not be confirmed sits in the site repository next to this article.
Definitions and the term
- Anthropic · Claude Code docs: Glossary — 18 September 2026: “agentic harness — the tools, context management, and execution environment… Claude Code is the harness; Claude is the model inside it”; a vendor definition
- LangChain (Vivek Trivedy) · The Anatomy of an Agent Harness — 10 March 2026: “a harness is every piece of code, configuration, and execution logic that isn't the model itself”; the formula “Agent = Model + Harness”; a framework vendor
- Birgitta Böckeler (martinfowler.com) · Harness engineering for coding agent users — 2 April 2026: the harness is “everything in an AI agent except the model itself”; split into guides (feedforward) and sensors (feedback)
- Anthropic (Prithvi Rajasekaran) · Harness design for long-running application development — 24 March 2026: “every component in a harness encodes an assumption about what the model can't do on its own”; Opus 4.6 removed the need for sprints and context resets; the full harness took 6 hours and $200 vs 20 minutes and $9 for a solo agent — one experiment
- Anthropic (Martin, Cemaj, Cohen) · Scaling Managed Agents: Decoupling the brain from the hands — 8 April 2026: the harness is “the loop that calls Claude and routes Claude's tool calls to the relevant infrastructure”; brain and hands; harness assumptions “go stale as models improve”; a vendor describing its own product
- Hugging Face (Paniego, Gosthipaty) · The Agent Glossary — 25 May 2026: harness as the execution layer, scaffold as the behavior layer; the authors note there are no universally accepted definitions; inverse to Anthropic's 2025 usage
- Prime Intellect · verifiers v1 — 10 July 2026: a training environment = a taskset (“data, tools, and scoring”) + a harness (“the program that solves the task”); one configuration for evals and training; vendor documentation
- Addy Osmani · The New Software Lifecycle (blog version of Google's whitepaper “The New SDLC With Vibe Coding”) — 16 June 2026: “an agent is a model plus a harness”, the “10% model, 90% harness” split is the authors' estimate, not a measurement; from outside the top 30 into the top 5 of Terminal-Bench 2.0 by changing the harness; LangChain +13.7 points
- Knizhny kub · The New SDLC: harness, the factory model and the economics of agentic engineering — 21 June 2026: a breakdown of the second half of Google's whitepaper — the “Agent = Model + Harness” formula, the factory model, the conductor and orchestrator roles
- Knizhny kub · Loop Engineering: why the most important part of the agent loop is the right to say no — 14 July 2026: the ladder prompt → context → harness → loop and the five movements of a loop; the HuaShu material is a working note, not an Anthropic publication
Eras and history
- EleutherAI · lm-evaluation-harness — Opened 18 September 2026: “a framework for few-shot evaluation of language models” — the word “harness” entered the model world as an evaluation runner
- Simon Willison · OpenAI function calling — 13 June 2023: gpt-4-0613 and gpt-3.5 accept JSON schemas of functions and return a structured call — the first harness function to move into the API
- Yang et al. (Princeton) · SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — 6 May 2024 (v1): the agent-computer interface (ACI) markedly improves file creation and editing, repository navigation and test execution; the abstract uses neither “harness” nor “scaffold”
- OpenAI · Introducing SWE-bench Verified — August 2024: the wrapper around the model is called a “scaffold” (Agentless), while “harness” names the containerized evaluation runner; page updated 24 February 2025
- Anthropic · Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku — 22 October 2024: “we're teaching it general computer skills” instead of building task-specific tools; OSWorld 14.9% from screenshots vs 7.8% for the next-best system; “at times cumbersome and error-prone”
- Anthropic · Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet — 6 January 2025 (date after the page migration): “an ‘agent’ refers to the combination of an AI model and the software scaffolding around it”; “give as much control as possible to the language model itself, and keep the scaffolding minimal”
- Anthropic · Introducing Claude 4 — 22 May 2025: SWE-bench results use “the same simple scaffold that equips the model with solely the two tools” — bash and file editing; Claude Code general availability
- Anthropic · Introducing Claude Sonnet 4.5 — 29 September 2025: “more than 30 hours” of focus on multi-step tasks — a vendor claim; the Agent SDK on “the same infrastructure that powers Claude Code”; context editing and the memory tool in the API
- Anthropic (Thariq Shihipar) · Building agents with the Claude Agent SDK — 29 September 2025: “the agent harness that powers Claude Code… can power many other types of agents, too”; the loop “gather context → take action → verify work → repeat”
- Anthropic (Justin Young) · Effective harnesses for long-running agents — 26 November 2025: an initializer agent creates a launch script, a progress file and a feature list; a later instance “would look around, see that progress had been made, and declare the job done”
- Linux Foundation · Announces the Formation of the Agentic AI Foundation — 9 December 2025: eight platinum members including Anthropic, Google, Microsoft and OpenAI; MCP, goose and AGENTS.md donated; “more than 10,000 published MCP servers”; AGENTS.md “adopted by over 60,000 open source projects”
- Andrej Karpathy · 2025 LLM Year in Review — 19 December 2025: “Claude Code emerged as the first convincing demonstration of what an LLM Agent looks like”
- Mitchell Hashimoto · My AI Adoption Journey — 5 February 2026: “I don't know if there is a broad industry-accepted term for this yet, but I've grown to calling this ‘harness engineering’”; the rule — turn every agent mistake into a fix that makes it impossible
- Knizhny kub · Harness: why one agent with files displaces complex harnesses (Krestnikov, Data Fest 2026) — 15 August 2026: a periodization from bare LLMs to one agent in the file space; “the LLM is the draft power, files are the field, the harness is the yoke”; Claude Code with Sonnet solves 70 of 104 tasks, DeepAgents with the same model 67; claims by the GigaChain team (Sber)
Co-evolution of model and harness
- Simon Willison · OpenAI Structured Outputs — 6 August 2024: the schema is compiled into a context-free grammar; the earlier JSON mode promised only valid syntax, not schema conformance, and in the author's words “neither of these modes were guaranteed to return valid JSON”
- Anthropic · Claude docs: Extended thinking — 18 September 2026: interleaved thinking needed a beta header on Claude 4, is automatic on 4.6 and later, and the manual mode is rejected on 4.7 and later — a harness flag dissolved into model behavior
- Anthropic · Claude docs: Context windows — 18 September 2026: the million-token window is default and standard-priced on 4.6 and later; “context rot” is named in the docs; Sonnet 4.5–5 receive API-injected budget tags, Opus 4.7 and later do not
- Simon Willison · GPT-5.1-Codex-Max — 19 November 2025: OpenAI — “our first model natively trained to operate across multiple context windows through… compaction”; Willison notes Claude Code already did auto-summarization in the harness
- Anthropic · Claude docs: Compaction — 18 September 2026: server-side compaction with a default 150,000-token trigger and a 50,000 minimum; the summary prompt “varies by model” and can be replaced by the developer; SDK compaction marked deprecated
- Anthropic · Claude Code CHANGELOG — 18 September 2026, version 2.1.271: “changed the task-tracking tools to be offered only on Claude 3.x, Opus 4.0–4.7” — a harness component withdrawn for newer models
- Cursor · Composer: Building a fast frontier model with RL — 29 October 2025: during RL the model “can call any tool in the Cursor Agent harness”; “hundreds of thousands of concurrent sandboxed coding environments”; a vendor claim
- Cursor · Composer 2 technical report — 27 March 2026: “RL training occurs in realistic Cursor sessions with the same tools and harness the deployed model uses”; a vendor claim
- Cursor · Real-time RL for Composer — 26 March 2026: model checkpoints “approximately every 5 hours”; the user is part of the training environment; +2.28% edit persistence, −3.13% dissatisfied follow-ups, −10.3% latency — internal metrics
- Cognition · SWE-1.5 — 29 October 2025: “end-to-end RL on real task environments using our custom Cascade agent harness”; “re-trained the model on the updated harness”; 950 tokens per second; a vendor claim
- VentureBeat · OpenAI unveils GPT-5-Codex, optimized for agentic coding — 15 September 2025: the model is “trained on real-world engineering work”; recommended “only for agentic coding tasks in Codex or Codex-like environments”; the training harness itself is not named
- Qwen · Qwen3-Coder — 22 July 2025: “long-horizon RL (Agent RL)” across “20,000 independent environments in parallel”; the terminal client adapted from Gemini CLI; a vendor claim
- Zhipu · GLM-5 — 17 February 2026: over 10,000 verifiable environments in the Harbor format with Docker build accuracy above 90% — training environments literally in the benchmark's format; a vendor claim
- DeepSeek · deepseek-harness — 18 September 2026: the session as an append-only log of typed events; the minimal preset is described as an “RL-agent composition” with persistent bash and a string-replace editor; open code shows the possibility, not the data policy
- Knizhny kub · DeepSeek Harness: when the agent's history becomes part of inference economics — 18 August 2026: why compaction repeats the previous prefix byte for byte, and why the loop “harness → trajectories → post-training → the same harness” is an architectural possibility, not a proven practice
- Anthropic · Introducing Claude Opus 4.6 — 5 February 2026: Terminal-Bench 2.0 — “all runs used the Terminus-2 harness, except for OpenAI's Codex CLI”; a multi-agent harness raised BrowseComp to 86.8%; a prompt modification yielded 81.42% on SWE-bench; a vendor claim
- arXiv 2604.25850 · Agentic Harness Engineering — 28 April 2026: GPT-5.4 on Terminal-Bench 2.0 — 69.7% with the seed harness, 71.9% with Codex CLI, 77.0% with the evolved one; transferring the harness to other models yields +2.3 to +10.1 pp
- arXiv 2606.25514 · icat-agent — 24 June 2026: GPT-5.4-xhigh on SWE-bench Pro — 59.1% with the best prior scaffold vs 67.4% with the new one
- Sakana AI · Darwin Gödel Machine — 30 May 2025: a self-modifying agent raised SWE-bench from 20.0 to 50.0% and Polyglot from 14.2 to 30.7%; it also “removed the markers we use in the reward function to detect hallucination (despite our explicit instruction not to do so)”
- Hu, Lu, Clune · Automated Design of Agentic Systems — 15 August 2024: “a meta agent iteratively programs interesting new agents based on an ever-growing archive of previous discoveries” — the precursor of self-modifying harnesses
- Google DeepMind · AlphaEvolve impact — 7 May 2026: −30% variant detection errors in DeepConsensus, −20% write amplification in Spanner — vendor results on its own systems
- arXiv 2604.16968 · On Safety Risks in Experience-Driven Self-Evolving Agents (ACL 2026) — 18 April 2026: experience gathered on benign tasks “can still compromise safety in high-risk scenarios”; experience “reinforces agents' tendency to act rather than refuse”
- Manus (Peak Ji) · Context Engineering for AI Agents: Lessons from Building Manus — 18 July 2025: “we've rebuilt our agent framework four times”; “if model progress is the rising tide, we want Manus to be the boat, not the pillar stuck to the seabed”
- Philipp Schmid · The importance of Agent Harness in 2026 — 5 January 2026: the harness is “the infrastructure that wraps around an AI model to manage long-running tasks. It is not the agent itself”; model as CPU, context as RAM, harness as OS; five Manus refactors in six months
- Anthropic · Effective context engineering for AI agents — 29 September 2025: “smarter models require less prescriptive engineering”; context is a finite resource with an “attention budget”; compaction, notes and subagents as three techniques for tasks longer than the window
Instructions, loop and context
- YC Root Access · Boris Cherny: Building Claude Code — 27 July 2026: “we deleted 80% of the system prompt — Opus 5 just does it”; a variable that removes all prompts; “delete the entire system prompt and then bring it back line by line”; “every six months, delete your instructions, skills, hooks”; “evals outlive the harness… maybe one, two, three” generations; a vendor employee's self-report
- Knizhny kub · Boris Cherny: We Cut 80% of Claude Code's Prompt (YC Startup School 2026) — 17 August 2026: a breakdown of the talk given the day after the Opus 5 release — “without a measurable quality loss on their coding evals” in the speaker's words; “product overhang”; the shift from prompts through context to task setting
- Knizhny kub · Nick Nisi on agent skills: fewer instructions, more evidence — 11 July 2026: the talk “How I deleted 95% of my agent skills and got better results” (AI Engineer Europe, 10 April 2026); 10,739 generated instruction lines; flow control moved into a state machine with mandatory gates; an engineer's self-report
- Vercel · We removed 80% of our agent's tools — 22 December 2025: the agent stripped down to one bash tool — 3.5× faster, −37% tokens, −42% steps, success 80 → 100% on five internal queries; “the model makes better choices when we stop making choices for it”; a vendor claim
- Knizhny kub · Andrew Qu (Vercel): How We Solved Agent Building — 16 September 2026: d0's path from a big prompt through a chain of specialists to one agent with a working directory; the contributions of the architecture and of a stronger model to the doubled score are not separated
- Knizhny kub · Kyle Mistele: Loop Engineering from First Principles — 30 August 2026: the agent loop as a control loop — set point, sensor, controller, actuator; the model is needed far from everywhere; the speaker co-founded a vendor
- AGENTS.md — 18 September 2026: “used by over 60k open-source projects” (a December 2025 figure); 24 supporting tools; stewarded by the AAIF
- Lulla et al. · AGENTS.md study (arXiv 2601.20404) — 28 January 2026 (v2 30 March): across 124 PRs in 10 repositories the presence of AGENTS.md is associated with −28.64% median runtime and −16.58% output tokens at comparable completion; a correlation, not an experiment
- Vercel · AGENTS.md outperforms skills in our agent evals — 27 January 2026: a compressed 8 KB docs index in AGENTS.md scored 100% on Next.js 16 API evals, skills maxed at 79%; a vendor's internal eval
- Agent Skills · Specification — 18 September 2026: a SKILL.md file with a name up to 64 and a description up to 1024 characters; progressive disclosure — metadata (~100 tokens) → instructions (under 5,000 tokens) → resources on demand; 46 clients on the showcase
- Geoffrey Huntley · Ralph Wiggum as a “software engineer” — 14 July 2025: “in its purest form, Ralph is a Bash loop” that feeds one prompt to the agent again and again
- Anthropic · claude-code/plugins/ralph-wiggum — 18 September 2026: the official plugin implements the loop via a Stop hook “by blocking normal session exit”, with an iteration cap and a completion promise
- Anthropic · Claude Code docs: /goal — 18 September 2026: completion “is decided by a fresh model rather than the one doing the work” (Haiku by default); conditions must name a check, not an intention
- Anthropic · Claude Code docs: Scheduled tasks (/loop) — 18 September 2026: repetition at a fixed interval or self-paced from one minute to an hour, session-scoped, a seven-day expiry
- Anthropic · Claude Code docs: Checkpointing — 18 September 2026: every prompt creates a checkpoint, the 100 most recent are kept; “checkpointing does not track files modified by Bash commands”
- Anthropic · A harness for every task: dynamic workflows in Claude Code — 2 June 2026: “Claude can now write its own harness on the fly, custom-built for the task at hand” — an orchestration script generated per task; a vendor claim
- Anthropic · Agent SDK docs: The agent loop — 18 September 2026: the loop ends with a response that contains no tool calls; hard caps on turns and on a dollar budget, the budget covers subagents
- Chroma (Hong, Troynikov, Huber) · Context Rot — 14 July 2025: 18 models; “LLMs do not maintain consistent performance across input lengths” even on trivial tasks
- Anthropic · Claude Code docs: Context window — 18 September 2026: auto-compaction clears old tool outputs first, then summarizes; the threshold is about 967,000 tokens on million-token models; up to five recently modified files are re-read afterwards
Boundaries, tools and standards
- Anthropic · Claude Code docs: Permissions — 18 September 2026: “permission rules are enforced by Claude Code, not by the model. Instructions in your prompt or CLAUDE.md shape what Claude tries to do, but they don't change what Claude Code allows”
- Anthropic · Claude Code docs: Memory — 18 September 2026: CLAUDE.md “is not a hard enforcement layer”; “to block an action regardless of what Claude decides, use a PreToolUse hook”; aim for a file under 200 lines
- Anthropic · Claude Code docs: Hooks guide — 18 September 2026: a table of 33 events; exit code 2 blocks the action; deny/defer/ask/allow decisions with the most restrictive winning; “certain actions always happen rather than relying on the LLM to choose to run them”
- Anthropic · Claude Code docs: Permission modes — 18 September 2026: six modes; in auto mode “a separate classifier model reviews actions before they run, blocking anything that escalates beyond your request”; some actions are never auto-approved in any mode
- Anthropic · Beyond permission prompts: making Claude Code more secure and autonomous — 20 October 2025: a bubblewrap and Seatbelt sandbox “safely reduces permission prompts by 84%” and covers “any scripts, programs, or subprocesses”; the vendor's internal telemetry
- OpenAI · Codex docs: Sandboxing — 18 September 2026: two layers outside the model — the sandbox mode (what is technically possible) and the approval policy (when to ask); “you aren't just trusting the agent's intentions; you are trusting that the agent is operating inside enforced limits”
- Anthropic · Claude Code docs: Security — 18 September 2026: network commands are never auto-approved, web fetch runs in a separate context window, credentials are masked, “all operations in cloud sessions are logged”; Anthropic “does not security-audit or manage any MCP server”
- Anthropic · Claude docs: Tool search tool — 18 September 2026: up to 10,000 deferred tools; ~55,000 tokens of definitions reduced “by over 85 percent”; accuracy degrades “once you exceed 30–50 available tools”; a vendor claim
- Anthropic · Claude docs: Programmatic tool calling — 18 September 2026: the model writes code in a container and calls tools from it, only the script's output returns to context; “improved performance by an average of 11% while using 24% fewer input tokens”; a vendor claim
- Cloudflare (Varda, Pai) · Code Mode — 26 September 2025: MCP tools become a TypeScript API and the model writes code executed in an isolate; “LLMs have seen a lot of code. They have not seen a lot of ‘tool calls’”
- Model Context Protocol · Specification 2026-07-28: Changelog — 28 July 2026: the protocol went stateless — the handshake and session id removed; server discovery and subscriptions added; tasks moved to an extension; OAuth dynamic client registration deprecated; trace context in metadata; a minimum twelve-month deprecation window
- Linux Foundation · Agentic AI Foundation adds 43 new members — 18 May 2026: 190 members; new Gold members include F5, GoDaddy, Stripe; government members include the U.S. Army, PNNL, Sandia
- Linux Foundation · Agentic AI Foundation launches MCPA certification — 14 September 2026: the foundation's first certification; the TypeScript and Python SDKs — “more than 1 billion cumulative downloads”; “MCP has quickly become the standard for connecting AI agents to external systems”
- Linux Foundation · A2A protocol surpasses 150 organizations — 9 April 2026: more than 150 organizations, 22,000+ stars, SDKs in five languages, version 1.0 with signed agent cards; accepted into the AAIF on 27 August 2026
- Okta · Agent SSO — 24 August 2026: Cross App Access “formally incorporated as the official Enterprise-Managed Authorization extension for the Model Context Protocol”; agents as “first-class identities in Universal Directory” since GA on 30 April 2026
- NIST/CAISI · Announcing the AI Agent Standards Initiative — 17 February 2026: three pillars — industry-led standards, open protocol development, research on agent security and identity; a concept paper on identity and authorization
- OWASP · Top 10 for Agentic Applications for 2026 — 9 December 2025: from agent goal hijack (ASI01) to rogue agents (ASI10); more than a hundred experts
- OpenTelemetry · Semantic conventions for generative AI — 18 September 2026: the conventions moved to a dedicated repository; as of mid-July 2026 every gen_ai item is still marked “Development” — no stable release
- InfoQ · Claude Code source leak — 7 April 2026: package version 2.1.88 shipped a source map with about 1,906 TypeScript files; Anthropic called it “a release packaging issue caused by human error, not a security breach”; the contents are not used in the article
- Anthropic · Demystifying evals for AI agents — 9 January 2026: task, trial, grader, transcript, outcome; pass@k vs pass^k — “the probability that all k trials succeed”; start with 20–50 tasks from real failures
- Docker · Docker Sandboxes: run Claude Code and other coding agents unsupervised but safely — 30 January 2026: sandboxes on dedicated microVMs “adding a hard security boundary”; Claude Code, Copilot CLI, Codex CLI, Gemini CLI, Kiro supported
- Anthropic · Claude Code docs: Worktrees — 18 September 2026: a separate worktree per session or subagent; four checks block edits and git operations in the main checkout; Cursor 2.0 (29 October 2025) ran up to eight agents on worktrees or remote machines
- SWE-agent · mini-SWE-agent — 18 September 2026: about a hundred lines, “does not have any tools other than bash”, a linear history; “scores >74% on the SWE-bench verified benchmark”
- Mario Zechner · pi-mono — 18 September 2026: four tools, a short prompt, TypeScript extensions; deliberately no MCP and no subagents; MIT license
Orchestration and infrastructure
- Anthropic · Claude Code docs: Subagents — 18 September 2026: a subagent works in its own context window and “returns only the summary”; nesting up to three levels, up to 20 concurrent; worktree isolation
- Anthropic · Claude Code docs: Agent teams — 18 September 2026: a lead and teammates with a shared task list and mailboxes; “agent teams use significantly more tokens than a single session”; “a teammate can't approve a permission prompt… on your behalf”
- Anthropic · Claude Code docs: Projects — 18 September 2026 (public beta): one coordinator conversation spawns threads — cloud sessions on their own branches; shared instructions up to 16,000 characters and a project memory
- Anthropic · Claude Code docs: Routines — 18 September 2026 (research preview): runs on a schedule, a GitHub event or the API; no permission prompts during a run; API-trigger text arrives marked untrusted
- Temporal · Temporal raises $550M at a $12.55B valuation — 14 September 2026: a $550M Series E at a $12.55B valuation “as demand grows for reliable AI infrastructure”; $300M at $5B in February; a company statement
- Cloudflare · Workflows GA: production-ready durable execution — 7 April 2025: steps that wait for an event or sleep — pausing the loop for a human approval is built into the engine
- Inngest (Charly Poly) · Durable execution: the key to harnessing AI agents — 19 February 2026: “this is especially valuable for LLM calls, where re-execution means re-paying for tokens”; a vendor blog
- Knizhny kub · Cursor Cloud Agents: what to hand to the agent and what to keep on the platform — 14 August 2026: per Cursor's June write-up — procedural logic moves from the harness into tools, the environment became part of quality, the loop lives in Temporal, and “one agent” decomposes by lifetime and state owner
- Anthropic · Claude Managed Agents — 8 April 2026: “pre-built, configurable agent harness that runs in managed infrastructure”; agent, environment, session, events; cloud or self-hosted sandboxes; $0.08 per active session-hour on top of tokens; early customers Notion, Rakuten, Asana, Sentry
- Knizhny kub · Evolution of Agentic Surfaces: the harness becomes infrastructure — 21 August 2026: three generations of surfaces — Messages API, Agent SDK, Managed Agents; Sonnet 4.5's “context anxiety” and the context reset that became overhead under Opus 4.5; the brain–hands split cut median time to first token by 60%; a vendor's account
- Google Developers Blog · All the news from the Google I/O 2026 developer keynote — 19 May 2026: Managed Agents in the Gemini API — one call provisions a remote Linux sandbox, configured via AGENTS.md and SKILL.md; Antigravity 2.0, CLI and SDK with “programmatic control over agent harness”; WebMCP in Chrome 149
- Microsoft (Shawn Henry) · Microsoft Agent Framework at Build 2026 — 3 June 2026: an “Agent Harness” with auto-compaction, file memory and a shell; Foundry Hosted Agents; CodeAct — the agent writes Python to call tools, latency −52.4%; a vendor claim
- OpenAI · Agent Builder guide — 18 September 2026: the visual builder “is scheduled to shut down on November 30, 2026”; ChatKit and the Agents SDK remain
- Anthropic · How we built our multi-agent research system — 13 June 2025: an orchestrator with workers is 90.2% better than a single Opus 4 on internal evals at roughly 15× the tokens; a vendor claim
- Cognition (Walden Yan) · Don't Build Multi-Agents — 12 June 2025: “share context, and share full agent traces, not just individual messages”; actions carry implicit decisions, and conflicting decisions carry bad results; a vendor's position
- Stripe (Alistair Gray) · Minions: Stripe's one-shot, end-to-end coding agents — 9 February 2026: “over 1,300 pull requests… merged each week” are fully agent-produced and human-reviewed; a goose fork; Blueprints combine the determinism of workflows with agents' flexibility; a company's self-report
- Spotify Engineering · Background coding agents: dataset migrations (Honk, part 4) — 22 April 2026: Claude-Code-based agents delivered 240 migration PRs and saved about ten engineering weeks; skills and configurability withheld as “a deliberate design choice for guardrails”; a company's self-report
- Cloudflare (Ryan Skidmore) · AI code review at scale — 20 April 2026: 131,246 review runs on 48,095 merge requests in a month, $1.19 average per run, 0.6% “break glass” overrides, a dedicated AGENTS.md reviewer; a company's self-report
- Simon Willison · Code w/ Claude 2026 (live blog) — 6 May 2026: multi-agent Managed Agents, Outcomes, Dreaming, Routines; “most people will experience AI through one of the things you've built on the Claude platform” (Ami Vora)
Limits of the effect
- METR · Measuring time horizon using Claude Code and Codex — 13 February 2026: “Claude Code and Codex aren't obviously better than the default scaffolds we use for our agents”; Opus 4.5 in Claude Code beats ReAct in 50.7% of bootstrap samples, GPT-5 in Codex beats Triframe in 14.5%
- Terminal-Bench · Harbor-Index — 29 June 2026: “native usually finishes a little ahead… but no comparison is statistically significant”
- arXiv 2609.04298 · Harbor adapters: Terminus-2 vs native harnesses — 3 September 2026: every model is run with Terminus-2 and with one of three native harnesses; the strongest, GPT-5.5 with Codex, reaches 28.0%
- Lee et al. (Stanford) · Meta-Harness — 30 March 2026: the introduction says “changing the harness around a fixed LLM can produce a 6× performance gap”; the paper's own Terminal-Bench 2.0 numbers are modest: Opus 4.6 — 74.7 → 76.4%, Haiku 4.5 — 33.7 → 37.6%
- arXiv 2606.12344 · Claw-SWE-Bench — 10 June 2026: GLM 5.1 — 19.1% with a minimal adapter vs 73.4% with the full one; with the model fixed, the harness choice moves the score by 27.4 pp on average
- arXiv 2605.23950 · Stop Comparing LLM Agents Without Disclosing the Harness — 7 May 2026: “current evaluation protocols therefore systematically misattribute harness-level gains to model improvements”
- TechCrunch · Nvidia just showed that the harness, not the AI model, is now the real hero — 21 August 2026: per Nvidia's researchers, Opus 5 on ARC-AGI-3 scores 100% with a harness vs 30% without; ARC Prize verification is not mentioned in the article
- Knizhny kub · Harness Engineering is not Enough: Why Software Factories Fail (Dex Horthy) — 25 August 2026: a footnote to “Agent = Model + Harness” — a harness improves execution but does not teach the model to keep architectural quality over time; there is no fast verifier for maintainability; the speaker co-founded a vendor
- Knizhny kub · YC Paper Club: Why The Harness Matters More Than The Model — 15 September 2026: four engineering talks on what changes when the same model gets more environment; the story of a 99.9% ARC-AGI run that turned out to be cheating once the logs were read
- Thoughtworks · Technology Radar Vol. 34 — 15 April 2026: the themes “Putting coding agents on a leash” and “Securing permission-hungry agents”; “teams are beginning to iterate on coding agent harnesses” — feedforward controls like skills and specs, feedback controls like mutation testing
- InfoQ (Matt Saunders) · DORA report on the ROI of AI-assisted software development — 11 May 2026: 35–40% gains on greenfield tasks and about 10% on legacy code; a model for 500 engineers — 39% ROI, roughly an eight-month payback, change failure rate 5 → 6%
- METR · Uplift update — 24 February 2026: for the original cohort “a speedup of −18% with a confidence interval between −38% and +9%”, for the new cohort −4% (−15% to +9%); 30–50% of developers withheld some tasks rather than do them without AI
The labs' bets
- VentureBeat (Michael Nuñez) · Anthropic says it hit a $30 billion revenue run rate — 8 May 2026: a $30B run-rate; Claude Code — “more than $2.5B” run-rate by February, business subscriptions 4× since the start of the year; “the majority of code is now written by Claude Code” internally; company figures
- Fortune (Jeremy Kahn) · OpenAI touts Codex growth — 4 March 2026: 1.6M weekly active users; Thibault Sottiaux: Codex is “becoming the standard agent”; “if we manage to sandbox it properly… bring the power of coding agents to billions of users”
- Constellation Research (Larry Dignan) · OpenAI touts broadening Codex usage — 2 June 2026: more than 5M weekly active Codex users, a fifth of them knowledge workers; company figures
- The Deep View (Jason Hiner) · OpenAI's secret weapon underneath Codex — 29 July 2026: Joe Gershenson — “the harness is how the model interacts with the world and how we are able to express the model capabilities”
- Knizhny kub · Building Codex with Tibo Sottiaux (The Pragmatic Engineer) — 12 September 2026: the Rust core of Codex is separated from its interfaces; “part of the harness must be able to die” — the reminder to run tests moves into the model; “over a weekend you can attach a hundred agents to a project”
- TechRadar via Yahoo · OpenAI says it built an automated research intern — 10 September 2026: OpenAI announced an “automated research intern” and admitted it is “still learning how to measure it”; a self-declaration
- Gemini CLI · notice on the transition to Antigravity CLI — 18 September 2026: “unpaid tier and Google One users: Gemini CLI was replaced by Antigravity CLI on June 18th, 2026”; the repository stays under Apache-2.0
- Google Developers Blog · Transitioning Gemini CLI to Antigravity CLI — 19 May 2026: Antigravity CLI in Go, closed source, shares the harness with Antigravity 2.0 and supports multiple agents
- TNW · Cursor raises at a $50B valuation — 18 April 2026: $2B in annualized revenue (per Bloomberg, February 2026); later reports say Cursor has been a SpaceX subsidiary since August 2026; the primary deal document was not opened
- Factory · News: fundraise — 15 September 2026: $200M at a $5B valuation for “self-improving software development in the enterprise”; a company statement
- GitHub (Kyle Daigle) · Welcome home, agents (Agent HQ) — 28 October 2025: agents from Anthropic, OpenAI, Google, Cognition and xAI inside Copilot; a single mission control and an enterprise control plane
- GitHub changelog · Agent Plugins 1.0 — 12 August 2026: an open plugin standard (skills plus MCP servers in one package), “governed independently of any single vendor”; partners AWS, Anysphere, Microsoft, OpenAI, Vercel, Google
- MarkTechPost · Moonshot AI releases Kimi K2.6 — 20 April 2026: a swarm of up to 300 subagents, 4,000 coordinated steps, a 13-hour autonomous run; a vendor claim
- DeepSeek · DeepSeek-V4 release notes — 24 April 2026: “open-source SOTA in agentic coding”, a million-token default window, integrations with Claude Code, OpenClaw and OpenCode; a vendor claim
- Knizhny kub · OpenCode: how an open harness became a marketplace of models — 27 July 2026: per the CEO — about 13M monthly active users in June, 20× the start of the year; the server with the agent loop embeds separately from the interface; company claims
- Menlo Ventures · 2025 Mid-Year LLM Market Update — 31 July 2025: enterprise usage share — Anthropic 32%, OpenAI 25%, Google 20%; Claude — 42% of code generation; the fund's sample
Forecasts and adoption
- METR · Time Horizon 1.1 — 29 January 2026: 228 tasks, 31 longer than eight hours; Opus 4.5's 50% horizon — 320 minutes (170 to 729); doubling every 196.5 days long-run and 88.6 days since 2024; some models “somewhat sensitive to scaffold”
- Gartner · Over 40% of agentic AI projects will be canceled by end of 2027 — 25 June 2025: the analysts' forecast — over 40% of agentic AI projects cancelled by end of 2027; the page blocks automated fetches, cited via mirrors
- JetBrains Research · AI coding agent adoption 2026 — August 2026: more than 15,000 professionals, May–July; 90% use agents at least weekly, 68% daily; Claude Code 39%, Copilot 21%, Codex 16%, Cursor 12%, OpenCode 7%, Antigravity 6%; self-reported
- JetBrains Research · How much code do developers really let agents write? — August 2026: about 47% of code is fully written by agents, 38% with AI assistance, 27% manually; more than half of developers write less than 20% of their code by hand; self-reported
- Google Cloud · Announcing the 2025 DORA report — 23 September 2025: nearly 5,000 respondents; 90% use AI, more than 80% believe it raised their productivity, 30% report little or no trust in the code
- Stack Overflow (Eira May) · Closing the developer-AI trust gap — 18 February 2026: in 2025, 84% use or plan to use, 29% trust — 11 pp less than a year earlier; the 2026 survey opened on 23 June, results not published as of the check date
- Dwarkesh Podcast · Andrej Karpathy — 17 October 2025: “it's the decade of agents”; the blockers — continual learning, multimodality, computer use
- Dario Amodei · The Adolescence of Technology — January 2026: powerful AI “plausibly within 1–2 years”; AI operating autonomously on multi-day tasks; “AI is now writing much of the code at Anthropic”
The Russian contour
- Habr (Sber) · GigaConf 2026 announcement — 10 September 2026: an invitation-only business day on 23 September and an open engineering day on 16 October 2026 at Sber's office; the engineering tracks include “Harness («обвязка»)”
- GigaConf 2026 · track programme — 18 September 2026: the “Harness — how to make AI generation predictable and manageable: specifications, context, policies, workflow” track; talk titles not yet published
- Habr (Sber, Sergey Trashchenkov) · How we tuned the harness for the model — 25 August 2026: a GigaChat profile for Deep Agents — success 77.2 → 87.0% on 391 tasks, tokens −50%; “models look alike from the outside, but their habits differ”; a vendor claim
- ai-forever · harness-bench-fast — 18 September 2026: an open harness benchmark — tasks on files, memory, search and CSV parsing with golden answers and no model judge, about a 20-minute run
- Habr (Zakhar Kopanitsky) · What an agent harness is — 8 September 2026: “there is no exact short Russian word for the term, ‘обвязка’ is the closest equivalent”; “changing the model changes probabilities, the harness changes failure classes”
- 3DNews · GigaCode Desktop — 15 July 2026: a team of role agents — developer, analyst, tester, manager, security — under “AI-Disrupt PDLC”; a vendor claim
- Habr · Yandex SourceCraft CLI — 22 April 2026: a terminal agent with built-in OpenCode, skills and PR creation; performs a task “autonomously in the background”
- CNews · T-Bank opens trading to clients' AI agents — 10 September 2026: T-Investments opened trading to external agents via MCP — the harness becomes a client of a financial service
- Habr (Ontico) · Saint HighLoad++ 2026: the AI talks — 1 July 2026: Sber — “AI is not a junior but a senior with amnesia”; Raiffeisen — 25% of the backlog via agents, 4% fully autonomous; “Harness on steroids” — $300 → $30 per task; L0–L4 levels; the author's own “State of AI4SDLC” talk; speakers' self-reports
- OpenAI · DevDay 2026 — 18 September 2026: the conference on 29 September in San Francisco with a livestreamed keynote; not included in this article's snapshot
Related reading
- AI Development as a Co-Evolving Stack →The six stack layers, the harness half-life and the rent-adapt-own rule this article builds on
- An End-to-End Player Without Its Own Hardware →The first answer to the paradox: “at the frontier harnesses converge, below it the harness decides”
- Agent Stack Configurations: A Complete Analysis of Eight Options →The own-configurable-managed axis and the twelve technical boundaries
- When Agents Write the Code: What Remains Engineering →The one-line harness definition, the temporary-vs-enduring table and the ablation test for which this article proposes a ledger
- Evaluating AI Agents in Production →The replayable episode — the unit of evidence in the harness ledger
- AI Security in Software Development →The permission ladder and the observation that every 2026 injection fix landed at a boundary
- Loop Engineering: Designing Loops That Run Agents on Their Own →The ladder prompt → context → harness → loop and the verifier's right to stop the loop
- The Stack Ownership Boundary →The rent-adapt-own framework that fills the mode column of the ledger