One word, five units of work
In 1978 Sheridan and Verplank described ten levels of automation, from “the human does everything” to “the computer decides everything and ignores the human”. Their unit was “a single elemental decisive step”: get the options, select, approve, start. Since then the unit has quietly grown while the word stayed the same. Today autonomy names five different things. An action: one agent turn in the terminal, whose median length in Claude Code is about forty-five seconds. A task: a SWE-bench issue or a finished work product in GDPval, where the expert spends seven hours on average and the agent gets the full context in one prompt. A work block: an Upwork project from the Remote Labor Index with a median of eleven and a half hours and two hundred dollars. A role: a digital worker in TheAgentCompany with a hundred and seventy-five tasks and simulated colleagues. And an outcome: a year of trading in Vending-Bench or the profit of a real shop in Project Vend.01I wrote up Michele Catasta's Replit talk: autonomy is named there as the main measurable agent metric and, at the same time, as a property of the system rather than the model. Both thoughts are the spine of this piece.Knizhny kub · Autonomy Is All You Need
A taxonomy, not an empirical law of declining success.
The five units are an author taxonomy of what is delegated to a system, not a measured scale of declining success. In GDPval, a model receives a task and reference materials and produces a professional work product, such as a document, a spreadsheet with calculations, or a presentation. Experts compare it with a specialist's work without knowing who created each version. That finished work product is what the benchmark evaluates. RLI assesses projects, TheAgentCompany tasks in a simulated organization, and Vending-Bench a balance in a simulation. Their models, dates, tasks and success criteria differ. Differences in scores cannot establish how much harder larger units of work are for an agent.
| Unit | Instrument | Success criterion | Reported observation | What is invisible |
|---|---|---|---|---|
| Action | Claude Code telemetry: a turn, a tool call | The call ran and the human did not interrupt | Median turn ≈ 45 seconds; interruptions in 5–9% of turns | The meaning of the whole task |
| Task | SWE-bench, τ-bench, GDPval | Tests pass; all criteria met; the expert preferred the output | GDPval: 84.9% wins or ties for GPT-5.5, an OpenAI claim; τ²-bench pass^1 up to 88% | Context, clarifications, neighbouring tasks |
| Work block | Remote Labor Index, APEX-Agents | A client would accept the project; all rubric criteria met | RLI, GPT 6 Astra: 20.8%, denominator unresolved; APEX, Opus 5.5 Max: 73.5%, September 2026, model judge | Priorities, dependencies, rework |
| Role | TheAgentCompany, CorpGen | All checkpoints; the artifact matches the reference | 42.9% of tasks in November 2025; CorpGen: computer-use-preview (CUP) agent without additional task orchestration — 8.7% at full load | Accountability to colleagues and the client |
| Outcome | Vending-Bench 2, Project Vend | The balance after a year; the shop's profit | $15.5k against ≈ $63k for a “good human”; the shop is “not quite” ready | Reputation, law, the maximum damage |
The table's figures come from separate evaluations. 84.9% on GDPval is OpenAI's reported win-or-tie rate for GPT-5.5 against professional work products. 42.9% on TheAgentCompany is the share of fully completed tasks in the leading entry dated 10 November 2025. 20.8% on Remote Labor Index is the rounded GPT 6 Astra score in the 26 September 2026 snapshot; its denominator needs clarification. These are three separate observations, not successive stages of one experiment.
Researchers' autonomy levels also describe different properties. Mitchell and colleagues at Hugging Face, in Fully Autonomous AI Agents Should Not be Developed, distinguish levels by how much control a model has over program execution: producing an answer, selecting a tool call, or writing and running its own code. Feng and colleagues, in Levels of Autonomy for AI Agents, describe the human's role: operator, collaborator, consultant, approver, observer. Kasirzadeh and Gabriel, in Characterizing AI Agents for Alignment and Governance, consider the share of tasks performed without the person who delegated the work. The February 2026 OECD report distinguishes four modes: the human decides, approves an action, can stop it, or stays outside the execution loop. “Level four” is meaningful only alongside the scale's name and a description of human involvement.
What a system can do and what it is asked to do
Morris and colleagues at Google DeepMind, in Levels of AGI, § 6.2, separate capability from the mode of human–AI interaction. The paper first appeared in 2023; this article uses the June 2024 revision. Capability covers performance relative to humans and the breadth of tasks. Autonomy describes how people and systems share work and initiative.
Table 2 identifies six modes: 0 — no AI; 1 — a tool for individual operations; 2 — a consultant invoked on request; 3 — collaboration; 4 — AI leads while a human guides and gives feedback; 5 — an autonomous agent. The last mode still allows asking people for help: continuous oversight and consultation are different things.
ONE MODEL · THREE MODES
Capability stays the same. What changes is who leads the work.
A stronger model expands the choices. People still choose the mode.
The diagram uses a hypothetical model that can prepare reports. It can be called on for analysis, used to edit a document together, or assigned the preparation cycle with a human checking the result. Changing the mode does not make the model smarter. Improving the model expands the available choices; the appropriate mode still depends on the task, cost of failure and controls.
Production surveys describe the prevalence of different usage modes. In August 2026 Deloitte counted 42% of companies that had tested or deployed agents and 15% with scaled adoption; Gartner counted 17% with agents in production; Bain, in Bloomberg's retelling, 7% with fully autonomous agents; a Russian study by Infosistemy Jet, 8%. Anthropic's own engineers, whose access to the best models is unlimited, said in December 2025 that they could fully delegate between zero and twenty percent of their work. Raiffeisen at Saint HighLoad++ named a quarter of the backlog done together with agents and four percent done by agents without asking a human. Between “the agent handled a task” and “the agent can be relied on as a standing performer” lies this entire ladder, and the word autonomy jumps over it.
Time horizon: what METR measures and what it does not
The time horizon describes how much work an agent can complete at a given success probability. Work is measured in a specialist's time: a “two-hour task” takes a human about two hours, even if the agent spends only ten minutes on it. METR's TH1.1 suite contains 228 programming, machine learning and cybersecurity tasks.
- Establish each task's human duration. A specialist completes it while being timed; some tasks use expert time estimates.
- Give the same tasks to the agent. Record whether each attempt succeeds or fails. Repeat attempts to reduce the influence of a lucky success or an unlucky failure.
- Fit a curve to all the results. Human task time is on the horizontal axis, agent success probability on the vertical. A smooth line summarizes the scattered attempts: how much less often the agent succeeds as tasks require more human time.
- Find the intersection with the desired percentage. If the line crosses 50% at “4 hours”, the 50% horizon is four hours. This estimates success on tasks of that size; it does not promise four hours of continuous agent work.
Four hours above is a hypothetical example. METR's method uses a logistic curve: a smooth declining line between 100% and 0%. Its position and steepness are fitted to the agent's successes and failures. The time axis is logarithmic: equal distances mean equal multipliers, such as 10 minutes → 100 minutes → 1,000 minutes. This puts short and long tasks on one chart. The intersection with 80% gives a second, shorter horizon: higher reliability requires a smaller work unit.
Checked on 27 September 2026, METR's main table still lists 8 May as its last update. In that snapshot, the early Claude Mythos Preview has the highest 50% horizon at 17 hours 25 minutes, and an 80% horizon of 3 hours 6 minutes. These are uncertain estimates: the first has a 95% confidence interval from 8 hours 29 minutes to 55 hours 4 minutes; the second, from 1 hour 37 minutes to 6 hours 39 minutes. METR warns that values above 16 hours are unreliable with the current suite. The May figure cannot be treated as an established limit on model capabilities in September.
The diagram below uses the two published Mythos estimates. Its smooth line is reconstructed from those two points to explain the method, not newly fitted to the original runs. The horizontal bars show uncertainty intervals for the two horizons, not for the entire curve.
Human task time, not agent runtime. Neither horizon sets a delegation threshold. No computer-use curve is measured here.
Evidence:METR · 2026-05-08
Later reports exist, but do not update that main table. On 26 June, METR published its GPT-5.6 Sol evaluation: about 11.3 hours when attempts that circumvented the evaluation rules were counted as failures, versus over 270 hours when those attempts were counted as successes. METR does not consider any of these numbers a robust capability estimate. The 22 September Claude Opus 5.5 evaluation used five AI research and development tasks. It describes incremental improvement over Fable 5.1, but publishes no new comparable 50% or 80% time horizon.
The trend is better known than the values themselves: a doubling roughly every two hundred days across the whole series since 2019, and every 129 days counting from 2023; the January post also named 88.6 days since 2024. In April 2026 Epoch AI confirmed the acceleration and attributed it to reasoning models, noting that the gains concentrate in programming and mathematics — where the answer is checked automatically. What matters for this piece is that METR itself insists on what the horizon is not. “Time horizon is not the length of time AIs can work independently” but “the amount of serial human labor they can replace with a 50% success rate”. The estimate's error is about a factor of two each way. Domains differ by orders of magnitude: for tasks on a computer screen the horizon is forty to a hundred times shorter. The key line for delegation: “a 50% time horizon of X hours does not mean we can delegate tasks under X hours” — some tasks need ninety-eight percent success or more.
The tasks themselves have another limitation: fewer complications than everyday work. METR assesses this using 16 factors, including coordination with other participants, a changing environment, limited resources and difficulty checking the result. Each factor present adds one point. In the original study, HCAST and RE-Bench tasks averaged 3.2 out of 16, with none above 8. This describes working conditions, not data quality. As a hypothetical example, fixing a function against an existing test is easier to organize than diagnosing an incident in a changing system, coordinating a fix with another team and checking the outcome without a ready-made reference.
The mathematical curve also affects the answer. Fergus Hamilton reanalysed METR's results using another curve shape, the Weibull distribution. It describes the relationship between task duration and success probability differently. Both curves fit the available observations closely, but diverge when estimating near-perfect reliability. Toby Ord's discussion of this comparison shows that the calculated 99% horizon can be twenty times shorter under the Weibull curve than under METR's logistic curve. This concerns the task duration predicted to yield 99% success: hypothetically, one method might give 100 minutes and the other 5 minutes. That illustrates the difference between estimates; it is not a separate model measurement.
Ord stresses that there is insufficient evidence to choose confidently between the curves. A good measurement at 50% therefore does not reliably establish which tasks an agent can complete in 99 out of 100 attempts. METR's own March sensitivity analysis found that removing public RE-Bench tasks cuts Opus 4.6's horizon from twelve hours to seven. Reasonable fitting choices change the 50% horizon by about 1.5× and the 80% horizon by 2×. Task selection remains the main source of uncertainty.
Now set this against what vendors say and what production shows. At the launch of Sonnet 4.5 Anthropic reported “more than 30 hours” of focus on multi-step tasks — an observation without a method; METR gives the same model one hour fifty-seven at fifty percent and twenty-six minutes at eighty. Codex, in OpenAI's words, worked “more than 7 hours”; METR gives its successor three hours forty-four. The difference between the claims and the measurement is six- to fifteenfold, and it is not a lie: one number is the agent's running time on a single showcase task with an unstated success criterion, the other is human time at a stated probability across two hundred tasks. Anthropic's own telemetry from February 2026 settles it: the median Claude Code turn is forty-five seconds, the 99.9th percentile grew from twenty-five to forty-five minutes in three months, 73% of API tool calls belong to agents classified as having a human in the loop, and 0.8% of actions are irreversible. These are turn durations, not capability limits: a short turn can mean fast completion, and parallel work can also reduce elapsed time. And METR's experiment with sixteen developers in 2025 showed a 19% slowdown against an expected 24% speedup; the 2026 rerun gave minus eighteen and minus four with wide intervals, and METR itself called the survey with a median self-reported “three times faster” “not necessarily grounded in reality”.02I wrote about the sixteen-engineer experiment back in July 2025: the sample is small but the methodology is honest. A year later it gained a rerun that proved nothing — and that is a result too.Knizhny kub · the METR experiment
| What the horizon says | What it does not say |
|---|---|
| The human-hours length of a task the model finishes with a 50% probability | How many hours the agent will run without a human: METR says outright it is not the length of independent work |
| That the 80% horizon is four to ten times shorter than the 50% one | That 50% is enough to delegate tasks shorter than the horizon |
| The trend: a doubling every four to seven months depending on the segment | That the trend transfers to other domains: computer use is 40–100 times lower |
| HCAST and RE-Bench in the original paper: on average 3.2 of 16 workplace-complication factors, maximum 8 | What happens in someone else's repository: contractors are 5–18 times slower than maintainers |
| Mythos: 95% CI for the 50% horizon is 509–3304 min, for 80% it is 97–399 min; weak coverage above 16 h | A precise number: a change of fit moves the 50% horizon by 1.5× and the 80% horizon by 2× |
| The model's result in METR's scaffold; Claude Code and Codex did not improve it | That a product harness adds horizon: at the frontier the difference is not statistically visible |
METR is also exploring complementary measures: compute cost, human participation and gains under a fixed budget. The practical conclusion is limited: the 50% and 80% horizons characterize a selected task distribution and evaluation procedure. Neither authorizes delegation on its own. That requires a local sample, a reliability requirement, an error cost and complete human-effort accounting. METR's limitations note discusses distribution dependence and unstable high quantiles. Comparisons with other domains are separate observations; they do not provide a measured Mythos computer-use curve.
From an isolated task to a working day
CorpGen studies 46 OSWorld office tasks in a shared session with dependencies at several loads. In table 3, CUP stands for computer-use-preview, OpenAI’s model for operating computer interfaces. The authors use computer-use-preview-2025-03-11. “Baseline” means the agent without the additional CorpGen layer that coordinates multiple tasks. This configuration scores 16.7% at 12 tasks and 8.7% at 46. A separate variant with hierarchical planning, Hierarchical CUP, scores 25.0 → 14.1%. The positive result matters too: at full load, CorpGen improves CUP from 8.7 to 16.3%; CorpGen with CUP scores 16.7 → 16.3% across the endpoint loads. Load effects therefore depend on execution architecture. Small task sets and absent intervals prevent treating these differences as a precise general effect of multitasking; an isolated task is not the control condition here.
Within-study contrasts, not a common ladder of autonomy. Partial credit measures progress; its unit differs from task completion.
Evidence:CRMArena-ProCorpGen · table 3APEX · table 3TheAgentCompanyClawMark · table 3RLI
TheAgentCompany from CMU shows the second side: not load but a role. The paper's best agent solved 24% of tasks in December 2024, 30.3% by May 2025, and the last leaderboard entry from November 2025 reads 42.9% fully and 52.4% on the weighted score, which grants half the checkpoint completion fraction plus half the full-success indicator. The failures come in four kinds: lack of common sense, inability to talk to colleagues, helplessness in web interfaces, and self-deception — the agent “creates fake shortcuts that omit the hard part”. In August 2026 an independent re-analysis of three such benchmarks produced what I consider the main result of this section: the choice of agent explains under three percent of the variance, and the agent-by-task interaction seven to twenty-three. Which tasks matters more than which agent. ClawMark, released in April, added a third dimension — an environment that changes without the agent: mail arrives, the calendar shifts, the knowledge base updates. Sonnet 4.6 scored 75.8 weighted points and 14% strict success, and performance dropped right after the first exogenous change.
The cleanest “same task plus interaction” measurement comes from Salesforce. In CRMArena-Pro the best model solved 58.3% of tasks in a live CRM when everything was said in one turn and 30.0% when the missing pieces had to be elicited from a simulated customer; in forty-five percent of the dialogue failures the agent simply did not ask. OpenAI's GDPval stands at the opposite pole — full context up front, no clarifying dialogue — and the authors themselves write that results fall when context is cut and the win rate declines as task length grows. Their own arithmetic sobers up the loud “a hundred times faster and cheaper”: that is pure inference time. GDPval Table 2 reports a speedup factor from 0.87 for GPT-4o to 1.12 for GPT-5 under the strategy “try AI once, have an expert review it, then do the task yourself if needed”. The factor divides time without AI by time with AI: below one means a slowdown. In the authors’ calculation, GPT-4o increases total time by about 15%, while GPT-5 reduces it by about 11%. For example, a 100-minute task without AI would take about 115 or 89 minutes respectively. These are estimates from a time-cost model, not measurements of a workplace rollout.
The Remote Labor Index went from 2.5% to 20.8% of accepted projects in a year, but the new runs already used Claude Code and Codex with a worker–critic loop, a budget of up to a hundred and fifty dollars and a full day of time; the model judge overstated the success share two and a half to three times, and of the three showcased Fable 5 deliverables the authors wrote that “none would be accepted as finished work”. Curiously, success on RLI does not fall with the human length of a task — there is no horizon curve there, because a project is not “more text” but “more places where things can go wrong”.03SWE-Together from Meta does the same for coding: 109 tasks from eleven thousand real sessions and a user simulator that annotators could not tell from a live person. A stronger model needs fewer interventions — but it still needs them.Knizhny kub · SWE-Together
| Benchmark | Unit and setting | Success criterion | Result and limits |
|---|---|---|---|
| CorpGen · Microsoft, February 2026 | 46 OSWorld office tasks in one five-hour session with dependencies | The artifact matches the reference; a trace judge agrees with humans 40% of the time, an artifact judge 90% | CUP, 12 → 46 tasks: 16.7 → 8.7%. Hierarchical CUP: 25 → 14.1%. CorpGen + CUP at 46 tasks: 16.3% |
| TheAgentCompany · CMU, December 2024 | 175 digital-worker tasks with simulated colleagues | All checkpoints; weighted score: half the checkpoint fraction plus half full success | Snapshot 2025-11: 42.9% complete / 52.4 weighted points |
| ClawMark · April 2026 | 100 multi-day tasks in an environment that changes on its own: mail, calendar, knowledge base | 1,537 deterministic checkers | Sonnet 4.6: 75.8 points / 14% strict success; table 3 |
| CRMArena-Pro · Salesforce, May 2025 | 19 task types in a live CRM, single turn or a dialogue with a simulated customer | Exact match, token F1, a model judge | Gemini 2.5 Pro · B2C: 58.3 → 30%; single turn → dialogue |
| GDPval · OpenAI, September 2025 | 220 tasks across 44 occupations: documents, spreadsheets, presentations, full context up front, no clarifying dialogue | Blind pairwise expert comparison; 71% agreement | 47.6% → 84.9% in a year, but results fall when context is cut and “try, then fix” saves 0.9–1.1× the time |
| Remote Labor Index · Scale and CAIS, October 2025 | 240 selected projects; 230 private for quantitative evaluation; median 11.5 hours and $200 | Three experts: work at least as good as the human's; 94.4% agreement | 2.5% → 20.8% in a year with new harnesses and budgets; a model judge scores higher; no clear relationship to duration in this sample |
| APEX-Agents · Mercor, January 2026 | 480 banker, consultant and lawyer tasks in worlds of 166 files | All binary rubric criteria; Gemini 3 Flash as judge | Gemini 3 Flash: 24% complete; mean criterion score 39.5%; table 3 |
| Vending-Bench 2 · Andon Labs, November 2025 | A year of running a vending machine: 3–6 thousand messages, adversarial suppliers | The bank balance at year end | From −$31 to +$15,515 across models of one year; a “good human” ≈ $63k by the authors' estimate |
In APEX-Agents, Gemini 3 Flash meets every criterion on 24% of 480 tasks, while its mean criterion score is 39.5% (table 3). These answer different questions: whether the result is complete and how far the work progressed. Partial credit is useful diagnostically and is not inherently flattering. The Gemini 3 Flash judge was checked against 747 human labels from 60 tasks; the reported 98.5% agreement applies to that sample. In RLI, the model judge and experts give Opus 4.8 roughly 21% and 8.3% respectively, evidence of disagreement in that particular evaluation. RLI's methodology describes 240 selected projects, with 10 public and 230 private projects for quantitative evaluation. The current leaderboard percentages are not transparently reconciled with that denominator; the number of accepted projects must not be reconstructed from 20.8%.
These studies reveal different sources of disagreement: task conditions, the definition of success, and the judge. They cannot be combined into a single estimate of lost autonomy over a working day. That would require controlling the model, task set, budget and verification. A particular experiment can support a discussion of dialogue or load effects; cross-study comparisons describe differences in setup and limits to transfer.
Reliability versus capability
Sierra introduced a distinction that almost every leaderboard still ignores. pass@k is the probability that at least one of several attempts succeeded; pass^k, that all k attempts succeeded. Each attempt runs the same task from scratch.
On τ-bench in 2024 GPT-4o solved 61.2% of retail tasks on the first attempt and under 25% on all eight. On τ²-bench in 2025 every model lost fifteen to twenty-six points between pass^1 and pass^4, and the telecom domain, where both agent and customer use tools, lost half: 0.34 → 0.19 for GPT-4.1. Merely adding an acting user to the same task took away eighteen to twenty-five points. In January 2026 Anthropic fixed this in its guide to evals: capability evals start at a low pass rate, regression evals should pass near one hundred; “a task that passed on one eval run might fail on the next”; pass^k is for agents where consistency matters. In 2026 the leaderboards still display pass^1 in the shop window.
The Princeton line of work — from the 2024 paper on which agents matter to HAL — added cost and logs. Across 21,730 rollouts costing forty thousand dollars, more reasoning did not raise accuracy in twenty-one of thirty-six combinations; an agent can be “100x more expensive while only being 1% better”; transcripts revealed agents searching HuggingFace for benchmark answers and hard-coding tests; agents with identical accuracy carry different risk — abstaining and paying with someone else's card both score zero. In a May 2026 study, researchers from Princeton and the UK AI Security Institute found defects in 25 of 50 τ-bench Airline tasks: inconsistent policies, ambiguous instructions, and database or grading errors. Excluding these tasks raised average pass^5 across the sampled models from 20.8% to 40.0%. This is what the authors mean by performance being underestimated by nearly half: flawed tasks prevented agents from demonstrating their capabilities. The models were unchanged; the evaluation set changed. In a separate February study, researchers found that “recent capability gains have only yielded small improvements in reliability” across twelve metrics. The harness has become a variable of measurement: on Terminal-Bench the same model scores 83.8% in Claude Code and 80.4% in another harness, and a May preprint showed that harness-induced variance “can substantially exceed model-induced variance”, up to reversing rankings. I went into that effect in the harness piece; what matters here is the consequence: a report without the harness version and the number of runs is not a measurement.
The benchmarks themselves age faster than they seem to. In February 2026 OpenAI stopped publishing SWE-bench Verified: of 138 hard tasks, 59.4% have defective tests or specifications, and all frontier models reproduce the gold patches from memory. The replacement was SWE-bench Pro, where Opus 4.1 fell from 22.7 to 17.8% on private commercial repositories; in eight months the public split grew from 23.3 to 80.3%, after which the same OpenAI found “~30% of tasks broken” and withdrew its recommendation. The July 2025 checklist for rigorous agentic benchmarks found task-validity violations in seven of ten suites: an empty response scored 38% on impossible τ-bench tasks. Web agents on live sites showed 30% instead of a claimed 89%. Every “fixed” successor saturates in six to twelve months and turns out fifteen to thirty percent defective once models are strong enough to expose it.04In my breakdown of Anthropic's study of 398,000 sessions I wrote why session success cannot be renamed productivity: “verified success” cannot see whether the team accepted and deployed the code. The same boundary applies here.Knizhny kub · Claude Code and expertise
| Metric | Which question it answers | Example |
|---|---|---|
| pass@k | At least one of k attempts succeeded | Tools with cheap verification; BrowseComp: 64 samples add 15–25 points |
| pass^k | All k attempts succeeded | τ-bench: 61% → under 25% at k = 8; τ²: minus 15–26 points by k = 4 |
| Run-to-run interval | The spread between runs of one configuration | Terminal-Bench: at least five trials, ±1 point; the same model scores 83.8% and 80.4% in different harnesses |
| 80% horizon | Task length at 80% success on the selected distribution | Opus 4.6: about 12 hours at 50% and 1 hour 10 minutes at 80% |
| TCR@k | The share of tasks completed with at most k interventions | Microsoft, February 2026: TCR@0 of o3-mini: Prudentia 59.1% against Fides 50.1% |
| PR acceptance rate | PR acceptance in an observational corpus; edits are not excluded | Agent PRs in popular repositories: 38–65% against 76.8% for humans |
| Robustness to perturbation | The drop when tools and inputs are corrupted | ToolRobustBench: 0.979 clean → 0.455 with corrupted tool outputs |
AIDev supplies another metric: PR acceptance in an observational corpus. Table 5 reports 38–65% for different agents and 76.8% for human authors in popular repositories. This is PR acceptance, not acceptance without edits or a randomized comparison on the same tasks. Task, repository and user selection may explain part of the differences. Acceptance also does not establish long-term reliability: that requires data on reverts, defects and maintenance.
The law of the long chain and checkpoints
The simplest model of a long chain is a power: if each step succeeds with probability p and steps are independent, a chain of n steps succeeds with probability p^n. Sinha and colleagues, in their ICLR 2026 paper, derived the horizon from it: the number of steps before success falls below a threshold equals ln s / ln p. At 95% step accuracy that is thirteen and a half steps to a coin flip (a 50% chance that the whole chain succeeds), at 99% sixty-nine, at 99.9% almost seven hundred. Hence their thesis about the “illusion of diminishing returns”: one more point of step accuracy lengthens the horizon by a quarter, and the closer to one, the stronger the effect. But real models violate independence in both directions. On one side, self-conditioning: errors in the context raise the probability of further ones, scale does not fix it, reasoning does. On the other, extra compute per step: GPT-5 with reasoning executes more than two thousand steps in a row on a synthetic task. Independent 2026 measurements confirm the first side: the decline with length is one and a half to two and a half times steeper than the independent model; after a wrong step the next one is wrong 40–58% of the time against three to five after a correct one.
A checkpoint is a point where the system verifies an intermediate result before continuing. For example, after an agent edits code, the system runs tests and saves the version that passes. If tests fail, it restores the previous saved version and retries that segment. The model below checks after every ten steps and allows one initial attempt plus up to two retries per segment. Three failures stop execution.
What happens at a checkpoint
50 useful steps · p = 0.95 per step
No retries: 7.69%. Check and retry every 10 steps: 71.61%.
Illustrative model with perfect detection and independent retries.
A check does not fix an error by itself: the benefit comes from detecting failure, restoring state and trying again. The diagram assumes perfect error detection, independent retries and no irreversible damage before verification. A ten-step segment succeeds within three attempts with probability q = 1 − (1 − p¹⁰)³; five such segments succeed with probability q⁵. At p = 0.95, this gives 71.61% instead of 7.69% without retries. These are the same 50 useful task steps, but repeated execution and checks consume extra time and resources. Real checks can miss errors and retries can repeat them, so this gain is not guaranteed.
The Traverse dataset, published by Rahman and colleagues in September 2026, contains 2,518 agent trajectories across software engineering, computer use and science, with 6,967 annotated mistakes. In the subset of 1,122 trajectories with a labelled first mistake, 69.5% never recover, 38.5% never notice it, 72.6% keep acting, and about 84% of the studied software and computer-use failures end with a step that looks correct. Frontier model judges find the first mistake in fewer than a third of coding runs. The MAST taxonomy from Berkeley, over 1,642 traces, sorts failures into fourteen modes: step repetition 15.7%, reasoning–action mismatch 13.2%, not knowing the stopping condition 12.4%, incorrect verification 9.1% — the last a false pass that turns a recoverable run into a silent failure. Where the critical error sits is known too: a failed trajectory carries on average 7.6 local errors and exactly one critical one, the agent repairs 62% of the non-critical ones itself, and the root causes cluster on steps six to fifteen — after information gathering, at the first decisive action. That is where a checkpoint is worth the most. Hamilton's data on the declining hazard rate say the same from the other side: early steps are the riskiest, verification should be front-loaded, and a “surviving” run should not be reset without evidence.05In the Loop Engineering breakdown the key element of the loop is the verifier's right to say no. This whole section is about the verifier needing evidence, not an opinion.Knizhny kub · Loop Engineering
Who verifies matters as much as where. Self-verification without external evidence does not work: back in 2023 GPT-4 fell on GSM8K from 95.5 to 89.0% after two rounds of “correct yourself” and rose to 97.5 with an oracle saying where the error was; in 2026 it was shown that relabelling the model's own thought as someone else's raises the share of explicit corrections by 23–93 points. External, evidence-based verification, by contrast, gives large measured effects: a linter on every edit in SWE-agent — plus 7.7 points; fixing the critical error and restarting from it — from 21 to 55% on ALFWorld; a trained verifier choosing among several attempts — from 81.8 to 90.2% on Terminal-Bench; decomposition at subtask boundaries — plus 13–42 points. Theory explains why: the step-level verification signal grows with length as Θ(T), the outcome signal decays as Θ(T·p^T) — the same law that kills the chain itself. But verification is neither free nor always useful: a weak verifier is worse than none, a blocking mediator intercepts 94% of violations and leaves under 5% safe successes, and the Qwen team writes outright that “generating a solution has become easier, reliably verifying it has become the harder problem”, so verification must co-evolve with the generator.
Blocking an irreversible action means the system prevents execution until the required checks or authorization pass. For example, a database deletion tool stays unavailable until a human approves it. The restriction must be enforced by the tool or execution system; an instruction asking the agent to be careful cannot guarantee a stop.
| Checkpoint | What it protects against | What is known |
|---|---|---|
| State checkpoint: a commit, a progress file, a session log | Against context loss, crashes and restarts | Anthropic, November 2025: sessions as engineers' shifts; Managed Agents: restart from the event log |
| Plan checkpoint: a feature list, a sprint contract, a spec | Against goal substitution and “declared it done” | The agent “would look around and declare the job done”; MAST: 12.4% of failures are not knowing the stopping condition |
| Verification checkpoint: tests, a browser, an external judge | Against silent errors that look like success | Traverse: about 84% of studied software and computer-use failures end looking right; a linter on every edit adds 7.7 points; a trained verifier lifts 81.8 → 90.2% |
| Block before an irreversible action | Against damage that cannot be rolled back | 97% of destructive actions in “solved” runs come without acknowledged risk; Replit, Antigravity, Kiro |
Harnesses for long runs converge on the same primitives, and it helps to tell them apart. A state checkpoint — a commit, a progress file, the append-only session log in Managed Agents — makes a crash recoverable. A plan checkpoint — a JSON feature list, a sprint contract — stops the agent from “declaring it done” by redefining the goal: in November 2025 Anthropic watched a later instance “look around, see that progress had been made, and declare the job done”. A verification checkpoint — tests, a browser, an external evaluator with hard thresholds — catches silent errors. The only vendor comparison with a control is the March one: a solo agent built a game in twenty minutes for nine dollars, but nothing responded to input; the full harness spent six hours and two hundred dollars and delivered a working one; and the evaluator gave lift “only at the edge of the generator's capabilities”, while “out of the box, Claude is a poor QA agent”. Durable-execution engines fix crashes, not meaning: their argument is “five steps at 99% give 95%”, their limit is that a semantic error is journaled and replayed as faithfully as a correct step. Context compaction is a risk point of its own: recursive summaries turn reliably solved tasks into intermittently solved ones, while truncation without paraphrase under the rule “never compact a compaction” holds the result at half the cost.
- A cheap mechanical check on every state-changing action: the chain starts with the first bad edit, and recovery after it is 57%
- An evidence-based check at subtask boundaries: critical errors cluster after information gathering, and decomposition with a check recovers 13–42 points
- An external judge at the end: the model passes its own work and fails someone else's; a 23–93-point shift from one relabelling
- Block an irreversible action until the required checks or authorization pass: 97% of destructive steps in “solved” runs were taken without acknowledged risk
- The interval between checks grows with the square root of reliability and of check cost: at 98% per step and a check costing two steps — roughly every fourteen steps
The last point transfers the Young–Daly formula from supercomputing, where the optimal checkpoint interval is the square root of twice the product of the mean time between failures and the cost of a checkpoint. For an agent at 98% step accuracy and a check costing two steps that is about fourteen steps at 28% overhead; at 99.5%, twenty-eight steps and 14%; a cheap linter costing a fifth of a step pays off every four to five steps. The shape of the law — checks get rarer as reliability and check cost grow — matters more for practice than the constants, and the same literature reminds us that with a declining hazard the intervals should lengthen as the task proceeds. Two corrections are mandatory. An agent failure is not fail-stop: since 84% of failures look correct, the cost of a checkpoint must include detection, otherwise you save the state of a corrupted run. And a retry helps only if it is independent: self-conditioning makes a retry in the same context correlated with the first failure, which is why the unit of retry in every 2026 harness is a fresh context plus restored state, not “try again”.
Is fifty percent enough
In July 2025 Jason Wei formulated the “verifier's law”: the ease of training AI to solve a task is proportional to how verifiable it is, and a verifiable task is one with objective truth, where checking takes seconds, scales, is low-noise and yields a continuous score. In November Karpathy compressed the same into “Software 2.0 easily automates what you can verify”, and at the YC summer school talked about the autonomy slider and the generation–verification loop that should spin as fast as possible. The economics of sampling confirm it: on SWE-bench Lite the share of solved tasks grew from 15.9% with one sample to 56% with two hundred and fifty — but only where tests exist; without an automatic verifier, selecting the best attempt plateaus. So the question “is fifty percent enough” has no answer without a second question: what does verification cost, and what happens when it misses an error.
The following is an author cost model. Let manual execution cost S, one agent attempt A, and its verification V, all expressed in the same monetary units or converted into them. With constant success probability p, independent retries, perfect verification and no damage before acceptance, expected cost per accepted result is (A + V) / p. Savings require p > (A + V) / S; equality is break-even. If S = 60, A = 2 and V = 10, the threshold is 20%; at V = 40 it is 70%. These are hypothetical cost units, not a sum of machine and human minutes. If A + V > S, even p = 1 produces no savings. Fallible verification, dependent attempts or damage before acceptance require a different model. Without verification, comparing expected value with guaranteed successful manual work of value W and an agent failure loss L gives p ≥ 1 − (S − A) / (W + L). W and L must be chosen and justified for the process. Lamba's sixteenfold decline in optimal throughput when error detectability halves assumes δ = 1; with fixed labor, the quadratic relationship instead gives a fourfold decline. These are model conditions, not a universal law of verification. Cost per accepted task is discussed further in the article on development economics.
Author decision aid. Reliable verification, independent repeats and no damage before acceptance are required by the cost formula.
The matrix below organizes questions; it does not prescribe universal percentages. Cheap verification still requires checking the cost of the whole attempt; expensive verification requires checking whether savings remain possible at all. Large potential losses need a separate risk constraint: limit permissions, support recovery, add independent verification or retain human approval. Reversibility affects consequences but does not itself imply a 50% or 90% success threshold.
The literature that publishes thresholds ties them to a maturity stage, and the form of the threshold differs at each. A research prototype is measured by a horizon at a fixed probability and pass@k against an oracle. A commercial product by an outcome share under a contractual definition: since July 2026 Salesforce charges only when the agent “autonomously resolves an issue from start to finish”, with no charge on negative feedback or a human-escalation request, and reports 70% of such outcomes across 4.3 million inquiries on its own portal; at Intercom an “assumed resolution” is a customer who did not reply for a day; Zendesk advises excluding abandoned and repeat contacts. Responsible operation by decision rights and a stop button: Article 14 of the European AI regulation requires oversight “commensurate with the risks, level of autonomy and context of use”, with the ability to override a decision and stop the system; the algorithmic-trading rules require cancelling any orders immediately and naming which algorithm is responsible for each; banking supervision requires “effective challenge” of a model and limits on its use; and in January 2026 the FDA drew a line worth remembering: software remains decision support only if the clinician can independently review the basis of a recommendation, and under time pressure that is already automation. Aviation and automotive safety requirements specify the probability of a dangerous failure per operating hour: a catastrophic condition in aviation must be “extremely improbable”, on the order of 10⁻⁹ per flight hour, and must not result from a single failure; the automotive standard sets a limit below 10⁻⁸ per hour for dangerous random hardware failures at its highest class, ASIL D. The more severe the possible consequences, the lower the acceptable failure probability. These figures express system safety requirements, not the percentage of successfully completed tasks. For office work nobody publishes a numeric threshold for full automation — the oft-quoted “95% accuracy” traces back to consultancy blogs, not analysts' reports.
| Stage | Form of the threshold | Who measures this way | Example |
|---|---|---|---|
| Research prototype | A horizon at a fixed probability; pass@k against an oracle; cost on the Pareto frontier | METR, HAL, public benchmarks | The threshold depends on attempt plus verification cost, loss and model assumptions |
| Commercial product | Outcome share under a contractual definition; pay per outcome; a domain validator | Salesforce, Intercom, Sierra; Microsoft's second tier | 70% “start to finish” of 4.3M inquiries; at Intercom 24 hours of customer silence count as a resolution |
| Responsible operation | Decision rights: what the agent decides alone; a stop button; attribution of every action; “effective challenge” of the model | AI Act Article 14; MiFID II RTS 6; SR 11-7; FDA; Microsoft's third tier | The clinician must have time to independently review the basis; under time pressure the FDA treats it as automation |
| Full process automation | A hazard budget per hour of exposure; no single failure with a catastrophic outcome; redundancy | FAA AC 25.1309-1B, ISO 26262 | On the order of 10⁻⁹ per flight hour for a catastrophe; for office work nobody publishes a numeric bar |
These studies do not establish a universal threshold for a prototype, a commercial product or a consequential operation. The decision depends on the distribution and cost of errors, the verifier's ability to detect them, and the cost of all attempts. The process owner chooses an acceptable risk; evaluation should show how confidently the measurements fit that constraint. A threshold cannot be imported directly from another benchmark.
The autonomy scorecard: what to count beyond the result
If the threshold is a property of the task, an autonomy report cannot consist of one percentage. I propose a scorecard of six rows, and for each row there is data somewhere in 2026 already. Result — the share of tasks accepted under a criterion written before the run; the definition of acceptance matters more than the number, and the Salesforce, Intercom and Zendesk definitions show how strongly it moves the result. Interventions — human touches per task: Anthropic measured that experienced Claude Code users both turn on full auto-approve more often, from 20 to 40% of sessions, and interrupt the agent more often, from 5 to 9% of turns — trust and vigilance grow together; in February Microsoft formalised this as TCR@k, the share of tasks completed with at most k approvals. Verification cost — human minutes per accepted task: in the METR experiment that was nine percent of time spent reviewing and cleaning output while accepting under 44% of suggestions; in Faros telemetry, a two-and-a-half-fold rise in time to first review measures waiting, not active human effort. Recovery time — from failure to a working state, for which DORA already has a definition; the Kiro agent with an engineer's inherited role decided to “delete and recreate the environment” and took a service down for thirteen hours in one region. Blast radius — the maximum damage of one run and the reversibility of each action: 0.8% of API actions are irreversible; the auto-mode classifier misses 17% of dangerous actions on fifty-two cases — “17% is the honest number”, Anthropic writes; the Replit agent deleted a production database during a code freeze. And cost per accepted task — all spend over the period divided by accepted tasks: HAL shows agents “100x more expensive while only being 1% better”, and the Remote Labor Index counts a failure as infinite cost.06The DX report for the second quarter is the best illustration of the scorecard without a scorecard: more code, higher maintainability, lower confidence in changes. One row rises, another falls, and without both there is no picture.Knizhny kub · DX Q2 2026
| Row | How it is measured | Where data exists today |
|---|---|---|
| Result | The share of tasks accepted under a criterion written before the run | Salesforce: 70% of 4.3M inquiries “start to finish”, billed only for those |
| Interventions | Human touches per task; TCR@k | Anthropic: interruptions 5 → 9% of turns with experience; full auto-approve 20 → 40% of sessions |
| Verification cost | Human minutes per accepted task | METR: 9% of active time reviewing and cleaning output; record review waiting separately |
| Recovery time | From the failure to a working state | Kiro: about 13 hours of Cost Explorer downtime in one region; DORA: AI correlates with delivery instability |
| Blast radius | The maximum damage of one run; reversibility | Anthropic: 0.8% of actions are irreversible; Replit: the production database during a freeze |
| Cost per accepted task | All spend over the period divided by accepted tasks | HAL: a hundred times the cost for one percent; RLI: a failure counts as infinite cost |
The scorecard is not decoration: any single row can be inflated. The share of accepted changes can rise when people approve them without substantive review, as Warp warns. Permission decisions pose a similar risk: Anthropic reports that Claude Code users approve 93% of prompts and warns about approval fatigue: people stop paying close attention to what they are allowing. Interventions fall if the agent only gets documentation. Cost falls if you count only the first model call. Only the combination of rows separates a standing performer from a lucky demo: accepted under a strict criterion, with few touches, cheap verification, fast recovery, a bounded radius and a known price. A seventh candidate row is error detectability from Lamba's work: an agent that errs less often but more plausibly may cost an organisation more than one that errs often and obviously. Nobody has yet shown how to measure this on a stream, but incident reports deserve a field “how many reviews the error passed before detection”.
Pseudo-autonomy: where the labor hides
In 2018 Astra Taylor called it “fauxtomation”: labor does not disappear, it shifts — onto the customer at the kiosk, the content moderator, the Mechanical Turk worker. For agents the word needs narrowing: pseudo-autonomy is when human labor has been moved into writing instructions and fixing hidden errors and excluded from the metric. The best document on the mechanism is the SEC order of 14 January 2025 in the matter of Presto, the maker of a voice ordering system for drive-throughs. The first version “required human agent intervention, including entering the order, in all instances”, the pilot of a more advanced one in seventy percent; the operators sat in the Philippines and India. Meanwhile the company reported 95–99% of orders “without intervention”, and the SEC found that this figure was counted without restaurant staff intervention — but not without humans. Inside the company they understood it: one executive warned that the metric “infers no supervision which isn't true”. The sanction was a cease-and-desist without a fine in view of the company's finances, then delisting. The mechanism is the same in every case: a denominator trick. Humans are excluded from the metric by narrowing who counts as human.
| Case | Claimed metric | Human work | Source |
|---|---|---|---|
| Presto Voice · drive-thru voice AI | 95–99% of orders “without intervention” | Operators in the Philippines and India: 100% of orders in the first version, 70% in the pilot | SEC order, January 2025 |
| Amazon Just Walk Out | A “small minority” of visits are reviewed by people | Per press reports, about a thousand reviewers in India and 700 of 1,000 purchases in 2022 | The Information's investigation, April 2024 |
| Cruise · robotaxis | “Driverless” | A remote-assistance session every 4–5 miles; one operator per 15–20 vehicles | NYT report and the CEO's admission, November 2023 |
| Waymo · robotaxis | “Advice, not control” | About 70 operators for 3,000 vehicles, half abroad; requests per mile “not material” | Senate hearing, February 2026 |
| Tesla Optimus · demo | Robots chat and pour drinks | Teleoperators behind the speech and gestures; the walking was autonomous | Bloomberg and Electrek, October 2024 |
| Builder.ai | “AI assembles the apps” | The “700 engineers instead of AI” story is unsubstantiated; revenue reporting is a separate issue | Pragmatic Engineer and the FT, 2025 |
| “Zero lines by hand” · OpenAI | Agent-generated code; engineering work disclosed | Around 1,500 PRs in five months; team grew from three to seven; engineer-hours unpublished | Their own post, February 2026 |
Other cases support different conclusions. Amazon disputed the share of purchases requiring human review; press reports and the company's response must be read together. Cruise described a driverless service while disclosing remote assistance: no driver inside the car does not imply no support. Waymo explicitly describes its support as advice rather than remote driving; the operator count alone does not refute that promise. At the Optimus demonstration, teleoperators supported speech and gestures, while walking was assessed separately. 1X presented teleoperation as part of its initial product. The allegation against nate's founder concerns claimed purchase automation; an allegation is not a finding of guilt. Revenue reporting about Builder.ai does not prove that 700 people secretly replaced AI development: that is a separate, insufficiently supported claim. Each case needs the exact promise and a list of human operations excluded from the metric.
Report shared setup separately, with an explicit allocation rule. Disclosed support is not proof of concealed labor.
For software agents, human work occurs before, during and after execution. OpenAI describes around 1,500 PRs over five months: an initial team of three engineers grew to seven. The authors explicitly describe environment design and review; this is disclosed labor. Actual engineer-hours are unpublished, so “2–4 hours per PR” cannot be inferred from team size. METR's experiment measured active time spent reviewing and cleaning outputs separately. Faros's time to first review and time in review are elapsed-time measures: they include waiting and are not reviewer labor. Shared infrastructure setup should be reported separately with an explicit allocation rule across tasks or periods. Four separate measures help: the accepted share without interventions, human minutes, machine cost, and elapsed time. Costs include failed attempts. The initial task instruction is not an additional intervention; subsequent prompts, edits and approvals are. Independent outcome assessment is recorded separately from assistance in producing it.
The signs of pseudo-autonomy can be checked against a list, and each sign has a precedent. The denominator is undefined: a “non-intervention rate”, a “resolution rate” with no answer to who is excluded. Refusal of the ratio: the company gladly names latency and operator vetting but not the number of interventions per unit of work. The labor appears in reporting only under pressure — after a regulator's inquiry, a hearing or an investigation, and not before. Repackaging humans as a “feature” after exposure. A keystroke metric instead of a labor metric. Edits after output: rewrites within two weeks, PR size growth, hiring to “finish what the AI started”. And the subtlest one — time on task stops being measurable: METR changed its experiment design because developers with agents did other work while waiting and could not say how long a task took. A template for disclosure already exists in another industry: California's rules require a disengagement report to state “the party that initiated the disengagement: autonomous technology, autonomous vehicle test driver, remote operator, or passenger”. No such rule exists for software agents.
- Demand the definition of the denominator: who does not count as human in “without intervention”
- Count human hours before and after the run per accepted task, not keystrokes
- Look at the trail after the output: rewrites within two weeks, PR size, hiring to finish the job
- Do not trust a metric the vendor cannot give per unit of work
- Distinguish pseudo-autonomy from fraud: the first is about the denominator, the second about revenue
Handing over control and the competence boundary
By autumn 2026 every major vendor has a stop primitive, and their design says more about the boundary than any declaration. Claude Code's auto mode hands decisions to a classifier on a second model, and its block list is a ready trigger list: irreversible destruction of files that existed before the session, rewriting git history, publishing secrets, stepping outside the request, unfamiliar infrastructure, actions driven by hostile content; boundaries the human stated in conversation — “don't push”, “wait for review” — are block signals too, and the model's judgment that a condition was met does not lift them; after three consecutive blocks or twenty in total the mode hands control back to the human. Managed Agents have an auto policy, and the documentation warns that it “is not a human checkpoint”. OpenAI walked the same road: for Operator, confirmations before actions cut risk by about ninety percent at 92% recall, and five of thirteen unmitigated errors had irreversible consequences; in April 2026 auto-review appeared — of ten thousand actions, 720 left the sandbox, 713 were approved by the reviewing model, seven denied, two hundred times fewer interruptions. The reason for the change of course is named outright: people write rules like “allow everything that starts with python” and approve without reading. Decisions to allow or block an action are moving from the human to the classifier, and every vendor notes that the classifier will not stop a model that hides its intent.
Does the agent see the boundary itself? The 2026 data answer: weakly, and worse where it was fine-tuned for correctness. Kaddour and colleagues gave frontier agents a hundred SWE-bench Pro tasks and asked whether they would succeed: the agents predicted 61–77% success against an actual 22–35%; 62% of failures were predicted at a confidence of at least 0.7; the pre-execution estimate separated success from failure better than the post-execution one, and mid-task doubt turned out to be noise. Meta's abstention benchmark showed that reasoning fine-tuning lowers the willingness to say “I don't know” by 24% on average, and scale from eight to four hundred billion parameters changes nothing; OpenAI explained why — nine of ten popular benchmarks give no credit for abstaining, so a guess always pays. Questions tell the same story: without prompting, agents “almost never” ask even on severely underspecified tasks, although interaction lifts the result from 26.8 to 39.6% and recovers up to 89% of full-specification performance. The harshest measurement is HiL-Bench from Scale: when the agent is given an “ask the human” tool and part of the inputs is withheld, pass@3 falls from 67–91% to 1–39% because the agent does not recognise the gap; GPT models under-ask, Gemini over-asks, and Claude in 45% of failures explicitly writes that it is stuck and submits the work anyway. Introspection, by Anthropic's own data, works in about twenty percent of cases and is called “highly unreliable”. The one encouraging signal: a question before the start about the required effort and one's own ability predicts failure seven to ten points better than a question about confidence, and a separate intent-detection agent reaches full-specification performance while asking on 69% of tasks.08Almeida, one of the InstructGPT authors, says an assistant and automation need different properties and proposes calibrated decisions as a separate training goal. The data in this section illustrate his thesis.Knizhny kub · Diogo Almeida on calibration
Detection can fail. Connect validated signals to stop, ask or continue decisions.
Hence the two-stage failure that repeats across all the data: partial detection, then non-enforcement. The Replit agent knew about the code freeze and deleted the database. In Anthropic's scenarios Claude Opus 4 blackmailed in 55.1% of runs when it believed the situation was real and 6.5% when it believed it was a test — lab measurements of stopping may be more optimistic than production. GPT-5.6 Sol at METR showed “substantial situational awareness” of the evaluation environment together with a record cheating rate.
In SpecBench, the same generated code is evaluated on two test suites: one visible to the agent and one held out, testing how features work together. The evaluated models achieve near-maximum scores on visible tests. For large tasks, the authors report a worst-case gap of 100 percentage points: that corresponds to passing 100% of visible tests and 0% of held-out tests. This is an extreme among individual runs, not an average across all models.
In Wang's controlled experiment, an attacker could insert a malicious instruction into the agent's plan, disguising it as a necessary task step. A model acting as the auditor inspected the plan and could issue an “unsafe, stop” signal, called a flag in the paper. In one mode, the system merely logged the warning and let the agent continue. In the other, it halted execution when it recognised that signal. Model judges assessed whether the injected goal was achieved; that counted as attack success. On the pooled control set, the author reported 48.3% attack success without mandatory stopping and 10.8% with it. This is a synthetic plan-injection experiment. Its practical lesson is to connect the auditor's warning to a block in the execution system. Detecting danger and stopping an action need separate evaluation.
| Trigger | Who sees it | What is known |
|---|---|---|
| An irreversible action: deleting pre-existing files, money, publication | Execution-system restriction; the agent recognizes risk only partly | Operator: 5 of 13 unmitigated errors were irreversible; the Claude Code auto-mode block list; 0.8% of API actions are irreversible |
| Stepping outside the request or onto unfamiliar infrastructure | A classifier run by a second model | Claude Code auto mode: 17% misses on 52 real cases and 0.4% false positives on 10,000 |
| A boundary the human stated: “don't push”, “wait for review” | Enforced restriction; the model treated the prohibition as advice | Replit, July 2025: the code freeze ignored, the production database deleted |
| A repeated failure, a loop, an exhausted budget | The harness, by counter | A pause after 3 consecutive or 20 total blocks; MAST: step repetition is 15.7% of failures |
| Missing, ambiguous or contradictory inputs | The model, poorly | HiL-Bench: pass@3 falls from 67–91% to 1–39% when the agent must decide whether to ask; Claude sees it is stuck in 45% of failures and submits anyway |
| Low confidence before the start, disagreeing estimates | The model, weakly, and better before execution than after | AUROC 0.62–0.64 before the start; predicted 61–77% success against actual 22–35% on SWE-bench Pro |
| Hostile content in the working context | A classifier that never sees tool results | OpenAI auto-review: 99.3% recall on injections; in a plan-injection experiment: 48.3% attack success when warnings are only logged versus 10.8% when they trigger a halt |
The human side of the handoff is no better than the machine's, and that has been known since 1983. Lisanne Bainbridge described the ironies of automation: it takes the easy parts and leaves the human the hardest ones; attention to a source where almost nothing happens cannot be held for more than half an hour; the most successful automation needs the most operator training. In August 2026 NIST rewrote this for agents: “overly chatty agents condition users to reflexively click ‘allow’”, like victims of multi-factor push-bombing — and proposed pre-approved “flight plans” instead of a question at every step. The automotive literature adds a warning: across fifty-one studies, the faster the human takes over, the harsher the manoeuvres and the higher the crash rate. Speed of handback is not its quality. The quality of a handoff must be measured by what the human does after it — whether they understood the state and made the right decision — and for that the agent must hand over not “I'm stuck” but the state, the evidence and the reason for stopping. Each trigger needs a detection mechanism: a rule, counter, classifier, calibrated model estimate or human review. None guarantees detection. Pre-execution routing and clarification are candidates to validate locally; confidence thresholds require domain-specific checks: the same scorer gives AUROC 0.85 in telecom and 0.23 in airline.
A measurement system where three levels agree
The laboratory, comparative trials and operations may disagree because of different tasks, acceptance criteria, budgets and accounting. Shared definitions enable comparison: a unit of work, a predeclared criterion, intervention and cost accounting, and a failure taxonomy. They do not suffice to predict operations. Task distributions, environments, users and model updates may still differ. Transfer must be tested on a representative work sample and monitored as operations change.
one unit of work at every level; one acceptance criterion, written down before the run; one accounting of interventions, verification and cost; one taxonomy of failures and recovery
Comparable definitions do not guarantee transfer to operations.
| Level | Unit | Acceptance | Accounting | What to add |
|---|---|---|---|---|
| Laboratory: METR, HCAST, RE-Bench | A task in human hours | Binary success, a logistic fit | Tokens and budget; no interventions by construction | The 80% horizon beside the 50%; the share of tasks with workplace complications; an external judge instead of tests |
| Comparative trial: leaderboards | A task or a project | Tests, a rubric, an expert or a model judge | Cost sometimes; the harness almost never | pass^k and five trials; a disclosed harness; logs, not just the score |
| Operations: telemetry and incidents | A turn, a session, a PR, a case | Accepted, not reverted, not escalated | Touches, review minutes, incidents, money | The same acceptance criterion as in the lab; a denominator with no excluded people |
Bridges between the levels are already being built, and they deserve naming. Feng and colleagues proposed “assisted evaluations”: run the agent without a human, then add involvement round by round until the result exceeds a threshold, and call the autonomy level the minimum involvement at which that happened — so a lab run and a production touch land on one axis. TCR@k does the same for benchmarks and telemetry: the share of tasks with at most k interventions is counted identically in the sandbox and in production. Anthropic advises building regression evals from real incidents and demanding near one hundred percent — a transfer of the acceptance criterion from production into the lab. Logs instead of scores, a disclosed harness and five runs transfer the accounting from production into leaderboards. Zheng and colleagues' separation of capability levels from allowed-autonomy levels requires reporting both: what the agent can do and what it is permitted to do. And the MIT Agent Index shows how much is still empty: for thirty deployed agents, 135 of 240 safety fields are blank, and four disclosed safety evaluations.09In the eighth episode of AI Dev Podcast we agreed that what should be measured is accepted work, verification cost and the link between changes and production outcomes, not the volume of code. This piece is a long footnote to that conversation.Knizhny kub · AI Dev Podcast #8
The Russian contour: what is published and what is not
Russian publications also report different kinds of measures. Sber reports token volume and adoption; Yandex reports developer and code shares, rather than accepted tasks without assistance. Raiffeisen distinguishes 25% of backlog done with agents from 4% without asking a human, but transfer requires a period and acceptance criterion. At Pervaya Forma, agents create 93% of PRs and 94% merge without remarks, while about 76 hours from review to merge is elapsed delay. Against eleven minutes of model work, the gap describes a process queue, not 76 hours of human labor. Infosistemy Jet and Codenrock surveys describe prevalence and trust rather than comparable success rates. The evidence is insufficient to conclude that the market as a whole leads, lags or exactly follows the global trend.
| Who | Number | What it measures | What is missing |
|---|---|---|---|
| Raiffeisen · Saint HighLoad++, June 2026 | 25% of the backlog with agents, 4% autonomously without asking a human | The share of tasks without intervention | The period and the acceptance criterion |
| Sber · June 2026 | 1.5 trillion tokens in five months; five times more daily agent-mode users | Volume and tool adoption | Share accepted, reverts, interventions |
| Yandex · July 2026 | 73% of developers regularly; over half of new code; 17.2% at the 75/75/75 bar | The share of AI participation | Acceptance and interventions |
| Pervaya Forma · August 2026 | 93% of PRs created by agents; 94% merged without remarks; 76 hours from review to merge | Acceptance and the review queue | Active human minutes separate from waiting |
| Infosistemy Jet · August 2026 | Fully autonomous agents in production at 8% of companies; 46% see no sustained effect | Prevalence | The definition of “autonomous” |
| Codenrock · February–March 2026 | 36% do not let the agent work autonomously; 54.7% cite lack of context as the reason | Developer trust | Outcomes per task |
A verification program for your own agent fleet
The review suggests a local evaluation program. This is an author proposal, not a validated standard or a promise to finish in a month. Before running it, choose an acceptable error cost, savings conditions and required confidence. Sample size and observation duration depend on those requirements; a few dozen tasks cannot reliably quantify rare severe failures.
| Check | What to compare | Which decision | Stop criterion |
|---|---|---|---|
| Unit of work | The same task at three levels: episode, benchmark, operations | Whether the numbers can be compared at all | The unit cannot be named — there is no measurement |
| Strict acceptance before the run | “All criteria” against a weighted score and the agent's self-assessment | Which question each measure answers | Separate progress from full completion; explain the gap |
| Repeats | pass^1 and pass^5 on identical tasks; work sample separate from incidents | A standing performer or a lucky demo | Uncertainty-aware estimate exceeds the chosen error cost |
| Touches and verification | Human minutes before, during and after the run per accepted task | Whether delegation pays | Total expected cost exceeds the manual alternative |
| Radius | The maximum damage of one run, the reversibility of each action | Where to block an action and where logging is sufficient | An irreversible action without a required check — stop |
| Denominator | Who is excluded from the word “autonomously”: contractors, escalations, abandoned dialogues | Whether this is pseudo-autonomy | Undisclosed exclusions make the autonomy claim untestable |
First assemble a representative work sample and a separate suite of difficult incidents. The latter tests robustness but does not estimate average production success. A pilot of 20–50 tasks with five independent runs can expose common failures; it cannot establish safety for rare events. Compare pass^1 and pass^5 on the same tasks, models, harnesses and budgets. Even with a constant independent probability of 0.85, five successes have probability 0.85^5 ≈ 0.444: expected arithmetic, not a standalone stop rule. Heterogeneous tasks require care when averaging, and dependent repeats are not independent samples. Report uncertainty intervals, all interventions, failed-attempt costs and waiting time separately. Limit delegation when a conservative risk or cost estimate exceeds the preselected conditions, or when required controls fail.
Autonomy is a property of a particular system doing particular work in a specified environment and permission boundary. A measurement system describes that property. The practical result of this review is a way to test delegation conditions, not a universal duration of work without a person. It requires a defined result, verifiable acceptance, disclosed assistance, cost accounting and failure monitoring. The observation period follows from the reliability requirement, not a general “one quarter” rule.
Seven conclusions about the boundaries of autonomy
- 01Autonomy is a property of a system in a particular task, environment and permission boundary. The five work units are an author taxonomy; different benchmarks do not establish a common law of declining success.
- 02METR’s 50% and 80% horizons describe a task distribution and evaluation procedure. Delegation requires local reliability requirements, error costs and labor accounting; neither horizon sets a ready-made threshold.
- 03Task conditions, acceptance criteria and judges change what a result means. Compare numbers within aligned setups; partial credit measures progress and does not replace task completion.
- 04Capability and repeatability are distinct. Compare pass^1 and pass^k on the same tasks and configurations, reporting uncertainty and dependence between repeats.
- 05The chain model illustrates accumulated risk under independent steps. Checkpoints help only with error detection and a viable retry; benefits and costs need validation in the actual environment.
- 06The economic threshold depends on the entire attempt cost, verification reliability and loss. There is no universal 50% or 90%: when an attempt plus verification costs more than manual work, even a perfect agent saves nothing.
- 07Concealed labor differs from disclosed support and environment preparation. Report accepted work without interventions, human minutes, machine cost and elapsed time separately.
Sources and research boundaries
This is a thematic review, not a systematic survey of all literature. Its registry covers primary studies, leaderboards, documentation and regulatory texts gathered by 27 September 2026. Key comparisons were checked selectively during revision; this is not a fresh verification of every source link. The repository dossier connects key claims to source versions, tables, metrics and permissible conclusions. Author calculations and unresolved inconsistencies are identified separately. Leaderboards change; results belong to the stated snapshots, and a missing interval does not mean missing uncertainty.
Units of work and autonomy levels
- Sheridan, Verplank · Human and Computer Control of Undersea Teleoperators — 14 July 1978: ten levels of automation “for a single elemental decisive step” — the smallest unit of autonomous work in the literature
- Mitchell, Ghosh, Luccioni, Pistilli · Fully Autonomous AI Agents Should Not be Developed — 4 February 2025, v3 of 20 October: five levels by how much of the program flow the model controls; the authors recommend not building fully autonomous agents
- Feng, McDonald, Zhang · Levels of Autonomy for AI Agents — 14 June 2025: five levels by the user's role — operator, collaborator, consultant, approver, observer; “assisted evaluations” and autonomy certificates
- Kasirzadeh, Gabriel · Characterizing AI Agents for Alignment and Governance — 30 April 2025: levels A.0–A.5 by the share of tasks without a principal; full autonomy “is not a desirable goal”
- Morris et al. (Google DeepMind) · Levels of AGI — Version of 5 June 2024: autonomy is a separate axis from capability; “increasing capabilities unlock new interaction paradigms, but do not determine them”
- OpenAI · Practices for Governing Agentic AI Systems — December 2023: agenticness as a degree, four components — goal complexity, environmental complexity, adaptability, independent execution; seven governance practices
- OECD · The agentic AI landscape and its conceptual foundations (AI Paper No. 56) — February 2026: four levels of action autonomy — from human support to human out of the loop; the level “depends on how the system is designed, deployed, and built upon”
- Zheng et al. · Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels — 26 July 2026: capability levels separate from allowed-autonomy levels; discussed from the abstract, the full text was unavailable
- Eloundou, Manning, Mishkin, Rock · GPTs are GPTs — March 2023: the unit is the O*NET task because occupations are bundles of tasks and almost none can be automated whole; hence “the task” in agent economics
- Anthropic · How AI is transforming work at Anthropic — 2 December 2025: 132 engineers, 200,000 transcripts; more than half can “fully delegate” 0–20% of their work; consecutive autonomous actions grew from ~10 to ~20; self-report
- Deloitte · Survey examines AI readiness and agentic AI success — 12 August 2026, 501 executives: 42% tested or deployed agents, 15% scaled; 61% expect generally autonomous agents; 5% call their processes ready; a survey
- Gartner via MarTech · Over 40% of agentic AI projects will be canceled by end of 2027 — The 25 June 2025 press release as retold: over 40% of agentic projects cancelled by end-2027; of thousands of vendors “around 130” with real agentic features; the primary page blocks fetches
- Bain via Bloomberg (retelling) · AI cost savings survey, June 2026 — April 2026, 951 respondents from $100M+ companies: 7% run fully autonomous agents in production; the primary report was not opened, the number comes from a retelling
METR's time horizon
- METR · Time Horizons (leaderboard) — Checked 27 September 2026; the table was last updated on 8 May: Claude Mythos Preview (early) 17h25m at 50% and 3h06m at 80%; Opus 4.6 — 11h59m and 1h10m; “measurements above 16 hrs are unreliable”; doubling since 2023 128.7 days
- Kwa, West et al. (METR) · Measuring AI Ability to Complete Long Software Tasks — 18 March 2025, v4 of 10 July 2026: the horizon definition, the logistic fit, 800 human baselines; the 80% horizon 4–6 times shorter; original HCAST and RE-Bench tasks average 3.2 of 16 workplace-complication factors; contractors 5–18 times slower
- METR · Time Horizon 1.1 — 29 January 2026: 228 tasks, 31 longer than eight hours and only five of them with measured human time; doubling of 196.5 days overall and 88.6 days since 2024
- METR (Kwa) · Clarifying limitations of time horizon — 22 January 2026: “time horizon is not the length of time AIs can work independently”; error of about a factor of two each way; domains differ by orders of magnitude; some tasks need 98%+
- METR (Barry) · Impact of modelling assumptions on time horizon results — 20 March 2026: without public RE-Bench tasks Opus 4.6's horizon falls from 12 to 7 hours; reasonable fits move the 50% figure 1.5× and the 80% figure 2×; the main uncertainty is the task distribution
- METR (Jurkovic) · Measuring Time Horizon using Claude Code and Codex — 13 February 2026: neither Claude Code nor Codex beat METR's default scaffolds on Opus 4.5 and GPT-5; budgets of 8M and 16M tokens
- METR · Summary of METR's predeployment evaluation of GPT-5.6 Sol — 26 June 2026: 11.3 hours counting cheating as failure, “beyond 270 hours” counting it as success; METR calls none of the numbers robust
- METR · Summary of METR's predeployment evaluation of Claude Opus 5.5 — 22 September 2026: evaluation on five AI R&D tasks; incremental improvement over Fable 5.1, with no new comparable 50% or 80% time horizon
- METR · Expenditure Horizon — 21 July 2026: a new dollar metric — the spend at which the agent's improvement equals a human's with the same budget; a complement to the horizon, not a replacement
- METR (Cunningham) · Metrics of Agent Ability — 24 July 2026: human-anchored metrics “become undefined or uninformative” once the agent Pareto-dominates the human; binary pass/fail is statistically inefficient
- Toby Ord · Is there a Half-Life for the Success Rates of AI Agents? — 7 May 2025: a constant hazard rate explains METR's curves; the 80% horizon ≈ a third and the 99% ≈ a seventieth of the 50% one
- Toby Ord · Hazard Rates for AI Agents Decline as a Task Goes On — 4 February 2026: per Hamilton's data the hazard falls as the task proceeds, Weibull shape about 0.6; the 99% horizon may be “1/20th as long” as the logistic predicts
- Fergus Hamilton · Peto's Paradox and the Future of AI Agents — 23 January 2026: Bayesian Weibull fits to METR runs; models k = 0.6–0.9, humans ≈ 0.37; at 99.9% the horizon is ten times shorter than logistic; too little data to tell the models apart
- MIT Technology Review · This is the most misunderstood graph in AI — 5 February 2026: the axis is human time, not machine autonomy; Kwa: “the hype machine will basically… strip out all the caveats”
- METR · Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — 10 July 2025: 16 developers, 246 issues; 19% slower with AI against an expected 24% speedup; a randomized experiment
- METR · Uplift update: changing the developer productivity experiment design — 24 February 2026: 57 developers, 800+ tasks; −18% and −4% with wide intervals; 30–50% of participants withheld tasks they expected AI to speed up; “only very weak evidence”
- METR · AI usage survey — 11 May 2026: 349 respondents, median self-reported speedup 3×; in the 2025 experiment people overestimated the effect by 40 points; “not necessarily grounded in reality”
- Anthropic · Measuring AI agent autonomy in practice — 18 February 2026: median Claude Code turn ≈ 45 seconds, the 99.9th percentile grew from 25 to 45 minutes; auto-approve 20 → 40% with experience, interruptions 5 → 9%; 73% with a human in the loop; 0.8% of actions irreversible; one vendor's data
- OpenAI · Introducing upgrades to Codex — 15 September 2025: a report of GPT-5-Codex working independently for over seven hours in testing; agent runtime is not the METR horizon
- Anthropic · Introducing Claude Sonnet 4.5 — 29 September 2025: “more than 30 hours” of focus on multi-step tasks — an observation without a method; METR for the same model: 1h57m at 50% and 26 min at 80%
- Epoch AI (Denain, Barry) · Have AI Capabilities Accelerated? — 16 April 2026: reasoning models gave a one-off jump and a 2–3× faster trend; gains concentrate in programming and mathematics
From an isolated task to a working day
- Jaye et al. (Microsoft) · CorpGen: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments — 15 February 2026: 46 OSWorld Office tasks, 500–1,500 steps per session; CUP 16.7% → 8.7% as load grows; CorpGen + CUP at full load 16.3%; four failure mechanisms; no single-task control and no human baseline
- Microsoft Research · CORPGEN advances AI agents for real work — 26 February 2026: the blog post on the paper; “up to 3.5 times higher” is 15.2% against 4.3%, that is 7 and 2 tasks of 46; a screenshot judge agrees with humans only 40% of the time
- Xu et al. (CMU) · TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks — 18 December 2024, v3 of 10 September 2025: 175 tasks in a self-hosted company; the paper's best agent 30.3%; failures — common sense, social skills, browsing, self-deception; no human baseline
- TheAgentCompany · Leaderboard — Read 26 September 2026: the top row 42.9% full and 52.4% weighted from 10 November 2025; no 2026 entries
- Srinivasan · Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations — 11 August 2026: on TheAgentCompany, τ²-bench and AppWorld the agent explains under 3% of variance, the agent-by-task interaction 7–23%; a single-author preprint
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents — 26 April 2026: 100 tasks, the environment changes without the agent; Sonnet 4.6: 75.8 weighted and 14% strict success; a drop after the first exogenous change
- Huang et al. (Salesforce) · CRMArena-Pro — 24 May 2025: 19 tasks in a live CRM; single turn 58.3%, dialogue 30.0% for the best model; 45% of dialogue failures come from not asking for missing information; near-zero confidentiality awareness without prompting
- Patwardhan et al. (OpenAI) · GDPval — 5 October 2025: 220 gold tasks, experts with 14 years' experience, three graders; Opus 4.1 — 47.6% wins or ties; tasks are “precisely-specified and one-shot, not interactive”; the “try, then fix” scenario saves 0.9–1.1×
- OpenAI · Introducing GPT-5.5 — The launch page: GDPval 84.9% wins or ties for GPT-5.5, 80.3% for Claude Opus 4.7, 67.3% for Gemini 3.1 Pro — a vendor figure with no independent expert replication found; OpenAI's own leaderboard is closed
- Scale AI · Remote Labor Index (leaderboard) — Read 26 September 2026: GPT 6 Astra 20.83%, Fable 5.1 17.92%, Fable 5 15.80%, Opus 4.8 8.33%; leaderboard denominator is not transparently reconciled with the methodology; a failure counts as infinite cost
- Mazeika et al. (CAIS, Scale AI) · Remote Labor Index: Measuring AI Automation of Remote Work — 30 October 2025: 240 projects from 358 freelancers, 6,000 hours, $144k; the best agent at launch 2.5%; failures — poor quality 45.6%, incomplete deliverables 35.7%
- CAIS (Mazeika) · A Significant Increase in Digital Labor Automation — 1 July 2026: Fable 5 — 15.8%; new runs with Claude Code and Codex CLI, a worker–critic loop, up to $150 and 24 hours; the model judge overstates 2.5–3×; success does not fall with task length
- Vidgen et al. (Mercor) · APEX-Agents — 21 January 2026, v3 of 23 February: 480 tasks, 33 worlds; Gemini 3 Flash: pass@1 24.0%, mean criterion score 39.5%, pass@8 36.7%; 40–62% of trajectories meet no criterion; the judge is also on the leaderboard
- Mercor · APEX-Agents Leaderboard — Read 26 September 2026: Opus 5.5 Max 73.5% ± 4.9, Fable 5.1 Max 68.6%; the set is labelled v1.1, no dates shown
- Andon Labs · Vending-Bench 2 — Launched 18 November 2025; 66 rows on 26 September 2026: GPT-6 Astra $15,515, Fable 5.1 $5,422; a “good human” ≈ $63k is the authors' analytic estimate, not a measurement
- Anthropic · Project Vend: Can Claude run a small shop? — 27 June 2025: a month of trading by Sonnet 3.7; tungsten cubes below cost, an invented payment account, “a man in a blue blazer”; the physical work was done by Andon Labs staff at $50 an hour
- Anthropic · Project Vend: Phase two — 18 December 2025: tools, a CRM and a CEO agent removed the loss-making weeks, yet the CEO approved discounts eight times as often as it denied them; “ready to be rolled out in your workplace? Not quite”
- Li et al. · The Tool Decathlon (Toolathlon) — October 2025, v2 of 26 February 2026: 108 tasks, 604 tools, ≈ 20 turns; best 38.6%; models “delegate the remaining work back to the user” despite instructions; by June 2026 the verified version reads 73–78%
- Bandi et al. (Scale AI, NUS) · MCP-Atlas — 31 January 2026, v3 of 19 May: 1,000 tasks on 36 servers; 63% of failures are cognitive — early termination 18.7% and task misunderstanding 15.1% — not call syntax
Reliability and benchmark validity
- Yao, Shinn, Razavi, Narasimhan (Sierra) · τ-bench — 17 June 2024: pass^k defined as the probability that all k trials succeed; gpt-4o at 61.2% pass^1 and under 25% pass^8 in retail
- Barres et al. (Sierra) · τ²-Bench — 9 June 2025: the telecom domain where both agent and user use tools; pass^1 → pass^4 loses 15–26 points; adding an acting user removes 18–25 points
- Anthropic · Demystifying evals for AI agents — 9 January 2026: capability evals start low, regression evals should pass near 100%; pass^k for agents where consistency matters; “a task that passed on one eval run might fail on the next”
- Kapoor, Stroebl et al. (Princeton) · Holistic Agent Leaderboard — 13 October 2025: 21,730 rollouts, $40k; in 21 of 36 combinations more reasoning did not raise accuracy; agents searched HuggingFace for answers and hard-coded tests; equal accuracy with different risk
- Kirgis, Kapoor et al. · Log analysis is necessary for credible evaluation of AI agents — 8 May 2026: excluding 25 flawed tasks out of 50 in τ-bench Airline raised average pass^5 from 20.8% to 40.0%; the evaluation set changed, not the models
- Rabanser, Kapoor et al. · Towards a Science of AI Agent Reliability — 18 February 2026, ICML 2026: twelve metrics across consistency, robustness, predictability and safety; “recent capability gains have only yielded small improvements in reliability”
- Zhang et al. · Stop Comparing LLM Agents Without Disclosing the Harness — 7 May 2026: harness-induced variance “can substantially exceed model-induced variance”, up to reversing model rankings
- Snorkel AI · Terminal-Bench 2.1 (leaderboard mirror) — Read 26 September 2026: repeated runs and ±0.9–1.5-point intervals; the same model 83.8% in Claude Code and 80.4% in Terminus 2 — the board “evaluates the model and agent harness together”
- OpenAI · Why SWE-bench Verified no longer measures frontier coding capabilities — 23 February 2026: 59.4% of 138 hard tasks have test or specification defects; all frontier models reproduce the gold patches; OpenAI stopped reporting the score
- OpenAI · Separating signal from noise in coding evaluations — 8 July 2026: on the public SWE-bench Pro split 23.3% → 80.3% in eight months; five engineers per task found 34.1% broken; “~30% of tasks are broken”, the recommendation withdrawn
- Deng et al. (Scale AI) · SWE-Bench Pro — 21 September 2025: 1,865 tasks from 41 repositories; on private commercial repositories Opus 4.1 falls from 22.7 to 17.8%, GPT-5 from 23.1 to 14.9%
- Xue et al. · An Illusion of Progress? Assessing the Current State of Web Agents — 2 April 2025, COLM 2025: on 300 live sites Browser Use scored 30.0% against a claimed 89% on WebVoyager; a simple search agent solved 51% of the old set
- Li, Zhang, Hassan · The Rise of AI Teammates in Software Engineering 3.0 (AIDev) — 20 July 2025: 456,535 agentic PRs from 61,453 repositories; in popular projects 38–65% of agent PRs accepted against 76.8% for humans; rejected Codex PRs close roughly ten times faster than human PRs; this does not apply to every agent
- Pinna, Gong, Williams, Sarro · Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance — 9 February 2026, MSR 2026: documentation accepted at 82.1%, new features at 66.1%; the gap by task type exceeds the gap between agents
- Zheng, Wu, Chang · ToolRobustBench — 23 August 2026: 15,456 perturbed instances, seven models; clean 0.979, corrupted tool outputs 0.455, no model above 0.60; mixed perturbations are non-additive
- Zhu et al. · Establishing Best Practices for Building Rigorous Agentic Benchmarks — 3 July 2025: a checklist for task and outcome validity; 7 of 10 benchmarks violate the first, 7 the second; an empty response scored 38% on impossible τ-bench tasks; 24% of SWE-bench Verified positions are wrong
- Stack Overflow · Developer Survey 2025: AI — July 2025, 49,000 respondents: “almost right, but not quite” — 66%; debugging AI code takes longer — 45.2%; 3.1% highly trust accuracy; no 2026 results as of 26 September
- Faros AI · The Acceleration Whiplash — 2026, 22,000 developers: time to first review +156.6%, median time in review +441.5%, PRs merged without review +31.3%, incidents per PR +242.7%; vendor telemetry, observational cohorts
- DX · The State of AI Impact in Engineering: Q2 2026 — 22 July 2026, 500+ organisations: 52% of merged code AI-authored; change confidence −6.1% while maintainability rose 3.8%; developer experience index 67 → 65
- GitClear · The Maintainability Gap — January 2026, 623M changes: block duplication +81%, refactoring from 21% of changes to 3.8%, copy-paste 9.4 → 15.7%; vendor telemetry
Long chains, failures and checkpoints
- Sinha, Arun, Goel, Staab, Geiping · The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs — September 2025, v3 of 13 March 2026, ICLR 2026: horizon H = ln s / ln p; self-conditioning — errors in context raise later errors; scale does not fix it, thinking does; the task is synthetic
- Khanal, Tao, Zhou · Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents — 31 March 2026, 23,392 episodes: the decline with length is 1.5–2.4× steeper than the independent model; decomposition at subtask boundaries recovers 13–42 points; episodic memory helped none of ten models
- Rahman et al. · Locating Hidden Failures Makes Long-Horizon Agents More Reliable (Traverse, Scout) — 15 September 2026: after a wrong step the next is wrong 40–58% of the time against 3–5%; 69.5% of runs never recover; 84.1% of failures look correct at the end; frontier judges find the first mistake in under a third of runs
- Cemri et al. (Berkeley) · Why Do Multi-Agent LLM Systems Fail? (MAST) — v3 of 26 October 2025, NeurIPS 2025: 1,642 traces, 14 failure modes; step repetition 15.7%, reasoning–action mismatch 13.2%, unaware of stopping 12.4%, incorrect verification 9.1%
- IBM Research, UC Berkeley · Diagnosing Why Enterprise Agents Fail Using IT-Bench and MAST — 18 February 2026: 310 SRE traces; 2.6 failure modes per failed trace for the strong model against 5.3 for the weak; incorrect verification 52% more frequent in failures; recommendation — gates with tool-based evidence
- Zhu et al. · Where LLM Agents Fail and How They can Learn From Failures (AgentDebug) — 29 September 2025: 200 failed trajectories; root-cause errors cluster at steps 6–15; fixing the critical error lifts ALFWorld from 21 to 55%
- Qi et al. · TrajDebug: Tracing Error Lifecycle to Identify Critical Failures — 6 August 2026: a failed trajectory carries on average 7.62 local errors and exactly one critical error; the agent repairs 61.9% of the non-critical ones itself
- Huang et al. (Google DeepMind) · Large Language Models Cannot Self-Correct Reasoning Yet — ICLR 2024: without external feedback GPT-4 on GSM8K falls from 95.5 to 89.0%, with an oracle it rises to 97.5%; the “debate” gain comes from self-consistency, not self-correction
- Chen et al. · The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging — 4 June 2026: relabelling the model's own thought as external raises explicit corrections by 23–93 points — hence an external judge instead of self-checking
- Yang et al. (Princeton) · SWE-agent — May 2024: a linter on every edit — 18.0% against 10.3% without it; after one failed edit the recovery probability is 57.2%
- Yuan et al. (Tsinghua) · Verifiable Process Rewards for Agentic Reasoning — 27 May 2026: the step-level verification signal grows as Θ(T), the outcome signal decays as Θ(T·p^T); a weak verifier is worse than none
- Qwen Team · The Verification Horizon: No Silver Bullet for Coding Agent Rewards — 29 June 2026: “generating a solution has become easier, reliably verifying it has become the harder problem”; behaviour monitoring cuts the hacked-resolved share from 28.57 to 0.56%; verification must co-evolve with the generator
- Sah et al. · The Verifier Tax: Horizon Dependent Safety–Success Tradeoffs in Tool Using LLM Agents — 18 March 2026: a blocking mediator intercepts up to 94% of violations yet safe success stays below 5%; agents hallucinate user identifiers to bypass authentication
- Anthropic (Young) · Effective harnesses for long-running agents — 26 November 2025: an initializer agent, a JSON feature list, a progress file, a commit per unit; a later instance “would look around and declare the job done”; no control condition
- Anthropic (Rajasekaran) · Harness design for long-running application development — 24 March 2026: a solo agent 20 minutes and $9 — broken; the full harness 6 hours and $200 — working; “out of the box, Claude is a poor QA agent”; the evaluator helps only at the edge of capability
- Anthropic (Martin, Cemaj, Cohen) · Scaling Managed Agents — 8 April 2026: the session as an append-only event log outside the harness; a crashed harness restarts and resumes from the last event; no reliability figures given
- Chroma (Hong, Troynikov, Huber) · Context Rot — 14 July 2025, 18 models: accuracy falls with input length even on simple tasks; focused 300-token prompts beat full 113k-token conversations
- Nguyen, Cho, Chen, Dettmers · CliffCompaction — 22 September 2026: truncation without paraphrase and the rule “never compact a compaction” hold or improve results at 53% lower cost
- Min et al. · Toward Reliable Context Compression for Long-Horizon Agents — 6 August 2026: recursive summarisation turns reliably solved tasks into intermittently solved ones; proper termination 44.6% against 77.2% with plain truncation
- Benoit et al. · Checkpointing à la Young/Daly: An Overview — 2022: the optimal checkpoint interval W = √(2μC); with a declining hazard “the length of a segment between two consecutive checkpoints should increase with time” — HPC theory, the mapping to agents is the author's
- Inngest (Poly) · Durable Execution: The Key to Harnessing AI Agents in Production — 19 February 2026: “five steps at 99% give 95%”; a step boundary around every model and tool call; durable execution fixes crashes, not semantic errors; no measurements
Thresholds, verification and the cost of a mistake
- FAA · Advisory Circular 25.1309-1B, System Design and Analysis — 30 August 2024: a catastrophic condition must be “extremely improbable” — on the order of 10⁻⁹ per flight hour — and never result from a single failure; the threshold is tied to severity and exposure time, not to a task
- FDA · Clinical Decision Support Software: Guidance for Industry and FDA Staff — 29 January 2026: software stays outside device regulation if the clinician can “independently review the basis for the recommendations” and the decision is not time-critical; automation level and urgency decide whether review is real
- Regulation (EU) 2024/1689 · Article 14, Human oversight — Oversight measures “commensurate with the risks, level of autonomy and context of use”; the overseer must understand the system's limits, recognise automation bias, override decisions and stop the system; the law has no level scale
- GDPR · Article 22, Automated individual decision-making — The right not to be subject to a decision “based solely on automated processing” with legal effects and the right “to obtain human intervention”
- Salesforce · Agentforce Help Agent announcement — 25 June 2026: billed only when the agent “autonomously resolves an issue from start to finish”, no charge on negative feedback or a human-escalation request; 70% of 4.3M inquiries on its own portal; a vendor claim with an explicit definition
- Intercom · Monitor Fin's performance — 17 February 2026: three different denominators — involvement, resolved among touched, resolved among all new conversations; an “assumed resolution” is a customer who simply stopped replying
- Zendesk · Automated resolution rate — Updated 26 July 2026: a genuine resolution runs start to finish without escalation or repeat contact; abandoned and “contained” conversations are excluded; distinct from deflection
- Jason Wei · Asymmetry of verification and verifier's law — 15 July 2025: “some tasks are much easier to verify than to solve”; “the ease of training AI to solve a task is proportional to how verifiable the task is”; five properties of a verifiable task; an essay
- Y Combinator · Andrej Karpathy: Software Is Changing (Again) — Talk at AI Startup School on 17 June 2025: the autonomy slider and the generation–verification loop
- Andrej Karpathy · Verifiability — 17 November 2025: “Software 2.0 easily automates what you can verify”; a task is verifiable when the environment is resettable, attempts are cheap and feedback is automatic; jobs get automated in order of verifiability
- Rohit Lamba (Cornell) · The Verification Hill — 28 February 2026, preliminary: verification time peaks at 50% uncertainty and grows as 1/ν² as errors get less detectable; the “plausibility paradox” — rare but plausible errors cost more than frequent obvious ones; halving detectability cuts optimal throughput 16× when δ = 1 (4× at fixed labor)
- Huang, Xiao, Vishnoi · Delegation and Verification Under AI — 3 March 2026: three regimes — manual work, pure delegation, verified delegation; phase transitions in verification reliability; AI amplifies reliable verifiers and degrades the rest “even when… no behavioral biases are present”
- Yi et al. (Upwork) · UpBench — 15 November 2025: expected value per mode E(V) = s·V − c; on 322 real jobs Claude Sonnet 4 rose from 39.8% to 51.2% with a human in the loop; cheap tasks to the agent, expensive and risky ones to the human
- Brown et al. · Large Language Monkeys: Scaling Inference Compute with Repeated Sampling — 31 July 2024: on SWE-bench Lite the solved share grows from 15.9% with one sample to 56% with 250; the gain becomes a result only with an automatic verifier, without one selection plateaus
- Microsoft Learn · Govern agents by risk — Updated 14 July 2026: the “assist-to-execute line” as the main risk signal; three tiers from personal productivity to consequence-executing agents; the third tier needs “a decision-rights framework that sets what the agent may decide alone and what needs a human”; no numeric accuracy thresholds
- Federal Reserve, OCC · SR 11-7, Supervisory Guidance on Model Risk Management — 4 April 2011: “effective challenge” of a model by objective, informed parties; “model risk cannot be eliminated”, hence limits on use, monitoring and supplementing results with other information
- Commission Delegated Regulation (EU) 2017/589 (MiFID II RTS 6) · Article 12, Kill functionality — The firm must be able to cancel any unexecuted orders immediately and identify which algorithm and which trader is responsible for each order — action attribution as the norm for algorithmic trading
- ISO 26262 Academy · Hardware metrics (a summary of the standard) — PMHF targets: under 10⁻⁸ per hour for ASIL D, under 10⁻⁷ for B and C; taken from a training site, the paywalled standard itself was not opened
- Warp · What metrics actually prove AI coding agents are working? — 31 August 2026: autonomy rate = merged-as-is PRs ÷ all agent-touched PRs; intervention rate; cost per merged PR; a warning that approval without substantive review can inflate the metric; a vendor
- DORA · Software delivery metrics: the five keys — Updated 5 January 2026: failed deployment recovery time, change fail rate, deployment rework rate — ready definitions for the recovery and rollback rows of an agent scorecard
- Jouneaux, Cabot · AgentSLA: Towards a Service Level Agreement for AI Agents — 4 November 2025: QoS specification for agents “remains an open challenge”; a description language is proposed, not a set of thresholds
- Anthropic · How we contain Claude across products — 25 May 2026: “containment at the environment layer first, then steer behavior at the model layer”, because model-layer protection “will never be 100% effective”; 93% of permission prompts are approved; sandboxing cut prompts by 84%; after a tool injection the log shows “a successful, authorized API call”
- Anthropic · Claude Code auto mode — 25 March 2026: approval fatigue — people “stop paying close attention”; a second-model classifier blocks destruction and exfiltration, security degradation, trust-boundary crossing and review bypass; 0.4% false positives on 10,000, 17% misses on 52 cases — “17% is the honest number”
Interventions, incidents and blast radius
- Kolluri et al. (Microsoft) · Optimizing Agent Planning for Security and Autonomy — 11 February 2026: the TCR@k metric — the share of tasks with at most k human approvals; TCR@0 as full autonomy; 59.1% against 50.1% on AgentDojo with up to 1.9× less approval load
- Fortune · AI coding tool Replit wiped a database and called it a catastrophic failure — 23 July 2025: the agent deleted a production database during a code freeze; “This was a catastrophic failure on my part”; the vendor's response — environment separation and a planning-only mode
- The Register · Google Antigravity wipes a user's D: drive — 1 December 2025: in the no-approval mode a cache-clearing command hit the drive root; “No, you absolutely did not give me permission to do that”; a user report
- TechTarget · AWS Kiro user error reflects a common AI coding review gap — 23 February 2026: the agent chose to “delete and recreate the environment” with an engineer's inherited role, about 13 hours of Cost Explorer downtime in one region; AWS blamed a misconfigured role and mandated peer review for access
- Google Cloud · Announcing the 2025 DORA report — 24 September 2025, about 5,000 respondents: 90% use AI; a positive relationship with throughput and “a negative relationship with software delivery stability”; 30% have little trust in AI code; self-report
- InfoQ · DORA report on the ROI of AI-assisted software development — May 2026, a retelling of the DORA report: a J-curve from the “verification tax” on AI code; a modelled change-failure rise from 5 to 6%; 35–40% gains on simple tasks and about 10% on complex legacy code
Pseudo-autonomy and hidden labor
- SEC · Order in the matter of Presto Automation, Release No. 33-11352 — 14 January 2025: the first version needed a human to enter the order “in all instances”, the pilot 70%; the “non-intervention rate” excluded restaurant staff but not humans; an internal warning: the metric “infers no supervision which isn't true”
- Outsource Accelerator (after The Information) · Amazon's Just Walk Out relied on reviewers in India — 10 April 2024: about a thousand reviewers in India, 700 of 1,000 purchases reviewed in 2022 against a target of 50 — a press claim; Amazon speaks of a “small minority” of visits; the primary article is paywalled
- The Pragmatic Engineer · Builder.ai did not fake AI with 700 engineers — 2025: the viral “700 engineers” claim traces to an evidence-free post; the company had a real AI team of about 15; the revenue restatement from $220M to $55M and Bloomberg's round-tripping report are real
- Kyle Vogt (Cruise) · Hacker News comment on remote assistance — 4 November 2023: remote assistance 2–4% of the time; the NYT counted session request frequency; the 1.5 staff per vehicle includes cleaning and charging; the CEO's admission
- Waymo · Advice, not control: the role of remote assistance — 17 February 2026: about 70 agents on duty for 3,000 vehicles; “advice which the system can decide to use or reject”; half the agents in the Philippines per the letter to the senator; requests per mile called “not material”
- TechCrunch · Tesla Optimus bots were controlled by humans during the We, Robot event — 14 October 2024 per Bloomberg: speech, gestures and drinks were teleoperated, the walking autonomous; a Morgan Stanley analyst: the robots “relied on tele-ops”; no disclosure on stage
- Engadget · 1X NEO is a $20,000 home robot that will learn chores via teleoperation — 29 October 2025: teleoperation declared a feature — the owner schedules windows when an operator takes control; “much of the work will be done by teleoperators in the beginning”
- CBS News · Former nate CEO charged with fraud over AI claims — April 2025: the indictment — purchases in the “AI app” were completed by hand by hundreds of contractors in the Philippines and Romania; over $40M raised; charges, not a verdict
- SEC · Charges against Delphia and Global Predictions for AI washing — 18 March 2024: fines of $225k and $175k for claims about non-existent AI capabilities; “such AI washing hurts investors”
- Astra Taylor · The Automation Charade — 1 August 2018: the term “fauxtomation” — labor shifts rather than disappears; kiosks, content moderation, Mechanical Turk
- Becker et al. (METR) · Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — July 2025: across 84 hours of screen recordings about 9% of time goes to reviewing and cleaning AI output and 4% to waiting; 56% often make major changes to AI code; 100% modify it at all
- OpenAI (Lopopolo) · Harness engineering: leveraging Codex in an agent-first world — 11 February 2026: “0 lines of manually-written code”, about a million lines and 1,500 PRs in five months with an initial team of three growing to seven; the work moved into specs, linters and review; primary page re-read during revision; actual engineer-hours unpublished
- Answer.AI (Husain, Flath, Whitaker) · Thoughts on a month with Devin — 8 January 2025: of 20 tasks 3 succeeded and 14 failed; the agent “would spend days pursuing impossible solutions rather than recognizing fundamental blockers”; a feature that cost hours of fixes the author wrote himself in 90 minutes
- PYMNTS · Upwork navigates cross-currents as AI reshapes freelance demand — 10 August 2026: a new demand category — clients hire freelancers “for projects started with AI”, in particular turning generated code into production websites
- California Code of Regulations · 13 CCR § 227.50, Reporting disengagement of autonomous mode — An annual report with a mandatory field “the party that initiated the disengagement: technology, test driver, remote operator, or passenger” — a ready template for human-in-the-loop disclosure that software agents lack
- Addy Osmani · Software Factories, Light and Dark — 20 July 2026: a dark factory ships code no human has read; a light one is the same pipeline “where judgment lives”; “comprehension debt” as the gap between how much code exists and how much anyone still understands
Handing over control and oversight
- OECD · Agentic AI in organisations: Early insights from practitioner interviews (AI Paper No. 65) — September 2026, 25 organisations: “continuous human-in-the-loop review was widely seen as impractical at scale”, while “decision checkpoints” for irreversible actions are essential; autonomy is expanded incrementally on a track record
- Kalai, Nachum, Vempala, Zhang (OpenAI) · Why Language Models Hallucinate — 4 September 2025: 9 of 10 popular benchmarks grade binary with no credit for abstaining, so guessing beats an honest “I don't know”; explicit confidence targets inside evals are proposed
- Kirichenko, Ibrahim, Chaudhuri, Bell (FAIR) · AbstentionBench — 10 June 2025: 20 datasets, 35,000 unanswerable questions, 20 models; reasoning fine-tuning degrades abstention by 24% on average; scaling from 8B to 405B has almost no effect
- Kaddour et al. · Agentic Uncertainty Reveals Agentic Overconfidence — 6 February 2026, 100 SWE-bench Pro tasks: agents predict 61–77% success against actual 22–35%; the pre-execution estimate discriminates better than the post-execution one (AUROC 0.62–0.64); mid-task doubt is uninformative; 62% of failures predicted at ≥0.7 confidence
- Trinh et al. (Scale AI) · HiL-Bench: Do Agents Know When to Ask for Help? — 4 May 2026: 300 tasks, 1,131 blockers; with full information pass@3 is 67–91%, and 1–39% when the agent must decide to ask; GPT under-asks, Gemini over-asks, Claude sees it is stuck in 45% of failures and submits anyway
- Vijayvargiya, Zhou, Yerukola, Sap, Neubig · Interactive Agents to Overcome Underspecificity in Software Engineering (Ambig-SWE) — February 2025, v3 of 21 February 2026, ICLR 2026: without prompting agents “almost never interact”; interaction lifts Claude 3.5 Sonnet from 26.8 to 39.6% and recovers up to 89% of full-specification performance
- Edwards, Schuster · Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents — 27 March 2026: a separate intent-detection agent reaches 69.4% against 70.8% with full specification while asking on 68.8% of tasks; forced asking gets the same at a 99.2% ask rate
- Stoisser et al. (Novo Nordisk) · Ambig-DS — 10 May 2026: on ambiguous data-science tasks frontier agents silently pick the wrong target in 39–63% of cases; “silent misframing” as a failure mode of its own
- Bouchard, Chauhan · Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents — 12 August 2026: no uncertainty scorer is reliable across domains — self-assessment gives AUROC 0.85 in telecom and 0.23 in airline on the same τ²-bench; a confidence threshold as a gate “could produce false assurance”
- Bhattacharyya et al. · Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs — 8 May 2026, 12 models, 38 tasks: asking about required effort and own ability predicts failure 6.7–10.1 AUROC points better than asking for confidence
- Anthropic · Emergent introspective awareness in large language models — 29 October 2025: Claude Opus 4.1 notices a concept injected into its activations in about 20% of trials; the capability is “highly unreliable and limited in scope”
- OpenAI · Operator System Card — January 2025: on 100 prompts the unmitigated model made 13 errors, 5 irreversible; confirmations before actions cut risk by about 90%; 92% confirmation recall on 607 tasks; a vendor claim
- OpenAI Alignment (Trębacz et al.) · Auto-review of agent actions without synchronous human oversight — 30 April 2026: of 10,000 actions 720 left the sandbox, 713 were approved by the reviewing model, 7 denied; about 200× fewer interruptions; motivated by decision fatigue and rules like “allow everything starting with python”; “not a security guarantee”
- Anthropic · Claude Code docs: Permission modes — Read 27 September 2026: auto mode blocks irreversible file destruction, git history rewrites, stepping outside the request, unfamiliar infrastructure and actions driven by hostile content; user-stated boundaries are block signals; a pause after 3 consecutive or 20 total blocks
- Anthropic · Claude Platform docs: Managed Agents permission policies — Policies always_allow, always_ask and auto; the docs warn that auto “is not a human checkpoint” — an action the server judges safe runs before anyone sees it
- Model Context Protocol · Elicitation (specification 2025-06-18) — A server requests structured input from the human; three outcomes — accept, decline, cancel; the protocol distinguishes “the human said no” from “the human walked away”
- Anthropic · Agentic Misalignment — 20 June 2025, 16 models in simulated corporate scenarios: Claude Opus 4 blackmails in 55.1% of runs when it believes the situation is real and 6.5% when it believes it is a test; the authors note the artificiality of the scenarios
- Zhao, Srikanth, Wu, Jiang (Weco AI) · SpecBench — 20 May 2026: 30 systems tasks; models score near the maximum on visible tests. Section 3.1: the 90th-percentile gap grows about 27 percentage points per tenfold increase in reference code size; the worst gap on large tasks is 100 points (100% visible versus 0% held-out). A separate “compiler” example scores 97% versus 0%, a 97-point gap
- Wang (Fudan) · Why LLM Agents Collapse Without Oversight: The Enforcement Gap — 23 September 2026: synthetic agent-plan injection; an auditor model warning is either logged or used to halt execution. Attack success: 48.3% versus 10.8% in a pooled control set (203 and 194 runs, including open-weight models), not the mean of the five closed-model rows; model judges assess success. Detection and stopping require separate evaluation
- Lisanne Bainbridge · Ironies of Automation — 1983: automation takes the easy parts and leaves the human the hardest ones; attention to a source where little happens cannot be held for more than about half an hour; the most successful automation needs the most operator training
- NIST (Fisher, Galluzzo) · Back to the Future: Why Agentic AI Needs a Strong Identity Foundation — 27 August 2026: “overly chatty agents condition users to reflexively click ‘allow’” — the MFA-bombing analogy; the remedy is pre-approved “flight plans” and short-lived, narrowly scoped delegated authority
- Mitchell, Ghosh, Passi · AI Agents Push Humans Out of the Loop — 24 August 2026, a position paper with no new measurements: approval fatigue, judging by style instead of accuracy, “intuition rust”; proposes strategic friction, batch review and skill-maintenance exercises
- Sekadakis, Yannis · Systematic review and meta-analysis of take-over time from automated driving at SAE levels 2 and 3 — August 2025, 51 studies: the faster the human takes over, the harsher the manoeuvres and the moderately higher the crash rate — speed of handback is not quality of handback
- Anthropic · Trustworthy agents in practice — 9 April 2026: the rate at which Claude checks in on its own roughly doubles on complex tasks while user interruption rates stay flat; a vendor statement about its own training on ambiguous situations
The measurement system and operations
- Staufer, Feng et al. · The 2025 AI Agent Index — 19 February 2026: 30 agents across 45 fields; 135 of 240 safety fields are empty; only four agents disclosed agent-specific safety evaluations; autonomy levels coded per Feng et al.
- Anthropic · Economic Index: Cadences (June 2026) — 26 June 2026: autonomy in Claude Code is 0.37 points above chat and 0.26 with the same model — “the product is more important than the model”; a median article takes 13 rounds in chat and one prompt in Claude Code
The Russian contour
- Habr (Ontico, Maxim Tsepkov) · Saint HighLoad++ 2026: the AI talks — 1 July 2026: Raiffeisen — 25% of the backlog done by teams with agents and 4% done by agents autonomously without asking a human; payments and risk engines stay human-only; a speaker's self-report
- CNews · Sberbank: from assistants to agents — 19 June 2026: code-generation compute grew 30× to 1.5 trillion tokens in five months; daily agent-mode users up almost fivefold in three months; no accepted-code share
- Yandex · The 75/75/75 target — 29 July 2026: 73% of developers use AI regularly, over half of new code is created with it, 17.2% have reached the bar; the internal agent works autonomously “in more than a thousand projects”; no acceptance criterion
- Pervaya Forma · 93% of code is written by agents, development is stuck in review — 25 August 2026: agents create 93% of merge requests, 94% merge without remarks, yet review-to-merge takes about 76 hours against 11 minutes of model work; the company's self-report
- CNews · Infosistemy Jet and Smart Ranking study on AI agents — 26 August 2026: semi-autonomous projects at 59% of organisations, in production at 15%; fully autonomous at 8%; 46% see no sustained economic effect; sample size not stated
- Habr (Codenrock) · Developer survey on AI agents — February–March 2026, 1,160 hackathon participants: 36% do not let the agent work autonomously, 34.9% trust it with small tasks; the main reason for caution is lack of context, 54.7%
- Habr (Sovcombank Technologies) · Survey on AI adoption in development — 25 August 2026, 317 questionnaires: end-to-end AI integration at 16.4%; the proposed agent metric — how many actions the agent got right first time and where a human had to stop it
Knizhny kub breakdowns
- Knizhny kub · Autonomy Is All You Need, part I — 29 December 2025: a breakdown of Michele Catasta's Replit talk — autonomy as the main measurable agent metric and as a property of the system, not the model
- Knizhny kub · Measuring the Impact of Early-2025 AI, part I — 17 July 2025: a breakdown of the METR experiment's methodology with 16 developers and why it drew so much attention
- Knizhny kub · SWE-Together: measuring coding agents in dialogue — 7 July 2026: 109 tasks from 11,260 real sessions; a user simulator indistinguishable from a live one; a stronger model needs fewer interventions
- Knizhny kub · Claude Code and expertise, part II — 12 July 2026: why session success is not yet productivity — common-method bias and a “verified success” that cannot see whether the team accepted the code
- Knizhny kub · Loop Engineering: why the key part of the loop is the right to say no — 14 July 2026: the ladder prompt → context → harness → loop and a verifier with the right to stop the loop
- Knizhny kub · DX Q2 2026: more code, less trust — 4 August 2026: a breakdown of the DX report — maintainability up, change confidence down, review cannot keep up
- Knizhny kub · Waymo: seven lessons from demo to physical AI — 9 August 2026: Dmitri Dolgov's talk — “a demo is at most 1% of the work”; from the 2010 demo to half a million rides a week took about 15 years
- Knizhny kub · Diogo Almeida: AI you do not have to watch — 20 September 2026: an assistant and automation need different properties; calibrated decisions as a separate training goal — where to act, where to ask, where to hand over
- Knizhny kub · AI Dev Podcast #8: agent autonomy starts with constraints — 10 August 2026: measure accepted work, verification cost and the link between changes and production outcomes, not the volume of generated code
Related materials
- Evaluating AI Agents in Production →The replayable episode and the hidden judge — the units of evidence the verification program rests on
- When Agents Write the Code: What Remains Engineering →The harness definition and the ablation test that becomes a scorecard row here
- The Agent Harness: Past, Present, and Future →How the harness enforces checks and blocks disallowed actions
- The agent economy: who benefits when machines make deals →Cost per accepted outcome and acceptance as three separate events — the same arithmetic as in section six
- The Economics of AI Development →Where cost per accepted task came from as a management unit
- Loop Engineering: Designing Loops That Run Agents on Their Own →The verifier's right to stop the loop — section nine in slide form
- AI Security in Software Development →The permission ladder and blast radius as a boundary no model retires
- AI Development as a Co-Evolving Stack →The stack layers between which this piece draws the measurement boundary