Anthropic: Who Will Build the Next AI (#AI)
If AI helps build the next, more capable model, it could also accelerate its own development. For us, that could mean less time between new tools arriving and another round of adapting our work to them. Eventually, people may have to manage research agents that do more and more work independently. And understand when to intervene.
In Anthropic's article, there is a chart that got me thinking; I wrote down a question about the trend's convergence by 2027. According to the company, Claude led 1% of research and engineering work in March 2026 and 26% in August. AI was at least collaborating with humans in more than 90% of the work. These percentages refer to a weighted basket of task categories within Anthropic. Here, “led” means a particular level of independence.
On June 17, 2026, Jean-Stanislas Denain, Joe Kwon, and Anson Ho proposed a scale from 0 to 5 in an Epoch AI publication. At the beginning of the scale, AI is either unused or helps with individual parts of the work. Further up, the distribution of initiative changes: - AL3: the agent does large chunks; a human continually directs it and assembles the result. - AL4: the agent receives a high-level task and does most of the work; a human supervises, course-corrects, and approves. - AL5: the task is completed end to end with minimal or no human involvement.
For example, at AL4, you can ask an agent to fix a broken data pipeline: it finds the cause, makes changes, tests them, and returns a report. The decision to deploy stays with the engineer. Anthropic's August snapshot contains no measured R&D categories at AL5.
Now the connection to 2027 becomes clear. Between March and August, the AL4 share grew by roughly five percentage points a month. If that pace continues, about 76% of the work would reach at least AL4 by summer 2027, and the linear calculation would hit 100% by the end of the year. This is my conditional extrapolation. It assumes the pace will hold and the remaining tasks will be comparable to those already delegated to agents. But they could include choosing a research direction, assessing an unexpected result, and deciding where to spend an enormous compute budget. An agent that implements an experiment may still depend on the person who worked out what to test.
Then the article introduces roughly 30,000 agents working simultaneously on one of Anthropic's internal platforms. At that scale, a human saying “I'll review everything” needs some fairly serious clarification :)
The company measures monitoring coverage, review latency, and the share of actions blocked or escalated. It reports 100% coverage of actions on that platform. Coverage shows that a check is built into the process; judging its quality requires knowing how many dangerous actions it misses.
I liked the detail about persistent agent identities: the history survives model changes, and messages are linked to their authors and original evidence. You can reconstruct who proposed a decision and where it came from. At the same time, a shared communication channel can spread a shared error: the ability to check a source still needs to be used.
Compute presents a similar problem. Spending is convenient to track, but the share of resources “for safety” depends on how work is classified and how efficiently it is done. That share can fall because the same amount of safety work requires less compute. Claude gathers evidence about automation, and Claude also evaluates it. A shared methodology and external verification matter for trusting those numbers.
It looks as though we may see more processes in which humans choose the direction, agents do the work, and other agents help check it. If that cycle starts reliably improving subsequent models, the acceleration could become much more noticeable. The ability to detect an error in time, reconstruct the chain of decisions, and stop the process would then become as important as the ability to start it.
P.S. I often turn the web versions of articles like this into PDFs so I can take my time reading them on a tablet and write notes directly on them as I go — it helps me think better
#AI #Agents #Research #Engineering #Metrics #Security
Files from the post
- Annotated_Measurements_for_understanding_the_pace_of_AI_development.pdfDownload PDF