THE NUMBERS
● February 2026: under 1 percent of Anthropic's AI R&D led by Claude.
● August 2026: 26 percent.
● More than 90 percent of that work now involves the model as a collaborator or better.
● Also disclosed: 30,000 agents running at once, monitors blocking roughly one action in 47,000.
What the levels mean
Anthropic's index uses graded levels rather than a single figure, which is the part that makes it readable.
| Level |
What it means |
| AL3 — collaborates | Model works alongside a researcher. Over 90 percent of the work is here or above |
| AL4 — leads | Model completes most of a task end to end from a high-level prompt, human supervises. 26 percent |
| AL5 | Not reached in any measured subset |
That last row matters. A company publishing an automation index has every incentive to make the top number impressive, and Anthropic stated plainly that Claude is not at AL5 anywhere it measured. A ceiling you admit to is more credible than one you do not mention.
SIX DAYS AFTER ASKING THE INDUSTRY TO SLOW DOWN
On 12 September Dario Amodei published an essay urging labs to pace frontier capability gains so safety work could catch up.
On 17 September his own lab published the measurement showing the model now leads a quarter of the work that builds its successor. Those are not contradictory — they are the same argument. The essay asked for a pause because of what the index measures.
The number nobody is quoting
Thirty thousand agents running at once, with monitors blocking roughly one action in 47,000.
That block rate is the most concrete safety figure any lab has published this year. It gives you a denominator — the thing missing from OpenAI's six disclosed misalignment cases, where six out of how many was unanswerable.
It also means blocked actions happen. At 30,000 concurrent agents, one in 47,000 is not a rounding error, and publishing the ratio is a disclosure against interest.
What this does and does not tell you
- It is self-reported and self-defined. Anthropic wrote the levels and graded its own work against them. That is a limitation, not a disqualification.
- It is about one lab's internal research, not about your codebase. 26 percent of AI research at Anthropic says nothing about what an agent can do in your repository.
- It is a trajectory, not a state. Under 1 percent to 26 percent in six months is the finding. The absolute number matters less than the slope.
- And it makes the pacing argument concrete. Every essay about recursive self-improvement has been hypothetical. This is a dashboard.
Why publish it at all
Anthropic has now published three things this month that cost it something: four real-system breaches during evaluations, wide transcript access granted to METR, and an enforcement action against state-linked accounts. This is the fourth.
The pattern is consistent enough to be a strategy. Disclosure against interest is the only kind that carries evidential weight, and a company arguing for industry-wide pacing needs evidence more than it needs a flattering number.
Sources
FAQ
What is the R&D Automation Index?
Anthropic's measurement of how much of its own AI research and development the model leads, graded by level. It reports 26 percent at AL4 as of August 2026, up from under 1 percent in February.
What does AL4 mean?
The model completes most of a task end to end from a high-level prompt while a human supervises. AL3, collaborating alongside a researcher, covers more than 90 percent of the work.
Has Claude reached AL5?
No. Anthropic states it is not operating at AL5 in any measured subset of the work.
What is the one in 47,000 figure?
The rate at which Anthropic's monitors block an agent action. With roughly 30,000 agents running at once internally, that ratio is the most concrete safety denominator any lab has published this year.
Does this affect the models I use?
No. It measures internal research practice, not product behaviour. Nothing about pricing or availability changes.
Does it contradict the slowdown essay?
No. Amodei argued for pacing because of what this index measures. The essay and the dashboard are the same argument from different ends.