Anthropic has published data showing its Claude model now leads more than a quarter of the company’s AI research tasks, alongside a proposal for standardised measurements that other labs could adopt.
Anthropic has announced three measurements designed to track how much of its AI research and development is now carried out by AI rather than humans — and the numbers are striking.
The San Francisco-based company says its Claude model now “leads” around 26% of its measured AI research and development tasks, completing most of each task end-to-end from a high-level prompt under human supervision. That figure sits at what Anthropic calls automation level AL4 on a scale borrowed from research organisation Epoch AI, which runs from AL0 — no AI involvement — to AL5, meaning fully autonomous AI with no human in the loop.
Claude hasn’t reached AL5 on any measured task. But the direction of travel is hard to miss.
From Near Zero to 90% in Six Months
The speed of the shift is what catches the eye. In February 2026, less than 1% of Anthropic’s measured AI research and development work involved AI at or above the “AI collaborates” level — roughly AL3, where AI handles large parts of a task under human direction. By August 2026, that figure had risen to over 90%.
That’s a dramatic change inside a single company in less than a year.
Anthropic is framing its three-metric framework as a voluntary transparency tool, not a legal requirement. The company has published its methodology and is encouraging other frontier AI labs to adopt similar measures, with the stated aim of giving policymakers and the public a clearer picture of how AI development actually works.
What the Three Measurements Cover
The first metric — the AI R&D Automation Index — tracks what share of Anthropic’s AI research is handled by Claude at each automation level. The 26% figure at AL4 is the headline number, but the broader picture is that AI is now involved at some meaningful level in nearly all the company’s measured research work.
Meanwhile, the second metric covers agent oversight. Anthropic says around 30,000 AI agents are active at any one time on its main internal platform, carrying out research and engineering tasks. Every action those agents take passes through monitoring systems before it’s executed — a safeguard designed to catch problems before they cause harm.
The third metric tracks how Anthropic splits its computing resources between safety work and general AI research. In a sample week running from 13 to 20 July 2026, roughly 6% of the compute used for AI research and development went to safety work. When looking only at AI-driven research and development, that share rises to about 12%.
Not Everyone Is Convinced
The metrics have drawn a mixed response. Developers and researchers who use Claude and similar systems have broadly welcomed the transparency — it gives them a clearer sense of how much AI is already doing substantive technical work, and how that work is monitored.
But critics have raised questions. The data is self-reported and covers only Anthropic’s own operations. There’s no independent verification, and no standardised way to compare these numbers against what other frontier labs are doing. Google DeepMind, OpenAI, and Meta AI all operate at a similar scale but have published no equivalent figures.
Some experts and civil society groups have also questioned whether 6% of compute for safety is enough, given the scale of the risks that frontier AI companies themselves acknowledge. And there’s a tension some observers have pointed to: if AI is already leading more than a quarter of the work that builds future AI, is that compatible with calls — including from some of Anthropic’s own leadership — for a more careful, coordinated pace of development?
Dario Amodei, Anthropic’s chief executive, has previously spoken about the need for the AI industry to take safety seriously as systems become more capable. The company’s new metrics are partly an attempt to show, in concrete terms, what that looks like in practice.
The Broader Policy Picture
No UK government body or regulator has yet adopted Anthropic’s three-metric framework as a required standard. The Department for Science, Innovation and Technology and the AI Safety Institute are both engaged in ongoing work on AI transparency and risk, but Anthropic’s specific measurements remain voluntary for now.
That could change. UK regulators are watching how frontier labs manage AI-led development closely, and voluntary metrics published by major companies often feed into later regulatory thinking. Whether the 6% safety compute figure, or the AL4 automation level, eventually appear in any formal UK guidance nobody knows yet.
For now, Anthropic is offering the methodology freely so that other organisations — labs, regulators, research bodies — can replicate or adapt it. Whether the rest of the industry takes up that offer will say a great deal about how seriously the sector treats transparency.
What This Means for Kent Residents
Anyone in Kent using AI-powered tools — whether through a local business, a school, an NHS service, or at home — is indirectly affected by how companies like Anthropic manage the safety and oversight of the systems they build. Organisations such as NHS Kent and Medway ICB or Kent County Council that are considering AI procurement now have a concrete set of questions they can ask vendors: how much of your system was built by AI, how are your agents monitored, and what share of your computing budget goes to safety? Anthropic’s framework won’t answer those questions for every supplier, but it gives local decision-makers a useful benchmark to work from.
Source: @AnthropicAI
Anthropic Unveils Three New Metrics to Track How AI Systems Build Future AI Models Quiz
5 questions