Anthropic wants to measure what happens inside frontier labs

Public model benchmarks only reveal part of the AI development process. They show what a system can do after it has been trained, but provide much less visibility into how frontier laboratories are increasingly using their own models to accelerate the next generation.

Anthropic is therefore proposing three categories of development metrics: how much AI R&D is performed by AI itself, how effectively internal agents can be monitored and interrupted, and how computing resources are allocated during model development.

The stated goal is to make a process currently visible primarily to a small number of private laboratories observable from the outside.

Claude now "leads" 26% of Anthropic's model R&D

The most striking number concerns automation of AI research itself. Anthropic uses an automation scale developed by Epoch AI ranging from AL0 to AL5.

AL0 represents work with no AI involvement. At AL3, AI collaborates with a human and can perform large parts of a task under close direction. At AL4, AI "leads": it can complete most of a task end-to-end from a relatively high-level instruction while a human remains in a supervisory role.

Anthropic estimates that Claude reached AL4 for 26% of its measured model R&D work in August 2026.

The increase has been rapid. Anthropic's published chart places the figure at roughly 1% in March, 3% in April, 12% in May, 14% in June, 22% in July and 26% in August.

More than 90% of research already involves Claude

The share becomes significantly larger when Anthropic includes work in which Claude collaborates without leading the task.

More than 90% of measured R&D is now at "AI collaborates" or above. That does not mean 90% of Anthropic's researchers have been replaced, nor that Claude independently performs 90% of development.

Much of this work remains directed, supervised or validated by humans. Anthropic also states that no measured category has reached AL5, the level representing fully autonomous AI operation without a human in the loop.

The important development is therefore not that full autonomy has already arrived, but how rapidly AI is becoming embedded throughout the process used to build future models.

How do you measure something this difficult?

To build its automation index, Anthropic first attempted to map the work actually being performed throughout its model development organization.

For each week in July, the company sampled 20% of employees from departments involved in the model R&D loop. Claude research agents analyzed work records and internal documentation to identify around 15,000 individual tasks.

Those tasks were then organized into a hierarchy containing 542 nodes and 378 terminal task categories, ranging from evaluation-platform troubleshooting to reinforcement-learning infrastructure and network policy.

Anthropic then weights these categories using human time spent on each area as a proxy for their importance and estimates the degree of automation for each one.

There is an obvious complication: Claude also helps evaluate Claude

Anthropic acknowledges an important methodological limitation. Claude is itself used during parts of the process that determine how automated Anthropic's work has become.

That creates the unusual situation of an AI system helping to evaluate the role that the same system plays inside the company developing it.

Anthropic compared model ratings with assessments from employees responsible for the relevant work areas. The company says exact model-versus-human agreement reached 59%, while model and human ratings were within one automation level 97% of the time.

Those results do not turn the metric into an objective ground truth. The distinction between AI that "collaborates" and AI that "leads" a task can still involve judgment.

Around 30,000 agents can operate inside Anthropic at once

Anthropic's second set of measurements provides an indication of how large its internal agent infrastructure has already become.

The company says approximately 30,000 agents were performing research and engineering work at any one time on its most-used internal platform during August.

These agents can work for extended periods and delegate tasks to other agents. That creates a different challenge from model capability alone: how do you oversee tens of thousands of semi-autonomous systems as the number and consequence of their actions increase?

Anthropic uses two monitoring layers. Online monitors examine actions before execution and can block or redirect them. Offline monitoring analyzes activity afterward to identify behaviors that may develop more slowly or require broader context.

One decision in roughly 47,000 was blocked

Anthropic says 100% of actions on the measured platform pass through the online monitoring system before execution and are also ingested for offline monitoring.

Across more than one billion decisions made by research and engineering agents during August, approximately 0.002% were blocked by the online monitor, equivalent to roughly one decision in every 47,000.

That figure requires careful interpretation. It does not mean one in every 47,000 decisions would necessarily have caused a severe incident. Monitoring systems can intervene on suspicious, erroneous or policy-incompatible behavior.

Likewise, a low blocking rate does not prove that all remaining decisions were harmless. The result depends directly on how effectively the monitors identify problematic behavior.

Around 100,000 transcripts are flagged each week

Offline monitoring operates on another scale. Anthropic says approximately 100,000 activity transcripts are flagged each week before being processed by multiple classifiers.

Most are resolved during that process, while roughly 50 of the highest-priority cases per week are escalated to human review.

Anthropic argues that developers could publish comparable figures for monitoring coverage, review latency and escalation rates, making it possible to assess whether oversight is keeping pace with agent deployment.

Anthropic is also publishing how much compute goes to safety

The third metric focuses on a more tangible resource: computing power.

Anthropic examined all of its compute use during the week from July 13 to July 20, 2026. Approximately 6% of compute allocated to AI R&D during that period was classified as safety work.

For AI-driven AI R&D specifically, that figure was approximately 12%.

Anthropic describes these estimates as deliberately conservative. Workloads that contributed equally to capability improvement and safety were counted as capability work rather than safety.

Compute is not a perfect measure of safety effort

Raw compute percentages also have limitations. Some safety research depends primarily on human analysis, experiment design or evaluation and may require far fewer GPUs than frontier model training.

A laboratory could therefore employ significant safety teams without those teams representing a large share of total compute.

Anthropic acknowledges this limitation. The main value of the metric is its ability to track change within one developer over time and, if common definitions emerge, make comparisons across laboratories.

The real issue behind the numbers is AI's development feedback loop

Anthropic is ultimately trying to measure something broader than Claude itself. When an AI helps write tools, run experiments, analyze results and modify the infrastructure used to train its successor, each generation can theoretically accelerate development of the next.

The extreme version of that process is often described as recursive self-improvement: a system capable of building or improving its successor with little meaningful human intervention.

Anthropic's own figures show that this stage has not been reached in the work it measured. None of the categories is classified as fully autonomous AL5.

But the figures also show why the concept is becoming measurable rather than purely speculative. Moving from around 1% of work led by Claude in March to 26% five months later creates a trend that can be tracked.

Anthropic wants independent evaluators to verify the process

The clearest weakness in these metrics is that they are initially produced by the same laboratory they are intended to make more transparent.

Anthropic therefore says it plans to embed independent third-party evaluators from multiple organizations and provide access to internal processes, systems and data comparable to that available to its own risk assessment teams.

Those organizations could verify safety practices, report incidents and monitor metrics like those Anthropic has now published.

If that model works, the significant change may not simply be that Anthropic releases more numbers. The pace of frontier AI development could become something external observers can inspect rather than something known primarily to the companies building the systems.

The most important benchmark may soon measure the development process itself

Traditional benchmarks will remain essential for measuring model capabilities, but they mostly examine the finished output of the development process.

Anthropic's automation index instead examines the engine driving progress: how much of the work needed to improve AI can increasingly be performed by AI.

If Anthropic continues publishing these figures, the evolution of that 26% figure may become especially revealing. Slow growth would suggest that important parts of frontier development remain difficult to automate. Rapid movement toward higher AL4 and eventually AL5 shares would indicate that more of the development loop can operate with reduced human involvement.

At that point, the question would no longer be only how much better the next Claude is than today's model. It would also be how much today's Claude accelerated the arrival of its successor.