Anthropic published a new set of proposed frontier-lab measurements on September 17.

The premise is straightforward. The industry has no shortage of evaluations describing model capability, but considerably less visibility into the process producing the next generation of those models.

How much of future AI development is already being performed by current AI? How comprehensively are internal research agents monitored? How much compute intended for AI R&D is explicitly being spent on safety work?

Anthropic proposes three groups of measurements as a starting point.

The headline figure is inevitably 26%

Anthropic calls its first prototype the R&D Automation Index.

According to the index, Claude was “leading” 26% of Anthropic's model R&D in August 2026.

“Leading” has a specific meaning here.

Anthropic uses Epoch AI's Automation Level scale, running from AL0, with no AI involvement, to AL5, where AI operates fully autonomously without a human in the loop.

At AL3, described as “AI collaborates,” a model can perform substantial pieces of work but remains under close human direction.

At AL4, or “AI leads,” the human can provide a relatively high-level objective and allow the model to perform most of the task end-to-end before supervision and final decision-making.

The 26% figure refers to AL4.

That does not mean 26% of Anthropic is autonomous

The distinction is critical.

Anthropic is not saying Claude independently builds 26% of its next model, or that 26% of researchers have become unnecessary.

The measurement covers a weighted basket of model-R&D tasks.

AL4 still contains a human supervisor. The person may no longer need to follow the process continuously, but remains responsible for review and decisions such as whether a change actually ships.

Anthropic explicitly says no measured subset of its AI R&D currently reaches AL5, the fully autonomous level.

AI involvement is already much broader than the 26% headline suggests

A second number gives a different view of the transition.

More than 90% of the work in Anthropic's measured basket is already at AL3 or above.

In other words, work without substantial Claude involvement is a relatively small portion of the R&D activity represented by the index.

The AL4 share has also moved rapidly.

Anthropic's chart places it below 1% in February 2026, then at 1% in March, 3% in April, 12% in May, 14% in June, 22% in July and 26% in August.

Those numbers come from Anthropic's own measurement system rather than an independent productivity benchmark.

Before measuring automation, Anthropic first had to define what its researchers actually do

That may be the most revealing part of the methodology.

A frontier lab does not necessarily have one clean master list called “all tasks required to build the next model.”

Anthropic therefore reconstructed the work from internal records.

For each week in July, it randomly sampled 20% of staff from departments participating in the model-R&D loop.

A Claude research agent examined Slack and internal documentation associated with those sampled employees and identified the work they had performed.

The process produced roughly 15,000 granular tasks.

Those tasks became a 542-node map of model development

Claude then organized the inventory into a hierarchical tree, moving from broad domains toward increasingly specific categories of work.

The resulting structure contains 542 nodes and 378 leaf categories.

Anthropic gives examples including evaluation-platform defect diagnosis, reinforcement-learning sandbox network policy and serving-incident postmortems.

That tree is frozen for subsequent measurements.

Monthly scores can therefore track whether the same broad basket of work is becoming more automated over time.

Human time provides the weighting

Simply counting tasks would produce a distorted index.

A minor scripting fix and a major training-pipeline project would otherwise each contribute one unit.

Anthropic instead uses person-time as an approximation of how important a category is to the overall R&D effort.

Each sampled employee contributes one unit of weight per week, divided across the work they performed.

Someone working on four tasks gives each 0.25 units. Someone working on ten gives each 0.10.

Category weights are then built from those person-time allocations.

Claude then judges how much Claude has automated

This is the obvious methodological vulnerability, and Anthropic explicitly acknowledges it.

For every category, a Claude agent researches how that work is performed, which tools are used and how much AI currently does.

A separate Claude judge then reviews that evidence and assigns an automation level.

Models from the same broader family therefore contribute both to the work and to measuring how much of the work has been automated.

A systematic blind spot in the judge could flow directly into the index.

Anthropic compared the AI ratings with human ones

Employees responsible for the relevant work areas independently rated automation without seeing the model's evidence or conclusions.

Anthropic reports exact model-versus-human agreement of 59%.

Human-versus-human exact agreement was only 35%.

More importantly, model and human scores landed within one automation level of each other 97% of the time.

That does not make the index objective.

It does demonstrate that boundaries such as the transition from “AI collaborates” to “AI leads” are fuzzy enough that human evaluators disagree substantially too.

A frozen basket creates another problem

Suppose old tasks become heavily automated while researchers use the freed time to invent entirely new kinds of work.

A fixed index could show dramatic automation without capturing the new human activity appearing around it.

Anthropic explored this by constructing another task tree from January 2026 data and comparing incoming work from February through July against that older basket.

It says it did not observe an increase in “novel” tasks at the level of analysis used.

The company still plans to rebuild and version the task basket periodically.

The second set of measurements asks who is watching the agents

Automation is only half of the picture.

If thousands of agents are performing longer stretches of work without continuous human attention, it also becomes important to measure whether those actions remain observable.

Anthropic therefore proposes publishing monitoring coverage, review latency and blocking or escalation rates.

As of August, approximately 30,000 agents were performing research and engineering work at any one time on Anthropic's most-used internal agent platform.

Anthropic stresses that this number covers that platform rather than every possible internal agent system.

Those agents encounter monitoring both before and after execution

The online monitor operates before an action occurs.

Anthropic says 100% of actions on the measured platform pass through it before execution.

The system can block dangerous activity or steer the agent away from a problematic action.

This layer is intended for events whose consequences could happen quickly and be difficult to undo, such as an attempt to move model weights outside approved systems.

Blocked actions are then scheduled for human review within one week.

About one in 47,000 decisions was blocked

Anthropic analyzed more than one billion decisions generated by research and engineering agents during August.

The online monitor blocked 0.002% of them, or roughly one decision in 47,000.

That rate is not equivalent to the underlying rate of dangerous behavior.

Some blocks can be false positives, while problematic actions could also evade detection.

Anthropic therefore argues that blocking rate should be published alongside monitoring coverage and independent evaluations rather than treated as a standalone safety score.

Offline monitoring reads everything again afterward

Anthropic says 100% of the same agents' activity is also ingested into an offline monitoring system after the actions occur.

That layer targets patterns that may develop more slowly and can be investigated retrospectively.

It generates roughly 100,000 flagged transcripts each week.

Multiple classifiers then remove false positives and perform initial oversight.

The highest-priority cases, approximately 50 per week, reach human reviewers.

Across the full stream, Anthropic says about one or two transcripts per thousand are flagged for additional review.

Every agent also has a persistent identity

Anthropic describes two design choices intended to make the system more auditable.

Each agent carries an individual identity and data records remain attached to it across model upgrades.

Actions can therefore be connected to a continuing agent history even when the underlying model changes.

Agents also communicate through an open shared messaging system instead of private exchanges that would be harder to traverse.

Messages retain their author's identity and connect back to transcripts, allowing behavioral monitoring to follow interactions between agents rather than looking only at each agent in isolation.

The third proposal measures something more physical: compute

Anthropic's final category concerns how accelerator capacity is allocated.

Terms such as “autonomy” require judgment. GPU and accelerator usage is at least closer to an infrastructure quantity that can be recorded directly.

Anthropic examined a snapshot of all its compute from July 13 through July 20.

It grouped workloads into broad categories and then measured how much AI R&D compute could be classified as safety work.

Six percent went to safety, under Anthropic's definition

During the measured week, approximately 6% of compute allocated to AI R&D was classified as safety work.

Looking specifically at compute used for AI-driven AI R&D, approximately 12% was classified as safety.

The two percentages have different denominators and should not be added together.

Anthropic also describes its classification as deliberately conservative.

If a workload advanced capabilities about as much as safety, it was counted as AI R&D rather than safety.

Safeguard classifiers are outside that 6%

Anthropic says compute used by its safeguard classifiers is excluded from those percentages.

The company describes that separate workload as being of a comparable magnitude.

That does not mean the safety percentage can simply be doubled.

The categories represent different functions and Anthropic intentionally reports them separately.

The company also stresses that compute is an imperfect proxy for overall safety effort because some safety research consumes large amounts of researcher time without requiring a large training run.

Even defining “safety compute” requires judgment

Some projects affect both capability and safety.

Scalable oversight research, for example, might make a future system easier to supervise while also improving the usefulness of a current model.

Anthropic wrote a detailed classification definition with boundary cases and used a Claude-based classifier to sort workloads.

For training and evaluation runs, the classifier reads metadata and associated code.

For research-agent inference, a variant uses session transcripts when those are available.

Anthropic says some cases remained genuinely difficult even after hours of human review.

The compute audit also uses sampling

The examined week contained almost 10,000 research training and evaluation runs.

Anthropic classified a sample of roughly 14%.

The sampling was weighted toward high-compute jobs so the final estimate better reflects where accelerator capacity went rather than how many individual runs existed.

The company says it wants to refine the system so an independent evaluator could re-run the classifier on random jobs and transcripts and verify both the classification and totals.

Anthropic wants these measurements to become externally useful

The company frames the proposal as a way to reduce the information gap between frontier labs, independent evaluators, governments and the public.

Anthropic also suggests that metrics of this kind could eventually support commitments or formal requirements, including testing windows before new models are used for more AI R&D or disclosures about safety-compute allocation.

Those are Anthropic's policy proposals, not current industry standards or generally applicable regulatory requirements.

The company says it plans to embed independent third-party evaluators from multiple organizations and give them access to internal processes, systems and data comparable to internal risk-assessment teams.

These numbers cannot yet rank Anthropic against other frontier labs

There is no shared measurement standard.

Another lab could define “AI leads” differently, build a different task basket, or draw the boundary between safety and capabilities elsewhere.

Two companies could both report 25% automation while measuring substantially different realities.

Anthropic explicitly identifies the lack of a common methodology as one of the main barriers to cross-lab comparison.

The proposal becomes much more informative only if definitions converge and external auditors can reproduce the underlying measurements.

The 26% figure may not be the most consequential part of the release

It is a striking snapshot of acceleration: Anthropic's own index moves AL4 work from effectively negligible levels early this year to more than one quarter of model R&D by August.

The more durable development may be the attempt to make that acceleration observable before the next model launch.

When the outside world sees a frontier lab only through successive model releases, it sees outputs but very little of the machinery creating them.

An automation index, agent-monitoring coverage and compute allocation would expose at least part of that production system.

For now, Anthropic has largely designed the measurements, supplied the data and run the models that classify them.

That makes the next metric unusually obvious: if these numbers are to become more than a detailed internal audit, their reproducibility and independent verification will need to be measured just as carefully.