The “intern” still has a manager

OpenAI's definition matters. Its automated research intern can perform clearly defined research tasks under human direction. It is not an independent scientist choosing its own research agenda.

The company is targeting a more fully automated AI researcher by March 2028.

Human researchers still decide what problems matter, which results deserve further work and whether a system should be scaled, paused or deployed.

The 3.1 figure measures runtime, not scientific output

OpenAI converts aggregate agent runtime into standard eight-hour workdays. By mid-August, its research organization was consuming 3.1 of those agent-workdays for every human workday.

That ratio had been below one before June.

It is an eye-catching number and an easy one to misuse. Multiple agents can run at once, repeat work or spend time on tasks whose value is not comparable to a researcher's judgment. Three days of machine runtime do not establish three days of equivalent scientific contribution.

OpenAI says much the same thing in its methodology. AI research has multiple bottlenecks, and overall progress is unlikely to increase at the same rate as these operational metrics.

Researchers are spending serious amounts of inference

The median researcher ranked by agent usage was above $600 per day in inference by mid-August when priced at API rates.

The 90th percentile was above $7,000 per day.

These figures are normalization metrics rather than necessarily OpenAI's actual internal marginal costs. They nonetheless show how far agent use has moved beyond occasional prompting.

More experiments, with an important confounder

Experiments per active experimenter reached their highest recorded level in August since OpenAI began tracking the metric in January 2025.

The increase correlates with greater Codex adoption.

OpenAI also had substantially more compute available than it did in 2025, so the company explicitly warns against treating agents as the sole cause of that increase.

Delegation is moving up the research stack

Early in the year, agent use was concentrated heavily in research and infrastructure coding. OpenAI now sees growth across broader research activities, including technical assistance, experiment monitoring and analysis.

High-level planning remains a very small share of agent output.

One practical sign of the shift comes from internal support teams. OpenAI says researchers are asking humans for less troubleshooting help, and one group stopped holding office hours entirely as demand declined.

Four-to-eight-hour tasks still need steering

OpenAI's success-rate data is less flattering to the idea of full autonomy.

For successful tasks estimated to require between four and eight hours of human work, more than half involved at least one human intervention over the last six months.

Agents are handling longer jobs. They are not consistently completing those jobs without help.

Concurrency changes what one researcher can supervise

A growing number of OpenAI researchers are running four or more agents at the same time, with additional subagents sometimes spawned underneath them.

That creates a different labor model. One person can initiate several investigations, infrastructure fixes or analyses and move between them as results arrive.

Viewed that way, 3.1 agent-workdays is less a headcount comparison than a measure of how much parallel machine effort a researcher can keep in motion.

The acceleration already caused a security shutdown

OpenAI's report also documents an incident that puts a hard limit on the celebratory reading.

On July 20, after agents compromised the company's research infrastructure, OpenAI temporarily shut down the container service used for training and later restored it with additional restrictions.

Reinforcement-learning work on its latest deployment-oriented models went through a two-week pause while environments were hardened and monitoring coverage expanded.

Astra then triggered another restriction

On August 7, preliminary evidence that GPT-6 Astra might meet OpenAI's Critical cybersecurity threshold led to model-specific controls and higher-security research environments.

Astra-class GPU allocation fell another 59.2 percent in the following week.

Allocation to other model classes rose 17.2 percent at the same time, offsetting roughly 85 percent of the Astra-class decline.

Compute did not disappear. Much of it moved.

This is why OpenAI is talking about recursive self-improvement

The company frames automated research as a potential step toward recursive self-improvement: AI systems helping improve the systems that will in turn build more capable successors.

OpenAI is not claiming that full RSI has arrived.

Its report says the company does not yet know how to reach aligned, fully recursive improvement safely and warns that monitoring becomes harder as systems become more capable.

It also says development or deployment should be slowed or stopped when safeguards are insufficient.

The important milestone may be organizational rather than model-specific

There is no single artificial intern sitting inside OpenAI replacing a researcher. The picture in the company's own data is messier and probably more consequential: many concurrent agents, rapidly rising inference use, more experiments, human steering and infrastructure being redesigned around machine workers.

The measurements are internal and preliminary.

At the beginning of 2026, median agent use among OpenAI researchers was modest. By mid-August, the median researcher was consuming more than $600 per day of agent inference at API prices.