Skip to content
AI Primer
breaking

OpenAI claims agents reached the automated research intern milestone

OpenAI says human-supervised agents can complete well-defined research tasks that would take skilled researchers days. Its internal measurements report 3.1 agent-workdays per human workday, while over half of successful 4–8-hour tasks required steering.

4 min read
OpenAI claims agents reached the automated research intern milestone
OpenAI claims agents reached the automated research intern milestone

TL;DR

  • OpenAI says its “automated research intern” can carry out well-defined, human-directed research tasks that would take a skilled researcher several days, as BorisMPower reported from the announcement.
  • The company’s 3.1-to-1 headline measures agent runtime against researcher workdays, not productivity-equivalent labor, a distinction rohanpaul_ai's summary preserved when quoting the figure.
  • Daily coding-agent use reached more than $600 at the median and more than $7,000 at the 90th percentile, according to reach_vb's roundup of OpenAI’s internal charts.
  • Longer autonomous runs remain fragile: more than half of successful tasks estimated at four to eight human hours needed at least one intervention, in WesRoth's summary of the results.
  • Researchers are increasingly supervising parallel runs, with the share running four or more concurrent workflows rising from about one-third in April to about three-quarters by mid-August, per rohanpaul_ai's workflow chart.

The official report publishes task-success data next to operational usage measures. Simon Willison flagged the late-July bend in agent spending and identified Astra as a guess, not a disclosed cause. OpenAI’s separate Astra safeguards post documents a July 20 safety-infrastructure pause and later security restrictions.

Automated research intern

OpenAI defines the milestone narrowly: a system working under human supervision on well-defined research tasks, including tasks that would take a skilled researcher a few days.

OpenAI called its measurements preliminary and said it was publishing both results and methods to encourage common measurement standards, in the statement TheRealAdamG linked.

An analysis from deredleritt3r extrapolated a December 2027 arrival from a claimed June completion, but OpenAI continues to state a March 2028 target for the more capable automated AI researcher.

Agent workdays

The 3.1 figure is a capacity measure. By mid-August, OpenAI counted 3.1 agent-workdays of runtime for every human researcher workday, while ai_for_success noted that agents were taking on longer, more complex research tasks.

OpenAI prices this usage at API rates rather than reporting an internal bill: the median researcher exceeded $600 per day and the 90th-percentile researcher exceeded $7,000. The distribution makes this less a single copilot rollout than a heavily skewed, compute-intensive workflow.

Four concurrent workflows

OpenAI’s concurrency measure tracks the share of researchers operating at least four agent workflows at once.

  • April 12: about one-third of researchers ran four or more workflows.
  • August 15: roughly three-quarters did.

Task horizons

OpenAI sorts agent assignments by estimated human completion time, then separates unattended successes, successes with intervention, and failures.

Across January through July, unattended success fell from 86% on tasks below 15 minutes to about 16% in the 64-to-128-hour bucket. The four-to-eight-hour band is the useful middle case: successful runs often existed, but a majority needed human steering.

A chart shared by scaling01 fitted July’s 50% success horizon at 4.7 hours, after 8.14 hours in June. Its own note says July had incomplete follow-up near the end of the observation window, and the seven monthly estimates do not establish a stable growth curve.

Code and experiments

OpenAI also published two activity measures:

  • Code changes per active contributor reached about 7 times the pre-2025 average.
  • Experiments per active experimenter reached 1.6 times the 2025 baseline.

These are throughput measures, not a count of successful discoveries or an external productivity evaluation.

Astra safety restrictions

The research-acceleration disclosure arrived with an account of model-development pacing. polynoamial highlighted that OpenAI had paired its internal usage figures with discussion of monitoring, alignment, security, and development restrictions.

The chart records three concrete details:

  • July 20: a safety-infrastructure pause and additional security requirements.
  • August 6 to 7: further Astra-class security restrictions.
  • July 20 through August 6: most Astra compute shown was allocated to testing the safety and security improvements.

Jakub Pachocki, OpenAI’s chief scientist, wrote in An Alien Mind that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer.

Research code

The internal usage curve was not uniform across the company. thsottiaux described research work as debugging dense Python systems whose behavior can hinge on both papers and subtle training dynamics, where the failure may sit in either the math or the system.

That post says engineering and product teams saw their large agent-use increase months earlier than research did.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR2 posts
Automated research intern3 posts
Agent workdays1 post
Task horizons1 post
Astra safety restrictions1 post
Research code1 post
Share on X