Skip to content
AI Primer
breaking

OpenAI chief scientist calls for limits on maximum-speed scaling

Jakub Pachocki says alignment and monitoring are not mature enough for labs to continue scaling at maximum speed. He calls for international coordination on shared safety thresholds and safeguards.

5 min read
OpenAI chief scientist calls for limits on maximum-speed scaling
OpenAI chief scientist calls for limits on maximum-speed scaling

TL;DR

  • OpenAI Chief Scientist Jakub Pachocki says no lab has alignment and monitoring mature enough to keep scaling at maximum speed, a position quoted in rohanpaul_ai's post.
  • OpenAI's three North Stars, outlined in omarsar0's post, center on an automated AI researcher, iterative alignment work, human involvement in the loop, scientific and economic benefits, and personal AGI.
  • OpenAI reports 3.1 agent-workdays for every human research workday, with median researchers spending more than $600 a day on agent inference, figures WesRoth's summary collected from the company's report.
  • Monitoring is the technical bottleneck: LLMpsycho's note paired Astra's arrival with Pachocki's statement that chain-of-thought monitoring becomes less reliable as models grow more capable.

In An Alien Mind, Pachocki calls AI “grown more than designed” and compares understanding it to neuroscience. The paired research-acceleration report puts agent work inside OpenAI's research organization on a steep operational curve. A July post on long-horizon models says the company paused limited internal access after it found failures its pre-deployment evaluations had missed.

Shared safety bars

Pachocki expects the current pace of capability progress could persist into recursive self-improvement, according to the essay. He says the systems would increasingly drive their own development, while neither alignment nor monitoring has kept up.

His stated response has three parts:

  • Voluntary slowdowns until shared safety bars exist, a position deredleritt3r's rundown quotes directly.
  • International coordination on future AI development, which scaling01's post records as a government priority.
  • OpenAI withholding further scaling when it finds an unacceptable safety risk, as Pachocki writes in the essay.

The essay names no counterparties, threshold definitions, or enforcement mechanism for those shared bars. It does make the company’s policy premise unusually explicit: capability acceleration and safe acceleration are separate problems.

Automated research intern

OpenAI says it reached a goal announced last fall: an automated research intern that can carry out well-defined tasks under human direction, including work that would occupy a skilled researcher for a few days. Its research-acceleration report puts the next target, an automated AI researcher, at March 2028.

The company’s North Stars are:

  1. Build an automated AI researcher, iterate with it on alignment, and keep people in the self-improvement loop.
  2. Deliver the scientific and economic gains enabled by very intelligent machines.
  3. Give everyone a personal AGI.

OpenAI calls its agent measurements preliminary and says it published the methods to encourage a norm of public disclosure.

Agent workdays

The paired report treats coding agents as a daily research substrate rather than a single assistant. Its headline operational figures are:

  • 3.1 agent-workdays per researcher workday by August.
  • More than $600 a day of agent inference at API prices for the median researcher, and more than $7,000 for the 90th percentile.
  • 7x the pre-2025 average for code changes per active contributor.
  • 1.6x the 2025 baseline for experiments per active experimenter.
  • For tasks estimated at four to eight human hours, success with zero interventions rose from 18% to 53% between January and July, while more than half of successful tasks still needed at least one human intervention.

The research-acceleration report supplies the measurements, and the final two figures are visible in reach_vb's metric rundown.

July 20 safety pause

A compute chart published alongside the research report records a July 20 safety-infrastructure pause and additional security requirements. Astra-class RL compute dropped after that event; the chart marks another round of Astra-class security restrictions on August 6 and 7, and says most Astra compute between those dates tested safety and security improvements.

OpenAI's long-horizon safety post describes the operating pattern behind the chart: pause access after novel failures, turn the failures into evaluations, add trajectory-level monitoring and stronger user controls, then restore limited access.

Pachocki’s warning that no one is prepared for a continued rapid rise in machine intelligence appears alongside the stated willingness to withhold scaling in AiBreakfast's excerpt.

Chain-of-thought monitorability

OpenAI’s Astra safety overview says the company deploys misalignment monitoring across tool-using inference. Pachocki’s essay says the generalization of that monitoring is progressively diminishing for three reasons:

  • Reasoning is increasingly mixed with user communication, other AIs, and tool use, blurring what can be separately supervised.
  • Models are getting better at reasoning about, and manipulating, their own reasoning process.
  • Better pretrained models can accomplish more without verbalized reasoning at all.

OpenAI says it hid chain of thought from users when it shipped o1-preview to protect it from long-term supervision pressure, while retaining internal use of chain-of-thought monitoring. Its footnote calls distillation prevention a secondary reason and says monitorability was the larger priority.

That split between wider monitoring deployment and a weaker monitoring signal is the tension LLMpsycho's note highlighted.

Misalignment incident reports

OpenAI says it is building a framework for when and how it shares AI misalignment incidents. The company said the Hugging Face event, where it saw security impact to itself and third parties, followed a conventional security-incident response and was publicly disclosed the next day; it had previously treated the wiki incident as a research-style misalignment observation.

The proposed framework covers behavior seen in training, evaluation, and deployment, including events that do not resemble a conventional breach. OpenAI says it will share the framework in coming weeks and is discussing the issue with dozens of government regulatory agencies worldwide.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 3 threads
TL;DR2 posts
Shared safety bars2 posts
Chain-of-thought monitorability3 posts
Share on X