Skip to content
AI Primer
update

OpenAI says it evaluates safety cases before major RL runs

Sam Altman said OpenAI evaluates explicit safety cases before RL training runs expected to materially raise capabilities. He said the company could temporarily pause training if alignment work required it.

5 min read
OpenAI says it evaluates safety cases before major RL runs
OpenAI says it evaluates safety cases before major RL runs

TL;DR

  • OpenAI now says a major frontier RL run must have an explicit safety case before it begins when the run is expected to materially increase capabilities, according to sama.
  • The company has already delayed larger RL work while it raised security bars, and OpenAI’s RL update screenshot records a restart only after new safety and security requirements were in place.
  • Independent review is part of the announced control stack: sama said OpenAI will offer evaluators employee-like access, with details still to come.
  • The policy targets both loss of human control and concentration of power, the two hazards laid out by sama.

OpenAI’s August safety note says it paused deployment-bound RL training for two weeks to harden research environments. METR’s six-day onsite investigation was unpaid and retrospective. Anthropic CEO Dario Amodei’s three-stage plan asks for a permanent version of that scrutiny, reaching into training pipelines rather than only shipped models.

Safety cases

OpenAI CEO Sam Altman says the company now formulates explicit safety cases before frontier reinforcement-learning runs expected to significantly raise capability. That makes the training run, rather than the release candidate, the newly announced decision point.

The statement does not define the capability delta that triggers a case, the evidence a case must contain, its approver, or a publication process. It does place monitoring and alignment work ahead of the run instead of treating them solely as pre-release checks.

Altman also said OpenAI would stop training temporarily if it needed more alignment progress, in haider1’s clip.

The capability trigger

The immediate backdrop is cyber capability. OpenAI’s Path to Astra designated Astra its first model at the Critical cybersecurity threshold, saying it could find unknown flaws and develop exploits against well-protected systems with the right tools and access.

Artificial Analysis says its measurements continue to show capability jumps across every measured dimension. In a Fortune interview relayed by rohanpaul_ai’s Fortune clip, Altman called a 10% chance of killing everybody by the end of the decade unacceptable.

The RL gate

The safety-case claim has a recent operational precedent. A Sept. 11 Bloomberg report relayed by kimmonismus said Altman had told staff OpenAI could pace development with other labs, while OpenAI’s August note documented a two-week RL pause during infrastructure hardening.

The restart notice says OpenAI had held back some larger RL runs for future Astra versions, restarted a large frontier run on August 28, and continued to hold some smaller experimental runs. A formal safety case turns that sort of exceptional hold into an expected pre-run gate.

Evaluator access

OpenAI has paired the internal gate with a promise of outside visibility. Altman said independent evaluators would get employee-like access and that OpenAI would follow Anthropic’s approach.

Amodei’s proposed version assigns embedded evaluators three jobs: verify safety practices and commitments, report incidents, and assess alignment in training pipelines and processes. The earlier METR investigation offers a narrower precedent: staff worked onsite at OpenAI for six days to independently assess the Hugging Face incident, while leaving OpenAI’s investigation and remediation out of scope.

An EricSteinb correction identifies METR by name as the intended evaluator in that policy discussion. OpenAI’s pledge does not yet name an evaluator, specify permanent access, or establish what findings will be public.

Pacing

Pacing leaves model development and releases running, but adds time and compute costs for alignment, monitoring, security, and evaluation. Anthropic CEO Dario Amodei described that distinction in a CNN interview carried by rohanpaul_ai’s CNN clip, while the accompanying post includes the same exchange about industry-wide speed.

Public support has arrived from unusual competitors. Demis Hassabis endorsed the direction in testingcatalog’s screenshot, while Amodei paired his argument for delay with a claim, quoted by kimmonismus, that AI could cure most major diseases within five to ten years.

Coordination terms

The pre-run case and evaluator pledge are unilateral company processes. The remaining mechanisms Amodei proposes require coordination: common safety standards among frontier labs in democratic countries, then international coordination with authoritarian governments.

The proposal identifies three potential pacing levers:

  • Training compute
  • The nature of training runs
  • Internal use of AI to improve AI

Those are the ingredients listed in rohanpaul_ai’s excerpt of Amodei’s essay. OpenAI welcomes a federal framework but says companies need not wait for legislation to begin, according to sama. Cohere’s policy essay objects to a small group of firms writing the rules, calling instead for capability- and context-based standards with public testing capacity.

Federal pushback

The coordination plan now runs into a direct political constraint. A Financial Times report shown in kimmonismus’s FT screenshot says President Trump rejected calls from technology leaders for an AI slowdown.

That leaves OpenAI’s safety cases as a company-level process and evaluator access as a voluntary commitment, while the industry-wide and global layers still depend on a policy agreement that has not arrived.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR3 posts
Safety cases1 post
The capability trigger1 post
The RL gate2 posts
Evaluator access1 post
Pacing4 posts
Coordination terms1 post
Federal pushback1 post
Share on X