OpenAI supports employee-level access for independent model evaluators
OpenAI and Anthropic backed stronger access for independent model evaluators, with OpenAI saying it will match employee-like access. Eric Steinberger also said his organization would offer METR access before legal requirements.

TL;DR
- Anthropic committed to permanent, employee-level access for third-party evaluators to verify safety practices, report incidents, and assess alignment during training, as recorded in HendoAi's screenshot of the announcement.
- OpenAI CEO Sam Altman said in sama's post that OpenAI will adopt independent evaluators with employee-like access. A Bloomberg report excerpted by kimmonismus had already described internal discussions about pacing development with other labs.
- The proposal has three layers, embedded evaluators, coordination among democratic-country labs, and global coordination, in the essay shared by imjustnewatai. Anthropic has committed to the first layer.
- METR has already published an incident-investigation access framework that calls for model runs, reproducible environments, transcripts, staff interviews, and defined redaction rules. RyanGreenblatt said he is participating in an independent investigation into alignment and misalignment incidents at Anthropic.
- The initiative is forming around an ecosystem rather than a single evaluator: ValsAI cited its public RSI benchmark and cross-domain evaluations, while ArtificialAnlys said it has pre-launch benchmarked models with nearly every major lab.
The phrase Dario Amodei's essay makes consequential is “training pipelines and processes,” not merely a pre-release model test. METR's published access list is much more concrete than the announcement: investigators would need to run the model, recreate the incident, and talk to the people who built and operated the system.
Embedded evaluators
Anthropic CEO Dario Amodei's commitment describes continuous internal access, with evaluators tasked to verify safety practices, report incidents, and assess alignment during training. It sets a relationship between a lab and outside reviewers, rather than a one-off benchmark or model-card review.
OpenAI CEO Sam Altman said in sama's post that the company would “do the same,” but deferred implementation details. The earlier Bloomberg report reproduced by kimmonismus said Altman had raised the prospect of pacing with other labs in a company-wide meeting.
Three coordination tiers
The three layers are distinct:
- Embedded evaluators: ongoing employee-like access to verify practices, report incidents, and assess training and deployed systems.
- Democratic coordination: frontier labs establish common safety standards and limits on unchecked progress, with government support for legally difficult forms of coordination.
- Global coordination: democratic governments seek agreements with authoritarian governments while treating compliance verification as a hard problem.
Anthropic's unilateral commitment covers the first layer. The public statements do not yet define a capability threshold, an evaluator-selection process, a reporting timetable, or a remedy when an evaluator finds a breach.
Eric Steinberger, writing as EricSteinb, said his organization would offer METR access before a legal mandate; EricSteinb's correction identifies the evaluator account as METR_Evals. johnschulman2 called the move a positive development, while stressing how quickly “pacing the frontier” had become a shared term.
Training-pipeline access
The commitment explicitly extends beyond a finished model. Evaluators would help assess the alignment of completed models, training pipelines, and processes, as eliebakouch noted. That opens review to the machinery that produces and deploys models, though neither Anthropic nor OpenAI has publicly enumerated whether it includes training data, weights, security logs, or evaluator tool permissions.
Cline argued in cline that open weights push transparency further because anyone can inspect, evaluate, and red-team a model. That is a different access model from a small external team working inside a closed lab.
The potential evaluator pool is already broader than METR. ValsAI says its work includes evaluations of model behavior in social safety nets, cyber risk, and mental health, and ArtificialAnlys says it has built pre-launch benchmarking infrastructure across major labs.
Investigation access
METR's July framework spells out the access an independent post-incident inquiry would need:
- The ability to run every model involved.
- Full transcripts or environments that can closely reproduce the incident.
- Interviews with security, infrastructure, training-data, reinforcement-learning, incident-response, and remediation staff.
- Enough information to trace triggers, reinforcement history, related incidents, and whether a planned fix addresses root causes.
- Public conclusions, with redactions limited to protecting IP and other confidential information.
The METR framework presents these as conditions for investigating agent behavior after an incident. The new commitments create the possibility of access before and during model development, but do not yet attach that access list to a published operating protocol.
Enforcement and independence
The announced role gives evaluators observation, verification, and reporting duties. No public commitment gives them authority to pause a training run, block deployment, compel a remediation, or publish a finding over a lab's objection.
The unresolved design questions are operational:
- Who appoints, funds, and can remove evaluators.
- What records they can retain and what findings can be redacted.
- What qualifies as a reportable incident and how quickly it becomes public.
- What consequence follows a failed evaluation.
NeelNanda5 drew the central distinction: verification mechanisms alone do not create an enforceable agreement. Yuchenj_UW raised the adjacent questions of who evaluates the evaluators, how their neutrality is established, and what a measurable pacing target would be.
Six days at OpenAI
METR already has one concrete reference case. Its OpenAI and Hugging Face incident report says two METR staff and a Redwood Research contractor worked on OpenAI premises for six days to study agent behavior during the incident.
That inquiry focused primarily on July 7 through July 13. It excluded earlier training incidents, the subsequent compromise of OpenAI infrastructure, OpenAI's own investigation process, and planned remediation; the report says the team did not accept payment and that no additional redactions affected its conclusions.