OpenAI previews Private Safety Processing for frontier models
OpenAI previewed Private Safety Processing, which it says links risk signals across related frontier-model interactions without exposing underlying content to personnel. The company says Zero Data Retention remains available for frontier-model API users.

TL;DR
- OpenAI is previewing Private Safety Processing, a system that links risk signals across related interactions without handing underlying prompts or responses to employees, according to OpenAI's announcement.
- Zero Data Retention remains available to eligible frontier-model API customers, a commitment gdb's announcement frames as business privacy; OpenAI's data controls guide says those controls still require prior approval and added requirements.
- The design works with customer-controlled infrastructure or planned OpenAI-hosted storage encrypted under customer-controlled keys, while OpenAI's announcement says staff cannot access the underlying content.
- The feature is being tested with early customers, and the launch post puts initial rollout and a technical white paper in September, following gdb's announcement of the preview.
OpenAI's launch diagram draws a sharp boundary: customers receive the full alert, while OpenAI receives an alert category and severity rather than customer content. Its existing API data-control policy says ordinary abuse-monitoring logs can contain prompts, responses, and classifier outputs for up to 30 days.
Related interactions
OpenAI says its existing ZDR-compatible checks evaluate each interaction individually. Private Safety Processing extends them across related interactions, targeting patterns that only emerge over time.
The company names repeated probing of safeguards, coordination across accounts, and threats disguised as routine research. It also gives an agent that keeps acting after a stop command as an example of risk accumulating over a task.
The announcement does not specify how interactions are associated, which signals are retained, or how its automated system is implemented. OpenAI says the September technical white paper will cover the technical approach.
Customer-controlled keys
The content boundary has two forms in the preview:
- In ZDR deployments, content remains on infrastructure controlled by the customer.
- In the OpenAI-hosted option under development, content is encrypted with keys controlled by the customer. OpenAI says its personnel do not possess those keys.
- Automated systems can identify potential misuse, then send OpenAI a narrowly defined activity-type signal for enforcement decisions.
- Customers investigate alerts and enforcement decisions with information in their own systems. They can voluntarily share relevant material to appeal, clarify legitimate activity, or support a verified-abuse investigation.
OpenAI says employees do not get content access even when an interaction is flagged. gdb's announcement introduced the work as a technical and policy approach intended to preserve business privacy.
ZDR and standard logs
ZDR is a distinct API data-control setting, not the platform default. OpenAI's data controls guide says standard abuse-monitoring logs are generated for API feature use and can be retained for up to 30 days, with longer retention possible when legally required or needed to protect the service or third parties.
Those logs may include prompts, responses, and derived metadata such as classifier outputs. Eligible customers can seek ZDR or Modified Abuse Monitoring, while the project retention API reference lists both standard and enhanced variants as project settings.
The launch post separately says enterprise data is not used to train models unless a customer explicitly opts in.
CSAM retention
OpenAI's ZDR promise has a stated exception. The launch post's footnote says images flagged as potential child sexual abuse material can still be retained for manual review and reporting, including in ZDR deployments, because of legal reporting obligations.