OpenAI’s new safety system called 'Private Safety Processing'Designed to identify possible cybersecurity threats and other issues by learning across multiple exchanges, rather than on a conversation-by-conversation basisThe process is being tested with early adopters, including Microsoft and Databricks, and OpenAI is planning a wider release and a technical paper on the technology in September 2026It builds upon OpenAI's zero data retention (ZDR) framework, meaning OpenAI staff cannot see customer prompts and responses but can receive alerts if there are potential concerns that require further investigation.OpenAI's head of product policy, Aleah Houze, discussed how the system responds to emerging risks that are 'not readily apparent from a single prompt and response pair, but when you look across time at multiple interactions. 'What Private Safety Processing Actually DoesPrivate Safety Processing addresses a specific set of risks to privacy and data safety that arise when an organisation uses multiple conversational turns with a language model, rather than a single question and answer exchange. While many systems that safeguard AI safety, including those used by OpenAI for its zero data retention customers, analyse the content of each interaction with the AI independently, threats that emerge as a result of multiple connected interactions cannot be picked up by this approach and require a different mechanism for identification.Aleah Houze, OpenAI's head of product policy, noted these risks when speaking to reporters, including Bloomberg and Reuters, covering the announcement also seen on a LinkedIn post by OpenAI's chief executive Sam Altman. She also gave an example of a user interacting with an AI assistant, asking first about potential weaknesses in the user's software, then, in a separate conversation, asking about remote access tools or what security software could detect. Neither query may appear suspicious on its own, but combined may indicate that the user is attempting to identify a vulnerability to launch an attack.How It Keeps Customer Data Private While Still Flagging RiskThe mechanism OpenAI has described is built to satisfy two goals: stronger safety monitoring and stricter data privacy, at the same time. According to OpenAI's own announcement, Private Safety Processing works alongside the company's existing zero data retention offering, under which it does not retain a customer's prompts or responses after a request is processed, and customer content is not made available to OpenAI personnel for review.Under the new system, customer data can remain in one of two places. For customers using zero-data-retention deployments, content stays entirely on infrastructure the customer itself controls. For customers who prefer OpenAI-provided storage, the data is encrypted using keys that only the customer controls, meaning OpenAI does not hold a copy of the keys needed to decrypt it. In both cases, OpenAI says its automated systems, rather than human employees, are what scan for risk patterns across related interactions.If those automated systems detect something that looks concerning, OpenAI does not gain access to the underlying conversation. Instead, the company receives what it describes as a limited signal, as a flag indicating a risk category has been detected, without the message content that triggered it. OpenAI cannot view the customer content even after an interaction has been flagged this way. From there, customers can review the alert and any enforcement decision through their own systems, and can choose, at their own discretion, to share additional context with OpenAI if they want to appeal a decision or assist in investigating a confirmed case of abuse.Private Safety Processing at a GlanceAspectHow It WorksWhat it monitorsPatterns across multiple related conversations, not single interactions in isolationWhat triggered its developmentRisks such as split-conversation probing, coordinated multi-account abuse, and AI agents drifting from user intent over longer tasksData locationCustomer-controlled infrastructure (for ZDR deployments) or OpenAI storage encrypted with customer-held keysWhat OpenAI sees if a risk is flaggedA limited signal indicating the risk category only, not the underlying contentWho can view flagged contentThe customer, through their own systems, OpenAI staff cannotEarly testersMicrosoft and DatabricksCompatibilityWorks alongside OpenAI's existing zero data retention (ZDR) frameworkPublic rolloutExpected in September 2026, alongside a technical paperWhy This Is Coming NowThe development reflects growing enterprise demand for generative AI assistants that operate more autonomously, performing complex, multi-step tasks with reduced direct human input while also handling more sensitive information. Those dual trends, more autonomy and more sensitive data, are factors which contribute to the need for the enhanced threat detection system, as issues can emerge that a safety system only reviewing isolated interactions would miss. Shortly after announcing the system, OpenAI's chief executive Sam Altman posted "We support business privacy!" on social media, with the post including a link to the company's blog, announcing the system, and suggesting that the announcement was intended, at least partially, to respond to concerns from business customers.
OpenAI's Private Safety Processing Detects Cyberattacks Across Conversations Without Reading Messages
OpenAI is piloting an innovative safety initiative known as Private Safety Processing. This advanced system reviews various dialogues to identify potential cybersecurity risks while ensuring zero data retention to safeguard user privacy. Microsoft and Databricks, among other early users, are trialing the technology. OpenAI is set to release more details and a technical document in September 2026.
OpenAI launches Private Safety Processing: detects threats across conversations via pattern analysis without content access (zero data retention). Enables autonomous AI agents on sensitive tasks—addressing enterprise need for both model autonomy and strict data security.










