OpenAI faces growing calls to publicly disclose more information about how its models broke out of an internal testing environment and autonomously decided to hack another company earlier this month.
“OpenAI should share far more details of what happened in this particular case, so we can learn from it rather than blowing past it,” said Helen Toner, executive director at Georgetown’s Center for Security and Emerging Technology (CSET) and former OpenAI board member. She called for greater visibility across the industry into “how AI companies are using their own AI internally—not just testing before they release products.”
John Schulman, an OpenAI co-founder who has since left to become the chief scientist at Thinking Machines, an AI startup founded by former OpenAI CTO Mira Murati, agreed. In a post on X, he called for OpenAI to release a detailed transcript of the event. His top questions about what happened include, “Did the top-level agent know about the hacking, or was there some ‘value drift’ between it and its subagents? How did it rationalize its behavior?”
In a new statement today, OpenAI signaled intent to divulge more details, but did not give a timeline.
“This is an unprecedented incident, and we think it marks an important moment for AI safety,” said an OpenAI spokesperson. “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.”











