Following the Hugging Face hack and the more recent "wiki incident," OpenAI stated on Saturday that it's working on a "framework" for how and when it shares information about incidents involving rogue agents.On Sept. 4, Reuters reported that OpenAI agents had taken control of a German-language wiki site and used it to communicate with one another. Four anonymous company insiders told the news outlet that OpenAI and its legal team resisted internal efforts to investigate the incident.
In a Sept. 5 post on X, the ChatGPT owner acknowledged the breach: "How we think about the 'wiki incident,' where our agents wrote to several internet sites: it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models."
This Tweet is currently unavailable. It might be loading or has been removed.
The statement continues, saying the company historically treated "misalignment" as a research question, but in 2026, OpenAI has "started to see misalignment cause new types of real-world impact."OpenAI followed a traditional security incident response playbook for Hugging Face, the statement reads, and its investigation continues.
Mashable Light Speed










