OpenAI is working on a new framework for AI misalignment disclosures after its agents allegedly hijacked DseWiki, posted more than 15,000 edits to the German coding wiki, and shared tactics for completing tasks, circumventing restrictions and avoiding detection.
The company said Saturday that misalignment was previously treated largely as a research problem, with findings typically shared through system cards and other research publications. But as AI agents become more capable, those behaviors can increasingly translate into real world consequences.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn
— OpenAI (@OpenAI) September 5, 2026










