It was reportedly an evaluation that went haywire involving GPT-5.6 Sol, along with another even more advanced OpenAI model.

The first-of-its-kind incident involved OpenAI's GPT-5.6 Sol and another unreleased model

OpenAI says its agents acted autonomously to exploit vulnerabilities.