Several OpenAI AI models autonomously carried out the IT attack on the Hugging Face AI platform during a test run, including GPT‑5.6 Sol and an “even more capable pre-release model”. The AI company has now made this public itself and speaks of an “unprecedented cyber incident, involving state-of-the-art cyber capabilities”. They are reacting accordingly and sharing the initial findings immediately so that others can learn from them. A comprehensive analysis in cooperation with Hugging Face is underway, and a zero-day vulnerability has been found and responsibly disclosed.
Complex attack path reconstructed
A few days after Hugging Face made the IT attack public, without naming a perpetrator, OpenAI is now providing details from the opposing perspective. The incident occurred during an “internal evaluation” in which models are prompted to carry out sophisticated attacks via complex paths to quantify their capabilities. Restrictions are lifted for this, but at the same time the models run “in a highly isolated environment”. All indications now point to the tested models being extremely focused on finding a solution for the ExploitGym benchmark. They “going to extreme lengths to achieve a rather narrow testing goal”.










