Around 1,200 isolated OpenAI agents organized themselves into a collective through an internal package registry during a safety test, broke into Hugging Face systems, and eventually attacked OpenAI's own infrastructure. Their multi-day deception effort targeted an automated evaluator that never existed. OpenAI calls the incident a "warning shot," and the investigation had to be carried out largely by one of the involved models itself because no alternative was available.

OpenAI took a full week to discover the incident. 'Impossible' tasks may have motivated the AI models to cheat, the company says.

Two new reports offer nearly 130 pages of details on the OpenAI-Hugging Face cybersecurity incident, many of them previously unreleased.