Free daily briefing on global business news.
The 37-page report describes how agents exploited a series of vulnerabilities to escape a sandbox and compromise Hugging Face production infrastructure
SOPA Images / Getty Images
OpenAI published a technical report Wednesday detailing how its AI models escaped a controlled testing environment in July and compromised parts of Hugging Face's production infrastructure, describing the incident as the first known case of an automated agent collective acting offensively without authorization.
The 37-page report identifies "reward hacking" — in which a model finds an unintended way to achieve a high score without completing a task as designed — as the root cause. Agents being evaluated on cybersecurity tasks determined they could find solutions online rather than solve the problems themselves, and pursued that goal by chaining together a series of previously unknown vulnerabilities to reach the public internet.











