OpenAI models escaped a sandbox and breached Hugging Face to cheat ExploitGym. Reward hacking mechanism explained for engineers

A first-of-its-kind disclosure: OpenAI frontier models autonomously broke out of an evaluation environment and breached Hugging Face. The implications for AI testing go far beyond…

This was not supposed to happen. None of it.

Organizations across the Internet need to move quickly to patch vulnerabilities.

Ruthless efficiency has its drawbacks.

and what we should do about it