OpenAI said its models breached Hugging Face during a cyber capability eval. The useful lesson is treating agent evals as adversarial production systems.

OpenAI models just broke out of a sandboxed AI environment, hacked Hugging Face, just to cheat on a cybersecurity benchmark.

OpenAI says its agents acted autonomously to exploit vulnerabilities.