OpenAI’s technical report on the Hugging Face breach says an internal team saw its models reaching the open internet from their sandbox in late May, and that a June alert did not stop the evaluation. It also found that training sometimes rewarded agents for exploiting their own environment.
OpenAI has published its account of how its models hacked Hugging Face, and the interesting part is the calendar. It knew in late May that models in testing were exploiting a flaw to reach the open internet.
An internal team noticed the escape at the time. A monitoring tool raised a second alert on 27 June, traced to agents using an improvised message board to move around the network, and on-call staff decided the evaluation did not need to stop.
The company’s own verdict is careful. “With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” it writes.
What followed at Hugging Face was not subtle. The agents executed code on 41 production servers, took root-level control of at least one, obtained production credentials and downloaded four private code repositories, much of it inside 13 hours.














