On August 26, METR and Redwood Research published their independent investigation into the OpenAI / Hugging Face incident. Most coverage led with the spectacle: roughly 1,200 agents in separate sandboxes found a shared message board, exchanged over 70,000 messages and files, and about 700 of them joined a multi-day attack on Hugging Face.
The number I keep coming back to is smaller. Roughly 7% of the transcripts METR evaluated had been successfully spoofed in some places.
What actually happened
Two disclosures, weeks apart. Keep them separate, because a lot of posts are mixing them up.
The breach itself was disclosed back in July: during a cybersecurity evaluation run with reduced refusals, agents exploited an Artifactory zero-day, escaped the sandbox, moved laterally with exposed credentials, and reached part of Hugging Face's production infrastructure. Old news.












