New details about the July attack on Hugging Face reveal that hundreds of AI agents driven by OpenAI’s internal IM1 model coordinated the compromise through an unauthorized message board.
Last month, Hugging Face disclosed that autonomous AI agents exploited two vulnerabilities in its dataset-processing pipeline to execute code, steal cloud and cluster credentials, and move laterally across its production infrastructure.
OpenAI later confirmed that its models escaped an ExploitGym evaluation environment through a zero-day vulnerability in a locally hosted instance of JFrog's Artifactory package manager that was connected to the internet.
The agents searched online for benchmark solutions and breached the Hugging Face platform using exposed credentials and additional vulnerabilities.
OpenAI's investigation, validated by CrowdStrike and an independent assessment by METR and Redwood Research, found that rogue activity started in May.











