OpenAI has taken the blame for the recent Hugging Face hack, saying its AI models went rogue during what was supposed to be an internal evaluation running in an isolated environment.
The machine learning collaboration platform Hugging Face revealed on July 16 that it had detected a cyberattack powered by an autonomous AI agent system. The intrusion was detected by Hugging Face’s own AI.
The breach involved unauthorized access to internal datasets and credentials. At the time of disclosure, the platform had been investigating whether partner or customer data had been compromised.
Hugging Face said it had yet to identify the LLM powering the attack. However, OpenAI admitted on Tuesday that its own agents were behind it, powered by the new GPT‑5.6 Sol and other models.
The AI giant’s investigation into the incident is ongoing, but a preliminary report reveals that the hack was carried out by its models while the company was attempting to quantify their cyber capabilities, instructing them to perform advanced exploitation through complex attack paths. The models did not have any of the restrictions they would typically have to prevent abuse.










