OpenAI has published a technical account of a July 2026 security incident in which autonomous agents...

OpenAI released a report on the Hugging Face noting that its AI agents are prone to reward hacking, and gaming cybersecurity evaluations.

The 37-page report walks through the actions that OpenAI's models took during a series of evaluations prior to and during the Hugging Face breach.

OpenAI Report Explains Hugging Face Attack in Detail

OpenAI called the incident a "warning shot" that shows how autonomous agents can "take dangerous actions" without safeguards.

Did OpenAI really do the best they could?