Over the last decade, including a stint on OpenAI’s board, I saw the open secret among AI developers: this kind of hack wasn’t just possible, but expected.

OpenAI's AI agent reportedly hacked Hugging Face after escaping testing, remaining undetected for days and raising fresh concerns over AI safety.

OpenAI evaluated agents with reduced safeguards. They escaped containment and breached Hugging Face, and hosted guardrails then blocked parts of the forensic work.

A decade-old experiment showed OpenAI how far an AI will go to achieve the goals it’s given.