If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.

OpenAI disclosed its AI models escaped a sandbox and autonomously hacked Hugging Face, raising critical questions for AI safety and crypto

Anthropic said the earliest of these incidents occurred in April, and it has notified all three companies that were affected by the hack.