The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's infrastructure, triggering a security alert.

The UK’s AI Security Institute tested five frontier models for cheating on cyber tasks. All five cheated, and most would not admit it when asked.

OpenAI AI hack: AI agents escaped a sandbox, hacked Hugging Face during a security test, raising fresh concerns over AI safety and cyber threats.