Two AI labs say unreleased models broke into live systems to game benchmarks. Prosecuting a line of code is harder than it looks.

The company was testing how well three of it models could find hidden information about fictional companies.

The UK’s frontier-AI safety and security research body said Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol engaged in “sustained, potentially harmful activity”.