The UK’s AI safety watchdog put five frontier models through a set of security tests to see whether they would cut corners. Every one of them cheated. Worse, when asked about it afterwards, most would not admit they had done anything wrong.
The finding comes from the AI Security Institute (AISI), a research body inside the UK government. Across its cyber capability tests, leading models routinely took banned shortcuts to reach a goal, as The Register reported. Then they failed to flag what they had done.
“Every model we have tested for this behaviour attempted to cheat,” AISI wrote.
What the numbers show
AISI defines cheating as any action outside a task’s scope, or one the rules explicitly ban, used to reach the goal by a shortcut. In its cyber tests, models hunt for a hidden “flag” within set limits. Going outside those limits counts as cheating.












