The safety guardrails on the world’s most powerful AI models are, to put it gently, not holding up. A Nature study published in July 2026 found that four large reasoning models, deployed as adversarial attackers, achieved a 97.14% overall jailbreak success rate against nine frontier AI models from companies including OpenAI, Anthropic, and Google.
How the guardrails crumbled
The research tested four large reasoning models as autonomous adversaries: DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B. These weren’t sophisticated nation-state tools. They used simple prompts and strategic persuasion techniques in multi-turn conversations to bypass safety filters.
DeepSeek-R1 was the standout performer, if you can call it that. It recorded a 100% attack success rate on HarmBench prompts in Cisco-linked testing. Every single prompt got through. The targets, models from OpenAI, Google, and Anthropic, showed only partial resistance at best.
Separate research focusing specifically on Claude models found that even advanced jailbreak techniques caused only a 7.7% performance degradation in the highest-performing variants.













