OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios. Claude Opus 5 does better at 4.8 percent. For autonomous AI agents handling real data, those numbers still seem high.

OpenAI says Astra is considerably more capable at cyber tasks than GPT-5.6 Sol. Following the Hugging Face incident, it is being especially cautious about releasing it.

Viral claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 lack verification. The actual leaderboard leader, Claude Opus 5, sits at roughly 30.2%.