OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump.

OpenAI says Astra is considerably more capable at cyber tasks than GPT-5.6 Sol. Following the Hugging Face incident, it is being especially cautious about releasing it.

The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.