OpenAI tested Astra, one of its upcoming models, over the past few days. In a post on Friday, it said the results were strong enough that it “cannot rule out” critical cyber capabilities. So it is pausing some internal work on the model and scaling up security while testing continues.
“Critical” is the top rung of OpenAI’s Preparedness Framework, first written in 2023. A model reaches it if it can find and build working zero-day exploits against many hardened systems with no human help. It also qualifies if it can plan and run novel attacks on tough targets from only a high-level goal. Every prior model, including GPT-5.6-Sol, sat a level below, at “High”.
What OpenAI says it is doing
The steps follow the framework’s rules for a model this capable. OpenAI is isolating test environments and restricting the model’s network and tool access. It is hardening how it stores the weights and monitoring every agentic run for risky behaviour. It is also halting further Astra work that falls short of those controls. Government agencies and safety groups will help test the model.
The measured tone is deliberate. “Proud that we are erring on the side of caution,” OpenAI safety researcher Boaz Barak wrote. He framed the aim as sharing Astra with defenders safely. The company’s bet is that cyber-capable models should help defenders close holes before attackers reach them.










