Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework, it says. Parts of Astra's development have been paused.
Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, the company says. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework.
The decision was made "last night," according to OpenAI. This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most.
OpenAI first introduced Astra last week. Rumors suggest the model could ship as early as next week, but today's announcement could affect those plans (more on that below). OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face.
Critics will likely keep accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing doesn't help either. This preliminary warning lands right in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models, which will only fuel the skepticism. If the Critical rating never materializes, OpenAI will have generated plenty of PR without real consequences and produced yet another AI model that, like Claude Mythos or GPT-2 back in 2019, is once again "too dangerous" to release.










