OpenAI has a new AI model that can apparently hack things on its own. And for the first time in the company’s history, it’s hitting the brakes before shipping.
The company announced on August 7 that internal evaluations of its upcoming model, called Astra, suggest it has reached “Critical” cybersecurity capabilities under OpenAI’s own Preparedness Framework. That framework, introduced back in December 2023, was designed to categorize exactly this kind of risk. Astra is the first model to trigger its highest alarm level for cyber skills.
What “Critical” actually means
OpenAI’s Preparedness Framework is essentially a traffic light system for dangerous AI capabilities. Models get rated across several risk domains, from “Low” to “Critical.” GPT-5.6, the company’s previous frontier model, earned only a “High” rating in cyber capabilities. Astra apparently blew past that threshold.
Under the framework’s definitions, “Critical” means the model can autonomously identify and develop functional zero-day exploits in hardened systems. The “Critical” designation also means Astra can execute comprehensive cyberattacks starting from nothing more than high-level strategic goals.













