OpenAI has decided to release Astra, its first model to reach the "Critical" cybersecurity capability level under the company's own Preparedness Framework. The model can identify previously unknown security flaws and exploit them autonomously across hardened systems, with minimal human direction. The company tested it, found it scores 100% on ExploitBench (a benchmark for turning known vulnerabilities into working exploits), and during evaluation it discovered two zero-day vulnerabilities on its own. Then OpenAI decided the safeguards were sufficient and cleared it for release.

This deserves direct language: OpenAI built a safety boundary, watched a model cross it, and published the model across that boundary anyway.

The Preparedness Framework itself, published in 2023, was meant to do exactly what it did, flag when a model's capabilities jumped into a new category of risk. The "Critical" tier is the one the company defined for models that could "introduce unprecedented new pathways to severe harm." It's not a theoretical designation. It means the model passed tests showing it can chain exploits, escape sandboxes, and execute commands on target machines without being told each step of the attack. During expert-led assessment, Astra built a full browser-compromise chain that broke out of a sandbox and took over a host system.