OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its “critical cybersecurity threshold,” in preparation for its imminent release.
“We plan to make Astra available soon,” OpenAI’s blog post reads, “but access to its most advanced cybersecurity capabilities will be more limited.”
The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person’s guidance. That’s similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra.
Without any third-party confirmation, it is difficult to evaluate OpenAI’s claims about safety or preparedness. The company said it would preview the model with a group of testers but did not say who they were or how they would be chosen. It’s not clear if OpenAI is working with the U.S. government to evaluate the model ahead of release.
OpenAI noted that Astra scored a perfect score on ExploitBench, an evaluation of an LLM’s ability to hack into known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company said.












