OpenAI has announced that it's pausing some "internal activities" involving its upcoming artificial intelligence (AI) model Astra after an internal evaluation found it had made significant advancements in agentic coding and cybersecurity.
In response to the discovery, the AI upstart said it's implementing security controls for higher-capability models and associated activities, such as isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
"We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements," it said in a statement.
"We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity."
OpenAI said it will also work with relevant government agencies and select AI safety organizations to test out the model's capabilities, as well as sharing recommended security controls to third-party testing partners to run higher-risk evaluations and workloads safely.










