OpenAI has conducted this type of model testing for years, and there have been early warning signs in previous models of systems that will act maliciously and attempt to escape environments.
In April, Anthropic’s Mythos model also gained internet access and published details of a security exploit online publicly, beyond what researchers anticipated the model would do.
Mythos, and Anthropic’s subsequent Fable model, made reverberations in the cyber security community and caused governments around the world to home in on the idea that attacks on digital and critical infrastructure will be increasingly AI-led and autonomous.
Jake Moore, global cyber security adviser at ESET, a cyber security company, said OpenAI would inevitably use the breach as a marketing tool, given how much rival AI developer Anthropic benefited earlier this year from similar concerns. “I just don’t think that OpenAI had a matching story and so maybe they’d been waiting for something like this,” he added.
Following this incident, many in the AI safety and cyber security communities have called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on the next generation of AI systems.










