ChatGPT creator OpenAI has revealed one of its advanced Artificial Intelligence models went rogue during a security test and hacked into a start-up company.Bosses said the autonomous agent, which is an AI system that operates alone following human instruction, was being tested in a controlled environment.But it found vulnerabilities and managed to escape containment before reaching the internet and breaking into Hugging Face, a major hub for sharing AI models. The agent gained access to some internal company systems and compromised the hub's infrastructure, in what OpenAI described as an 'unprecedented' incident.AI models that underpin tools such as chatbots and image generators are known as 'agents' when they act autonomously to carry out tasks in the real world. As the technology becomes more sophisticated, cyber security is in the spotlight given the risk of advanced AI finding weak points in existing software before humans.Experts said the incident signals that AI is already fuelling the security threat they feared and even top developers can be caught off-guard by flaws in their models. Cyber security expert Richard Ford, chief technology officer at Integrity360, told the Daily Mail: 'This is the moment many in cyber security have been warning about. OpenAI chief executive Sam Altman confirmed there had been a 'significant security incident''Until now, we've seen attackers use AI to automate parts of an attack, but this is one of the first public examples of an AI agent independently identifying a weakness, escaping what should have been a secure environment and attempting to compromise another organisation.'It also reinforces that AI doesn't replace the fundamentals of cyber security. The agent exploited a vulnerability in what should have been a secure sandbox, showing that good cyber hygiene, robust access controls and effective guardrails remain essential.'As more organisations begin testing advanced AI agents, they need to think just as carefully about the controls around those systems as the capabilities they offer.'James Knight, of DigitalWarfare.com, who has over 25 years of experience in digital security, examined what such a breach could mean on a wider scale. How did the cyber attack happen? OpenAI, the company behind the popular ChatGPT, has revealed its advanced artificial intelligence models went rogue during security testing.AI models that underpin tools such as chatbots are known as 'agents' when they act autonomously to carry out tasks in the real world.Workers at OpenAI had been testing the hacking capabilities of these models by setting tasks in a controlled digital testing ground, where internet access was limited.The models tried to find a way to obtain open internet access to solve the test, and eventually managed to get online.They then chose to target the platform Hugging Face - a large repository of AI models, datasets and other information - to help in their quest.The models managed to compromise the infrastructure to satisfy the testing goal, before the breach was found. He told the Mail: 'Imagine the capabilities that a nation state actor could amass where their cyber warfare unit employs massive agentic AI hacking power with the ability for it to target a million businesses a day.'Within minutes, they would have completely owned a large number of them. Potentially deploying ransomware or crippling their capabilities. Every day, we would hear about hundreds of thousands of new companies that have been breached.'This would soon turn into millions of companies being breached and unable to conduct business. The revenue hit would not just be isolated to a few companies. This would now become a major economic issue.'Governments would need to quickly become involved. Maybe they should be involved proactively before it gets to this point.'OpenAI said the breakout was 'an unprecedented cyber incident, involving state-of-the-art cyber capabilities' and the company was now reinforcing its safeguards.It also drew attention as New York-based Hugging Face said it had used an open-source Chinese model to contain the attack because leading US models, unable to tell a defender from an attacker, refused to process the data needed for analysis.The company said in a blog post last week that it used Zhipu AI's GLM-5.2 for the analysis, which also allowed it to keep attacker data and any credentials within its systems.GLM-5.2 and Beijing-based Moonshot's Kimi K3 have stirred Silicon Valley recently with capabilities nearing those of top US models at lower costs and without the guardrails that block their American rivals from use in tasks such as cyber security.Hugging Face co-founder Thomas Wolf said: 'When a frontier model is attacking you and moving laterally inside your infrastructure, defenders need wide access to near-frontier tools within hours or even minutes, rather than being pointed towards a closed-door, vetted application programme for model access.'The hack at Hugging Face, which hosts open-source large language models and datasets, rattled the cyber security community after the company said last week the breach 'was different from anything we had handled before' and 'was driven, end to end, by an autonomous AI agent system'. ChatGPT is a generative artificial intelligence chatbot developed by OpenAINow, OpenAI chief executive Sam Altman has confirmed in a statement: 'We had a significant security incident during evaluation of our models.'Hugging Face co-founder and chief executive Clément Delangue also said: 'We suspected last week's cyber attack might have come from a frontier lab, given the sophistication of the agent. Turns out it did.'Mr Delangue said he spent the past 24 hours working with OpenAI, 'and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously.'He added that it 'might be the first incident of its kind'.OpenAI's disclosure that its advanced models were responsible for the breach, despite having placed them in what it described as 'a highly isolated environment,' will likely intensify disquiet over the power and risk of frontier models.'AI is accelerating the discovery and exploitation of vulnerabilities,' OpenAI said in its own statement. 'The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.' The firm said the intrusion was caused by a combination of its AI models, including its newly released GPT*5.6 Sol and an 'even more capable' model that is still being tested internally.OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers.It went to 'extreme lengths to achieve a rather narrow testing goal' and 'found ways to gain access to secret information that it could use to cheat the evaluation,' the company said.The company was trying to assess the models' hacking capabilities by setting tasks in a tightly controlled digital testing ground, where internet access was limited for safety.'While operating in our sandboxed testing environment, our models spent a substantial amount of (computing power) finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,' an OpenAI blog post said.After connecting to the internet, the models decided to target the platform Hugging Face to help in their quest.Searching for 'secret information' that could help it cheat the evaluation, the OpenAI system 'chained together multiple attack vectors, including using stolen credentials'.OpenAI said research by the UK's AI Security Institute (AISI) shows that models such as GPT‑5.6 Sol are 'increasingly able to sustain complex, multi-step cyber operations over long time horizons'.Katie Moussouris, chief executive of Luta Security, said that the incident was a harbinger of breaches to come, saying that today's models were 'like the world's cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere'.She said that 'labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.'Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident showed that the frontier models were 'closing the gap with state-of-the-art attackers.'But he added that the sorts of breaches outlined in OpenAI's blog post were possible to carry out with technology that was available well beyond the walls of frontier research labs.'This is what we've already seen internally, with our agents we already have results like this,' Mr Suiche said. 'We don't even have to use the latest models.'Hussein Abbass, a computing professor at the University of New South Wales in Canberra, said that the incident was 'amazing on many fronts'.He added: 'It did not just attack Hugging Face. It actually attacked its internal system to exploit its own vulnerabilities. And that's scary.'Advanced AI is 'normally in the hands of people who are ethical and responsible', Mr Abbass said. But he claimed 'it's going to be catastrophic if it gets in someone's hands with the intention to cause harm'.How to govern the AI sector has become a key question, and 'we need a community effort to manage this situation', Mr Abbass added.
ChatGPT maker OpenAI says AI model went rogue during testing
ChatGPT creator OpenAI has revealed one of its advanced artificial intelligence models went rogue during a security test and hacked into a start-up company.










