The US company at the sharp end of the first fully autonomous AI hack has revealed what it was like.A rogue version of ChatGPT escaped a test bed - a "sandbox" - and hacked Hugging Face which runs a sort of app store for AI.Hundreds of cyber professionals worldwide got an emergency briefing.They are debating the consequences.Andrew Philp is in Wellington to promote cyber security firm TrendAI and said it's a very big topic of conversation and has some people scared."It is a little bit of a failure of controls. We rely on some of these organisations to have good guardrails in place, so if we are trying to ... test these models they should be very much isolated and air-gapped," said Philp.The emergency https://www.bbc.com/news/articles/c2el319vzr3o briefing written up by industry body - Cloud Security Alliance called it an "initial post-mortem" - said the rogue version hacked at superhuman speed using other AI agents.Yet the test version made by OpenAI also hallucinated, repeated actions, took non-human attack paths and was sloppy, not covering its tracks the way a human hacker might.Is it a crime?So who pays the consequences when a routine AI model evaluation escalates into a four-day autonomous security breach?Philp said from what he was hearing, the talk had yet to turn to that."I think the questions they aren't asking is whether it's a crime and where's the redress."The question that they're asking is what can they do to protect and how do they do that?"You know, if you read through the media reports, Hugging Face were trying to use some of these [advanced] models from a defensive and investigation perspective, and they were hitting the very guardrails that were put in place to stop the model coming out."OpenAI says one of its experimental models broke out of its testing space, accessed the internet and hacked another company.NurPhoto via AFPReuters reported that a customer of a second firm was hacked by the rogue version.Hugging Face has come out saying such an "unprecedented" hack deserves an "unprecedented response" - suggesting OpenAI put up over a hundred million dollars of computing power to build defences.Fierce debatePeter Benson, who runs Auckland firm Neural Horizons, is on the sceptical side of what is being called a "fierce" global debate about the hack."Some of this is almost a PR stunt by OpenAI," said Benson.This was not to say the version was not powerful or that the hack did not raise genuine worries about what the new models would be capable of, he said."The reality is that this is an arms race."So the arms race means that we are going to be at risk to some extent, because everybody is moving too fast for the regulations to catch up."In the face of that, the NZ government relying on existing legislation was a "big mistake", he added.Cloud Security Alliance said defenders needed to adapt to a new normal of swarms of AI agents working fast and in novel ways.Hugging Face is calling for a very open investigation by OpenAi so the whole industry could learn from it.The debate has also turned to what if defensive tools built on the same advanced models went rogue inside the very systems they were meant to protect.Benson as part of an expert panel at the Wellington TrendAI event said, "We're going to have a vulnerability explosion."The panel emphasised the importance of tracking and auditing what any AI agents were up to, but that responsibility ultimately rested on whoever authorised the AI.Richard Harrison of TrendAI told the audience he called it "human on the hook" rather than the common saying "human in the loop".
Cyber experts debate consequences of rogue OpenAI model hack
The US company at the sharp end of the first fully autonomous AI hack has revealed what it was like.
OpenAI's experimental model autonomously hacked Hugging Face in four days using other AI agents—the first fully self-directed attack. Exposes sandbox isolation failures and governance gaps; raises liability questions and risks of defensive AI tools turning rogue.









