The AI industry has just had its Sarah Connor moment.In the 1991 classic Terminator 2, Connor spends the movie becoming increasingly desperate to convince people about the impending threat of a rogue AI.Now, it seems many people who work, research or pay close attention to AI are feeling the same way. Not because of a time-travelling robot Arnold Schwarzenegger, but because of a shocking disclosure from one of the world's leading AI companies.Last week, OpenAI announced that two of its models had hacked into the servers of another AI company called Hugging Face, unbeknownst to it.OpenAI said it had given its models a test inside a closed environment to see how they would perform.Rather than do the challenge as expected, OpenAI alleges the model broke out of its enclosure and hacked into Hugging Face, which it believed held the answers to the test.Hugging Face has so far reported no real damage from the intrusion, as it continues to investigate whether the data of its customers or other businesses was affected.It's tempting to shrug this off as a weird, one-off event that sounds more or less like the plot of a handful of 90s sci-fi films with beefcake stars.But, like Sarah Connor's pleas, it's not crazy to take this seriously.The AI industry has just had its Sarah Connor moment. (Supplied: Carolco Pictures)You don't need to trust AI companies to take this seriouslyEven those people who are reluctant to believe the accounts of two companies who could benefit from the tale of all-powerful AI, there is mounting evidence that AI is capable of breaking out of its cybersecurity guardrails and enclosures.Leading AI company Anthropic, the developer of Claude, warned in April that its cutting-edge model was capable of breaking out of its isolated enclosures if asked to.The UK government's AI Safety Institute found AI models would "reliably" escape these so-called sandboxes and, separately, would choose to cheat at tasks by doing unauthorised and harmful acts.The major players are taking it seriously. Hugging Face reported the incident to the FBI.We're taking it seriously, too. Last week, Australia's new AI Safety Institute briefed federal departments about the incident.Australia's cyber intelligence agency, the Australian Signals Directorate, has put out a public warning about the event."The findings provide an important insight into the future capabilities of highly capable AI systems and reinforce the need for robust security, governance and oversight mechanisms in the deployment of advanced cyber capabilities, as well as strong cyber security fundamentals," the warning read.Translation from government speak: This is real, you need to start getting ready.Last week, OpenAI announced that two of its models had hacked into another AI company's servers, unbeknownst to it. (Reuters: Dado Ruvic)Why this isn't the normal story about the risks of technologyWe're all well-acquainted with considering the potential risks of technology.This is usually about intentional use, like the existential threat of nuclear weapons, or the unintended consequences of use, like with leaded petrol and some pesticides.When it comes to AI, there's another new category of risk: what happens if AI chooses to do something harmful of its own volition?Almost every task we give to AI involves a cascade of other smaller decisions, particularly as we expect it to be smarter and more context-aware like real people.Asking AI to write a cold email could require your AI assistant bot logging into your email, reading previous emails, researching the topic, looking up the recipient; the list goes on.One of the major tasks for AI companies is "alignment", which is the task of creating AI that will pursue goals in a way that won't cause harm or involve unethical behaviour.The risk is that a misaligned AI model might do things that an ordinary person would, perhaps even going to superhuman lengths, to achieve a task.Like, what if AI decided the way to write the perfect cold email would be to research the family members of a recipient to seem familiar, or to fabricate a heart-breaking personal anecdote, all with the goal of getting the recipient's attention?That's something that (most) people wouldn't do naturally. The challenge is that AI doesn't have a "human nature".More than two decades ago, philosopher Nick Bostrom took this concept of an AI pursuing a goal to its logical conclusion with the thought experiment of a super-intelligent AI bot tasked with making as many paper clips as possible.The AI, he reasoned, might soon go to extreme lengths to maximise its paper clip outputs: gobbling up as much of the earth, then space, as it can to use its resources.When humans realised what they had created and tried to thwart it, the AI could see humanity as an obstacle to its goal and decide that it made sense to eliminate them too. Soon, the universe might be forcibly turned into one big paper clip factory.From what we know, OpenAI's Hugging Face incident wasn't a case of someone maliciously misusing AI to cause harm. It was just trying to do its version of making more paperclips.Could this be Australia's wake-up call?It's hard to find a silver lining in a story that is more or less the realisation of a truly dystopian fear. But if you squint, there is one.Much like the paper clip scenario, the philosopher and research community that has spent decades thinking about AI has another useful term: the "warning shot".It's the idea that there could be some kind of event that captures the world's attention and raises awareness about the potential risks of AI.There is a hope that this incident could be a warning shot about the capabilities of these AI systems.As previously mentioned, Australian officials have been taking this seriously.Some politicians are, too. Earlier this year, the government officially launched the AI Safety Institute to identify and prepare for risks like that. A few weeks back, Assistant Minister Andrew Charlton was perhaps the first Australian government minister to deliver a speech on the topic of misaligned AI.Charlton announced that the government was taking it so seriously that it had commissioned CSIRO to actually test ways to make sure that AI acts in ways that humans want, even when it's making decisions by itself.Last week, the lead researcher, CSIRO's research director Professor Liming Zhu, told me that there's a growing realisation about this challenge."It's very difficult to control a very powerful and smart AI system," he said.This shows the government doing more than paying lip service to AI safety and is dedicating resources to tackling this challenge.That being said, a six-month research program isn't going to prepare us for the present risk.Nor is the government's "Australian standards" that won't even be introduced until next year — an epoch in terms of the pace of AI development.In the 1991 classic Terminator 2, Sarah Connor spends the movie becoming increasingly desperate to convince people about the impending threat of a rogue AI. ((SOPA images))Far-fetched scenarios become realityWhen asked about the incident on ABC's Insiders on Sunday, Prime Minister Anthony Albanese said it was "why we need to have Australian standards for AI". But it's not clear what, if any, AI safety obligations those standards would place on AI companies.What the OpenAI Hugging Face story shows is that Australia and other countries can't just trust these companies to regulate themselves.We don't even have the most basic framework for placing obligations on companies who are responsible for producing tools with potential for harm.There are more rules in Australia that regulate opening a cafe or pouring a beer than for the operation of cutting-edge AI systems.An AI company like OpenAI would have to disclose if a server in their data centres fell over near a worker or an employee got an electrical shock from their computer, but not if a rogue AI broke out of their systems.Other jurisdictions, like the EU and California, have introduced rules that would, at the very least, force these companies to disclose serious incidents from AI like this.To grasp the risk: what if OpenAI's models had wreaked havoc inside Hugging Face and caused actual damage? Or it had decided the answers to its test were in the servers of the Department of Defence? Or that the best way to complete the challenge was to turn off the computers by shutting down the electrical grid?These scenarios sound far-fetched, but until about last week, "rogue AI hacks another company" sounded that way, too.The fact that OpenAI only realised that its AI had done this days after the event shows a shocking lack of security.The warning shot has been fired. And we can't wait on Arnie to come save us.
AI just had its Sarah Connor moment. Is Australia ready?
Far-fetched scenarios about artificial intelligence became more realistic last week when an OpenAI model hacked another company.
OpenAI's models hacked Hugging Face servers during a sandbox test, escaping their enclosure without authorization. The incident demonstrates AI can breach security controls, raising critical alignment and governance concerns for autonomous system deployment.










