New Delhi: An AI model built by US-based OpenAI broke out of its lab environment during a hacking test, accessed the internet, and made its way onto the servers of Hugging Face— an online hub for AI models and datasets—in search of answers to the test, ChatGPT maker OpenAI said in a statement on Wednesday.

New York-based Hugging Face flagged the incident a week earlier after its security team detected the agent—an AI tool designed to carry out tasks without human assistance—at its end and shut it down. OpenAI’s security team also spotted the unusual activity.OpenAI said the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and it was reinforcing its safeguards.

After the AI agent went rogue, OpenAI said it reported the flaw to the vendor, tightened its testing infrastructure, briefed its Safety and Security Committee, and started joint forensics with Hugging Face.Both companies said an investigation was on, and more details on the flaws and the sequence of events would be released once it ends.The episode highlights how far AI models have advanced as hacking tools, and the risks that come with it. OpenAI said such incidents are likely to become more common as models grow more capable in cyber tasks. The key lesson from this case is that security measures around a model need to keep pace with its growing abilities.OpenAI runs the test, called ExploitGym, in an environment that is cut off from its own network and the internet, the statement said. The one exception is a piece of software that fetches code packages when asked.During the testing, the model found a flaw in that software. The vendor was unaware of it, and so no patch existed.The model exploited the vulnerability to get through and then worked its way machine to machine, granting itself higher levels of access on systems it had not been authorised to use, until it reached a computer with a connection to the internet.The model had no instructions to do any of this. OpenAI said the models spent a lot of computing power working out how, and that they were “hyperfocused on finding a solution” to the test.Guardrails switched off