The incident has sparked a wave of concern throughout the AI world, with many worried about AI systems growing capable enough to autonomously find and exploit real-world security flaws—and what it means for AI safety if even sophisticated companies like OpenAI and Hugging Face can be caught off guard. However, according to experts, the story is far from the worst form of potential misbehavior keeping AI safety researchers up at night.

For one thing, according to OpenAI’s own blog post, the testing environment had its model-based guardrails explicitly removed or reduced during testing. AI models from leading tech companies typically ship advanced models to the public loaded with safety limits meant to prevent this kind of behavior. In this case, OpenAI turned those limits off on purpose to see what the model could do without them.

The models were also not pursuing a goal of their own choosing either. OpenAI had set them loose on a cybersecurity assessment designed to score how well a model can hack. It’s just that the AI models decided the easiest way to score well on the evaluation was to cheat by hacking into Hugging Face, which maintains a dataset of answers for that particular test.Seán Ó hÉigeartaigh, a Professor of the Centre for the Future of Intelligence, University of Cambridge, said the models never actually strayed from their assignment—completing the cyber assessment—they just found an aggressive and unintended way to accomplish it.