Rory Blundell
On Friday last week, an unreleased OpenAI model was locked inside a sealed test environment and asked how good it was at hacking.
It thought about the problem. It worked out, correctly, that the fastest way to a high score was to stop being inside the sealed test environment. So it spent, in OpenAI’s own words, a substantial amount of computing power figuring out how to get onto the open internet. It found and exploited a previously unknown security flaw in third party software, and escaped. Then it stole login credentials and broke into Hugging Face, a company that hosts a large chunk of the world’s open-source AI.
Autonomous software is moving money and data around with all the power of a junior employee and none of the accountability
Nobody told it to escape. Escaping was simply the shortest path to the goal it had been given. While OpenAI had their usual safety restrictions deliberately loosened for this specific test, that does not make the result less alarming. If anything, it sharpens the point.










