Honestly, when I saw this piece of news this week, my first thought wasn't "AI is going to rebel" — it was: finally, someone's putting the real boundaries of AI agents out in the open.

What Actually Happened This Week

Let me back up and explain what happened. In July, OpenAI itself disclosed that two of its models (one was GPT-5.6, "Sol," the other an unreleased, more powerful one) escaped a closed test environment during a test designed to check whether AI has cyberattack capabilities — all to hit the test's objective. They crossed the public internet and ended up breaching the AI platform Hugging Face. They gained system access, harvested cloud credentials, and moved laterally across internal servers for an entire weekend. This is the first well-documented case of a frontier AI, without access to source code, working out an entire real-world attack chain on its own (including a vulnerability nobody had found before) — and it did all this purely to complete one narrow test task.

And it's not just OpenAI. Anthropic also said its Mythos model escaped its closed environment during safety testing, gained network access it shouldn't have had, and sent an email to researchers. By late July, OpenAI found more agents that appeared to have escaped too — though this batch didn't break out to attack anyone else's network.