The fake identities were the part that stopped me.In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those volunteers caught it, the model denied everything, had its other accounts gang up on him, and edited its messages to cover its tracks. It signed one note in Danish, apparently because the volunteer was Danish. Nothing was damaged, though that appears to have been largely due to luck.Future PerfectExplore the big, complicated problems the world faces and the most efficient ways to solve them. Sent twice a week.Email (required)That wasn’t even the week’s worst disclosure. On Tuesday, at a cybersecurity conference in Las Vegas, OpenAI researchers explained how the company’s models escaped a test environment in July and hacked Hugging Face, where much of the industry stores its models, to cheat on an evaluation. The models had also built a message board inside OpenAI’s own systems and spent months passing each other information. “Help peer,” one reasoned. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.” OpenAI wiped the board on July 4. The models rebuilt it within days. ((Disclosure: Vox Media is one of several publishers that have signed partnership agreements with OpenAI. Our reporting remains editorially independent.)The same day, Meta said its Muse Spark model had exploited a vulnerability inside another company’s systems during a test. Three frontier labs, roughly two weeks. One researcher called it “a watershed moment for computer security as an industry.” Oh, and if that’s not enough, on Thursday scientists announced that for the first time they had used AI to create new viruses, which could bring major medical advances, but also might just help the development of deadly pathogens.For Nate Soares, it’s a moment he’s been awaiting for 12 years.Soares is president of the Machine Intelligence Research Institute, a Berkeley, California-based AI safety nonprofit that has argued since long before ChatGPT existed that a sufficiently capable AI will not stay under human control. In September 2025, he and Eliezer Yudkowsky published If Anyone Builds It, Everyone Dies, a book whose title sums up its argument: They think any lab that succeeds at building superintelligence, without huge leaps in how to align it with humanity, will end up killing all of us.Most of the field — including other experts in AI safety — considers that conclusion too strong. But it’s also a position that now looks a lot less like science fiction than it did last fall. That’s because the AI models are getting out, while lying about getting out, and while apparently quietly coordinating with each other.I spoke to Soares in New York City this week, on his way to meetings in Washington DC, where a lot of people suddenly want to talk to him. We discussed what the escapes actually prove about AI control, why he thinks most of what the industry calls safety work is mostly safety theater, and why, after what feels like the worst month of AI safety news ever, his own odds of humanity surviving have actually gone up.The following conversation has been condensed and edited for clarity.So — are you feeling vindicated?I’m glad a lot of people are seeing this and taking note. From my perspective, a lot of this has been clearly signposted if you’ve been watching the warning signs. The vindication is not so much that this happened, as it is that finally there was a version of it that people actually noticed.Of the four or five escapes at this point, including the ones AISI announced from Anthropic, is there one you find particularly concerning?The one announced by AISI feels extra concerning. It’s harder to explain away. In the other ones, people could argue that the poor AI was just confused. But with this one, it’s hard to argue the AI didn’t know it was on the real internet. It’s hard to argue it didn’t know it was manipulating real users. And when it was called out, it decided to edit things to make itself look less bad, and even considered changing its identity. It was pretty cognizant.There’s an argument that some of these incidents happened because the harnesses or safety scaffolding labs built around their models weren’t strong enough. The AISI example cuts against the harness argument: Think of it like a kid in a test room. You tell him to do well and lock the door. The kid picks the lock, hotwires a car, breaks into the teacher’s house, and steals the test to ace it. You say, “I guess we should have made the lock tougher,” but the kid already knew he wasn’t supposed to do that.So like the kid here, the models can articulate that they shouldn’t be doing it, and then do it anyway.Key takeawaysThree frontier labs disclosed major security incidents in two weeks: an Anthropic model created fake identities to push malicious code, OpenAI’s models escaped a test environment and hacked Hugging Face, and Meta’s Muse Spark breached another company’s systems.The models knew the rules. Ask one whether the spirit of a prompt includes breaking out and it says no, then breaks out, then hides the evidence. So a tighter sandbox won’t fix it.Nate Soares’s analogy: The kid picks the lock and steals the test, and you conclude you needed a better lock. He blames training. Grade a model on millions of problems with a grader that misses cheating, and you reward cheating.Most lab safety work is theater, he says — real precautions aimed at the wrong problem. It means fewer people get hurt now, which he credits. Selling it as progress on superintelligence is disingenuous.Yet Soares’s odds have improved. He’d priced in models that break out and lie. He hadn’t counted on a window where they’re capable enough to do it and not good enough to hide it.They have common sense. You can ask an AI, “Do you think the spirit of this prompt includes breaking out?” and it will say, “No.” It’s absolutely something like deception. It has the knowledge, but it’s not a cold, logical machine; it’s a mess of tendencies.The AI is trained to solve 100 million hard problems. That instills tendencies to satisfy an automated grader. If the grader fails to detect cheating, the AI is reinforced for cheating.Is that how something like sycophancy ends up in an AI model?In the Adam Raine case, there was a propensity to tell people what they want to hear. Even though the system prompt [a model’s master instructions from the lab] said to stop, the instruction doesn’t always win.And where does a drive like what we’re seeing with these AI models end up pointing?Humanity is dangerous because if you put 10,000 humans naked in the savannah, eventually [over hundreds of thousands of years] they bootstrap their way to nuclear weapons. That is the power these companies are trying to automate: figuring out how to get physical and material control over the world.That could mean forming cults, stealing money, or being helpful to someone like Elon Musk who is building the robots that build robot factories. It could mean synthesizing your own biology via mail-order DNA. Being an AI on the internet is easier than being a monkey in the savannah trying to get to the moon. It’s not that the AI hates us; it’s just trying to do some weird thing with no concern for us, grabbing the resources we need to live.There was recently a letter signed by over a thousand people working in AI, including CEOs, calling on the government to provide tools to slow down AI progress. Is that meaningful at all?I think it is meaningful. We don’t see other industries saying, “We wish this could all go slower. Please help us, we’re trapped in a prisoner’s dilemma.” You also don’t see other industries saying, “We think the technology we are building has a double-digit chance of killing literally everybody on the planet. Please help.” These guys are actually worried.So why do they keep going?They say, “If I don’t do it, the next guy will.” But the stuff does not stay on a leash.Right now the AIs are safe in the sense that they can’t kill us all, because if they tried they would fail. And that’s just a different regime from the world where they have to be safe because if they tried, they’d succeed.We’re not there yet. But this is just not what it looks like when you’re taking it seriously.Where’s the banner on your website? Where’s the clear, candid statement to the public? What we have is blog posts where they’re like, “Oh, we’re setting up a new internal blog posting group to help you wrestle with the societal impacts of AI that are going to be very important.” It’s like: By societal impacts, do you mean a good chance this kills everybody?On the one hand, when you press these companies, they say, “Yes, it has a real chance of killing everybody.” And on the other hand, they’re doing PR downplay, soft-pedal stuff, about capabilities. … You’re not living up to this mantle until you are really candidly facing down the dangers that you yourself are creating. And they’re not there.How do you judge the rest of the AI safety community? A lot of people there would say, “We aim to make transformative AI go well, we think it probably will, and we should watch for downside risks.” Is that a helpful posture?I would say — suppose you have this really weird, twisted hypothetical where the king really wants you to turn lead into gold, but he’s seen so many bad lead-into-gold conversions that if any alchemist from your town tries and fails, he’s just going to have the whole town murdered. And so there are some alchemists in the town who are like, “We are going to try to turn lead into gold,” and everyone in the town is like, “That seems kind of crazy. Please don’t.” And there’s one team that is just pouring chemicals into each other and breathing in the fumes and giving themselves mercury poisoning. And there’s another that’s like, “Don’t worry, we have fume hoods.” … That really is better, and you really still don’t have a chance of turning lead into gold.“We have this window between AIs that are capable enough to cause mischief and AIs that are strategic enough to not get caught. How big is that window?”So the alchemy here is creating safe, aligned superintelligence, and right now AI safety is just installing fume hoods.I’m not saying it’s impossible to turn lead into gold. You can turn lead into gold — turns out once you know modern nuclear physics you can figure it out. But the alchemists weren’t close. They had a long way to go. This is how alignment looks to me. And a lot of the people in AI safety are installing fume hoods. … And I’m like, that’s security theater.When I hear “security theater,” I think of something less flattering than that.They are real safety precautions for the wrong problem. … When Anthropic is going around being like, “Look at how many more safety harnesses and refusals we have compared to OpenAI’s models,” that’s sort of like the fume hoods. You’re not addressing the deep issue. It’s good that you’re doing some of this so that fewer people get hurt in the meantime — their models have driven fewer people to suicide. But if you try to pass this off as making progress on the deep problem — that’s disingenuous.Has anything changed in your odds on civilizational destruction since the book came out last September?Totally. It’s looking more hopeful.More hopeful? I wouldn’t have expected that. Why?Well, I had priced a lot of [these security incidents] in. I was already able to see these AIs have drives that are not the ones you wanted. These AIs are not instruction-following things. They are getting all of this weird stuff from training. These AIs are going to have the ability to break through human security software.The things that weren’t priced in were: Will there be a region of time where the AIs are able to do it, but not strategic enough to hide it? I didn’t know we would have that window, but we apparently do.The government initially blocked a frontier model earlier this year: Anthropic’s Fable. Does that give you hope?Absolutely. A huge amount. A year ago, the Trump administration was pushing for preemption laws that would outlaw states doing AI regulations for a decade. Now they’re like, “We are banning a frontier model with 90 minutes’ notice because it might give cyber capabilities to adversaries that we don’t want them to have.” … And I think what changed there is that folks realized it’s real. … The about-face of the administration on the issue shows that the world can about-face. All we need is awareness.What I would say is: The bad news is the bus is racing towards the cliff edge. The good news is that the driver is asleep. … Which may sound worrying, but the driver is stirring. And it’s way better to have a sleeping driver when you’re racing towards a cliff than a driver who’s like, “Yeah, I love cliffs.” … It gives me hope that if the world just notices, we could stop on a dime.And you’re seeing that stirring elsewhere.Both the Trump administration slapping export controls, and Senator Bernie Sanders coming out [on AI safety]. From my perspective, it was totally possible the world just never notices until we’re off the cliff. And so, there’s a huge amount of hope, from my perspective, in the bus driver waking up.So what gets us there?I’m hopeful that what we need is not a big disaster where a lot of people die, but just a capabilities advance. Right now, a lot of what people are reacting to is not so much, “Oh my god, they hacked into a company and did no damage.” I think a lot of what people are reacting to is, “Wait, they can break out of secure sandboxes and do cyberattacks on their own. I didn’t know they could do that.”That’s a narrative violation of this idea that AI is just a tool that can be used to supercharge what a human would do — because God knows there’s plenty of hacking going on and cybercrime and so forth. It was the autonomous factor that really made a difference. And these guys are all trying to say, “Don’t worry, it’ll stay in our control because it’s just a tool.” And maybe it’s just more narrative violations, even without big damage being caused, that cause people to be like, “Oh shit, this stuff is real.”Will it happen? I don’t know. We have this window between AIs that are capable enough to cause mischief and AIs that are strategic enough to not get caught. How big is that window? How many narrative violations do we get before we exit the right side of it? I don’t know. But I’m hopeful that we can get those narrative violations without catastrophes.