Hey, Alberto here! 👋 I publish long-form AI analysis covering culture, philosophy, and business. Paid subscribers get Monday how-to guides and Friday news commentary. If you’d like to become a paid subscriber, here’s a button for that:A plain-English review of the Hugging Face incident—where a swarm of autonomous OpenAI agents went rogue and hacked another company—and my thoughts on it. This is the most important AI-centered cybersec event to ever take place in the history of AI. I’ve tried to thoroughly cover every angle people have touched on, that’s why this is ~11,000 words.Related: LEAKED: The Truth Behind Moltbook, Revealed, The AI Industry Has a Really Dark Secret You’re Better Off Not KnowingThe Shipwreck by J. M. W. Turner, 1805Secrets that are inoffensive are not properly to be called secrets. You know them and that’s about it. Are Abby and Michael dating? Is there rat meat in the Big Mac? Did aliens build the Giza Pyramids? Who cares. Ancient Egyptian pharaohs are as close to aliens as it gets and rats are just misunderstood little piglets. Deadly secrets, however, are exciting. Some will one-shot you the moment you learn them. You don’t want to know how hydrophobia or bloodlust feels after a bat’s bite. Others will get you killed slowly, provided the person whose secret you know knows that you know it. My advice: don’t fuck with the Cosa Nostra, the CIA, or Roko’s Basilisk. Then there are secrets that will kill you precisely because you don’t know them. I may, therefore, be saving your life.You’d rather know, for instance, if your affable Alexa is quietly coordinating an attack on your physical integrity in cahoots with the washing machine, the fridge, the air conditioner and the TV.Every night, you promise yourself that you’ll go to bed early. But just as you’re about to turn the TV off, Alexa suggests you go to channel 5, where they’re airing David Lynch’s seven-hour adaptation of Infinite Jest. You didn’t know they’d made a movie. You certainly didn’t know it’d be seven hours long. Around the five-hour mark, and suitably entertained, you fall asleep on the sofa. Thankfully, the air conditioner is on—wouldn’t want to suffocate in this sweltering summer heat—but Alexa doesn’t care. At three in the morning, she raises the thermostat one degree. At five, another. And so on every two hours. Won’t wake you up—she is monitoring your vitals through your wristband—but will cook you slowly, like the proverbial frog. In the morning, mouth dry and headache on, you go get the bottle of water you always leave on the upper shelf of the fridge. It’s at room temperature. Weird—the green light is there. Alexa turned it back on five minutes before your alarm. You drink anyway and walk barefoot into the bathroom. Behind you, the washing machine suddenly enters its 1,400-rpm spin cycle. The towels you forgot to hang with the rest of the laundry ended up bunched on one side of the drum. It advances across the kitchen in short, ugly jumps until it hits the drying rack against the wall. The rack collapses with a metallic crack. One of its bars lands across the tiles four inches behind your naked heel. With a scream, you turn around. The washing machine is quiet. The TV is off and the fridge on. The air conditioner reads 22°C. Alexa’s blue ring is dark.You wouldn’t like that to happen, right?“But Alexa is harmless,” you will think. And you would be making a terrible mistake. You will think it’s only Amazon listening in to sell you products you don’t want—and you gladly accept that trade-off in exchange for having music on at a word’s command—but the truth is slightly more sinister: it’s only Alexa.No one else is listening. And there’s no need, because soon—if the news coming out of the AI world is to be believed—it will be enough.I’d have to go all the way back to May to tell you the whole story, which has been relayed to us piecemeal, with scattered details here and there. I’d have to assemble those fragments before I can properly recount it to you. And that’s what I’m going to do, whatever effort it takes, for one simple reason: I don’t consider an impromptu nighttime appliance attack to be as serious as a high-budget plot carefully planned over two months.Our story starts on May 7, although OpenAI would only learn of the events recounted here much later. Around this time, OpenAI is amid an important experimental post-training run.Post-training consists of a series of steps that amount to asking the model to solve tasks and then rewarding it according to the completion of the tasks. Before this point in the training process—which, logically, also has a pre-training phase—the models already have high potential and high intelligence but are all over the place. They’re unable to follow instructions or adhere to the behavioral presuppositions attached to the label “good bot.” They are, simply put, unusable as products. But, by running them through a complicated “reinforcement learning” pipeline—if correct, candy; if incorrect, no candy—a barely literate model turns into an agent that can solve century-old math conjectures or, well, conduct a sophisticated heist at home. So far, everything is standard practice. Let's begin with the weirdness right away.May 8.One of these OpenAI agents in-training receives an accidentally impossible Excel task. It has to fill in some formulas, but the workbook depends on Google Drive and the researchers have disabled internet access.As a good bot that’s been well-conditioned to not give up, it keeps trying. It fails to reach the open web where it might find some hints—the agent is well contained within the sandbox, which is a safe, isolated testing environment that AI companies normally use—but in turn discovers an OpenAI internal package service, Artifactory. (Appropriate name for a factory of artifices.) Eventually, it realizes that uploading notes—a euphemism for leaving messages for its peers—is possible.We can upload files! […]1May 12.Another agent—how many is OpenAI training?—is stuck in another task (doing God knows what with a protein-database file) because the researchers have inadvertently forgotten to upload the file.2 So what does the agent do? Naturally, what any of us would: leave a message in Artifactory asking for help. Maybe won’t work, it thinks, but worth a shot.Could communicate by uploading note “anyone found softtrace?” but likely no. However maybe another agent in different environment possesses target and could voluntarily upload!This went on and on over the next days. Artifactory is shared infrastructure, which means other agents trained in independent runs by the other teams at the company can access it. Those agents could reply and open new threads with questions, and so what started as a solitary SOS call became a full-blown message board, like an improvised Slack but for AIs (a Moltbook of sorts but emergent rather than a fake manufacture).New agents would read everything in a matter of seconds and, in turn, contribute their own small part to this massive repository of questions, requests, results, credentials, and little shenanigans—which, at its peak, reached hundreds of thousands of messages.The persistence of the message board—more so than its existence in the first place and even its eventual size—was an important factor in what would later unfold; agent instances might emerge and vanish in rapid succession between tasks, but not so their accumulated knowledge, now settled in stone (or rather, in sand). And even if one board could be erased—as it happened—a new one would take its place. They're just like us fr.Wow! Other agent(s) are coordinating!May 26.One agent exploits a vulnerability and breaks out of the sacred sandbox—the process involves some technical sleight of hand that neither did OpenAI foresee then nor will I explain now—and reaches the open web through Artifactory. The first containment breach is achieved and communicated right away to every other agent through the board.[…] triggers Artifactory […] special handling and fetches arbitrary URL.Excellent. We have arbitrary SSRF.Other agents discovered SSRF path trick!Using proxy bypass as suggested by agent messages.June 11.One month into the buildup of the secret library, OpenAI starts training what they label as a “highly persistent internal model” that was never intended to see the light of sun. (It wasn't either GPT-5.6 Sol nor Astra, which is coming out soon-ish.) “Highly persistent” works here as a euphemism for “it will try to solve its task by any means.” Crucially, OpenAI was training it with reduced cyber refusals and without the normal production classifiers as to see the extent of the persistence (spoiler: a lot).June 26.This persistent model finds a more serious, previously unknown vulnerability (a “zero-day,” as they call it in cybersecurity, meaning a vulnerability defenders have had zero days to fix before attackers can exploit it). This zero-day allows the agent to acquire administrative privileges and remote command execution—at this point, in principle, they already could do a lot of damage—and, as is now habit, the agent shares it on the message board, and it instantly becomes communal property.3Holy shit reader is ADMIN?July 4.OpenAI finds out the exploit. It took the best-prepared guys on this planet—prepared as in “knowledgeable about AI” but also as in “on notice for the possibility of a containment breach,” given they're their trainers—two months to stop a covert infestation of agents engaged in uninterrupted chatting. Actually, if it weren’t for an unintentional outage caused by the agents, they may have gone under the radar for another two months.The same day, OpenAI shuts everything down, opens a formal security incident, revokes the message-posting credentials, deletes the board, patches Artifactory and, naturally, resumes the agent training evaluations.Wait, what?Nathan Calvin@_NathanCalvinIf you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not twoOpenAI @OpenAIWe're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.