Storytelling about risk is a tricky subject. We tend to reach for metaphors we hope improve understanding. Instead, we may just be adding to the problem.
“Going rogue” is the latest iteration—a way of explaining the failings of Anthropic and OpenAI models that have been found acting well beyond the safe limits that should have been set for them. Actual rogues tend to skulk in corners and nick your smartphone. “Going rogue” in the world of cybersecurity and hacking has far more dangerous implications.
AI models also now “escape” as if prisoners attempting to find their way to the forbidden outside world. When they provide erroneous answers to even simple questions they do not malfunction, they hallucinate, a far more human sounding affliction.
This is convenient for the technology companies behind the likes of Claude and ChatGPT. If the general public believes that the agents are somehow autonomous in their actions, human responsibility for what is being built dissipates.
“The anthropomorphic language we use (e.g. ‘going rogue’) makes the challenge of control much harder than it already is,” Anil Seth, professor of cognitive and computational neuroscience at the University of Sussex, posted on Bluesky this week. He had just appeared on the BBC News program, The World This Weekend, to discuss “preparing for the future of increasingly intelligent AI.”













