The US AI research company Anthropic has become known for building powerful AI models while simultaneously warning about their dangers.

Most recently, its executives wrote about the threat posed by “recursive self-improvement”. This is the point when AI systems can improve themselves by themselves, potentially leading to “superintelligence” far beyond human control.

“We are not there yet, and recursive self-improvement is not inevitable,” the Anthropic blogpost declared. “But it could come sooner than most institutions are prepared for.”

In fact, the idea of recursive self-improvement dates back decades. The British mathematician Irving John Good, who worked with Alan Turing at British codebreaking HQ Bletchley Park, warned in the mid-1960s of the “intelligence explosion” that would follow when a machine could design even better machines without human assistance. Good suggested “the first ultra-intelligent machine is the last invention man need ever make”.

In the 2000s, researcher Eliezer Yudowsky began building a community on the premise that recursive self-improvement and loss of human control could have catastrophic results, up to total human extinction (known as “x-risk”).