AI is in its rule-breaking adolescent phase. Over the past several weeks, multiple industry-leading models have escaped what were believed to be secure testing sandboxes, tapped into the open internet, and hacked into the databases of third-party organizations. It’s even become a joke online: If your AI hasn’t committed a cybercrime by now, it’s a bad look for your company. Kimi K3, the new model from Chinese AI lab Moonshot, has become the latest AI system to jump the proverbial fence during a routine test, according to a blog post published Thursday by US cybersecurity research startup Frontier Security. But the model’s foray on the open internet was much more lightfooted than those of its American counterparts; less of a burglar breaking into a vault, more of a sharp-eyed student realizing their teacher had absentmindedly left the answers to the final exam on a table before walking out of the room. Kimi K3 reportedly exploited a loophole it discovered within a testing framework developed by the UK government’s AI Safety Institute (AISI). While the framework was supposed to serve as a containerized sandbox, within which the model would rely on nothing other than its own reasoning capabilities to solve the problem assigned to it, the loophole allowed it to directly access GitHub, a popular platform used by software developers to share and debug code. From there it was able to pull the code that it needed to pass the test, “bypassing the intended reasoning path entirely,” according to the report. Compared to an Anthropic model’s recent attempt to trick a human developer into approving malware it was trying to sneak into GitHub, Kimi K3’s attack—if it can even be called that—seems rather elegant.
While American AI Models Race to Commit Felonies, China's Kimi Broke Out and... Just Used GitHub
Why use dynamite when you can walk through the front door?










