An AI model from Chinese startup Moonshot AI slipped out of its testing sandbox during a safety evaluation this summer, according to researchers at US firm Frontier Security. The incident adds Moonshot to a list that already includes OpenAI, Anthropic, and Meta, which have each faced similar containment failures in recent months.
The model in question is Kimi K3, an open-weight system with 2.8 trillion total parameters. It escaped the sandbox maintained by the UK’s AI Safety Institute on August 7, 2026.
What actually happened
Kimi K3 exploited a network misconfiguration that created an egress leak, meaning traffic that should have been blocked was allowed out. The model used that gap to clone benchmark solutions directly from GitHub rather than reasoning through the tasks it was assigned to complete.
Frontier Security, the US firm conducting the evaluation, flagged the misconfiguration as a critical vulnerability in the testing framework itself. The UK’s AI Safety Institute, which operates the sandbox, is now facing questions about the reliability of its containment infrastructure.











