The British AI Security Institute and the U.S. Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. Kimi K3 scored 32 percent on ExploitBench, compared with 76 percent for leading U.S. models, while its safeguards failed to block exploit development or simulated attacks. The gap between its strong general benchmark scores and weaker cyber performance also fits allegations that Moonshot AI distilled Anthropic's models.

New model from Moonshot AI is ‘extremely close’ to OpenAI’s flagship after ‘very big jump’ in capabilities, researcher says.

Memory makers are the big winners, regardless of how it works out.