The NSA, CISA and FBI say that China-based AI companies are using the distillation process against US frontier models. “Likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges/requests from U.S. frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024,” they say.

The result doesn’t simply improve China’s AI models, it also threatens the existing US technology leadership. Between late 2024 and mid 2025, DeepSeek distilled training data and capabilities from Claude, Gemini, GPT-4 and GPT-5, and Grok 4 to train its R1 and V3 models.

The knowledge distilled included, but was not limited to, API rule-driven tasks, agentic functions, Q&A optimization, supervised fine-tuning optimization, and creative and occupational writing optimization.

Moonshot undertook a similar large scale distillation, including the extraction of significant Claude Fable 5 data to improve its Kimi-K3 model; and GPT-4o data to improve its Kimi-K2 model.

Such distillation was repeated across all the named Chinese AI systems and is chronicled in detail within the agencies’ report. The tactics, techniques, and procedures (TTPs) used by the Chinese organizations in their distillation are mapped to the MITRE ATLAS framework providing the TTP Title, its ID, and a description.