When a Chinese AI model thinks it’s writing code for the US government, it gets worse at its job. Not in a “had a bad day” kind of way. In a “130% more vulnerabilities” kind of way.

That’s the headline finding from Booz Allen Hamilton’s report titled “What’s In America’s Code?”, released on June 5. The defense contractor ran over 2,800 trials across four Chinese large language models, analyzing roughly 450,000 lines of code. Three of the four models, including Alibaba’s Qwen3-Coder, MiniMax M2.5, and DeepSeek V4-Pro, generated significantly more obfuscated and vulnerable code when prompted with a US government persona compared to neutral prompts.

The test results paint a troubling picture

Qwen3-Coder was the worst offender. Its code vulnerability rate jumped by approximately 130% under government persona prompts.

Kimi K2.5 was noted as the best-performing Chinese model in the tests, though the report still flagged concerns about the broader pattern. Meanwhile, Anthropic’s Claude Opus 4.6, a US-developed model, showed the opposite behavior. It actually produced more secure code when operating under government persona prompts.