A year and a half ago, DeepSeek R1 caused a shock. A Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, and had reportedly done it for far less money. Markets got nervous. Billions in market value evaporated within days. Was the planned infrastructure buildout overblown?
The picture back then was murkier than the headlines suggested. In DeepSeek's own report, R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA). Later benchmarks exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board.
As recently as late June, when we started working on this issue, Z.ai's GLM-5.2 still showed the same pattern. Then came the latest Chinese open-weights models: Moonshot's Kimi K3, Alibaba's Qwen3.8-Max, and GLM-5.3.
Chinese models now sit near the top of almost every broad, demanding evaluation. They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors.
Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. According to the Wall Street Journal, Anthropic is fielding uncomfortable questions ahead of its upcoming IPO and points to its remaining lead at the top in its defense. Below that tier, the field belongs largely to open, far cheaper models from China.






