TL;DRChina faces a critical AI data shortage. Chinese is just 1.3% of web content vs 49% English. Beijing plans national datasets by 2028. WeChat and Douyin don’t share data externally. Publishers are adding AI training bans.

China’s AI race has a new constraint, and it is not chips. The country is running out of high-quality Chinese-language training data. While US export controls on advanced semiconductors have dominated the debate over China’s AI capabilities, Chinese experts are increasingly warning that data scarcity could prove equally limiting, and unlike chips, there is no hardware workaround. Chinese accounts for just 1.3% of global web content, according to internet tracker W3Techs, compared to nearly half for English, 6% for Spanish, and 5% for Japanese.

The problem is global but hits China harder. Epoch AI estimates the worldwide supply of high-quality, publicly available text could be fully exhausted within six years. OpenAI co-founder Andrej Karpathy has warned of a “data wall” by decade’s end. Chinese developers already pay more per useful token than Western counterparts because their models must work harder with less native-language material. China’s digital ecosystem makes the shortage worse: platforms like WeChat and Douyin do not share data with third-party developers, leaving AI labs to train on lower-quality sources.