If you have shipped an AI feature in the last year, you have probably felt this without naming it: real, usable training and test data is getting harder to find, even as every roadmap assumes AI will just keep improving. That is not a feeling. It is a dated constraint researchers call the data wall, and 2026 is the year it stopped being theoretical.
The Wall Is Not a Metaphor
Epoch AI's research projects that high quality, human generated public text will be largely exhausted for frontier training purposes somewhere between 2026 and 2032. Publishers are blocking AI crawlers more aggressively, licensing costs are climbing, and the open web has already been mined about as thoroughly as it can be.
The response has been fast. By some estimates, 30 to 60 percent of training tokens in recent frontier model runs are already synthetically generated, not scraped. Gartner projects synthetic data will make up roughly 75 percent of all data used in AI development by the end of 2026, up from about 1 percent in 2021. Microsoft Research Asia's SynthLLM uses graph algorithms to recombine high level concepts from existing corpora into new synthetic examples, specifically to cut dependence on scraping more of the web.






