The dominance of AI-generated content on the Internet is a major concern, and these percentages will steadily rise as AI tools improve and gain wider adoption

The worst fears of computer science experts about the possible deterioration of original writing, content quality and diverse writing styles have come true, with researchers finding that the share of AI-written content has grown sharply since Large Language Models (LLMs) arrived in 2022.AI experts have expressed concerns that this could affect future AI systems, as their training data would contain far more synthetic material, potentially degrading its quality over time. According to the Pew Research Centre’s analysis of nearly half-a-million English language webpages, about 10 per cent of the sample (10,000 webpages collected in July 2026) reflects signs of AI fingerprints. It points out that the share of AI-generated content could be much higher if you exclude previously existing texts. When researchers filtered the data to look strictly at content published after the public release of ChatGPT, they found that over one-third of those recently published pages were likely written, or substantially edited by AI.AI detectionThe report, however, maintains that AI detection models aren’t perfect. They sometimes misclassify human-written documents as showing signs of AI authorship and vice-versa. Another study, The Impact of AI-Generated Text on the Internet, by researchers from Imperial College London, Internet Archive and Stanford University, found that by mid-2025, roughly 35 per cent of the newly published websites were classified as AI-generated or AI-assisted, up from zero before ChatGPT’s launch in late 2022.PJ Narayanan, the former Director of International Institute of Information Technology-Hyderabad (IIITH), felt that the dominance of AI-generated content on the Internet is a major concern, and these percentages will steadily rise as AI tools improve and gain wider adoption.“While humans can only create content at a fixed, limited pace, individuals and companies utilising powerful AI systems can generate data at an exceptionally high rate, practically without limits,” he said. Citing the Pew study, he said that about 35 per cent of the recent content contained significant AI elements, a figure that continues to grow.Degrading Quality“One direct consequence will be on future AI systems, as their training data will contain far more AI-generated material, potentially degrading their quality over time. More subtly, people interacting with online content may unconsciously alter their writing styles to mimic AI,” he said.AI analyst Kashyap Kompella said that as AI-generated content grows, “the problem is no longer access to text; it is access to original, grounded, high-quality human knowledge”. He added: “The next constraint is not data quantity, but data provenance [the chronological record of information origins]. AI systems will increasingly need to know whether information came from real human activity, an expert, a proprietary workflow, or another AI system.”He said this would raise the value of private and proprietary datasets. Google’s proposed purchase of Spirit Airlines’ internal business data for $10 million is a useful signal. Real workplace communications, decisions and workflows are becoming valuable training assets because they cannot be easily recreated from the public web.“We are moving from a data-labelling economy to a data-creation economy. The first AI data industry paid people to annotate existing information. Going forward, we will pay experts to create new problems, examples, scenarios, workflows and evaluations specifically for AI systems,” said Kompella.Published on August 25, 2026