Major publishers are increasingly restricting AI web crawlers from accessing their journalism, changing how AI companies can obtain high-quality news content for model training and related applications. The shift is not a single coordinated policy, but a sustained pattern across established outlets that gives publishers more control over whether, and on what terms, their work is used by AI systems.
The New York Times, CNN and ABC were among the outlets reported to have blocked OpenAI's GPTBot in 2023, according to the Guardian's reporting on the crawler restrictions. The Guardian subsequently adopted its own GPTBot block. Reporting and robots.txt indicators also point to restrictions at the BBC and other traditional publishers. The common issue is access to content for machine-learning uses, which is distinct from the long-standing question of whether search engines can index a page.
The scale of the trend matters. The Reuters Institute found that by late 2023, roughly 48% of leading news sites across ten countries blocked OpenAI's crawlers. Its analysis also found that legacy publishers were more likely to block than newer outlets. That does not mean every publisher has adopted the same rule, or that every bot is treated alike. It does show that unrestricted web crawling is becoming a less dependable route to premium news data.







