Major publishers have increasingly restricted AI crawler access to their websites, turning a once largely technical robots.txt decision into a strategic question about training data, licensing, and control of digital content. The shift does not mean publishers have adopted one common position on artificial intelligence. It does mean that AI companies can no longer assume that publicly reachable web content is available for automated collection.
One early, high-profile example came in September 2023, when Guardian News & Media blocked OpenAI's GPTBot. The Guardian's report on its GPTBot block described the decision as part of a wider response by news organizations to AI systems trawling their work. Contemporaneous reporting also identified restrictions by other major outlets, including The New York Times, CNN, and ABC. Reuters Institute analysis from that period found that a significant share of leading publishers had blocked AI crawlers, while major publishers had not reversed blocks affecting OpenAI or Google crawlers.
The most important development is therefore not a single publisher's policy. It is the emergence of a more deliberate publisher-led framework for deciding whether, and on what terms, AI systems can access journalism and other professionally produced material.







