The pitch from data broker ISBNdb is blunt: the world’s best AI training data is sitting on a shelf. The catch is that getting at it means slicing the spine off millions of books, then never admitting which lab paid for the job.
AI has a pollution problem of its own making. So much of the web is now machine-written that models risk feeding on their own exhaust. One company thinks the antidote is sitting on a shelf: old, printed books, published before the chatbots arrived.
404 Media reported that ISBNdb now sources physical books in bulk for AI labs to scan into training data. The firm calls itself the world’s largest book database. “The world’s best AI training data is sitting on a shelf,” its site says. Books, it adds, are “dense, edited, authoritative.”
The value is in the date. ISBNdb argues that books printed before 2022 predate the large language model era, so they cannot contain AI-generated text. That matters because of model collapse. It names the documented decline that sets in when models train on the synthetic output of earlier models. Each generation ends up a little worse than the last.
The 💜 of EU techThe latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!The open web offers no such guarantee. A fast-growing share of online text is now machine-made, part of the same slop flood the labs helped create. A pre-2022 print run, by contrast, is a fixed, human-authored record that nobody can quietly rewrite.









