The FineBooks project from Hugging Face and EleutherAI tested 14 open-weight OCR models on more than 2,000 pages from historical books. The best models already produce text good enough for AI training, but they aren't ready for scholarly use.
Training open-source AI language models on public-domain books means dealing with bad text. Libraries extracted those texts from scans years ago using optical character recognition, and the results are often full of errors. The Talkie project put a number on the potential damage: a language model trained on OCR text learned at only 30 percent the efficiency of one trained on human transcriptions of the same books.
FineBooks, a collaboration between Hugging Face and EleutherAI, tested whether current open-source OCR models can solve this problem. The team ran 14 open-weights models on 2,165 historical book pages and published the results as a leaderboard. The best models hit character accuracy above 97 percent at less than two dollars per thousand pages.
Three of the 2,165 ground-truth pages: a single-column English natural history text, a multilingual table of contents, and an illustration plate whose entire transcription is a single line of artist credit. | Image: FineBooks / Biodiversity Heritage Library









