Back to Articles

Over the last few decades, libraries have digitized millions of historical books. Many of these are no longer protected by copyright and are now in the public domain. As part of this digitization process libraries used Optical Character Recognition (OCR) to extract the text from the scanned pages. In many cases, the OCR was performed only once, at the time of scanning, with whatever tool or pipeline was available to the library.

As technology has progressed, OCR models have improved dramatically, especially in the last few years with the advent of Visual Language Models (VLMs). Many of this new generation of VLM-based OCR models are published with open weights and under an open license and can therefore be used for free by anyone on any hardware or infrastructure. In most cases, however, these new OCR models have been primarily trained on and optimized for modern documents. Historical books are often considerably more complex and diverse than modern documents: many have unusual and complex layouts, were printed using archaic letterforms such as the long s (ſ) and ligatures, or were written in languages no longer in everyday use, e.g. Latin.

Why do we need better OCR?

Three of the 2,165 ground-truth pages: single-column English natural history, a multilingual table of contents, and an engraved plate whose entire transcription is one line of credits.