Storia: Old OCR text cripples language model training, and FineBooks wants to fix that at scale — Warptech Lab News