Treating OCR Output as a Language (TOOL) – Improving OCR Output with Seq2Seq Translation

Thomas Asselborn, Magnus Bender, Ralf Möller, Sylvia Melzer · Annals of Computer Science and Information Systems · 2025

Optical Character Recognition (OCR) systems are frequently used to digitise text, but often produce noisy results, especially with historical, poor-quality or multilingual data.Despite advances in OCR technology, post-processing remains a significant bottleneck.We propose TOOL (Treating OCR Output as a Language), a new approach that understands OCR correction as a machine translation task.By treating noisy OCR text as a language in its own right, TOOL employs sequenceto-sequence models like Marian to translate it into clean, standardised text.This method is scalable, model-independent and language-flexible.We demonstrate this approach by translating "OCR German" to Standard German from around 1871 to the present day, improving accuracy at the token level by using matched training pairs of OCR output and base text.

Read the paper · More papers on PaperTik