Supervised OCR Post-Correction of Historical Swedish Texts
Dana Dannélls, Simon Persson · Digital Humanities in the Nordic and Baltic Countries Publications · 2020
Current approaches for post-correction of OCR errors offer solutions that are tailored to a specific OCR system. This can be problematic if the postcorrection method was trained on a specific OCR system but have to be applied on the result of another system. Whereas OCR post-correction of historical text has received much attention lately, the question of what role does the OCR system play for the post-correction method has not been addressed. In this study we explore a dataset of 400 documents of historical Swedish text which has been OCR processed by three state-of-the-art OCR systems: Abbyy Finereader, Tesseract and Ocropus. We examine the OCR results of each system and present a supervised machine learning post-correction method that tries to approach the challenges exhibited by each system. We study the performance of our method by using three evaluation tools: PrimA, Sprakbanken evaluation tool and Frontiers Toolkit. Based on the evaluation analysis we discuss the impact each of the OCR systems has on the results of the post-correction method. We report on quantitative and qualitative results showing varying degrees of OCR postprocessing complexity that are important to consider when developing an OCR post-correction method.