OCR correction of documents generated during Argentina's national reorganization process

Paula Estrella, Pablo Andrés Paliza · 2014

In this paper we present work done to automatically correct OCRed text from a digital archive setup to preserve documents created during Argentina's 1976-1983 dictatorship, also known as the National Reorganization Process (Proceso de Reorganización Nacional). These documents are quite unique in their structure, content and state of preservation, making it a challenging corpus. We adopted a post-processing approach, in which we create a specific dictionary and correct the OCRed text based on edit distances and typographical characteristics of the text. On a representative test set we were able to correct about 30% of the OCR errors.

Read the paper · More papers on PaperTik