Improvement Optical Character Recognition for Structured Documents using Generative Adversarial Networks

José David Bermúdez Castro, Smith Washington Arauco Canchumuni, Cristian Muñoz, Fábio Corrêa Cordeiro, Antônio Marcelo Azevedo Alexandre, Marco Aurélio C. Pacheco · 2021

The Optical Character Recognition (OCR) models have improved a lot in the last years thanks to the advance of machine learning techniques. However, there is still a great challenge when images are too distorted, such as old scanned documents, where the characters are blurred or/and with irregular backgrounds. Document digitization is a crucial task in many enterprises to get all their information available. For industries with long life cycles, such as energy and mining, old reports are relevant. This paper proposes a methodology to improve the performance of OCR algorithms. It uses conditional generative adversarial networks (GANs) for improving the quality of rendered text in images, which maximizes the performance of the OCR algorithms. We created a synthetic dataset, which emulates most of the problems presents on scanned documents, to train the network. Besides synthetic data, we tested our methodology on real images collected from USA journals and Brasilian thesis. The results showed relevant improvements, in particular when evaluated images were too distorted. Besides, we performed experiments with text in English and Portuguese, demonstrating the efficiency of our work in different languages. The proposed methodology overcomes most of the problems that impact the quality of Tesseract, the OCR algorithm used in this work. Our results demonstrated the capability of the proposed OCR system over other algorithms in low-quality documents. Metrics of errors showed a decrease of up to 50% regarding the best results without using our methodology.

Read the paper · More papers on PaperTik