The Application of Deep Convolutional Denoising Autoencoder for Optical Character Recognition Preprocessing

Christopher Wiraatmaja, Kartika Gunadi, Iwan Njoto Sandjaja · 2017

The process of converting physical documents into digital texts generally requires a scanner tool to obtain high-quality document images. These high quality images will be read by OCR software to get digital text results. The weakness of this method is that OCR software requires a high quality document with low blur noise and no parallax in the image to have high accuracy. We developed an application to increase the document image quality with the help of Deep Convolutional Denoising Autoencoder, afterwards read by the OCR application. The final product of this program is a digital text converted from a document image which has been taken from a smartphone. There is an increase in accuracy using this application by 26.68% in a blurred image compare to standard Tesseract OCR and outperformed Simple OCR in average accuracy testing.

Read the paper · More papers on PaperTik