Selecting Automatically Pre-Processing Methods to Improve OCR Performances
Quang Anh Bui, David Mollard, Salvatore Tabbone · 2017
In this paper, we propose an approach that automatically selects suitable document pre-processing algorithms to increase OCR performances. We first provide an experimental evaluations protocol to study effects of document pre-processing methods on different OCR engines for document images that have different type of distorsions. We remark that, when distortions on the document image is unknown, a pre-processing methods does not always improve but sometimes decreases the OCR performance. We conclude that the effectiveness of a pre-processing algorithm depends on the nature of the OCR and type of distorsions. In the context that distortions on the document and information about OCR system's mechanism are unknown, we propose an automatic pre-processing selection method based on a convolutional neural network with 15 layers and where the last layer contains neurons representing our different pre-processing algorithms. Experimental results show the effectiveness of our approach to improve OCR performances in a mobile-captured document images framework.