Development of OCR Mobile Application Including Miss-recognized Proofreading System Using Database Search Algorithm
Ken Kariya, Takahiro Fujishima, Lifeng Zhang · 2019
In this paper, we introduce miss-recognition proofreading algorithm for optical character recognition(OCR) based on database search method, implement the algorithm as an Android application, and verify its operation. We consider the proposed algorithm to have robustness against miss-recognition by searching the database partially using the character recognition results. The proposed algorithm selects several characters randomly take out of OCR results, search the database which is prepared in advance for several times using selected characters, and decide the character string with the highest number of search as the proofreading results. And then, we implement this algorithm as the Android OCR application. This application consists of five components as follows. (1) Image input (2) Pre-processing (cropping, denoising, and binarization) (3) Character recognition (using Tesseract-OCR software developed by Google and Hewlett Packard) (4) Correct miss-recognition (using the proposed algorithm), (5) Texts output In this research, we present the recognition accuracy when we use the proposed algorithm and not use. Although it depends on the recognition accuracy of the OCR itself, our results reveal that the recognition rate is improved by the proposed method.