Development of Training Data for Optical Character Recognition using Deformed Printing Characters

Ken Kariya, Takahiro Fujishima, Lifeng Zhang · 2018

In recent years the Internet and IoT devices have broadly spread. As a result, the information which was printed on paper has been converting to digital data actively. Especially, the use of character recognition in industrial applications is increasing. When we convert printed characters to character codes, we manually input them to the computer or use a technology called Optical Character Recognition(OCR). OCR software does various image processing operations internally before doing the actual OCR. It generally does a very good job of this, but there will inevitably be cases where it isn't good enough, which can result in a significant reduction in accuracy. Noise can make the text of the image more difficult to read. uneven color of the background in the printed document can make binarization difficult. In addition, character deformation can also happen. The character deformation cause the difference between the training data for character recognition included in OCR software and the actually read printing characters. This difference cause incorrect character recognition. Therefore, we propose denoising and binarization system for OCR, and create training data using deformed printing characters. And then, we verify the character recognition accuracy when using the training data.

Read the paper · More papers on PaperTik