Tesseract OCR for Hindi Typewritten Documents
Jaspreet Kaur, Vishal Goyal, M Vinay Kumar · 2021 Sixth International Conference on Image Information Processing (ICIIP) · 2021
The development of Optical Character Recognition (OCR) in Indian text is an active area of research today. The presence of a large number of characters in a set of alphabets, their complex combinations is a major challenge for the OCR designer. This OCR is giving very good results even for Hindi documents that are machine printed. But, it is not giving good results for typewriter typed Hindi documents. Even there is no OCR that is available to date which is capable of recognizing text from typewriter typed Hindi documents. This paper introduces an automated training data framework, provided only with labelled text images, thus removing the need for manually generated text. In contrast to the training images and images text as ground-truth, this approach is based on the random, rule-based generation of meaningless text in an image file and their ground-truth text file. In this paper we describe a dataset of typewriter type Hindi documents and Ground truth (GT) typewritten Hindi text images paired with their transcription. Typewriter typed documents can be incorporated into the repository of Plagiarism detection tools because text cannot be recognized by any OCR. Thus, the functionality of the existing OCR needs to be extended for recognizing typewriter typed documents.