Improving the accuracy of tesseract OCR engine for machine printed Hindi documents
Jaspreet Kaur, Vishal Goyal, M Vinay Kumar · AIP conference proceedings · 2022
The development of Optical Character Recognition (OCR) of the Indian text is an active area of research today. The presence of many characters in a set of alphabets, their intricate combinations, and intricate grapes are a challenge for the OCR designer. This paper focuses on improving the efficiency of the Tesseract OCR in the Hindi language. This paper introduces Google’s Tesseract Optical Character Recognition software. An optical Character Recognition (OCR) is the process of identifying and converting the text into images in pixels in a computer-friendly presentation. The present work aims to improve the accuracy of the Tesseract 4.0 OCR engine. The OCR engine will extract the text from the image. Some information has in the extracted text is calculated to provide the decision to accept. The test results in Hindi, taken from a random sample of books, show the characters and the accuracy of the words.