Assessment of Optical Character Recognition Techniques for Hindi Language

Matuli Das, MADHU BAJAJ · International Journal of Innovative Research in Science Engineering and Technology · 2019

Optical Character Recognition(OCR) is one of the actively researched areas in both industry and academia because of its potential applications.OCR is a widely used technology to recognize text inside images, such as scanned documents and photos.OCR technology is used to convert virtually any kind of images containing written text (typed, handwritten or printed) into machine-readable encoded text data.It is a common method of digitizing printed texts so that they can be electronically edited, searched, stored more compactly, displayed on-line, and used in machine processes such as cognitive computing, machine translation, text-to-speech and text mining.OCR engines are used to read typed (machine printed) characters.At present, there are several OCR engines available for usage.One such OCR engine is Tesseract.Tesseract is an optical character recognition engine that can be used with various operating systems.It can recognize over 100 different languages.It does various image processing operations internally before carrying out the actual OCR.Tessaract performs a reasonably good OCR, but there are certain cases where it is not good enough, which can result in a significant reduction in accuracy.The present work involves using different predefined functions and features of MATLAB and various techniques available in MATLAB for improving the efficiency and performance of Optical Character Recognition (OCR) for the Hindi language.The results of this work are subjected to comparison with an existing OCR engine.

Read the paper · More papers on PaperTik