A Study to Recognize Printed Gujarati Characters Using Tesseract OCR
Milind Kumar Audichya · International Journal for Research in Applied Science and Engineering Technology · 2017
Optical Character Recognition (OCR) is a widely-known technique to recognize theprinted text using computer with the help of various peripheral devices.Research works for OCR of many languages scripts is in process and many languages are still far away.Gujarati script is one of the least focused script in research area of OCR as compared to other scripts.A wellknown Open Source OCR Engine called Tesseract which is already used for the recognition of numerous scripts, can be used to recognize printed Gujarati characters from digital images.This paper is trying to enlighten the use of Tesseract to recognize Gujarati characters with the help of already available trained data for Gujarati Script.