Multifont multisize Gujarati OCR with style identification
Ami Mehta, Ashishkumar Gor · 2017
In recent years, OCR (Optical Character Recognition) technology has been applied throughout the entire spectrum of industries, revolutionizing the document management process. Document Analysis (DA) is a pre-requisite in every OCR task. Detection of style is one of the most important issues in DA. Style identification is the cause for accuracy of the text recognition system in the OCR. An approach is introduced for multifont, multisize OCR with style identification for printed Gujarati script. Here, scanned image is segmented into lines and words followed by style identification (Bold, Italic, or Normal) on segmented words. Then characters are segmented from words. For identifying characters uniquely Histogram of Oriented Gradients (HOG) and Chain Code Histogram (CCH) are used. Individual classifiers are used for recognizing bold, italic and normal style characters using SVM classifier. At the end post processing is carried out to finalize the class label. Approach is tested on various documents of different fonts (LMG-Arun, Gujarati-saral, thesis fonts) and sizes (12, 14, and 16). Total 34 consonants(to), 3 vowels, 6 modifiers and combination of modifiers with consonant are consider for recognition purpose. Overall recognition accuracy of all types of document is comparable.