An Exhaustive Font and Size Invariant Classification Scheme for OCR of Devanagari Characters
Manoj Kumar Gupta, Chellapilla Vasantha Lakshmi, Madasu Hanmandlu, C. Patvardhan · International Journal on Natural Language Computing · 2015
Main challenge in any Optical Character Recognition (OCR) system is to deal with multiple fonts and sizes.In OCR of Indian languages, one also has to deal with a huge number of conjunct characters whose shape changes drastically with fonts.Separating the conjunct characters into its constituent symbols leads to segmentation errors.The proposed approach handles both the above listed problems in the context of Devanagari script.An attempt is made to identify all possible connected symbols of Devanagari (could be a consonant, vowel, half consonant or conjunct consonant henceforward shall be referred as a basic symbol) in the middle zone without segmenting the conjunct characters.On observing 469580 words from a variety of sources in our study, it is found that only 345 symbols are used more frequently in the middle zone and cover 99.97% of the text.They are then classified into 16 different classes on the basis of structural properties which are invariant across fonts and sizes.To validate the proposed classification scheme, results are presented on 25 fonts and three sizes.