Zone-based hybrid feature extraction algorithm for handwritten numeral recognition of two popular Indian scripts
S. V. Rajashekararadhya, P Vanaja Ranjan · 2009
India is a multi-lingual multi-script country, where eighteen official scripts are accepted and there are over hundred regional languages. In this paper we propose a zone-based hybrid feature extraction algorithm scheme towards the recognition of off-line handwritten numerals of two popular south-Indian scripts. The character centroid is computed and the character/numeral image (50×50) is further divided in to 25 equal zones (10×10). An average distance from the character centroid to the pixels present in the zone column is computed. This procedure is sequentially repeated for all the zone/grid/box columns present in the zone (10 features). There could be some zone column having empty foreground pixels. Hence feature value of such zone column in the feature vector is zero. This procedure is sequentially repeated for the entire zone present in the numeral image (250 features). Similarly we extract zone centroid coordinates as features. The numeral image is divided into 50 equal zones (5×10). The zone centroid is computed. This procedure is sequentially repeated for the entire zone present in the numeral image (100 features). There could be some zone having empty foreground pixels. Hence feature value of such zone in the feature vector is zero. Finally 350 such features are extracted for classification and recognition. The nearest neighbor and the support vector machine classifiers are used for subsequent classification and recognition purpose. We obtained 97.75 % and 93.9 % of recognition rate for Kannada and Tamil numerals respectively using nearest neighbor classifier. We obtained 98.2 % and 94.9 % of recognition rate for Kannada and Tamil numerals respectively using support vector machine classifier.