Feature extraction from printed Persian sub-words using Haar wavelet transform

Samira Nasrollahi Dizajyekan, Afshin Ebrahimi · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2011

This article presents a novel set of shape descriptors which are especially well-suited for the recognition of printed Persian sub-words based on their holistic shapes. The descriptor set is derived from the wavelet transform of a sub-word's image. The proposed algorithm is used to extract features from 87804 sub-words of 4 fonts and 3 sizes. To evaluate the feature extraction results, this algorithm was used to obtain recognition rate for a set of sub-words in a printed Persian text document. Features of an unknown sub-word are extracted and compared with all sub-words features in the dictionary and the desired sub-word is identified. In this stage to increase the recognition rate, dot features of the unknown sub-word are used as the second feature and compared with dot codes of 10 last sub-words in before stage and the sub-word with maximum similarity is extracted as correct recognized sub-word.

Read the paper · More papers on PaperTik