A New Method for Identification of Partially Similar Indian Scripts

Rajiv Kapoor, Amit Dhamija · 2012

In this paper, the texture symmetry/non-symmetry factor has been exploited to identify the Indian scripts. Biwavelants have been proposed to obtain the script texture using third order cumulant and bispectra. As the Indian scripts are partially similar to each other, in order to identify them, the samples must include more number of dissimilar characters. The features of individual lines are added repeatedly to enhance the dissimilarity until it reaches to a saturation level which in turn is used to compute a confidence factor i.e. amount of confidence attained in identifying a particular script sample. This variation in confidence factor also gives an estimate of the optimum sample size (number of lines) required for expected results. Cumulants are sensitive to the script curvatures and therefore are most suitable for the partially similar Indian scripts. The double discrete Fourier transform of third order cumulant gives bispectra which estimates the factor of symmetry/non-symmetry in terms of the quadratically coupled frequencies. The envelope of bispectra (biwavelant) obtained using wavelet (db8) provides an accurate behavior of the script texture; which along with Newton-Raphson technique is used to classify the Indian scripts. Various classifiers have been tested for script identification and out of them SVM gives the best results. The method successfully identified the 8 Indian scripts like Devanagari, Urdu, Gujarati, Telugu, Assamese, Gurmukhi, Kannada, and Bangla with desired accuracy.

Read the paper · More papers on PaperTik