A Two‐Stage Approach to Handwritten Indic Script Identification

Pawan Kumar Singh, Supratim Das, Ram Sarkar, Mita Nasipuri · 2017

Development of monolingual optical character recognition (OCR) systems for Indic scripts has already realized a noteworthy achievement. But in a multilingual country like India, it is quite obvious that a particular document may include text words written in more than one script. Therefore, in such and environment, selection of a particular OCR engine for text recognition through a script identification system is of pressing need. Script identification from handwritten document images is still a less explored research area. This chapter proposes a two-stage word-level script identification technique for eight popularly used handwritten scripts in India, namely, Bangla, Gurumukhi, Oriya, Telugu, Devanagari, Urdu, Malayalam, and Roman, using texture-based features. In the first stage, discrete wavelet transform (DWT) is applied on the input word images to extract the most representative information, whereas in the second stage, we exercise radon transform (RT) to the output of the first stage. Finally, a set of 48 statistical features is computed from each word image. The present technique is then evaluated on a database of 16,000 handwritten word images using multiple classifiers. Based on the statistical significance tests, it is concluded that a support vector machine (SVM) attains the best identification accuracy of 97.69%. Experimentation results also ensure that the present technique performs better than conventional methods.

Read the paper · More papers on PaperTik