Handwritten Script Recognition at Line Level - A Multiple Feature Based Approach
G. G. Rajput · 2013
90 Abstract— Automatic recognition of scripts present in a multiscript document has a variety of practical and commercial applications in banks, post offices, reservation counters, libraries, etc. In a country like India, a printed/handwritten document consisting of English script and a regional script is quite common. Such a document is termed as bi-script document. In this paper, a multiple feature based approach is proposed to identify the script type from a bi-script document. Features are extracted using Gabor filters, Discrete Cosine Transform, and Wavelets of Daubechies family. The classification is done using k-nearest neighbor classifier and SVM classifier. Experiments are performed on nine popular Indian scripts along with English script. An average recognition accuracy of 96.7% is obtained using SVM classifier. motivated us to design a robust system for script identification from handwritten bi-script documents at line level. To discriminate between printed text lines in Arabic and English, three techniques are presented in (4). Firstly, an approach based on detecting the peaks in the horizontal projection profile is considered. Secondly, another approach based on the moments of the profiles using neural networks for classification is presented. Finally, approach based on classifying run length histogram using neural networks is described. Further this has been extended with Water Reservoirs to accommodate more scripts rather than triplets. Using the combination of shape, statistical and Water Reservoirs, an automatic line-wise script identification scheme from printed documents containing five most popular scripts in the world, namely Roman, Chinese, Arabic, Devnagari and Bangla has been introduced (5). This has been