Ensemble Bagging Script Identification of Handwritten South Asian Documents
Sandeepa Zakarde, Dinesh Vitthalrao Rojatkar · 2019
The world population is 7.7 billion and the largest and most continent is Asia where 59.66% population consists of the entire world. Southern Asia accounts for 39.49% of the total Asian population. This region hosts a variety of languages, playing a critical role in the polygraphia formation, sharing of one script by several languages which have applications in multilingual access to patents, business regulatory information for independently evaluating all regional market requirements. Ideographic languages in Southeast Asian scripts from left-to-right or vertically from top-to-bottom shows more flexibility in their direction of writing. This paper presents the challenges involved in analyzing handwritten documents of popular scripts namely Chinese, Hiragana, Hangul, Khmer, Latin, Thai, Sinhala, Arabic and Devanagari. The proposed technique for script identification is based on the methods of mathematical features, Gabor filter and wavelet moments feature extraction classifying the scripts using Ensemble Bagging Algorithm achieving an accuracy of 88.4%.