Recognition Of Pre-Segmented Characters In Printed Bilingual Gujarati-English Text
Shailesh Chaudhari, Ravi M. Gulati · Zenodo (CERN European Organization for Nuclear Research) · 2017
Character recognition is always a challenging task for OCR of multilingual documents. Most of the script identification work is done on paragraph/block, line, and word level, but no significant work is done at character level. Generally English language is interspersed with most of the regional documents in India. In this paper a system is proposed to address such a diverse situation at character level in the context of Gujarati and English language used in Gujarat state, a western part of India. Experiments are carried out on pre-segmented multi font and multi-sized characters of both Gujarati and English. In this work features are extracted using Discrete Cosine Transform (DCT). The input character is identified and recognized as either English or Gujarati character. As it as multi-class problem, multi-class SVM has been used for classification and recognition. From the experiment promising results have been reported. To assess the global recognition accuracy 10-fold cross validation is used. The average recognition rates obtained are 98.53, 89.62%, 89.57%, and 91.56% using kNN, SVM_Linear, SVM_Pol3nomial, and SVM_RBF classifiers respectively.