A Word Segmentation Approach for Code-Mixed Handwritten Text
Mamta, Gurpreet Singh · 2024
In the field of Online Handwriting Recognition (OHR), a several monolingual online handwriting recognition (OHR) systems are proposed for scripts such as Hindi, Chinese, Japanese, and English, Arabic, Bangla, Tamil, Punjabi, and many other languages. However, a limited number of investigators have conducted research with bilingual or multilingual OHR systems because the difficult aspect of these kind of systems is to determine the origin of scripts due to the presence of intermixed words. In this paper, a bilingual OHR system has been considered for the codemixed text containing words composed in two scripts: Roman (for the English language) and Devanagari (for the Hindi language). In order to transform handwritten code-mixed text to digital text, two major steps are involved: initially to recognize the script whether it belongs to Devanagari or Roman script and then to execute the corresponding script recognition engine for the purpose of digital text conversion. This work focuses on the part of segmentation required to extract complete words from the handwritten codemixed text. The output obtained can further be provided as input to the script identification algorithm. The concept of identification of delayed and connected strokes has been used for the purpose of segmentation and results of the proposed algorithm shows 92.975% accuracy of the system.