Effective Compound Character OCR for Printed Devanagari Script
Ansh Mathur, Harshit Kumar Gupta, Prem Prakash Vuppuluri · 2024
Devanagari script serves as the foundation for three prominent languages: Hindi, spoken by over 520 million, Marathi by over 83 million, and Nepali by over 14 million individuals. It holds the distinction of being the most widely used Brahmic script globally, employed by countless people for reading and writing in their daily lives. Devanagari has a separate class of characters called compound characters which are the conjunction of two consonants. Recognition of printed OCR has been widely explored, however, the use of compound characters in Devanagari makes the problem significantly challenging. This paper focuses on optical character recognition of these compound characters. The task is quite challenging since several of these compound characters are very similar and hence difficult to distinguish. There are a total of 955 such unique characters in Devanagari. We prepared a dataset of these characters, preprocessed them and trained on a linear SVM model. An independent set of test instances was developed as part of the work. A dictionary was also prepared for these characters and complete OCR was performed on these compound characters. The proposed model was tested extensively on the test instances, and obtained an average accuracy of about $84.5 \%$.