Optical Character Recognition for Handwritten Telugu Text
Buddaraju Revathi, B. N. V. Narasimha Raju, Kaki Surendranath, K. Mounika, G Sravani · 2023
Optical Character Recognition (OCR) is a process of digitalization of scanned documents. OCR for Telugu language is challenging due to intricacy of characters with almost similar structure. As Telugu has spaces between the characters in a word, character level training decreases the complexity in training dataset and enhances the character recognition rates. The major confronts for developing Telugu OCR especially for handwritten text is character level segmentation which depends of a person’s handwritten style. To deal with Unicode training problem that arises for most of Indic scripts, we have converted label to a string during training and after recognition they are inverse mapped. For feature extraction we have used ResNet-50, which captures the small variances in the character there by increasing the recognition rates. The character level recognition rates of 93.5% are achieved and word level recognition rates of 84% are achieved.