Document Text Analysis and Recognition of Handwritten Telugu Scripts
Deepa R.N Ashlin, Y. Vijayalata, Atul Negi · 2022
Handwritten text recognition is an open problem of great interest in the area of automatic document image analysis. Handwriting recognition has been studied for a long time with only few practicable results when written on paper. The tasks related to text recognition becomes complex when it comes to Indic languages due to large size of character set. One such language is Telugu where the character list involves 150 unique characters. This paper addresses the procedure to recognize the Telugu handwritten document and convert the text into machine understandable language. The model is built on the ResNet-18, DBSCAN clustering algorithm for detecting the words in the hand written text image. The CNN, RNN and CTC. The proposed model is then evaluated using the IAM Telugu handwritten image dataset and few government school handwritten notes. The proposed model achieves an accuracy of 72.6% with a character error rate of 11.4%.