Improved Handwritten Character Recognition for Sinhala Language based on Convolutional Neural Networks

W.V.S.K. Wasalthilake, T. Kartheeswaran · 2022 IEEE 7th International conference for Convergence in Technology (I2CT) · 2022

Automating handwritten character recognition (HCR) for Sinhalese is still an emerging field, as Sri Lanka is the only country in the world that uses Sinhala as the national language. The Sinhala alphabet includes 60 characters that appear to be fairly complex when compared to many other languages. Although approximately 25 research studies have been conducted since 1990s to recognize handwritten Sinhala characters, there is still a large room for improvement. We have investigated two different methods for handwritten Sinhala character recognition: (i) total pixel length-based method which was newly derived by assuming the differences among characters in pixel length and (ii) Convolution Neural Networks (CNN) based classification, to overcome the limitations in the primitive methods. Multiple preprocessing techniques were used in the pixel length-based method while only a resizing technique was used in the CNN based method. The total pixel length and CNN based methods were tested using a dataset of 55 characters, each with 33 samples. The result of the pixel length method is imperfect as distinct characters may hold the same length. Therefore, we abode this method for further processing in future with more improvements. But the proposed CNN classification produces an accuracy of 85.71% and when compared to current studies, this can be considered as the best result achieved for 55 characters among total 60 characters, with nearly 10% of accuracy improvement.

Read the paper · More papers on PaperTik