Damaged Sinhala Handwritten Character Location-Identification using Neighbour Mapping
L. B. V. Lakshitha, H. K. I. L. Madhuwanthi, L. Ranathunga · 2023
In Sri Lanka, the Sinhala language is the major language among other languages. Identification of the damaged Sinhala handwritten characters from the degraded Sinhala handwritten document is the primary goal of this research. Although mean-based threshold and character width and height are used to classify the damaged and non-damaged characters in English, they cannot apply to the Sinhala language because Sinhala characters consist of different widths and heights. Therefore, classifying the damaged and non-damaged characters in a degraded Sinhala handwritten document is di cult. The gradient-based clustering technique is the main approach for clustering the character segment pixels based on their gradient values. Boundary drawing around the damaged zone cluster approach leads to identifying the damaged locations of the character foreground areas and classifying the character as damaged or non-damaged. Thinning-based neighbouring pixel mapping technique determines that a character has exactly damaged end tips. The accuracy of the output images is calculated using a confusion matrix by considering both damaged and non-damaged characters. The results show that the boundary-overlapping approach outputs 98.5% accuracy while the thinning-based neighbouring pixel mapping approach provides 93.6% accuracy.