Recognition of Tamil handwritten characters using Scrabble GAN
N. Sasipriyaa, P. Natesan, E. Gothai, G. Madhesan, E. Madhumitha, K.V. Mithun · 2023
Recognition of hand-or machine-printed documents is now a crucial component of applications. A commonly used method for converting a printed or handwritten file into its equivalent text format is Optical Character Recognition (OCR). Handwriting recognition is a major research area in pattern recognition and attracts extensive research in OCR. In Southern part of India, Tamil is the primary language with the longest continuous literary heritage. OCR is used to extract characters from scanned digital input images and convert them into formats that can be edited by machines. It is quite challenging to recognize handwritten Tamil characters due to variations in font size, orientation angle and writing styles, bigger category sets, and misunderstanding over the similarities between handwritten characters. In this study, we suggest a deep learning-based technique for handwritten Tamil character recognition. Text documents printed on paper require time-consuming character correction and reprinting. Here, a model incorporating scrabble GAN with CNN is used to increase the dataset. New training instances are created from the old instances using Scrabble Generative Adversarial and Networks augmentation to improve the accuracy of detecting the Tamil handwritten characters. The primary goal of the research work is to expand the handwritten characters and recognize them. Hyperparameters such as kernel size, activation function, padding, step size and layer count are adjusted to enhance CNN-based handwriting recognition. In order to distinguish between characters, CNN evaluates their differences in shapes and features.. Dataset consists of nearly 6,000 images and the 80% of the images are used as training images and 20 % of the images used as testing images and obtained the accuracy of 96.23%.