Devanagari Character Recognition using Statistical and Transform based Features
Duddela Sai Prashanth, Priya R. Kamath · 2024
Optical Character Recognition (OCR) on hand-written characters is much more complex when compared to printed characters. Complexity increases for Indian scripts like Devanagari because of its complex structures. Many researchers developed datasets for Devanagari characters, but they are either unavailable or limited in their labs. More than 200 subjects of different age groups are considered, and a dataset of 38,750 images of Vowels and Numerals is created, publicly available in Mendeley. Experiments are conducted on the dataset using 1. Statistical Features like Zoning, Vertical and Horizontal profiling, Histogram of Oriented Gradients (HoG) and 2. Transform-based Features like Discrete Cosine Transform (DCT), Discrete Wavelet Transform with DCT and Radon Transform with DCT. An accuracy of 97% is achieved using DCT for Numerals and 86.9% using HoG using Cubic Support Vector Machine (SVM). This study presents an improved recognition accuracy using unique Feature extraction methods for the Devanagari character by applying DCT on Radon Transform. DCT features to solve many pattern recognition problems. Radon transforms, denotes the project data on the 2-dimensional image, and calculates the 1 Dimensional Fast Fourier Transforms on the sum of the column data and Discrete Cosine Transform (DCT) features to solve many pattern recognition problems. DCT is applied on the Radon Transforms of the Devanagari character and generates unique features to improve the accuracy. With Cubic-SVM, a maximum accuracy of 98.7% for Numerals and 89.6% for Vowels. True Positive Rate (TPR), False Negative Rate (FNR), Positive Predictive Values (PPV), and False Discovery Rate (FDR) are parameters used for performance measures for the quality of the proposed method.