Hindi Character Recognition Using Segmentation, Feature Extraction, and Classification Techniques

Piyush Ranjan, Sandeep Kumar, Saurabh Tamta · 2025

This research gives a vast methodology for Hindi character recognition using methodologies of segmentation, feature extraction, and classification. The research of this paper is conducted systematically with a step-by-step process for data collection followed by preprocessing, segmentation, feature extraction, and classification. Performance evaluation follows this process. A dataset of handwritten and printed Devanagari script characters is used, including standalone characters, conjuncts, and diacritical marks. To improve the accuracy of recognition, data preprocessing techniques such as normalization, grayscale transformation, noising removal, and binarization are used. Feature extraction using a 3-layer CNN includes capturing geometric, structural, and textural features as well as stacked dense layers for classification. ReLU activation functions as applied to the hidden layers while SoftMax is used at the output layer of classifier's algorithm. The performance of the suggested approach is analyzed by using precision, recall, FI- score, and confusion matrices for both training and test data. From the results, one can find that the proposed approach works efficiently with high precision, recall, and FI-scores for all classes of characters. The confusion matrices also reflect the effectiveness and generalization capability of the model while learning from unseen data. This research reports the ultimate application of the deep mechanism of learning and CNNs to achieve 97.5% accurate and effective Hindi character recognition.

Read the paper · More papers on PaperTik