Gujarati Optical Character Recognition Using Efficient Text Feature Extraction Approaches
Avani Samir Bhuva, Dhirendra S. Mishra · Informatica · 2025
India isthe most populous country, with 22 official regional languages. Retrieving information from these regional languages is a challenging task. Approximately 62 million people worldwide speak the Gujarati language. This research paper aims to understand and extract the meaningful text features of the Gujarati text from the OCR Gujarati dataset. This research focuses on extracting meaningful text features from the Gujarati OCR dataset, which comprises 23,100 samples generated using the TERAFONT-VARUN font and augmented with horizontal/vertical shifts and rotational transformations. This study explores three levels of text feature extraction: Mid-level features using the Integrated Shape Numeric Encoding Approach (ISNEA) and Fusion of Region Geometric Features (FRGF), Mid-high-level features via the One-Bit Frequency Count Approach (OBFCA), and High-level features through a deep learning-based CNN model. The extracted features were stored in a structured Gujarati Text Feature Vector Dictionary. ISNEA struggles with characters containing maatra’s modifiers, with 87.5% on standardized OCR images. OBFCA resolves the maatra’s issue by row-wise binary frequency computation, yielding 90.11% accuracy. FRGF significantly outperforms ISNEA and OBFCA, with 93.5% accuracy using Eccentricity as a single feature and 92.75% using Eccentricity + Perimeter as a fused feature. The Euclidean distance and cosine similarities were also used to measure the similarities between the extracted text features. Comparative analysis against existing methods confirms the superiority and robustness of the proposed approaches in Gujarati OCR feature extraction.