Assessment of CNN Models for Optical Character Recognition of Modi Lipi Script
Mangesh Ingavale, Jayamala Kumar Patil · 2025
The preservation of ancient scripts like Modi Lipi, historically used for writing languages such as Marathi, is crucial for cultural heritage. This study assesses the effectiveness and performance of standard Convolutional Neural Network (CNN) models in the Optical Character Recognition (OCR) of Modi Lipi script. Nine prominent CNN architectures—VGG16, VGG19, ResNet-50, ResNet-101, DenseNet121, InceptionV3, MobileNet-V3, EfficientNet-B0, and EfficientNet-B1—were evaluated on a comprehensive dataset comprising 575,920 images of numerals, vowels, and consonants. Performance metrics including accuracy, sensitivity, specificity, and F1 score are analyzed across three classification tasks. Advanced models like EfficientNet-B1 and DenseNet121 demonstrated superior performance due to their innovative architectures capable of capturing complex script patterns. The findings highlight that while deeper and more complex models offer higher accuracy, they require more computational resources. This assessment underscores the viability of standard CNN models for Modi Lipi OCR and provides insights for future developments in script digitization efforts.