An Ensemble Based Stacking Architecture For Improved Bangla Optical Character Recognition
Tareque Bashar Ovi, Adil Ahnaf · 2022
Bangla handwritten character recognition holds an important position since it is the world's seventh most widely spoken language. However, the task is challenging as there exists wide range of individual writing styles and structural similarities across characters. Deep learning models like CNN have been employed to categorize Bangla characters although the accuracy hasn't improved much. In this research, the comparison of accuracy between the use of single machine learning model to stacking of the models has been shown by ROC curves. Stacking of Random Forest, Extra Trees, and XGBoost models based on ensemble and feature extraction has been used in this research that outperformed all current models in the Bangla OCR with 99.98% accuracy on the CMATERdb 3.1.3.3 dataset.