Towards building a Bangla text recognition solution with a Multi-Headed CNN architecture
Md. Majedul Islam, Avishek Das, Ibna Kowsar, AKM Shahariar Azad Rabby, Nazmul Hasan, Fuad Rahman · 2021 IEEE International Conference on Big Data (Big Data) · 2021
Bangla is among the ten most popular languages in the world by the number of speakers. The task of Bangla recognition is quite challenging than other languages because of the existence of graphemes of multiple single characters, and diacritics of vowels and consonants. The purpose of this study is to develop an innovative large-scale Bangla OCR solution based on character-level recognition. Two types of documents were used to test our method: handwritten and printed. In addition, our method was applied to the handwritten documents as well as three subdomains of the printed domain: computer-composed, letterpress, and typewritten documents using our proposed attentionbased multi-headed CNN architecture. Extensive testing shows that our method provides state-of-the-art performance on both handwritten and printed texts.