A Novel Deep Learning Character-Level Solution to Detect Language and Printing Style from a Bilingual Scanned Document

AKM Shahariar Azad Rabby, Md. Majedul Islam, Nazmul Hasan, Jebun Nahar, Fuad Rahman · 2020

Bangla is one of the world's most widely-spoken languages, but few languages (or "script") automation solutions have been reported for it. To build an OCR system, it is very important to detect the language and type of printing style to run specific character recognition and segmentation modules. This paper presents a novel solution to automatically detect the language (Bangla vs English in terms of the script), and printing style (printed vs handwritten) from any given bilingual scanned document using multiple deep learning models.

Read the paper · More papers on PaperTik