An absolute Optical Character Recognition system for Bangla script Utilizing a captured image

Md. Ruhulamin Siddique, Md. Ashiq Mahmood · 2021

Character recognition from a captured image is a significant field of research because there are 230 million native speakers in Bangladesh and India. In addition, there are many signboards, billboards, and many other image sources that contain Bangla Script. Since mid-1980, researchers started to recognize Bangla characters from scanned images. However, they already tried different kinds of methods to identify characters and examine the performance of recognition. This paper focuses on developing an eclectic OCR system that can recognize and extract Bangla text. This recognition process predestines captured images by digital camera or scanner containing Bangla scripts. Preprocessing steps include binarization, segmentation, noise cleaning, scaling characters by font size, skew detection, and correction. Freeman chain code represents a character from the image after feature extraction from a scaled character. A multilayer feedforward neural network-based recognition scheme is constructed to recognize and classify the unknown character and samples. We concluded that the success rate is approximately 99% in identifying characters and demonstrating the Unicode text from experimental results.

Read the paper · More papers on PaperTik