Segmenting bangla text for optical recognition

Md. Abdus Sattar, Khaled Mahmud, Humayun Arafat, Amrin Zaman · 2007

One of the important reasons for poor recognition rate in optical character recognition (OCR) system is the error in character segmentation. Existence of different type of characters in the scanned documents is a major problem to design an effective character segmentation procedure. In this paper, a new technique is presented for identification and segmentation of Bengali printed characters. This paper focuses on the segmentation of printed Bengali characters for efficient recognition of the characters. Our Line segmentation success rate is 99.7 % for 1000 lines, we have tested. Our Word segmentation success rate is 99.8 % for 4900 words tested. From the experiment we noticed that isolated characters fall into isolated group in 99.50 % cases. Most of the errors come from connected characters and characters having τ in front of them as segmenting τ we take the help of width. From the experiment we noticed that most of the errors came from components having multi-touching points between two characters.

Read the paper · More papers on PaperTik