Elimination of Splitting Errors in Printed Bangla Scripts
Mumit Khan · 2009
Accurate and robust character segmentation is a significant challenge in Bangla optical character recognition (OCR). The two main errors in segmentation are joining and splitting errors. To solve the problems of joining errors, several algorithms have been proposed in the literature, with varying degrees of accuracy. Few solutions have been proposed to handle the splitting error issue; however, the accuracy of these proposed solutions were not measured. In an actual implementation of the proposed techniques, we observe the presence of over segmented units. In this paper, we present a dissection based splitting error elimination method which solves the problem of over segmentation under a wide range of document images. Our methodology performs its tasks in two stages: we first concentrate on the careful clipping of the matraa (headline) and put our effort in keeping the pixel information of the units intact which are sensitive to splitting errors. In the second stage, we apply several rules based on the feature information of the units in a word. The combined performance of these two stages results in success rate of 99.93% in eliminating the splitting errors.