Cut classification for segmentation
Thomas A. Bayer, U. Kressel · 2002
In optical character recognition (OCR) and document analysis, many reading errors are not caused by inadequate classifier power, but by segmentation errors. In particular, merged characters are a major remaining problem. An efficient and powerful method of determining cut hypotheses for the segmentation of merged characters is presented. The method is based on a classifier deciding for each column of the character image, whether it represents a cut hypothesis or not. Since in the training phase the classifier is adapted by a sample set consisting of images of merged character patterns, the decision rules are created automatically rather than being man-made heuristics. The results obtained from a large test set show that a high recognition rate can be achieved with a reasonable computational effort.>