Segmenting merged characters

Thomas A. Bayer, Ulrich Kreßel, M. Hammelsbeck · 2003

In optical character recognition (OCR) and document analysis many errors are not caused by inadequate classifier power, but by segmentation errors. Besides broken characters, merged characters constitute the major remaining problem. This paper presents an efficient method for segmenting merged characters. The algorithm combines a statistical adapted cut classifier and a search algorithm, which employs further experts for selecting the proper cut positions from a set of hypotheses. It is designed especially for proportional fonts and even succeeds, if the characters are in italics font style.>

Read the paper · More papers on PaperTik