On the segmentation of multi-font printed Uygur scripts

A. Ymin, Yoshikazu Aoki · 1996

In many OCR systems, character segmentation is a necessary preprocessing step for character recognition. It is an important step because incorrectly segmented characters are not likely to be correctly recognized. The most difficult case in character segmentation is cursive scripts. Uygur character is a cursive script. This paper presents the problem of segmenting the Uygur characters in various fonts and size in printed scripts. The technique for the segmentation is presented as following: line separation, word separation, segmenting the word into isolated characters consists of the two step's algorithms, topological segmentation, and quasi-topological segmentation. Topological segmentation is based on tracing the outer contour of a given word. Quasi-topological segmentation is based on the decision to section a character on a combination of feature-extraction and character-width measurements. Our approach relies on the feature of characters and fonts and profile models.

Read the paper · More papers on PaperTik