Printed Uyghur Texts Segmentation

Jin Jian-Ming · Zhongwen xinxi xuebao · 2005

Uyghur is spoken in Xinjiang Uyghur Autonomous Re gion of China, which adopts Arabic script to write. As a cursive script and othe r characteristics, it is very difficult to do text segmentation and recognition. In this paper, a method, which hybrid horizontal projection and connected compo nents analysis, based on connected components classification is proposed to do t ext line segmentation and word segmentation of Uyghur texts. And then, the basel ine position of each word is estimated. All candidate character segmentation poi nts are found out by calculating the distance between word contour and baseline. Finally, over-segmen ted characters are merged according to rules. Experiment shows that the characte r segmentation accuracy has achieved 99%.

Read the paper · More papers on PaperTik