A Markov language model in Chinese text recognition

H.-J. Lee, C.-H. Tung, C.-H.C. Chien · 2002

A two-stage Chinese text recognition system is presented. In the first stage, a Chinese character is first segmented nonuniformly into 10 strips horizontally and vertically. Then three statistical features, viz. crossing counts peripheral background area and contour line length are extracted to form a 60-dimension feature vector. A feature matching method based on the city-block distance metric is employed to select N nearest neighbors as the candidates for each input character from the reference template base, which consists of 5,401 frequently-used Chinese characters. In the second stage, a 3-part-of-speech (tri-POS) Markov language model is employed to extract the most promising characters from all candidate characters in an input sentence. The dynamic programming method is applied to find the most promising sentence hypothesis whose part-of-speech sequence has the maximum likelihood to occur among all of the candidate sentences for an input sentence. The tri-POS contextual information is estimated from a tagged corpus.>

Read the paper · More papers on PaperTik