Experimental study of hidden Markov model based part-of-speech tagging for Chinese texts
Sun Maosong · Journal of Tsinghua University(Science and Technology) · 2000
The technique of part of speech tagging plays an important role in many applications of Chinese information processing. A large scale manually annotated Chinese corpus and a number of well conducted experiments were used to identify the following points of the hidden Markov model based part of speech tagging scheme for Chinese texts. The results are: ① The Bigram model is better than the Trigram model in terms of the performance cost ratio. ② An annotated corpus of about 70000 words tokens would be sufficient for training the Bigram model, to produce system performance of about 93% tagging accuracy for ambiguous word tokens and 97% tagging accuracy for all word tokens in the texts. ③ The Bigram model can be suited to different application domains quite well. These conclusions will facilitate the development of practical part of speech tagging systems for Chinese texts.