Chinese Part-of-speech Tagging Based on Full Second-order Hidden Markov Model

Huang De-gen · Jisuanji gongcheng · 2005

This paper describes an extension to the hidden Markov model for Chinese part-of-speech tagging using second-order approximations for both contextual and lexical probabilities, as well as the traditional Viterbi algorithm is extended. The model makes use of more contextual information than standard statistical models. A smoothing algorithm based on the linear interpolation algorithm is introduced to solve the sparse data problem. The new full second-order HMM is proved to improve Chinese part-of-speech tagging accuracies and disambiguation accuracies over current models.

Read the paper · More papers on PaperTik