Chinese Word Sense Disambiguation Based on Hidden Feature Extraction and CRF Model
Yin Huang · Journal of Guizhou University · 2013
The disambiguation model is built in the traditional methods of Chinese word sense disambiguation by observing dominant features,such as the context information and part of speech. We found the grammatical structure and semantic information hidden in those words also lead to ambiguities by analysis in-depth the reason of producing ambiguity. We can consider this information into the disambiguation model. Because the collocation information between words is summarized in the How Net,we extracted hidden semantic features of the collocation information from the training corpus by the How Net. Then,we combined it with the characteristics of the dominant context to do the word sense disambiguation by using conditional random field( CRF) method. Finally, we did the experiments of the word sense disambiguation and verification of its effects. We found that the method used in this paper improved accuracy of word sense disambiguation by compared with the traditional CRF.