Chinese Term Recognition and Extraction Based on Hidden Markov Model
Yonghua Cen, Zhe Han, Ji Peipei · 2008
Motivated by the probabilistic characteristics of syntax compositions especially POS (part of speech) matching of Chinese textual information and the inner structures of most unlexicalized Chinese domain terms, a system framework to recognize and extract domain-specific Chinese terms based on hidden Markov model (HMM) was proposed and implemented. The system learns the HMM parameters by the input training corpus with words roughly segmented and POS tagged by the ICTCLAS system developed by Chinese Academy of Sciences and term boundaries manually labeled. Based on HMM with the learned parameters knowledge, the system conducts term boundaries labeling for Chinese textual information from different domains and recognizes terms according to these boundaries. The system shows good performance, and the terms recognized can be treated as candidate terms for false-eliminating and optimizing combining with other parameters such as mutual information and domain dependency.