Combining segmenter and chunker for Chinese word segmentation
Masayuki Asahara, Chooi Ling Goh, Xiaojie Wang, Yūji Matsumoto · 2003
Our proposed method is to use a Hidden Markov Model-based word segmenter and a Support Vector Machine-based chunker for Chinese word segmentation.Firstly, input sentences are analyzed by the Hidden Markov Model-based word segmenter.The word segmenter produces n-best word candidates together with some class information and confidence measures.Secondly, the extracted words are broken into character units and each character is annotated with the possible word class and the position in the word, which are then used as the features for the chunker.Finally, the Support Vector Machine-based chunker brings character units together into words so as to determine the word boundaries. MethodsWe participate in the closed test for all four sets of data in Chinese Word Segmentation Bakeoff.Our method is based on the following two steps: