Word-Level Chinese Named Entity Recognition Based on Segmentation Digraph
Hong Gao, Degen Huang, Yuansheng Yang · 2006
This paper presents a statistic method to recognize word-level Chinese named entities based on segmentation digraph. The main idea is to generate named entity (NE) candidates according to their internal characteristic, and those NE candidates with high confidence are added into the segmentation digraph of a Chinese string as vertices along with lexical word candidates. Bigram model and trigram model are used when ambiguities occur to evaluate each path of segmentation digraph. The shortest path is selected as the optimal segment of the Chinese string and NEs that are recognized just in it. The performance of our method was evaluated on the corpus of Peking University, and the results show the method is simple and effective.