Identification of Chinese Unknown Word Based on Decision Tree
Chun Yuan · Zhongwen xinxi xuebao · 2004
Unknown words can cause segmentation mistakes in the automatic word segmentation processing of large Chinese texts. Meanwhile the recognition of unknown words is a difficult point in word segmentation processing. This article suggests the recognition of unknown words as a question of classification first, that is, the segmentation fragments, upon the segmentation processing, are divided into two categories as combination (combining an unknown words) and segregation (segregating to two single character words). Then, decision tree is used to solve this problem of classification. Six aspects are summarized from the Corpus and the modern Chinese morpheme database: front position formation probability of former character, back -end position formation probability of latter character, former character freedom, latter character freedom, mutual information and single character words co -occurred probability. Training set is constructed using these as attributes. And lastly, the decision tree is produced using C4 5 algorithm. After word segmentation processing, some unknown words have been recognized, but there are still some segmentation fragments usually. In this case our method should be used. For an open test, its recall rate is 69 42%; its precision is 40 41%. Experimental results show that shis recognition method based on decision tree is worth to continue to study in the future.