Algorithm to recognize unknown Chinese words based on BBS corpus
Hailing Xiong · Jisuanji gongcheng yu sheji · 2010
To deal with the problem of low efficiency of Chinese unknown word segmentation,a new method based on BBS corpus is presented.Network spider is used to download BBS web pages to build a corpus and this corpus is updated periodically in order to obtain a strong limitation.The new statistic MD(constructed by the mutual information function and duplicated combination frequency function) is used to segment the corpus to generate a candidate word list.By comparing candidate words list and the previous lexicon to recognize the unknown words,and added these unknown words into the basic lexicon.Experiments showed that the proposed method effectively improve the efficiency of identification of unknown words.