A Chinese Word Extraction Algorithm Based on Information Entropy

Junfang Zeng · Zhongwen xinxi xuebao · 2006

Targeting at extending the dictionary for word segmentation so as to improve its accuracy,this paper presents a high-frequency Chinese word extraction algorithm based on information entropy.We firstly transform noisy words and characters to separators,thus a text can be viewed as a Chinese string collection isolated by separators.Then we compute the frequencies of all the substrings of these Chinese strings.Finally,we judge whether each substring is a word by computing its information entropy.Preliminary experiments show that this simple algorithm is effective in extracting high-frequency Chinese words,with the accept rate up to 91.68%.

Read the paper · More papers on PaperTik