Chinese Word Segmentation Based on Contextual Entropy

Jin Huang, David M W Powers · Institutional Repositories DataBase (IRDB) · 2003

Chinese is written without word delimiters so word segmentation is generally considered a key step in processing Chinese texts. This paper presents a new statistical approach to segment Chinese sequences into words based on contextual entropy on both sides of a bigram. It is used to capture the dependency with the left and right contexts in which a bigram occurs. Our approach tries to segment by finding the word boundaries instead of the words. Experimental results show that it is effective for Chinese word segmentation.

Read the paper · More papers on PaperTik