Chinese word segmentation based on conditional random fields with character clustering

Liping Du, Xiaoge Li, Chunli Liu, Rui Liu, Xian Fan, Jianing Yang, Dayi Lin, Mian Wei · 2016

Chinese word segmentation plays an important role in Chinese text mining. It is the foundation of automatic relation extraction and identification in Chinese information processing. In this paper, we propose a method for Chinese word segmentation based on conditional random fields (CRF) with character clustering. For the character clustering, we firstly use the Skip-Gram model to obtain character embedding from a raw corpus (without word delimiters). We then apply two different clustering algorithms, K-means and Brown clustering algorithm, to get the clusters of character embedding. The effect of different numbers of dimensions of character embedding, the number of clusters, and different clustering algorithms have been studied. We verify our method using the 4th CCF Conference on Natural Language Processing and Chinese Computing (NLPCC2015) Weibo text segmentation task. Our system achieves an F-score of 95.67% and an out of vocabulary (OOV) rate of 94.78%. The result shows that clustering character embedding based on character representation can improve the performance of Chinese word segmentation on short text.

Read the paper · More papers on PaperTik