A comparative study on Chinese word segmentation using statistical models

Wenchao Meng, Liu Lianchen, Anyan Chen · 2010

Recent years, character based approaches to Chinese word segmentation task are developed, which show great success. In this paper, a detailed comparison among different statistical models are done. Three models (HMM, MEMM and CRF) are considered. First different tag sets are chosen to evaluate the models' precision and efficiency. Then HMM and MEMM are compared with the similar features. At last different features are compared to measure which feature contributes most to Chinese word segmentation. Finally some suggestion is given for developing Chinese word segmentation systems.

Read the paper · More papers on PaperTik