An improved embedding matching model for Chinese word segmentation
Xiaolong Deng, Yingfei Sun · 2018
To date, various deep neural network based models have been extensively applied in Chinese word segmentation (CWS) task, however, some of these models either consume so much time to train, or perform weakly without additional dictionary or training corpus. In this paper, we proposed an improved embedding matching model for CWS, which achieved 0.4% and 0.5% improvement of F1-Measure respectively on PKU and MSR dataset compared to the original model. Also, our proposed model achieved 0.77% improvement of F1-Measure and 2.59% improvement of weighted F1-Measure on the NLPCC dataset. The improved model outperforms most of previous models on the F1 measure of segmentation performance, and consumes less time to train and test. After improvement, the model can be a better choice in practical use for its exceedingly fast convergence and excellent performance.