Research on The Word Segmentation Model Construction Based on CNN+BiLSTM+HMM
Xuemei Sun, Bin Wen, Rong Fu · 2021
In natural language processing, Chinese word segmentation is a common and key task. For a great quantity of chinese natural language processing tasks, word segmentation is the basis for completing the follow-up work. At present, the most commonly used word segmentation methods are machine learning and deep learning methods. However, many models have problems such as the inability to identify unregistered words, complex models, and greater dependence on manual processing features. Therefore, in this paper, we propose a word segmentation model based on CNN+BiLSTM+HMM. First, the semantic representation of the word is obtained by using CNN, and then the context information is obtained through BiLSTM. Finally, the HMM is used to classify predict and output the optimal label sequence. This approach can not only improve the accuracy of high score words but also effectively retain longdistance sentence information. The MSRA corpus provided by Bakeoff 2005 to conduct experiments in this paper. The experiment shows that the word segmentation model based on CNN+BiLSTM+HMM is 1.92% higher than the BiLSTM word segmentation model inaccuracy and 1.41% higher than the BiLSTM+HMM word segmentation model. Respectively, which proves that the segmentation model based on CNN+BiLSTM+HMM has certain advantages.