A study of Chinese word separation based on bi-directional LSTM

Yin Zhou · 2022 International Conference on Data Analytics, Computing and Artificial Intelligence (ICDACAI) · 2022

In this paper, the method of machine learning and depth learning is applied to the field of Chinese word segmentation, and the traditional method of word segmentation based on machine learning is improved to improve word segmentation effect. In this paper, the vector corpus is quantized, and the injected LSTM adds the context relation in the language to the vector, which provides sufficient context information for the next conditional random word segmentation, thus improving the accuracy of the word segmentation. Compared with the general neural network, the advantage of LSTM is that it can preserve the dependency information of context. Compared with ordinary loop neural network, LSTM can not keep the gradient dispersion and gradient explosion to keep long distance dependence information, so as to better support word segmentation. This paper validates the proposed model on the corpus provided by the university of Beijing language, and tests and compares the effect of the traditional model on the same data set. The experimental results show that the fusion of the depth of learning characteristics and the shallow machine learning characteristics of the Chinese word segmentation compared to the traditional machine learning word segmentation, probability model word segmentation, word annotation word and word segmentation effect has improved to a certain extent. In the corpus of Peking University, our experimental results achieved 92.80% F value, which was 1% higher than that of conventional neural network segmentation method.

Read the paper · More papers on PaperTik