Short text similarity computation method based on feature expansion and Siamese network

Xinyuan Niu, Wenguang Zheng, Yingyuan Xiao, Qian Wang · 2021

Text similarity computation issues is a widely studied problem in natural language processing (NLP). Short text similarity computation is a new and more challenging problem, which cannot be effectively solved by using previous regular text similarity computation approach. The main reason is that, a short text generally contains limited number of words and fewer features can be extracted. In this paper, we propose a short text similarity computation method based on feature expansion and Siamese neural network. Firstly, a latent Dirichlet allocation (LDA) based model is constructed to expand the features of a short text. Then, deep features are extracted by using Siamese neural networks model which contains both convolutional neural networks (CNN) and Bi-directional long short-term memory (BiLSTM). Finally, the similarity of two short texts can be achieved by computing the Manhattan distance between generated feature vectors of these two texts. Experimental results show that, based on the data set of Ant Financial NLP Challenge, our method achieves higher accuracy and F1 score.

Read the paper · More papers on PaperTik