Sentences similarity analysis based on word embedding and syntax analysis
Xinchen Xu, Feiyue Ye · 2017
The calculation of sentence similarity is widely used in various areas of natural language processing (nlp) technology. At present, bag of word is the most common method to calculate sentence similarity. This method only takes into account the word level information, ignoring the impact of word ambiguity and the semantic information contained in the internal structure of the sentence. In this paper, we propose a sentence similarity calculation method based on word embedding and syntactic analysis. By means of machine learning word embedding to express the word in order to solve the word ambiguity problem, and using the dependency parser to analyze the internal grammatical structure of the sentence, we are able to identify the relationship between corresponding components in two sentences more accurately in the calculation of sentence similarity. In this paper, the correlation degree is calculated on the SemEval dataset and some better results are obtained.