HEP: Heuristic Similarity Calculation based on the Minimum Entropy Principle
Jinhua Yang · 2021 7th International Conference on Computer and Communications (ICCC) · 2021
Automatically calculating the semantic similarity of two short sentences is one of the fundamental tasks in natural language processing. There are several models used to evaluate semantic similarity. However, most methods in unsupervised learning are still traditional models. In this paper, we propose a new approach that is in line with the law of nature. We use a minimum entropy weighted model based on the heuristic algorithm to perform this task. Instead of individual weigh vector for each sentence, searched a global weight that applies to all sentences. We choose accuracy and area under the curve (AUC) as evaluation metrics. By calculating the similarity of sentence pairs on different datasets, our minimum entropy weighted model outperforms previous traditional unsupervised models in accuracy and the area under the curve (AUC). This paper optimize the threshold selection in calculation and obtain an excellent experimental results in more general cases. All in all, this paper provides an excellent weighting method for word vectors to form sentence vectors.