Research on Key Technologies for Text Similarity Calculation Based on Small Datasets
2024
Text similarity is an important criterion for achieving text retrieval.Compared to neural networks that require a large amount of data for training, character based similarity calculation is more suitable for similarity calculation in small datasets.In this paper, we propose to optimize the calculation formula of Jaccard similarity coefficient in view of the missing judgment of the existing character based text similarity calculation algorithms such as cosine distance and Edit distance.Firstly, the denominator of the original Jaccard similarity coefficient was modified to avoid the impact caused by the text length in the database in the original formula.At the same time, the distance between intersecting characters was recorded for weight calculation to improve the accuracy of the calculation.Then, by recording user matching results, the number of similarity calculations is reduced to achieve the effect of improving running speed.Finally, different weights and compensation mechanisms are set for the input text based on its order, achieving differentiation in the order of input text at the lexical level, further improving the accuracy of similarity calculation.The experiment shows that the algorithm performs well in terms of accuracy in similarity calculation.