SEMANTIC SIMILARITY ANALYSIS USING TRANSFORMER-BASED SENTENCE EMBEDDINGS
Bohdan M. Pavlyshenko, Mykola Stasiuk · Electronics and Information Technologies · 2025
Background . Transformer-based models have become central to natural language processing, demonstrating state-of-the-art performance in semantic similarity assessment, a task critical for various applications. These models capture detailed relationships between text, advancing the ability to gauge semantic relatedness. Materials and Methods. The performance of sentence embedding models, including all-mpnet-base-v2 , all-MiniLM-L6-v2 , paraphrase-multilingual-mpnet-base-v2 , bge-base-en-v1.5 , all-roberta-large-v1 , a ll-distilroberta-v1 , LaBSE , paraphrase-MiniLM-L3-v2 , bge-large-en-v1.5 , was assessed across different dataset sizes with two datasets. The following preprocessing steps were applied to the datasets: lowercasing, removing stop words, cleaning from special symbols and numbers, and lemmatization. Cosine similarity scores with negative values, indicating semantic dissimilarity, were treated as equivalent to a human-annotated similarity score of 0, and non-negative cosine similarity values were scaled to the 0-5 range. Metrics such as R 2 , MSE, RMSE, MAE, Spearman’s Correlation Coefficient, and Kendall's Tau were used for evaluation. Results and Discussion. Models’ performance generally improves with increased data. Evaluation of sentence embedding models revealed performance variations. all-roberta-large-v1 showed strong accuracy with high R 2 values and low errors. BAAI/bge-large-en-v1.5 excelled in capturing semantic relationships, demonstrating high Spearman’s and Kendall's Tau coefficients. all-MiniLM-L6-v2 demonstrated the fastest embedding generation. BAAI/bge-base-en-v1.5 presented the lowest accuracy. Processing times generally increase with data size. Conclusion. This study highlights a trade-off between accuracy and efficiency in sentence embedding. Model selection depends on balancing these factors to align with specific application needs. In cases when requiring high accuracy should favor all-roberta-large-v1 , while those prioritizing speed would benefit from all-MiniLM-L6-v2 . BAAI/bge-large-en-v1.5 is most suitable for tasks demanding semantic understanding of text details. Keywords : semantic similarity, sentence embeddings, transformers.