Enhancing Cosine Similarity: A Unified Approach to Reduce Common Entity Bias in Sentence Similarity
Amit Kumar Gupta, Aadya Gupta, Sumukh Gupta, P. Tarun, Sushant Jain · 2025
Cosine similarity is widely used for measuring semantic similarity between sentences by calculating the cosine of the angle between their embedding vectors. Despite its widespread application, it often miscalculates, overestimates similarity when common Named Entities (NER) are present among text corpus, as these entities inflate the score without adequately reflecting real intents of the sentence. In this paper, first, the existing statistical and AI based approaches were evaluated for measuring semantic similarity between text corpora. Since the results were not acceptable, an enhanced cosine similarity framework was proposed, which addresses these challenges by using weighted scoring mechanism applied on Named Entity Recognition (NER), sentence intent analysis. Dataset is curated from FAQs of various products and user queries of initial deployment of LLMs. Logistic regression was used to learn optimal weights for these components using annotated data. Experimental results demonstrate improvement from 65% to 85% classification accuracy with proposed method.