Semantic Similarity for Text Comparison between Textual Documents or Sentences

Anshul Modi, Yuvraj Singh Dhanjal, Anamika Larhgotra · 2023

The inclusion of semantic information into any similarity metric boosts its effectiveness and gives findings that are interpretable by humans for further inquiry. A similarity calculation approach only centered on word properties inside the text sometimes provides less precise findings. This document gives three methods that aims to focus on textual terms and incorporate semantic information into their feature vectors, thereby computing semantic similarities. These strategies are founded in both corpus and knowledge-based approaches, namely: cosine similarity using tf-idf vectors, cosine similarity employing word embeddings, and soft cosine similarity utilizing word embeddings. Among these three, cosine similarity utilizing tf-idf vectors is deemed the most proficient in spotting similarities among succinct texts. The texts found by this technology are easily comprehensible and can be easily employed in other applications for important information retrieval.

Read the paper · More papers on PaperTik