A comparative analysis of different approaches to lexical and semantic document similarity

Pratik Kanani, Devang Shah, Riya Bihani, Khushi Chavan, Aniket Kore, Nilesh Madhukar Patil · 2024

With a boom in digitalization of academic research, the number of scholarly publications has been growing exponentially. This has resulted in logistical problems. It is now impossible for a single person to keep up with all scientific literature. But recent advancements in artificial intelligence (AI) and natural language processing (NLP) paved the way for understanding and interpreting all such texts with document text similarity. For the use case presented, we aim to provide a comprehensive review of several algorithms in this work. On self-scraped literary articles, our study applies unsupervised machine learning approaches. The pre-processed dataset undergoes word embedding algorithms such as TF-IDF vectorization, LDA topic modeling, BERT sentence transformer encodings, GloVe Word embeddings, and Doc2Vec embeddings. Cosine similarity and Euclidean distance are then used to calculate the pairwise percentage of similarity between the articles. The results display the top-k papers most similar to the query document along with the percentage of similarity. This chapter presents the combinations of these text representation algorithms and similarity metrics. Furthermore, performance, challenges, and issues are also analyzed.

Read the paper · More papers on PaperTik