An Approach for Measuring Unstructured text Document Similarity using LDA-BERT Embedding Model
Ashwini V. Zadgaonkar, Akhilesh Dhoot, Ishan Chinchghare, Pranay Rahangdale, Gaurav Iyer, Prateek Bhattacharya · 2024
This research work proposes an innovative method for measuring text similarity of unstructured PDF documents using a hybrid approach that combines Latent Dirichlet Allocation (LDA) and Bidirectional Encoder Representations from Transformers (BERT) embedding models. The proposed methodology initially employs LDA to extract topic distributions from the documents and subsequently employs BERT to create contextualized embeddings of these topics. These resulting embeddings are then utilized to compute similarity scores between pairs of pdf documents.To assess the effectiveness of the proposed approach, experiments are conducted using a diverse datasets of scientific research papers, legal documents, and news articles. Our findings reveal that the proposed method surpasses conventional techniques such as Cosine similarity and Jaccard similarity in terms of accuracy and robustness. Notably, our approach demonstrates exceptional performance in detecting cases of plagiarism, particularly when the plagiarized content has been rephrased or paraphrased.Overall, this research underscores the potential of the LDA-BERT embedding model to enhance the precision and efficiency of text similarity measurements, with wide-ranging applications across document analysis, plagiarism detection, and information retrieval.