Text Similarity Based on Modified LSA Technique

Ahmed T. Sadiq, Khudhair J. Kadhim · 2015

The most applications or practice areas which are interrelated with text mining such as information retriever (IR), clustering, text summarization, automatic answer grading, machine translation and automatic essay scoring and other. All of them are depends on how find similarity distance between a pair or more of texts as a major process. This paper proposes two approaches are focuses on the problem of the semantic similarities between texts in English language by using Latent Semantic Analysis (LSA) technique. It's trying to enhance the process of finding the semantic similarity distance between texts and making it more adaptable for both long documents and short sentences. The two proposed approaches are using the same style that used in Knowledge-based measures, where derived semantic relationship on terms level from the semantic space, and thus calculate the similarity between the two texts fully. Evaluation results on three different data sets show that the two approaches gives good results comparison to human judgment equal to 76% and outperforms several competing methods which are used for detecting Plagiarism in texts, where the proposed system achieves 92%.

Read the paper · More papers on PaperTik