Documents Similarities Algorithms for Research Papers Authenticity

Izzat Mahmoud Alsmadi, Zakaria Issa Saleh · 2012

Abstract —Studying documents similarity can have several fields of applications. In this paper, we focused on evaluating documents’ similarity to predict possible plagiarism in research papers. We evaluated the usage of several document similarity algorithms such as: Cosine, Dice, Manhattan, Euclidian, etc. We also tried several approaches of selections for the length of characters or words as a baseline for the search algorithms. Preprocessing steps were necessary to remove several types and categories of stop words that may bias the similarity measurement algorithms. Some of the algorithms are developed to search through local file and others are developed to search through the Internet. Results showed that there is a great deal of trade off between the two conflicting criteria: accuracy and performance. Keywords- Text mining, plagiarism; documents similarity, and string searc. I. I NTRODUCTION The amount of possible plagiarism in research publications may vary in its seriousness. Plagiarism can be in wording through copying statements or paragraphs from other research papers. It can be also in copying ideas (i.e. semantic plagiarism) where methods for comparing text or statements similarity may not work well in discovering such plagiarism. In this area, there are conflicting opinions on the levels under which a paper can be classified as plagiarism or not. Double publication is another related problem whether the same author is possibly publishing the same or similar idea in more than one research publication channel. Languages are also barriers that put challenges on detecting possible plagiarism where some authors may publish the same paper two times in two different languages especially in journals that are exclusively published in one language. Documents similarity can have several areas of applicability. Besides, inspecting possible plagiarism which is the focus of this paper, identifying similar documents can be used to improve searching facilities by keeping fewer documents to search for or within and giving the users less to browse through. File synchronization is another important application for users who keep files on several machines (e.g., work, home, etc). Document similarities can be either on language based or syntactic. However, similarity can be more complex to include semantic similarity despite the fact that words and statements may not be the same or similar.

Read the paper · More papers on PaperTik