Developing Monolingual English Corpus for Plagiarism Detection using Human Annotated Paraphrase Corpus Notebook for PAN at CLEF 2015

Salar Mohtaj, Habibollah Asghari, Vahid Zarrabi · 2015

In this paper, we describe an approach to create monolingual English plagiarism detection corpus for the task of text alignment corpus construction in PAN 2015 competition. We propose two different obfuscation methods to fragment obfuscation for creating the cases of plagiarism. The first method is an artificial obfuscation which consists of variety of obfuscation strategies such as synonym substitution, random change of order, POS preserving change of order and addition/deletion. The second obfuscation method is a simulated obfusca- tion, in which the SemEval dataset is used for creating the cases of plagiarism by using pairs of sentences with their similarity scores.

Read the paper · More papers on PaperTik