Developing Bilingual Plagiarism Detection Corpus Using Sentence Aligned Parallel Corpus Notebook for PAN at CLEF 2015
Habibollah Asghari, Khadijeh Khoshnava, Omid Fatemi, Heshaam Faili · CLEF (Working Notes) · 2015
Plagiarism detection is the process of locating text reuse within a suspicious document. The plagiarism detection corpora are used for evaluating plagiarism detection systems. In this paper, we present a bilingual Persian- English plagiarism detection corpus. We provide our corpus for the task of text alignment corpus construction in the PAN 2015 competition. Our approach is based on parallel corpus sentences. We have used a Persian-English sentence aligned parallel corpus in a combination with Wikipedia articles to create our corpus. Paired sentences in parallel corpus have a similarity score between 0 and 1. We have used similarity scores to establish the degree of obfuscation for constructing the plagiarism cases.