Creating A Corpus of Plagiarised Academic Texts
Mark Stevenson, Paul David Clough · 2009
Plagiarism is a serious problem in higher education and generally acknowledged to be on the increase (McCabe, 2005). Text analysis tools have the potential to be applied to work submitted by students and assist the educator in the detection of plagiarised text. It is difficult to develop and evaluate such systems without examples of such documents. There is therefore the need for resources that contain examples of plagiarised text submitted by students. However, gathering examples of such texts presents a unique set of challenges for corpus construction. This paper discusses current work towards the creation of a corpus of documents submitted for assessment in higher education that contain examples of simulated plagiarism. The corpus is designed to represent the types of plagiarism that are found within higher education as closely as possible. We describe the process of corpus creation and some features of the resulting resource. It is hoped that this resource will become useful for research into the problem of plagiarism detection.