DCU@FIRE-2014
Debasis Ganguly, Gareth J. F. Jones · 2015
We investigate an information retrieval (IR) based approach to source code plagiarism detection. The standard method plagiarism detection by extensively checking pairwise similarities between documents is not scalable to large collections of source code documents. To make the task of source code plagiarism detection fast and scalable in practice, we propose an IR based approach. In this method each document is treated as a pseudo-query which retrieves a list of documents which are potential candidate for containing plagiarised material in decreasing order of their similarity to the query. A threshold is then applied on the relative similarity decrement ratios to create a set of documents as potential cases of source-code reuse. Instead of treating a source code as an unstructured text document, we explore term extraction from the annotated parse tree of a source code and also make use of a field-based language model for indexing and retrieval of source code documents. Results confirm that source code parsing plays a vital role in improving the plagiarism prediction accuracy.