Scalable Source Code Plagiarism Detection Using Source Code Vectors Clustering
Michal Ďuračík, Emil Kršák, Patrik Hrkút · 2018
Nowadays, the plagiarism is a growing problem due to a lot of easily accessible resources on-line. New algorithms are constantly being developed, but there are not currently many systems, that could be used for successful plagiarism detection in large source files databases. Aim of our work is to deal with plagiarism on a large scale. This paper describes our new scalable approach to the detection of plagiarism in source code in the academic environment. The aim of the algorithm is to search for plagiarism in a huge number of source code files. An incremental clustering approach is applied to achieve modularity and scalability of the algorithm. The paper also details structures of data persistence and methods of searching for source code snippet matches. In addition, we present some results of this approach on real student submissions and compare the results with other detection systems.