Machine Learning Techniques for Multisource Plagiarism Detection
Akhil Eppa, Anirudh Murali · 2021
Taking someone else’s work and claiming it as your own is termed as plagiarism. Plagiarism is a concerning issue in every field of education. There are various tools to detect plagiarism and help maintain the necessary integrity. Specifically, multisource plagiarism is when one copies from multiple resources and mixes them and presents it and hence is challenging to detect accurately. This paper deals with detecting multisource plagiarism in the specific category of C programming assignments. The approach used in this paper is to compare programs at the function level using an attention-based model to extract features from the program. By making comparisons at the function level, it is possible to detect plagiarism from multiple sources by performing a one-to-many mapping instead of the regular one to one mapping. The DBSCAN clustering algorithm is then used to find groups of similar submissions based on the extracted features, using point density in the vector space to cluster submissions. This approach brings a new dimension to multisource plagiarism detection, which is not handled by popular existing tools.