Auto Clustering Source Code To Detect Plagiarism Of Student Programming Assignments in Java Programming Language
Yusni Amaliah, Wilem Musu, Suprianto Suprianto, Muhammad Noer Fadlan · 2021 3rd International Conference on Cybernetics and Intelligent System (ICORIS) · 2021
Informatics among students who are always interacting with computers that facilitate the practice of plagiarism is given the facility to copy and change the text (copy and paste) and connection facilities that allow to access other people's work freely via the internet, the practice of plagiarism is often done. Plagiarism detection is divided into three major processes, namely preprocessing, plagiarism detection and calculation accuracy using precision, recall, and F-measure. In preprocessing is done with the help of library compilation techniques ANTLR (Another Tool for Language Recognition) Furthermore, the similarity distance calculation using tree edit distance algorithm, to calculate the order of the strings using the technique ngram. Clusters are thought to be similar to be grouped into several clusters using Agglomerative Hierarchical Clustering with average linkage method. the results of the calculation accuracy of the system and manually counted using the precision, recall, and f-measure. The result of edit distance algorithm is good enough in calculating the similarity source code, as well as the calculation ngram to measure the change in the sequence string in the source code. It is characterized by the value of similarity that can determine the cluster group that included plagiarism, as well as with the process of grouping using hierarchical clustering methods. Calculation of accuracy using the Precision, Recall and f-Measure to prove the detection process is almost the same as the detection system manually, and the search for the best threshold value.