A Code Similarity Detection Algorithm Based on Maximum Common Subtree Optimization

Zhikai Lin, Lin Li · Proceedings of the 2020 4th International Conference on Electronic Information Technology and Computer Engineering · 2020

The code similarity detection is different from the traditional text duplication checking. The former has a lot of the same syntax content in the code. There are two code duplication detection algorithms. One is realized by extracting and counting characteristic attributes, which can result in a lack of the logical relationship between code structures. The other is realized by abstracting code into a string, tree structure or graph structure, which can lead to a lack of codes' semantic features. To rectify these deficiencies, an optimization algorithm based on maximum common subtree is proposed. First of all, the structural information based on the largest common subtree is extracted to calculate the structural similarity of codes. After this, the semantic information based on the longest common subsequence is extracted to calculate the semantic similarity of codes. Finally, the semantic similarity and structural similarity are assigned different weights using TF-IDF algorithm. Experimental results show that, the optimization for code similarity detection based on the maximum common subtree is able to reduce the non-plagiarism similarity of codes and keeps more feature information than the traditional code duplication checking algorithm.

Read the paper · More papers on PaperTik