Code Similarity Detection Technique Based on AST Unsupervised Clustering Method
Yifan Xie, Wenan Zhou, Huamiao Hu, Zhicheng Lu, Mengyuan Wu · 2020
Code similarity detection is one of the important means to maintain the healthy development of the software eco-environment. Early code similarity detection method is mainly based on the code attribute measurement technology, this method can quickly detect more obvious similar parts on small-scale code datasets, but when faced with large-scale datasets or code segments with similar syntax and semantics, the detection performance is not effective. The new code similarity detection technology based on tree-structure representation and machine learning can adapt to different data changes, thus providing new solutions. This paper proposes a new detection method that combines the abstract syntax tree (AST) structure and unsupervised clustering, this new method can improve the detection efficiency of code clustering analysis, because it not only considers the syntax and semantic information of the code, but also reduces the workload of code data preprocessing through unsupervised clustering.