Code Similarity Detection Based on Siamese Network

Yi Wu, Wei Wang · 2021

At present, with the continuous expansion of software scale, problems such as plagiarism, clone and reuse in software code become increasingly prominent, and the study of code similarity plays an important role. The existing studies have problems such as inaccurate representation of code semantic information and insufficient acquisition of word vector feature information. To solve the above problems, this paper proposes a code similarity calculation model based on a deep learning framework. This method firstly expresses the source code semantics. Secondly, it uses the Siamese network to extract semantic feature information. Finally, it utilizes the cosine distance to calculate the similarity of feature vector in high-dimensional space. Experiments have proved that our method has better performance in terms of precision, recall and F1compared with the baseline method. For this reason, our method can effectively obtain code semantic information and improve the performance of code similarity measurement.

Read the paper · More papers on PaperTik