Cross-language code clone detection via flow-enhanced graph attention network
Mengyao Hu, Jia Yang, Weiqi Zhou · The Computer Journal · 2025
Abstract Code clones are similar code fragments at the syntactic or semantic level, commonly seen in software development. Excessive cloning harms maintainability and may introduce persistent bugs. We analyze cross-language code clone detection at the accurate semantic level. Most existing clone detection approaches target single-language environments and focus mainly on syntactic similarity. However, complex software systems are often developed using multiple programming languages, resulting in semantically similar cross-language code clones. These clones pose challenges beyond the capabilities of current detection tools. In this paper, we propose a novel flow-enhanced graph attention network approach, called FEGAT, to effectively detect cross-language code clones at the semantic level. First, we design a flow-enhanced code graph using abstract syntax tree along with the added control and data flow edges. Then, we input this code graph into the pre-trained model CodeBERT to learn the initial flow-enhanced node representation with semantic information. Third, we design FEGAT to learn flow-enhanced graph representation of cross-language codes from their semantic information and detect clones by computing the similarity score. Finally, we conduct experiments on the AtCoder and CodeChef datasets to evaluate the performance of FEGAT in terms of precision, recall, and F1-score. The experimental results demonstrate that FEGAT outperforms existing cross-language code clone detection tools.