Nuanced Code Clone Detection Through LLM-Based Code Revision and AST Graph Modeling

Chunguang Li, Jessada Konpang, Adisorn Sirikham, Yan Wang · IEEE Access · 2025

Detecting semantically equivalent but syntactically diverse code clones (Type-4) remains challenging for traditional AST- or token-based approaches. We propose a clone detection framework that couplesLLM-based code revisionwithAST graph modelingandGraph Attention Networks(GAT), trained via a joint objective that includes a contrastive loss to align embeddings of semantically equivalent fragments. Concretely, we use an LLM to generate syntactically altered but functionally identical variants, thereby augmenting training data withnuanced code clones—subtle edits such as identifier renaming, control-flow restructuring, or statement reordering that preserve program behavior while confusing purely syntactic matchers. On Google Code Jam (GCJ), the proposed method attains Precision 0.99, Recall 0.98, and F1 0.985, outperforming the strongest baseline (F1 = 0.97) by +1.5 percentage points (relative ≈ +1.6%). On BigCloneBench, it achieves Precision 0.97, Recall 0.96, and F1 0.97, improving over the best baseline (F1 = 0.93) by +4.0 pp (relative ≈ +4.3%). For the hardest WT3/T4 category, the framework reaches F1 97.4 versus 93.6 for a strong GNN baseline (+3.8 pp). Ablations indicate that (i) LLM-based augmentation and (ii) GAT-based AST encoding are both critical to performance. In addition to accuracy, we discuss computational considerations (augmentation/training cost), portability beyond Java, and potential extensions toward dynamic clone detection. The results suggest that marrying LLM-generatednuancedvariants with graph-based program representations yields robust Type-4 clone detection while remaining scalable to practical codebases.

Read the paper · More papers on PaperTik