Leveraging IR based sequence and graph features for source-binary code alignment

Zhouqian Yu, Wei Zhang, Tao Xu · 2024

Code similarity analysis is a versatile technique that can be applied across various domains, including code clone detection, code search, malware detection, patch analysis, and vulnerability search. The core of source-binary code similarity analysis lies in effective modeling of the both of source code and binary code. Existing researches have overlooked the fact that code has both text and graph structure features, so they usually only use one of these features for modeling. In this paper, we propose a method that combines sequence features based on intermediate representation and program graph features derived from intermediate representation to enhance the capture of program semantics. Our approach involves the parallel embedding and learning of LLVM IR and the derived program graphs, leading to the extraction of the final code feature representation. Furthermore, we employ a triplet-loss network to identify code characteristic disparities and produce similarity rankings. The experimental results have demonstrated that our approach outperforms other models on the baseline, achieving a 6% improvement in accuracy for source-binary code similarity tasks.

Read the paper · More papers on PaperTik