CrossLMD: Cross-Language Malicious code Detection
Zhengyuan Xu, Xiaohui Han, Wenbo Zuo, Xuejiao Luo, Zhiwen Wang, Xiao-Ming Wu · 2023
Current malicious code detection models usually mainly focus on variants of popular programming languages and source codes, while ignoring the inherent semantic information in source codes. These approaches present significant challenges in the face of the vast amount of multi-programming language code that exists in repositories, open-source communities, blogs, and forums. To address this problem, we propose CrossLMD, a cross-language malicious code detection model capable of identifying malicious code across various programming languages using a single model. CrossLMD leverages large-scale pre-trained models to understand semantic dependencies between code written in different languages. We built a neural network classifier on top of the pre-trained model to fine-tune the pre-trained model for cross-programming language malicious code detection tasks. Our experimental results demonstrate that CrossLMD can effectively identify malicious code written in never-encountered programming languages, exhibiting superior performance compared to baseline methods. Our proposed model provides a more comprehensive solution for malicious code detection, taking into account the inherent semantic information of source code, regardless of language or syntax changes. Utilizing a single model to detect malicious code in multiple programming languages simplifies the detection process and improves efficiency and cost-effectiveness.