Prefix Tuning for Few-shot Malware Classification with Supervised Contrastive Cross-Entropy Learning
Zhengming Yuan, Yaoxiang Yu, Yuxia Wu, Siwei Huang, Bo Cai · 2024
The accurate detection of malware is of paramount importance in today’s society, where the number of computers is vast. Traditional methods for classifying malicious code face three main challenges. Firstly, previous deep learning models encounter difficulties in handling excessively long input sequences due to their structural and computational resource limitations. Secondly, the rapid changes and rapid propagation of malicious samples pose certain challenges in collecting a sufficient number of samples and ensuring the accuracy of sample labels. Thirdly, in scenarios with limited sample availability, prior approaches failed to consider the augmentation of the model’s generalization performance when dealing with a limited number of samples. To address the first challenge, we propose a novel method that segments malicious samples into subroutines to extract opcode sequences. These sequences are then converted into two-dimensional opcode sequences and fed into a pretrained model. Additionally, we transform the relevant features of malicious sample opcodes, registers, and assembly language comments into one-dimensional vectors as prefixes. This approach guides the model to better learn sample feature representations, ultimately enhancing the classification performance. To solve the latter two challenges, we introduce a combination of supervised contrastive learning and cross-entropy loss, utilizing support set computation to generate prototype vectors. This enables the model to effectively tackle emerging variations of malicious code and enhance its performance in few-shot scenarios. Experimental results demonstrate that our method outperforms existing malicious code classification models and achieves outstanding performance in few-shot experiments.