Proof: Pre-Training Model of Malware Family Classification Based on Active Defense

Shuai Guo, Hao Liu, Zhiyong Zhang, Shen Su, Zhihong Tian · IEEE Transactions on Sustainable Computing · 2025

Malware acts as a critical component in network attack and defense mechanisms. As Malware threats become more complex and diverse, timely automatic malware classification is urgently needed. Take into account consistency and system resource consumption, malware classification should be compatible with other security devices to form a comprehensive defense system. In this study, we proposed a pre-training framework Proof based on malware behavior captured by honeypoints inside and outside the protected system, and the framework’s performance in the categorization of malware families is evaluated. To describe malware, a structure called Malware Behavior Instruction Set (MBIS) was designed by selecting a subset of all behaviors captured by honeypoints. Then, a self-supervised pre-training model MBI2vec is applied to learn the internal mode of malicious code to guide the downstream malicious code family classification model composed of Bi-LSTM and attention mechanism. Finally, the Proof framework evaluated 846 malware samples from five malware families that we captured in the real world, and the f1-score of the classification result was 0.9554.

Read the paper · More papers on PaperTik