A Self-Supervised Contrastive Learning-Based Method for Portable Executable Malware Classification

Timothy Miskell, Ting-Li Huoh, Yan Luo, Peilong Li · IEEE Access · 2026

Malicious software presents significant risks to computer systems, networks, and sensitive data, making malware detection a critical cybersecurity challenge. Labeling malware data not only requires expert knowledge but is also time-intensive. Given the constantly evolving threat landscape, an ever increasing amount of malware remains unlabeled, and those samples that are labeled may be incomplete or inconsistent. In this work, we introduce a self-supervised method using contrastive learning to perform static analysis and classification of malicious Portable Executable (PE) files, reducing the dependency on labeled data. We also develop data augmentation techniques that generate multiple augmented views and design PE-specific augmentation operators to be used during self-supervised learning such as shuffle, encryption, and compression based on the IMAGE_SECTION_HEADER. Our method is built upon raw PE byte sequences extracted from a large-scale publicly available dataset, SoRel-20M, which contains 20 million PE samples. Utilizing a two-stage framework consisting of a self-supervised contrastive learning pre-training phase followed by a supervised fine-tuning phase with limited amounts of labeled data, our model learns label-free invariant representations of the PE structure and as a result outperforms a traditional supervised Convolutional Neural Network (CNN), achieving a macro-averaged F1 score of 78.6% with only 10% of the labeled data. In addition, our method only requires the first 1 KB of header data, whereas the supervised baseline requires 1 MB of the underlying PE header.

Read the paper · More papers on PaperTik