Leveraging Large Language Models for Malware Detection Through PE Header and Library Analysis

Matthew Maximillian Tane, Nicholas Daniel Wijaya, Hidayaturrahman Hidayaturrahman · Procedia Computer Science · 2025

In 2024, ransomware attacks reached unprecedented levels globally, with over 5,400 incidents reported. Malware attacks continue to cause critical damage, causing high business and organization losses. As time goes on, our reliance on Large Language Models for everyday tasks continues to grow due to its high speed and accuracy. In this research, we investigate the accuracy and effectiveness of Large Language Models as malware detectors by training 4 open-source models—BERT, GPT-2, TinyLlama and Qwen 2—to classify malware and benign samples on PE file headers and libraries extracted from an imbalanced combination of the BODMAS and Dike Dataset, where 98.35% of the samples are malware and 1.65% are benign. Bayesian optimization was used to find optimal parameters. TinyLlama performed best with an accuracy of 99.69% and a balanced accuracy of 91.69%. All models achieved accuracy and balanced accuracy greater than 99.6% and 90.3% respectively. TinyLlama, whose attention weights and gradients were analyzed, was able to recognize structural patterns and identify unusual headers and libraries within the input text. In general, the model relied more on libraries and structural patterns compared to the section headers. In this paper, we demonstrate large language models’ relatively high performance in detecting malware despite an imbalanced dataset.

Read the paper · More papers on PaperTik