Research on APT Malware Detection Based on BERT-Transformer-TextCNN Modeling

Jian Zhang, Shengquan Liu, Zhihua Liu · 2024

In recent years, with the development of the Internet, APT malware detection remains an important issue for society. The lengthy nature of API sequence text leads to polysemy issues with distant repeated words. Traditional deep learning models must improve extracting the long-term dependencies of entire API sequences. During an APT attack, attackers combine benign and malicious sequences to conceal their malicious intent, resulting in localized malignancy. In response to these issues, the paper proposes A BERT-Transformer-TextCNN model for attack detection. BERT encodes API sequences, addressing the polysemy issue of distant repeated words. The Transformer-encoder extracts long-term dependencies, resolving the global feature issues of long sequences, and TextCNN is applied to capture local sequence information to solve the problem of localized malignancy. Based on actual APT malware samples, BERT encoding achieves higher precision than traditional encoders. Employing a Transformer-Encoder to extract long-term dependencies yields higher precision than using BiLSTM to extract long-term relationships. When adding features related to local feature correlation, the precision increased by 3.39%.

Read the paper · More papers on PaperTik