A Transformer Embedding-Based Model for Malicious DGA-Generated Domain Detection

Suleiman Y. Yerima, P. Vinod, Khaled F. Shaalan · 2024

Domain Generation Algorithms (DGA) techniques are frequently used by cybercriminals to make malware more evasive by generating thousands of fake domain names that prevent established security solutions from effectively discovering or disrupting attacks. The proliferation of DGA has increased the difficulty of distinguishing genuine domains from fake ones, thus necessitating the investigation of new approaches to improve detection efficiency. Therefore, in this paper we propose a method for detecting malicious DGA domain names based on the Generative Pre-Trained Transformer (GPT) Large Language Model (LLM). Our approach utilizes GPT to learn dense vector representations with the aim of enabling more effective detection of the malicious domains by machine learning algorithms. The system is implemented and evaluated using a publicly available dataset and compared to widely-used techniques such TF-IDF, Bag-of-Words, N-grams, word2vec as well as to two variants of BERT. The results of our experiments demonstrates the superiority of the GPT-based technique over the others, with the best performance recorded by the XGboost classifier achieving the highest accuracy of 93.1% compared to the other machine learning algorithms.

Read the paper · More papers on PaperTik