A Malicious Domain Name Detection Method Based on Variational Autoencoder
Donglin Ma, Xichuan Wu · 2024
Malicious domain names are a critical area of focus within cybersecurity research. The detection and analysis of these domains are crucial for thwarting network attacks and preventing fraud, thereby safeguarding user security and data privacy. Addressing the issues related to the sparse occurrence of malicious domain names within specific Domain Generation Algorithm (DGA) families, and the challenges in detecting them in other families, this study proposes a novel method for identifying malicious domain names using a Variational Autoencoder (VAE). This approach begins by preprocessing the domain name under investigation, eliminating extraneous and irrelevant details, and normalizing the data. Subsequently, the Word2vec technique is employed to convert the domain name into a vector, which is then fed into the VAE network. The model's optimal weights are determined by optimizing the objective function through backpropagation. Next, the mean and variance derived from sampling the latent space are utilized to compute the reconstruction probability. A threshold-based classifier is subsequently applied to categorize the domain names into either legitimate or malicious categories. The method is trained and tested using datasets from Cisco and the 360 open lab. Experimental outcomes demonstrate that this method achieves a high level of detection accuracy, confirming its robust performance and effectiveness in recognizing challenging or infrequently occurring malicious domain families.