Evaluating Machine Learning Techniques for DNS Traffic Classification Using Three Publicly Available Datasets
Matee Witawasiri, Pongsarun Boonyopakorn · 2024
The Domain Name System (DNS) is a fundamental service for accessing the Internet. Following the COVID-19 pandemic, online usage through Internet networks has increased, as has cybercrime, particularly through new domain registration. As of June 2024, there had been 751 million registrations. In the past, researchers have studied DNS protection and security with various protocols, including DNSSEC, DNSCrypt, DoT, DoH, and DoQ. These protocols protect against hackers but also make it more difficult to verify their usage by administrators. The introduction of machine learning has helped to scan and isolate the transmission of data between benign and malicious DNS. Researchers have extensively studied and applied popular machine learning models to detect malware in DNS transmission, algorithms, domain creation, and DNS theft detection. In this study, we evaluated the accuracy, specificity, F-score, and learning time of various well-known machine learning models, including probabilistic, linear, decision tree-based, ensemble, and neural network models. We performed our analysis with Altair AI Studio on three publicly available datasets. We discovered that LightGBM outperformed almost all other methods. The neural network is unable to surpass the performance of the ensemble, and it also requires a significant amount of time for model training. In the field of machine learning, you cannot apply the model or solution universally; it depends on the dataset that requires analysis. Machine learning is not the only weapon in system security. It should complement other security measures to enhance efficiency.