Performance Comparison of Feature Selection Methods for Machine Learning Models on DDoS Attack Dataset
Nicole Kayambu, Sanjeev Kumar · 2025
Distributed Denial of Service (DDoS) attacks have posed a major threat to the stability and security of computer networks and Internet. Detection and mitigation of DDoS attacks remain challenging due to the methods in which such attacks are launched. Machine Learning (ML) and Artificial intelligence (AI) are powerful tools that can be used to develop effective Intrusion Detection Systems (IDS) for the purpose of detecting and mitigating DDoS attacks. ML and AI models, however, require different features to be observed to create an efficient model. In this study, four different feature selection methods were used to determine the efficiency of selected features and to determine the effect they have on model accuracy-based on DDoS attack data from the CIC-DDoS 2019 dataset. The feature selection methods explored in this paper were Extra Trees Regressor, Decision Tree Regressor, Mutual Information, and Analysis of Variance. For this study, 200,000 samples of attack data were used. To address the data imbalance of this dataset, Synthetic Minority Oversampling Technique (SMOTE) was used in combination with Edited Nearest Neighbors (ENN). The model used to test the different feature selection methods is the Random Forest (RF) scheme, which is common among ML DDoS detection and mitigation applications. The model was evaluated using metrics such as precision, recall, F1-score, balanced accuracy, and the AUC score for each of the feature importance and selection methods. In this paper, the performance of these models was obtained and analyzed alongside their run times. Analysis of results showed a 75% decrease in testing times when using the top 5 features to train the RF model with SMOTE+ENN.