An Experiment on Feature Selection Using Logistic Regression
Raisa Islam, Subhasish Mazumdar, Md. Rakibul Islam · 2024
This study investigates feature selection using L1 and L2 regularization methods associated with logistic regression (LR) by leveraging its coefficient-based feature ranking. This research aims to optimize the feature set, enhancing model explainability and performance. The CIC-IDS2018 dataset was selected for the experiment, partially due to its huge volume and the inclusion of problematic classes. The research undertakes a detailed analysis, initially excluding one of the problematic classes and subsequently including both. Feature ranking was performed first with L1 followed by L2 regularization; and thereafter comparing performance of LR with L1 (LR+L1) against LR with L2 (LR+L2) by varying the feature set sizes for each ranking. Through comparative analysis, the outcome reveals no significant discrepancy in accuracy upon finalizing the feature set. Adopting a synthesis approach, the research selects features common to both L1 and L2-derived sets, and this optimized set was tested on more complex models such as Decision Tree and Random Forest. Results indicate a marginal mean accuracy reduction, by 0.8% and 0.6% respectively, while significantly reducing the feature set by 72%, regardless of the incorporation of the problematic class. Additionally, the confusion matrix is reported to facilitate the calculation of standard metrics: accuracy, precision, recall, and F1-score.