A Comparative Study on the Impacts of Data Leakage During Feature Selection using the CIC-IoT 2023 Intrusion Detection Dataset

Seshu Bhavani Mallampati, Hari Seetha · 2024

The growing adoption of devices based on the Internet of Things (IoT) has led to a significant problem in terms of security. To ensure the integrity of networked systems inside IoT contexts, it is crucial to use Intrusion Detection Systems (IDS). Feature selection is a vital component in optimizing the effectiveness of IDS. Its primary objective is determining the most relevant attributes contributing to precise threat identification. However, incorporating feature selection in the analysis raises the potential concern of data leaking. The phenomenon of data leaking, sometimes referred to as pattern leakage occurs when the training data includes information on the target variable, but equivalent data is not accessible during the model’s prediction phase. Therefore this research is conducted to analyze outcomes with and without data leakage on CIC-IoT2023 dataset. In the proposed work we used Gini index (GI) based feature selection to select the optimal features. Then the model was trained and tested by using traditional Machine Learning (ML) and Deep Learning (DL) models such as Decision Tree (DT), Extra Tree (ET), Light Gradient Boosting Machine (LGBM), Extreme Gradient Boosting Machine (XGBM), Multi-layer perceptron (MLP), Deep neural networks (DNN), Gated neural networks (GRU) and Recurrent neural networks (RNN). It transpires when data originating from the testing part is improperly applied in the training process, resulting in overfitting and exaggerated accuracy values and therefore feature selection after data splitting mitigates this issue in IoT networks.

Read the paper · More papers on PaperTik