Generalizability Assessment of Learning‐Based Intrusion Detection Systems for IoT Security: Perspectives of Data Diversity
Zakir Ahmad Sheikh, Narinder Verma, Yashwant Prasad Singh, Sudeep Tanwar, Abdulatif Alabdulatif · Security and Privacy · 2025
ABSTRACT Machine learning (ML) and deep learning (DL) models have become vital tools in Intrusion Detection Systems (IDS), yet their effectiveness depends heavily on the quality and distribution of training data. This study investigates the impact of dataset size and dataset balance on the performance of ML and DL models using the CIC‐IDS 2017 dataset. Five subsets (20%, 40%, 60%, 80%, and 100% of the dataset) were created to assess learning models across varying dataset sizes. Four models, including Random Forest (RF), Artificial Neural Network, Convolutional Neural Network (CNN), and CNN+Long‐Term Short Memory (CNN+LSTM), were trained and evaluated on these subsets, focusing on precision, recall, and F1‐score. To test model generalizability, a synthetic dataset of 20 million over‐sampled samples was generated using Synthetic Minority Oversampling Technique, followed by manual under‐sampling to create a balanced dataset of 1.5 million samples with approximately 100 000 samples per attack class. Upon generalizability assessment of already trained models on the synthetically generated datasets, CNN+LSTM consistently outperformed other models in generalizability but utilized more time for training and testing in each case. The RF showed the weakest generalizability performances but was the fastest in both training and testing scenarios. Moreover, to evaluate the importance of the dataset in general and the balanced dataset in particular, we have also considered the NSL‐KDD dataset and evaluated all four learning models for multiple classifications and binary classification. Our results highlight the importance of the dataset, balanced dataset, and structure of learning models.