Maximizing the Reliability of Machine Learning Based Invasion Detection Systems on a Modern, Unbalanced Dataset
D. David Neels Ponkumar, M. R. Arun, M. Hrushikesh Bharadwaj, G. Nagendrababu, A. Divijesh Reddy, R. Saravanakumar · 2025
More computers are networked and included into our everyday life as the Internet gains popularity. Common and less obvious server weaknesses might be leveraged by hackers to get into networks and launch intricate attacks. Popular computer security technology Intrusion Detection Systems (IDs) use machine learning on an already collected dataset. The datasets may not include the most recent information as they were produced on different networks during short times. They cannot sustain an onslaught with their bloat and lack of sufficient info. With inconsistent and obsolete information for infrequent attack pathways, modern intrusion detection systems are less effective. Five machine-learning intrusion detection models are presented in this work. Adaboost (AB), Random Forest (RF), K Nearest Neighbour (KNN), Decision Tree (DT), Gradient Boosting (GB) are used. To enhance Intrusion Detection Systems (IDS), the most recent security dataset CSE-CIC-IDS-2018 substitutes more regularly used older datasets. One chose an imbalanced dataset. The synthetic data model Synthetic Minority Oversampling Technique (SMote) helps to lower the imbalance ratio. While reducing false alerts and missed incursions, the method improves attack type-specific system performance. Small classes raise their data counts to match the average data size. Experimental data show that the advised approach significantly improves rare intrusion detection.