Comparative Analysis of Machine Learning Models for Real-Time Network Threat Detection

Sahil Varma Penmetsa · 2024

This research explores the efficacy of various machine learning models in detecting network threats, focusing on their performance across known, similar, and new attack datasets. I evaluated models including Random Forest, XGBoost, Logistic Regression, and Gaussian Naive Bayes, utilizing different preprocessing pipelines such as scaling, PCA, and correlation-based feature selection. Our findings reveal that while most models achieve near-perfect performance on known and similar threats, their ability to generalize to new threats is limited, with Gaussian Naive Bayes showing the highest F1 score on new attacks. Furthermore, Logistic Regression and Gaussian Naive Bayes models demonstrate the fastest detection times. The study highlights the importance of preprocessing techniques and model selection in enhancing threat detection systems. Deployment strategies for these models are discussed, emphasizing potential online and offline applications. Despite achieving significant insights, the study acknowledges the limitations posed by dataset size and computational resources, suggesting areas for future improvement.

Read the paper · More papers on PaperTik