Shallow and Deep Learning Based Approaches for Malicious URL Detection: A Comprehensive Performance Evaluation

Sujay Dinakar, K. Deepak · 2024

In the realm of cybersecurity, the relentless surge in cyber threats, particularly emanating from the proliferation of malicious URLs, poses a formidable challenge. This study delves into the pressing need for robust detection mechanisms to thwart these threats effectively. Through meticulous dataset curation involving 30,000 URLs, employing a diverse array of extraction methods, and leveraging advanced machine learning models this research comprehensively evaluates malicious URL detection performance. Notably, the ensemble Random Forest model emerges as the top performer, showcasing its pivotal role in cyber threat classification. Additionally, the incorporation of novel dataset creation methods and effective data preprocessing techniques, such as the Synthetic Minority Over-Sampling Technique (SMOTE), underscores the significance of this research endeavor. The study’s findings underscore the efficacy of Random Forest in accurately discerning between benign and malicious URLs. Moreover, insights gleaned from model assessment metrics, including precision, recall, F1 score, and confusion matrices, shed light on the models’ generalization and stability. Overall, this research contributes to fortifying digital environments against evolving cyber threats, underscoring the pivotal role of machine learning in bolstering cybersecurity defenses.

Read the paper · More papers on PaperTik