Detecting Malicious URLs using Machine Learning

C Sivakumar, P. Lakshmi Sagar, K. Swetha, K. Lakshmipathi Raju, Gundluru Naveen Kumar, M N Nanditha · 2025

Identifying malicious URLs is an essential cybersecurity function since the timely and correct detection prevents catastrophic security breaches. Although approaches such as BERT (Bidirectional Encoder Representations from Transformers) have proven useful in deciphering intricate text patterns, they still fall short when it comes to the computational resources and scalability needed for real-time detection. This research utilizes advanced machine learning techniques— Extreme Gradient Boosting (XGBoost) and Light Gradient Boosting Machine (LGBM)—to create a light yet efficient model for detecting malicious URLs. A comparative evaluation of the two algorithms is presented based on performance metrics such as accuracy, precision, recall, and F1-score. Also, this work identifies the challenges brought about by varied and intricate URL sets and underscores the significance of sophisticated feature engineering methods in distinguishing between benign and malicious URLs. Hyperparameter tuning and cross-validation are used for model robustness improvement. Large-scale experimental results exhibit dramatic gains in detection accuracy with low computational overhead, and hence, real-time deployment becomes possible. The results of this research can be applied to other methods, providing scalable and secure cybersecurity solutions that make browsing safer.

Read the paper · More papers on PaperTik