WebGuardML: Safeguarding Users with Malicious URL Detection Using Machine Learning
Md. Shamiul Islam, Nabila F. Rahman, Julkar Naeem, Abdullah Al Mamun, Farjana Akter, Sumaiya Jahan, Sabila Rahman, Samin Yasar, Md. Mushfiqul Haque Omi · 2023
The prevalence of cyber security threats stemming from rogue URLs is widespread. A malicious URL contains phishing and spam to initiate assaults. Such websites attract innocent people who fall for frauds like identity theft and financial loss. Such threats must be identified and addressed promptly. The main goal of this research is to create a machine-learning-based malicious URL detection system. The method uses Extreme Gradient Boosting, Categorical Boosting, Decision Trees, Random Forest, Extra Trees, Gradient Boosting and others to detect hazardous URLs. The process begins with dataset collection, extracts domain attributes, path components, length features and addresses class imbalance. The best model is chosen after lengthy training and testing, considering accuracy, precision, recall and F1-score. Hyperparameter tuning and ensemble techniques are optional optimization procedures. The process of testing serves to assess the performance of a system in real-world scenarios. After deployment, the system is monitored and maintained to keep up with emerging dangerous URL trends. Overall, this research improves cyber security by automating URL protection.