Shielding Cyberspace: A Machine Learning Approach for Malicious URL Detection

P Anusri, Swati Sneha, Prabhu Natarajan, Thangavel Palanisamy, Senthil Kumar Thangavel · 2024

Amidst the evolving landscape of cybersecurity threats, adept hackers deploy sophisticated strategies to exploit end-to-end technology and capitalize on human vulnerabilities. Hacking techniques including social engineering, phishing, and pharming are used, and manipulating Uniform Resource Locators (URLs) has become a central hub for malevolent activity. The purpose of the research is to determine how well various models work with sophisticated machine learning techniques to identify dangerous URLs. A sizable dataset merged from Mendeley, Phinh Tank and Kaggle is used in the study. By including four distinct types of URLs—malware, benign, phishing, and defacement—the dataset broadens the scope of the research. Lexical features are seamlessly integrated into the proposed model, which adeptly combines regression and classification techniques. Among them, the XGBoost Classifier is particularly noteworthy as a high performer, demonstrating robustness and interpretability with an outstanding accuracy of ${96.6\%}$. This research adds distinctive perspectives to the discussion of the Random Forest Classifier’s function and emphasizes the significance of customized model selections in tackling the complex problems associated with malicious URL detection in the context of cybersecurity.

Read the paper · More papers on PaperTik