Harnessing Language Models and Machine Learning for Rancorous URL Classification

Prabhuta Chaudhary, Ayush Verma, Manju Khari · 2024

As the internet expands exponentially, the need for robust network security has become increasingly critical. This increase in online activity has led to a rise in cyber threats, making individuals and organizations more vulnerable to security breaches. These threats include rogue websites, malware, and Trojan horses, which have evolved to become more difficult to detect, more automated, and more complex. These attacks and data breaches often involve malicious URLs. Therefore, identifying these harmful URLs is essential to assess network threats. This study focuses on using machine learning transformer-based models like BERT, RoBERTa, and XLNet and proposes a comparative study to classify URLs into benign, defacement, phishing, and malware categories. These models are chosen for their robustness, ability to handle the complexities of URLs, strong performance in various natural language processing (NLP) tasks, and transfer learning capabilities. The research utilizes a Kaggle dataset of 651,191 unique URLs for analysis. The models achieve high accuracy, precision, recall, and F1-score rates. Notably, the BERT model achieved the highest accuracy at 86.75%, while RoBERTa attained the highest precision score of approximately 87.32%, with identical recall rates for all models. This study contributes to safeguarding users from URL threats and enhancing their online experiences.

Read the paper · More papers on PaperTik