A Comparative Study on Detecting Phishing URLs Leveraging Pre-trained BERT Variants

Chanchal Patra, Debasis Giri, Tanmoy Maitra, Bibekananda Kundu · 2024

In the cybersecurity sector, phishing attacks are one of the new security concerns that have received a lot of attention lately. Phishing efforts that try to steal confidential data or spread malware frequently use malicious universal resource locators (URLs). Accurately identifying phishing URLs is crucial. There are several methods available now for identifying phishing websites. Nevertheless, as attackers might adapt their strategies to evade recently implemented detection techniques, phishing website detection continues to be a research focus. Our study presents an enhanced model for detecting phishing URLs, which is based on optimizing the BERT model family to identify phishing URLs with more precision. BERT variants such as ALBERT, DistilBERT, and RoBERTa are used in the phishing website detection model presented in this research. These models perform significantly better than existing detection techniques and demonstrate impressive accuracy. In the experiment, we found that the transformer-based BERT model achieved the highest accuracy, precision, recall, and F1-score of 99.29%, $\mathbf{9 9. 2 6 \%}, \mathbf{9 9. 2 7 \%}$, and $\mathbf{9 9. 2 7 \%}$, respectively. As per the findings, out of all the models that are now accessible, this one performs the best.

Read the paper · More papers on PaperTik