Phishing URL detection by leveraging RoBERTa for feature extraction and LSTM for classification

K. S. Jishnu, B. Arthi · 2023

A serious cybersecurity threat is phishing attacks, which use bogus URLs to fool users into disclosing critical information. These attacks affect human vulnerability and potentially result in significant data breaches and financial damages. Phishing attack prevention is vital for protecting individuals and organizations from falling for fraudulent schemes and maintaining internet security. This study introduces a novel method for phishing URL detection employing the RoBERTa transformer-based model for feature extraction and the LSTM for classification. RoBERTa extracts semantic and contextual information from URLs during the feature extraction step. By encoding the URLs into contextualized embeddings, RoBERTa successfully learns to represent the URLs in a way that captures their complicated meaning and surrounding context. The LSTM layer accurately categorizes URLs by capturing their sequential relationships using the collected features. The dataset is extensive, with 3,00,000 URLs. After comprehensive training and testing, the proposed system successfully differentiates between legitimate and phishing URLs with an accuracy of 97.14%. The results highlight the importance of incorporating LSTM for classification and RoBERTa for feature extraction in phishing URL detection. The findings of this study broaden phishing detection techniques and offer practical countermeasures.

Read the paper · More papers on PaperTik