An Interpretable Fine-Tuned BERT Approach for Phishing URLs Detection: A Superior Alternative to Feature Engineering
Yi Wei, Masaya Nakayama, Yuji Sekiya · 2024
Phishing is a kind of cybercrime that deceives online users into disclosing confidential information, leading to identity theft and financial loss. Attackers typically use a wide range of vectors to spread phishing links, further increasing the reach and effectiveness of these malicious campaigns. As this threat continues to grow, artificial intelligence strategies have emerged as a promising solution for phishing detection in recent years. However, traditional methods that require substantial manual feature engineering have been observed to be dataset-dependent and lack generalization to unknown or newly evolving phishing attacks. To address these limitations, this study proposes a fine-tuned BERT model-based phishing detection approach utilizing a novel tokenization method and providing interpretable outputs. By thoroughly searching the latest publicly available datasets, the generalization capabilities of both feature engineering methodologies and the proposed approach are tested. The experimental results demonstrate its effectiveness, significantly surpassing feature engineering, and ensuring the approach remains robust in generalizing across different datasets.