Generative Phishing URL Detection Based on Large Language Model
Bo Zhou, Jia Liu · 2025
Phishing is a form of cyberattack in which the attacker pretends to be a trustworthy entity to trick users into providing sensitive information. The most common form of phishing is URL-based phishing, which uses fake URLs to trick users into entering sensitive information, leading to serious security risks. The core of the phishing URL detection task is to accurately identify forged URLs and prevent potential threats. Although existing detection methods have achieved certain results, there are still some shortcomings, including only giving labels without giving reasons for judgment, and not being able to handle complex and changeable URLs. To alleviate these deficiencies, and inspired by the deep semantic understanding capabilities of large language models, we propose a phishing URL detection model based on large language models. In this work, we transform the traditional classification paradigm in phishing URL detection into a generation paradigm to increase the interpretability of model judgment and improve the detection performance. At the same time, by doing so, we also fill the gap of underutilization of large models in phishing URL detection. We have conducted sufficient experiments, and the results show that the proposed model is superior to existing models in terms of detection accuracy and other metrics. At the same time, the proposed model requires less training data and also shows better generalization ability.