Rethinking Phishing Detection: How Dataset Quality Affects Model Generalization

Pedro Afonso, Eva Maia, Ivone Amorim, Isabel Praça · 2025

Phishing remains a pervasive cybersecurity threat, prompting the development of numerous detection models and the creation of various public benchmark datasets for evaluation. However, despite the apparent diversity of these datasets, their quality and consistency remain largely underexamined. In this paper, we conduct a comprehensive analysis that jointly evaluates the generalization ability of phishing URL detection models and the integrity of widely used datasets. Our findings reveal critical data quality issues, such as high internal duplication, inconsistent labeling both within and across datasets, substantial URL overlap, and severe domain overrepresentation, that undermine model evaluation and encourage overfitting to dataset-specific artifacts. To quantify the impact of these flaws, we perform cross-dataset generalization experiments comparing a baseline character-level CNN with the state-of-the-art URLNet architecture. Results show that despite URLNet’s complexity, it offers limited generalization advantage under realistic evaluation settings. These insights highlight that the limitations of current benchmark datasets, rather than model design alone, form a significant bottleneck in advancing robust phishing detection.

Read the paper · More papers on PaperTik