AutoPhish: A Grammar-based AutoML Approach to Learn Classifiers for Phishing Detection
João Guilherme Miranda, Mateus L. S. D. Barros, Tapas Si, Carlo Marcelo R. Silva, Pericles B. C. de Miranda · 2025
The increasing sophistication of cyber threats—particularly phishing—demands advanced and adaptive detection mechanisms to protect users and organizations. Traditional defenses struggle to keep pace as phishing techniques such as Clone Phishing, Spear Phishing, DNS-Based Phishing, and Man-In-The-Middle attacks evolve. Recent research has extensively applied machine learning (ML) models to phishing detection, emphasizing the importance of attribute selection and classifier optimization. While approaches using rule-based systems, ensemble models, and artificial neural networks (ANNs) have shown promising results, the reliance on static datasets and generalized models limits their effectiveness in dynamic, real-world scenarios. This article proposes AutoPhish, a novel phishing detection approach based on Grammatical Evolution (GE). GE is a grammar-driven genetic programming method capable of generating customized and optimized classifiers. AutoPhish implements a single objective GE aiming to produce classifiers which maximize F1−score. Our method is evaluated in realistic settings with imbalanced datasets and compared to traditional machine learning algorithms using performance metrics such as accuracy, recall, precision, and F1−score. The results show that classifiers evolved using AutoPhish consistently outperform most baseline methods, demonstrating strong potential for practical deployment. This study underscores the value of evolutionary computation in cybersecurity and advances the development of adaptive, high-performance phishing detection systems.