PHILDER: Lightweight Framework for Intelligent Phishing Detection on Resource-Limited Devices

Guilherme Dantas Bispo, César Augusto Borges de Andrade, Gabriela Mayumi Saiki, Raquel Valadares Borges, André Luiz Marques Serrano, Geraldo P. Rocha Filho, Vinícius P. Gonçalves · IEEE Access · 2025

The increasing sophistication of phishing attacks represents a significant challenge to cybersecurity, particularly in the context of fraudulent emails crafted to deceive users and steal confidential information. This paper proposes a deep learning–based approach for the automatic detection of phishing in emails, the PHILDER (PHishing Intelligent Lightweight DEtection on Resource-limited devices) model. The study employs a dataset composed of real emails extracted from PhishTank and the SpamAssassin Public Corpus, spanning both legitimate and fraudulent samples. To address the natural imbalance between spam and phishing emails, two strategies were applied: undersampling, which reduces the number of legitimate emails to balance the classes, and oversampling using SMOTE (Synthetic Minority Over-sampling Technique) to generate new synthetic samples of the minority class. Five models (ALBERT, DistilBERT, MobileBERT, MiniLM, and TinyBERT) were trained: one using the original dataset, one with undersampling, and a third with oversampling. The experimental results show that the oversampling approach had to be discarded due to computational cost issues, the tendency to overfit, and inferior performance compared to the other two approaches. The main contributions of this paper include the application of computational efficiency metrics for model evaluation, which not only consider traditional metrics but also enable assessment in real-world applications under hardware computational constraints, and the exploration of different data composition strategies for training. The use of PHILDER (TinyBERT combined with data balancing techniques) demonstrates a promising approach for the automatic detection of phishing in emails, contributing to strengthening digital security against social engineering attacks.

Read the paper · More papers on PaperTik