Phishing Email Detection Through Machine Learning and Word Error Correction
Deeksha H Kulal, Leul Shiferaw, Quamar Niyaz · 2025
Phishing is considered as one of the effective fraudulent activities on the Internet. Numerous machine learning (ML) based models have been implemented for detecting phishing emails using publicly available datasets (e.g. Nazario, Millersmile). These datasets exhibit noticeable flaws, such as poor grammar structure and incorrect word usage, which are learned as critical features by ML models. However, the grammatical quality and overall structure of phishing emails have been significantly improved with large language models (LLMs). This shift could reduce the effectiveness of traditional grammatical cues for identifying phishing emails, as the improved grammatical quality makes these emails appear more legitimate. In order to handle this challenge, we pose the following research question: Can an ML-based phishing detection model, enhanced with word correction and splitting techniques, effectively identify phishing emails? To answer this question, we implement a phishing detection system considering misspelled word correction and splitting of combined words during the data processing stage and employing state-of-art natural language processing techniques. To enhance the robustness of models, datasets from diverse sources and timelines are used for model training and deployment.