Natural Language Processing-Enhanced Machine Learning Framework for Comprehensive Phishing Email Identification
Siva Theja Kopparaju, Cesar Chavarriaga, Erick Galarreta, Sajal Bhatia · 2024
Phishing emails remain a significant cybersecurity threat, bypassing traditional rule-based detection methods. This paper proposes a novel Machine Learning (ML) and Natural Language Processing (NLP) based approach for enhanced phishing email detection. The core of the model utilizes Term Frequency-Inverse Document Frequency (TF-IDF) for text vectorization and a Multinomial Naive Bayes (MNB) classifier. This combination facilitates the identification of semantic cues and enables dynamic adaptation to evolving phishing tactics. The model prioritizes a balance between sensitivity and specificity, ensuring accurate detection of phishing attempts while minimizing false positives. The proposed PED (Phishing Email Detection) model achieves an accuracy of ${9 7. 2 5 \%}$ demonstrating a high overall accuracy in correctly classifying both legitimate and phishing emails. High F1-scores for two email categories are indicative of the proposed model’s robust performance in balancing precision and recall. The research offers a practical solution for cybersecurity professionals, contributing to a more secure email environment.