GPT-4 Meets TF-IDF: A Hybrid Approach for Detecting Spam Emails Using Machine Learning

Sherif Elmeligy Abdelhamid, Benjamin Davis, Dang Khoa Le · 2025

Spam emails, designed to flood email inboxes with unwanted, irrelevant, or harmful content, have evolved into a persistent cybersecurity threat. Modern spam emails span multiple categories, including advertising spam, which promotes products or services; scam spam, which attempts to deceive recipients into revealing sensitive information; and malicious spam, which carries phishing links or malware. Traditional techniques usually detect spam emails by analyzing specific word patterns or applying a rule-based filtering system. However, spammers now employ more sophisticated language to evade detection, making it harder to recognize spam emails. Traditional methods may overlook subtle contextual and behavioral indicators critical for distinguishing spam from legitimate emails. Recent advancements in large language models present new opportunities by capturing complex patterns related to urgency, requests for sensitive information, and deceptive linguistic structures. This study explores a hybrid approach that combines TF-IDF (Term Frequency-Inverse Document Frequency) with GPT-4 (Generative Pre-trained Transformer) derived features to enhance spam detection, especially with classifiers that benefit from diverse feature sets. We evaluate this combined feature set across multiple machine learning classifiers, including K-Nearest Neighbors, Decision Tree, Naive Bayes, and Random Forest. Our findings reveal that this hybrid approach significantly improves model performance, with Random Forest achieving high accuracy (99.86%), precision (98.46%), and recall (100%). Notably, other classifiers, like K-Nearest Neighbors and Decision Tree, benefit from the added contextual and behavioral cues. The results demonstrate the importance of a multifaceted approach for detecting spam emails, where linguistic patterns and contextual features together can capture the deceptive tactics used by spammers. This research demonstrates the value of combining term-based methods with advanced language model features, providing a novel, scalable, and effective solution for enhancing email security systems against diverse spam threats.

Read the paper · More papers on PaperTik