A Comparative Analysis of Count-Based and Inference-Based NLP Models in Spam Email Detection Task

Longhua Zou · 2024

In the era of digital communication, distinguishing spam from legitimate emails has become a crucial task. This study provides a comparative analysis of count-based and inference-based Natural Language Processing (NLP) models in spam email detection. Focusing on the TFIDF-SVD-RF model and the Word2Vec-BiGRU-NN model, the research utilizes a real-life email dataset. The count-based model achieved an accuracy of 99.10% and an F1 score of 96.58%, while the inference-based model reached an accuracy of 98.39% and an F1 score of 93.92%. Through a detailed evaluation of methodology, text representation, and classification accuracy, the study reveals significant insights into the strengths and weaknesses of each model. Conclusively, this research underscores the importance of model selection in NLP tasks, offering guidance for future developments in spam email detection and similar text classification challenges.

Read the paper · More papers on PaperTik