Comparative Analysis and Optimization of Spam Filtration Techniques Using Natural Language Processing

Aditya Kumar Mehta, Suresh Kumar · 2024

Spam emails, messages, and content remain a persistent nuisance in the digital landscape, demanding effective and efficient classification methods. This paper presents an exploration of several techniques for spam classification using Natural Language Processing (NLP) techniques, we aimed to optimize the performance of spam filtration algorithms Through rigorous experimentation and analysis. In Spam classification, the spam Text is converted into a machine-readable form by NLP preprocessing tasks such as Tokenization, stemming, Lemmatization. After preprocessing the dataset is converted into a transformed text on which we will train our model to detect spam and ham messages using various classifiers. We trained our dataset on different Machine Learning (ML)classification models like Naive bayes, Random Forest, Extra trees classifier. We were able to achieve 96.7% accuracy and 100% precision using Extra trees classifier and 97.1% accuracy using multinomial Naïve bayes model. We also explored spam classification using pretrained ‘BERT’ model Which can predict spam or ham messages with contextual understanding, our spam detector model using BERT could classify spam messages with 86.0% accuracy on UCI spam messages dataset. This research contributes to the ongoing efforts to combat spam by presenting a comprehensive exploration of optimization techniques in the context of NLP-based spam classification.

Read the paper · More papers on PaperTik