Enhanced Spam Classification Model Through Text Processing and Fusion DL Models
S Kamalesh, N Niveatha, C M Yogita · 2024
The growing usage of communication via digital technology has resulted in excessive information spam to users and organizations, which is a major challenge. To tackle this problem, we present a hybrid model for spam detection that employs Multinomial Naive Bayes (MNB), Gated Recurrent Unit (GRU), and Long Short-Term Memory (LSTM) networks as its components. The ensemble model is based on traditional machine learning and deep learning techniques in order to improve the performance of spam classification. In The proposed model, raw text already goes through processes such as normalization, stop word removal and TF-IDF mathematics for better presentation and structure of data in readiness for the training of a model. The MNB model provides a baseline classification which is simple and fast, moreover effective for the spam classification. The temporal dependencies and context within the text is captured using GRU and LSTM networks. Additionally, to enhance classification performance, a majority voting approach is integrated at the final stage to combine the outputs from all the three models. It was observed that accuracy, precision, recall and F1 score demonstrate that the proposed approach ensemble approach performs better than the individual classifiers overcoming the challenge of spam detection to a very great extent. Other representative metrics such as ROC and confusion matrix, also confirm performance improvement of the proposed approach hence making it accessible and practical for real time systems on spam detection. The model proposed can accurately scale and implement a spam filtering mechanism geared towards spam that is constantly changing thus ensuring maximum detection and minimally false positives.