Enhancing Security in SMS by Combining NLP Models Using Ensemble Learning for Spam Detection with Image Steganography Integration
Aditya Kumar, C. Fancy · 2023
Spam SMS messages are a prevalent problem in today's world and have become a source of annoyance for users. This research study proposes a novel approach to detect SMS spam using Natural Language Processing (NLP) and creating multiple models and apply ensemble learning on them. The proposed method includes pre-processing the data with NLP techniques such as stopwords removal, stemming, and lemmatizing, extracting relevant features from the text data, and transforming it into numerical representations. The performance of the proposed method was evaluated on a real-world dataset and compared to traditional machine learning algorithms such as Naïve Bayes, SVM. The multiple classifiers are trained on the transformed data, and an Ensemble Learning algorithm is applied to combine their predictions to obtain a more accurate result. The resulting model KNR gave higher accuracy than traditional models. The custom model was integrated into a custom Image Steganography tool which can hide textual data into an image using 1 bit Lease Significant Bit (LSB) technique and while decrypting a hidden message, the custom model would detect if the textual data is spam or not. The purpose of the Image Steganography tool is to provide users with a secure method of communication. For example, sender can hide sensitive data they wish to send to the receiver by hiding it inside an image and the receiver decrypting it to get the original message. If a malicious actor finds out about this communication method and tried to send spam links or messages by hiding it in an image and sending it to its victim, the custom spam detection model will detect it.