Advancing Bilingual SMS Spam Detection: A Comparative Exploration of Machine Learning and Deep Ensemble Techniques
Hafsa Sultana, Jamal Uddin Tanvin, Nusrat Sharmin, Mohammad Shamsul Arefin, Md. Mahbubur Rahman · 2024
Spam messages are increasingly becoming a major issue in digital communication, creating significant challenges for both users and organizations. With the rise in text-based interactions, especially in low-resource languages like Bangla-English SMS, there is a growing need for more effective spam detection systems. This paper presents BESpam, a unique bilingual dataset that contains 5,730 Bangla-English SMS messages, specifically designed to address this challenge. In the literature, various machine learning and deep learning techniques align with feature extraction techniques from natural language processing used for spam SMS detection. Still, no studies have been found that focus on comparing deep and machine learning approaches. As a further contribution, we propose an ensemble model that combines advanced transformer architectures with FastText to improve the accuracy of spam classification. Our results demonstrate that this ensemble approach outperforms both traditional machine learning models and individual transformer-based methods. The stacking ensemble achieved an accuracy of 99%, exceeding the performance of standalone models such as mBERT (98%) and XLM-RoBERTa (98%), as well as traditional classifiers such as logistic regression and SVM, both of which also scored 98%. This work provides insight analysis of machine learning and deep learning in terms of bilingual SMS spam detection, in addition to advancing spam detection methodologies.