Bilingual SMS Spam Detection Using Deep Ensemble Learning
Hafsa Sultana, Jamal Uddin Tanvin, Fairooz Tasnia, Nusrat Sharmin · 2024
Spam messages have become a pervasive issue in digital communication, posing significant challenges to users and organizations alike. As the volume of text-based interactions continues to grow, particularly in low-resource languages, e.g. Bangla-English SMS messages, the need for effective spam detection systems has never been more critical. This paper introduces BESpam, a novel bilingual dataset comprising 5,730 instances of Bangla-English SMS messages aimed at addressing this gap. We explore the development of an ensemble model that leverages advanced transformer architectures alongside FastText to enhance spam classification performance. Our findings indicate that this proposed ensemble approach outperforms traditional machine learning methods and standalone transformer models, demonstrating the efficacy of combining contemporary natural language processing techniques with established feature extraction methods. The ensemble stacking method achieved a 99% accuracy, surpassing both the individual performances of mBERT (98%) and XLM-RoBERTa (98%), as well as traditional models like Logistic Regression and SVM, both of which scored 98%. This research not only advances the field of s pam detection but also serves as a vital resource for future studies focused on multilingual text processing, particularly in regions where SMS spam remains a pressing concern.