A Exploring Hybrid Classifiers Through Stacking and Voting Ensembles for Robust Multilingual Spam Classification
P.R. Sudha Rani, P.L.R. Kameswari · 2025
Spam detection in multilingual environments is a challenging task due to the diversity in language structures, vocabularies, and contextual nuances. This paper explores the effectiveness of hybrid classifiers using stacking and voting ensemble models to enhance spam classification across multiple languages. Using both TF-IDF and Word2Vec for feature representation, the work processes multilingual datasets, including English, Hindi, German, and French. To address class imbalance, SMOTE is applied, ensuring balanced and efficient training datasets. Later several ML models applied for four languages separately. In the next phase, top-performing classifiers from each language are identified and combined to form ensemble models. Two ensemble models namely stacking ensemble classification and voting ensemble classification applied. The experimental results shown that the ensemble approaches significantly improve classification accuracy and generalization, providing a robust solution for multilingual spam detection.