Optimizing SMS Spam Detection: Comparative Analysis of Hybrid Voting Ensembles and Bi-LSTM Networks with Stratified Cross-Validation
Arifur Rahman, Shahriar Parvej, Kazi Saeed Alam, H. M. Abdul Fattah · 2024
SMS spam has considerable impacts, affecting users and service providers alike, leading to a substantial trust deficit between both parties. This study aims to add to the expanding knowledge base in SMS spam detection, emphasizing the potential of various classifier algorithms, both soft and hard voting ensemble methods, and a customized Bi-LSTM (Bidirectional Long Short-Term Memory) model as effective tools in spam SMS classification. The classifier algorithms encompass Random Forest, Support Vector Machine, CatBoost, LightGBM, Decision Tree, XGBoost, Logistic Regression, and K-Nearest Neighbors. These models were employed to evaluate the benchmark dataset SMS Spam Collection sourced from the UCI Machine Learning repository. Additionally, we employed 5-fold Stratified cross-validation, ensuring a more thorough assessment of the model’s performance and reducing the risk of overfitting by validating the model on various representative subsets of the data. After cross-validation, the top three classifier algorithms were selected based on their performance metrices. Subsequently, both soft and hard voting ensemble were employed on these algorithms. The customized Bi-LSTM model was trained for 100 epochs as the loss curves showed minimal change after this point. This model was compiled using the adam optimizer (learning rate=0.0001) along with the binary crossentropy loss function. Prior to cross-validation, the Bi-LSTM model attained an accuracy score of 99.30%, the highest among all proposed models. Following cross validation, the soft voting ensemble on top three models achieved the highest accuracy rate of 93.77% among all recommended models, with significant precision, recall, F1-score, MAE, MSE, RMSE, RAE, RRSE, and AUC metrics of 93.50%, 93.77%, 93.51%, 6.23%, 6.23%, 24.96%, 4.30%, 8.28%, and 94.96%, respectively. The hard voting ensemble using the top three models achieved accuracy, precision, recall, F1-score, MAE, MSE, RMSE, RAE, and RRSE rates of 93.49%, 93.30%, 93.49%, 92.89%, 6.51%, 6.51%, 25.52%, 4.30%, and 8.25% respectively, placing third overall. On the other hand, LightGBM model secured the fourth-best performance, achieving accuracy, precision, recall, F1-score, MAE, MSE, RMSE, RAE, RRSE, and AUC results of 93.30%, 93.55%, 93.30%, 92.26%, 6.70%, 6.70%, 25.88%, 31.00%, 78.75%, and 92.80%, respectively.