Multilingual Spam Email Classification Using LSTM with Convolutional Block Attention Module
Md Masum Rana, Mahmudul Hasan, Ashis Kumar Mandal, Muhammad Abu Rayan, Syed Shahir Ahmed Rakin, Md. Mehedi Hasan Jony · 2025
Email is a widely used communication medium for formal information sharing, but spam emails pose significant challenges by wasting time, consuming bandwidth, introducing security risks such as phishing and malware, and causing financial and reputational damage. Effective spam email classification is crucial for enhancing cybersecurity. However, most existing studies focus on a single language, limiting scalability. To address this limitation, we propose a multilingual spam email classification approach using a bilingual dataset (Bangla and English). Our method introduces a hybrid model, LSTM-CBAM, which integrates Long Short-Term Memory (LSTM) with the Convolutional Block Attention Module (CBAM). The model’s performance is evaluated using both a single train-validation-test split and 5-fold cross-validation. LSTM-CBAM achieves 96.43% accuracy on the Bangla dataset and 99.78% on the English dataset in the train-validation-test split. Under 5-fold cross-validation, it attains average accuracies of 94.78% and 99.81% for the Bangla and English datasets, respectively, outperforming existing machine learning and deep learning models. Additionally, explainable AI techniques, such as Local Interpretable Model-Agnostic Explanations (LIME), enhance the interpretability of the model, providing valuable insights into its decision-making process. The results demonstrate the effectiveness and scalability of the proposed multilingual approach for spam email classification.