Harnessing BERT for Advanced Email Filtering in Cybersecurity
Avishek Majumder, Tanjim Mahmud, Tikle Barua, Nusratul Jannat, Md. Faisal Bin Abdul Aziz, Dilshad Islam, Rishita Chakma, Mohammad Shahadat Hossain, Karl Andersson · 2025
In the realm of digital communication and cybersecurity, the identification and filtering of ham and spam messages pose significant challenges due to the overwhelming volume of unsolicited and unwanted emails. This paper presents an in-depth analysis of various ML and DL techniques for efficient and robust spam detection systems, critical for enhancing cybersecurity defenses. Seven ML algorithms-RF, LR, SVM, XGBoost, GB, NB, and KNN-along with four deep learning models-CNN, LSTM, BiLSTM, and RNN-are evaluated for their effectiveness in spam classification. Additionally, we fine-tuned the BERT model, achieving a ground breaking accuracy of 99.37%, surpassing the 99.14% accuracy of the Bidirectional and Auto-Regressive Transformers (BART) model reported in recent research. Using a publicly available dataset of labeled ham and spam messages, the models were trained and tested, and their performance was assessed based on accuracy, precision, recall, and F1 score. The results demonstrate the superiority of the BERT model in spam detection, setting a new benchmark for cybersecurity. The success of BERT is attributed to its advanced capability to capture intricate patterns and contextual information, essential for distinguishing legitimate messages from spam, thus bolstering cybersecurity efforts.