Explainable Email Spam Detection: A Transformer-based Language Modeling Approach
Mohammad Amaz Uddin, Md. Mahiuddin, Ali Asgar Chowdhury · 2024
Email spam, also known as junk email, is any unwanted or irrelevant message that is delivered to a large number of recipients without the recipient's consent. Email communication is undoubtedly efficient and convenient, but there are serious risks associated with its rise in unwanted spam emails. These risks include resource waste and potential cybersecurity breaches. It also serves as a backdoor for malware infections, phishing attempts, and other cyber threats. Transformer-based models have a positive effect on the cybersecurity sector in order to address those cybersecurity concerns. In this research, we use a Bidirectional Encoder Representations from Transformers (BERT) variant, the DistilBERT model, to detect spam emails. For this purpose, we utilize the benchmark Enron spam email dataset, applying several preprocessing techniques to address the noisy data. Our model demonstrates high accuracy, achieving 99.18%. Additionally, to interpret the model's decision-making process, we used Transformers Interpret. This explainability tool provides insights into how our proposed model works and makes its decisions throughout the experiment.