Analysis and classification of spam email using artificial intelligence to identify cyberthreats

Francisco Jáñez Martino · 2023

In this Thesis, we propose new models, methodologies, approaches and datasets to analyze and identify rising cybertreats in spam emails.Motivated by our collaboration with the Spanish National Institute of Cybersecurity (INCIBE), we focus our efforts on developing applications and conducting studies to improve the earlier detection of these risky and harmful emails.Several of the contributions presented in this dissertation are planned to be incorporated in tools developed by INCIBE to launch more detailed and earlier warnings to organizations and citizens about potential risks associated with spam emails.Our approach heavily relies on the application of Natural Language Processing, as well as Machine and Deep Learning techniques, mainly centred around supervised learning methods.on Chernyavskiy et al. ( 2020)).We created a novel dataset called Persuasion Sentence in Spam Emails (PerSentSE) containing annotated sentences based on binary, i.e., persuasion or not, and multilabel classification.For the multilabel approach, we considered eight persuasion techniques: Appeal to authority, Appeal to fear/prejudice, Doubt, Exaggeration or minimization, Flag-waving, Loaded Language, Name Calling or Labeling and Repetition.We collected spam emails from the Bruce Guenter repository.Lastly, our objective was to create an intelligent system capable of detecting potentially risky spam emails for both individuals and organizations.We created Spam Email Risk Classification (SERC-4K), a novel dataset encompassing spam emails classified in two categories based on the potential risk for users due to their content, low and high risk, as well as a continuous value from 1 to 10.The dataset is composed of two subdatasets, one with spam emails shared by INCIBE (SERC-I) and another collected from the Bruce Guenter repository, Spam Archive (SERC-BG).SERC-I contains English and Spanish emails, while in the case of SERC-BG almost all of them are written in English.Firstly, our approach attempted to extract potentially worthy features from headers, text, attachments, URLs and protocols (56 features in total).Then, the sets of features along with three popular Machine Learning classifiers were evaluated resulting in Random Forest as the highest classifier-performance (0.914 of F1-score).Regarding regression approach, the Random Forest Regressor achieved the lowest MSE (0.579).Our work also included a feature evaluation to determine the importance of each feature and set.In the design of our methodologies, we have considered the influence of the dataset shift, as well as the spam domain is and adversarial environment.Our email processing sought to overcome some spammer strategies such as image-based spam and hidden text.

Read the paper · More papers on PaperTik