Digital Forensics and Machine Learning to Fraudulent Email Prediction
Norah Al-Ghamdi, Tahani Alsubait · 2022
E-mail is widespread in the modern commercial environment, providing an appropriate and efficient method for communication. However, today’s e-mail security threats are multiplying at an unprecedented rate. Sending phishing, spoofing, spam, and scam e-mails attempting to gain access to victims’ personal or financial information is a common way e-mail is used to commit a crime. The criminal activity needs to be combated through digital forensics. Unfortunately, cyber events are becoming significantly challenging, and human capabilities are limited. Using the SeFACED dataset, this research proposes a content base, E-mail multi-classification, into four different classes: Normal, Fraudulent, Threatening, and Suspicious, using four primary Machine Learning algorithms, namely Naïve Bayes (NB), Support Vector Machine (SVM), Logistic Regression (LR), and Random Forest (RF). In addition, the Term FrequencyInverse Document Frequency (TF-IDF) and Word2vec are also used as feature extraction techniques to compare their results. The findings show that the best accuracy is achieved with the RF, LR, SVM model by the TF-IDF feature extraction with an accuracy of 95%. In addition, the result for classification with TF-IDF outperformed Word2vec classification.