Machine learning forensics to gauge the likelihood of fraud in emails
K. Uma Maheswari, S. Nikkath Bushra · 2021
Digitization leads to unarguable high speed processing but our data is becoming the soft target for many types of cybercrimes. Even though digital forensic investigation is growing to recover volatile and non-volatile data from suspicious locations, manual investigation for evidence acquisition and proving the admissible evidential artifacts is still tedious and time consuming. Automation in this field of digital forensic investigation is the need of the hour to reduce large interactions with humans and to avoid slowing down of investigation process in the pace of growing digital crimes. The human-like behavior depicted by the Machine Learning (ML) to support automation in digital forensic investigation and to help forensic investigators at various stages of digital investigation is the motivation for the proposed work. This paper proposes a machine learning forensics prototype to predict the fraudulent emails in advance to avoid victimization of cyber-crimes. The model applies supervised machine learning technique through an algorithm of Enhanced Forensics Fuzzy C-Means Clustering (EFFCM). The fraudulent components are identified and chain of custody is reported in the training and testing phases performed on Enron email dataset through Weka tool in Cloud environment. The results are comparatively presented in the variations of accuracy in training and testing data partitions and in email application. The prototype presented in this paper can greatly aid the automation in digital forensic investigation for the prediction of fraudulent emails.