Fraud & Anomaly Detection: Using Fine-tuned OCSVM Algorithm and visualization of the enhanced results using Machine Learning Techniques
Kaushiv Garg, Kanwarpartap Singh Gill, Priyanshi Aggarwal, Ramesh Singh Rawat, Deepak Banerjee · 2024
Diversion from traditional methods of fraud detection by tackling the issue of insufficient class labels in corporate data is very essential. In contrast, the use of the One-Class Support Vector Machine (OCSVM) method is employed for the purpose of unsupervised anomaly detection. This study investigates a dataset including credit card transactions conducted by cardholders from Europe during the month of September in the year 2013. The data indicates a significant discrepancy, with just 0.172% (492 out of 284,807) of the transactions being classified as fraudulent. The dataset used in this research comprises numerical variables that have undergone pre-processing by Principal Component Analysis (PCA). Nevertheless, due to confidentiality concerns, the initial attributes of the dataset have not been revealed or accessed, with the exception of two variables: 'Time', denoting the duration of each transaction in seconds, and 'Amount', indicating the monetary amount associated with each transaction. In precision when there is a lack of class labels, the 'Class' property is used to distinguish between fraudulent transactions and legitimate transactions. Given the significant disparity in the allocation of these categories, it is deemed suitable to use the Area Under the Precision-Recall Curve (AUPRC) as a metric for assessing accuracy. This study presents a new approach to anomaly detection that addresses the issue of data balance and utilises the One-Class Support Vector Machine (OCSVM) technology. This study focuses on the difficulties related to the management of datasets that exhibit significant imbalances in the context of fraud detection. Specifically, it addresses the obstacles that arise when class labels are not readily available.