One-Class Classifier Performance: Comparing Majority versus Minority Class Training

Joffrey L. Leevy, John Hancock, Taghi M. Khoshgoftaar, Azadeh Abdollah Zadeh · 2023

Our study evaluates the impact of training OneClass Classification (OCC) algorithms on the majority class compared to training them on the minority class, using large and big datasets that are highly imbalanced. It is important to note that class availability can present a significant challenge in model training. In reality, there may be situations where one class is readily obtainable within a reasonable time frame while another class is not. Our task involves detecting instances of fraud in the Credit Fraud Detection Dataset and our Medicare dataset derived from Medicare Part D data and List of Excluded Individuals and Entities (LEIE) data. The Credit Card Fraud Detection Dataset has real-world transaction content as well as a significant class imbalance, making it suitable for use as a benchmark for credit card fraud detection. In addition, it is the only publicly available large data for credit card fraud analysis. Part D is big data, allowing researchers to analyze national trends and patterns in prescription drug usage and expenditures. The algorithms used in the study are One-Class Gaussian Mixture Model (GMM), OneClass Adversarial Nets (OCAN), and One-Class Support Vector Machine (SVM). Their performance is measured with the Area Under the Precision-Recall Curve (AUPRC) and Area Under the Receiver Operating Characteristic Curve (AUC). Our results indicate that OCC produces better results when models are trained on the majority class.

Read the paper · More papers on PaperTik