A Comparison of Low-Shot Learning Methods for Imbalanced Binary Classification
Preston Billion-Polak, Taghi M. Khoshgoftaar · 2024
The modern tasks of few-shot, one-shot, and zero-shot learning - or collectively Low-Shot Learning (LSL) -at first glance are quite similar to the long-standing task of class-imbalanced learning; specifically, they both aim to learn classes for which there is little labeled data available. Despite this similarity, and the shown effectiveness of LSL techniques in complex many-class image datasets, a previous literature review of the recent work in this overlap found that LSL methods are rarely applied to “classical” imbalanced tasks, i.e., those that are binary and tabular. To fill this gap, we select two of the papers in this survey that provide open-source models (each representing one of the two major approaches to LSL, optimization-based and similarity-based), train and thoroughly evaluate them on a traditional credit card fraud dataset, and compare their performances to each other. Our evaluation methodologies improve on those of the original works, by implementing an Area Under the Precision-Recall Curve (AUPRC) measurement, utilizing ten runs of five-fold cross-validation to ensure fairness, and conducting statistical analyses to confirm the statistical significance of our results. We find that, while our chosen optimization-based model outperforms somewhat in terms of Area Under the Receiver Operating Characteristic Curve (AUC-ROC), the similarity-based model vastly outperforms in terms of our primary metric, AUPRC, and shows promise for future research on similar datasets. To the best of our knowledge, we are the first to directly compare the performance of these two LSL approaches, and the first to examine a similarity-based model on this credit card fraud dataset.