Improving Classification Performance on Rare Events in Data Starved Medical Applications

Wan D. Bae, Angelo Alfonso, David Stanko, Lili Hao, Linh Le, Matthew Horak · 2023

Predicting rare events is a critical task in many medical applications. Recent improvements in sensor technologies and superior computing power have enabled machine learning techniques to provide valuable data insights pertaining to patients’ symptoms and risk factors. However, for individual-level analysis, class imbalance and small sizes of training datasets for classification models are the main roadblocks to the full utilization of machine learning for prediction of rare events in data starved medical applications. While several techniques have been proposed in this field, little exists in the literature about their performance in data starved contexts. In this study, we propose a novel extension to synthetic minority oversampling techniques, called "Average Neighbor Vector Oversampling (ANVO)", which maintains the data distribution within the minority class but increases data diversity and keeps similar class boundaries. We systematically compare it with state-of-the-art data augmentation techniques using synthetic and real asthma patient datasets. Further, we present a transfer learning framework that integrates oversampling techniques and focuses on retraining a large data model with a small amount of specialized training data and is therefore well-suited to data starved contexts. Results from our experiments demonstrate that the proposed solutions improved the performance of classifiers with a significant increase, 23.7%-61.6%, in sensitivity compared with baseline models.

Read the paper · More papers on PaperTik