A Novel Data Reduction Technique for Medicare Fraud Detection with Gaussian Mixture Models

John Hancock, Taghi M. Khoshgoftaar · 2025

This study addresses the issue of fraud within the Medicare program, which results in losses amounting to billions of dollars annually. By leveraging one-class classifiers (OCCs), we demonstrate a data reduction technique in the task of fraud detection, utilizing Medicare Part D and Part B insurance claims data. Our research introduces a novel data reduction technique that allows for the training of one-class Gaussian mixture models (GMMs) on a significantly reduced subset of the full dataset. This approach maintains high performance in fraud identification while optimizing computational efficiency. We demonstrate that models trained on as little as 20% of the original data yield performance, in terms of area under the receiver operating characteristic curve (AUC) and area under the precision recall curve (AUPRC), comparable to models trained on the complete training data. Our findings suggest that it is feasible to efficiently classify highly imbalanced datasets, such as those encountered in Medicare fraud detection, using one-class GMMs with a substantial data reduction. This research contributes to the field by presenting a scalable and effective methodology for fraud detection in large-scale healthcare datasets, showcasing the potential of machine learning in enhancing the integrity of public health insurance programs.

Read the paper · More papers on PaperTik