Healthcare Provider Summary Data for Fraud Classification

Justin M. Johnson, Taghi M. Khoshgoftaar · 2022

Fraud, waste, and abuse are spreading throughout the healthcare industry and costing patients and taxpayers billions of dollars. Fortunately, electronic medical records and publicly available data sources like the Centers for Medicare & Medicaid Services (CMS) have enabled data mining and machine learning techniques that can help automate the detection of healthcare fraud. In this study, we explore the application of healthcare provider summary data for the purpose of fraud detection. We leverage the latest CMS Part B Summary by Provider big data sets to curate two new labeled data sets for supervised learning. The two new data sets are compared to a popular baseline data set from related works using six runs of cross validation with two popular ensemble learners, multiple complementary performance metrics, and statistical tests. Classification results show that the proposed provider summary features are good indicators of healthcare fraud. A two-way analysis of variance test and 95% confidence intervals show that the new features yield significantly better performance on the fraud detection task when used to enrich existing data sets. Finally, feature contributions are measured with Shapley values to illustrate the top 20 features that contribute to fraud estimation.

Read the paper · More papers on PaperTik