Enhancing Medicare Fraud Detection: Random Undersampling Followed by SHAP-Driven Feature Selection with Big Data
Qianxin Liang, Richard A. Bauder, Taghi M. Khoshgoftaar · 2024
SHapley Additive exPlanations (SHAP) is a method used to explain the output of machine learning models. SHAP provides a unified measure of feature importance through SHAP values and serves as a feature selection tool for handling Big Data. This paper presents a study on optimizing feature selection using SHAP for Medicare fraud detection by applying the Random Undersampling (RUS) technique. Our approach aims to mitigate the big dataset's complexity, stemming from class imbalance, size, and high dimensionality, by employing RUS, followed by the integration of the SHAP model within a feature selection framework. The SHAP model is integrated using algorithms including LightGBM, XGBoost, CatBoost, and Decision Tree. To evaluate the effectiveness of our approach, we use the Area Under the Precision-Recall Curve (AUPRC) as the primary evaluation metric to measure the performance of the classification model, utilizing a Random Forest algorithm. Our experiments utilize the Medicare Part D dataset from the Centers for Medicare and Medicaid Services (CMS). Our primary objective is to investigate whether applying RUS with SHAP-based feature selection leads to measurable improvements in binary classification performance. Our findings indicate that the feature subset generated by applying RUS before selecting features using SHAP outperforms those created without the RUS enhancement. This approach not only enhances model performance, but also improves efficiency by reducing computational demands.