Health Insurance Fraud Detection: The Role of Feature Engineering and Preprocessing Techniques
Yuechi Chen, Chuqing Zhao, C L Nie · 2025
This study leverages feature engineering and preprocessing techniques to enhance machine learning performance in health insurance fraud detection. Using real-world healthcare fraud data, we compare the performance with and without Synthetic Minority Over-sampling Technique (SMOTE) for each of the logistic regression, decision trees, and ensemble techniques to measure the ability to detect fraud. The selected evaluation metrics are accuracy, recall, precision, and F1-score. Results show that feature interaction significantly increases model performance, with Shapley Additive exPlanation (SHAP) analysis validating the feature importance. We also notice that, while smote improves recall for certain models, ensemble methods inherently handle class imbalance. This study highlights the critical role of feature engineering in fraud detection, and demonstrates that engineered features contribute more to performance gains than SMOTE alone. The code used for this study has been open-sourced and is available at GitHub Repository for future research and benchmarking1.