Classification of Health Insurance Fraud Risk with Machine Learning

Muhammad Kent Al-Ghazi, Ryan Bertrand, Muhammad Dzul Qarrnayn Destra, Alexander Agung Santoso Gunawan, Karli Eka Setiawan · 2024

Health insurance is a useful service that can help its users gain lifesaving medical aid when they are in need. However, health insurance is also exploitable to insurance fraud through the falsification of information to increase the amount of reimbursement and cause massive loss of funds to the insurance provider. We propose the usage of machine learning to accurately determine potential health insurance fraud. The objective of conducting this research is to determine which features are the most important to determine healthcare insurance fraud. This research used a dataset provided in Kaggle titled Healthcare Provider Fraud Detection Analysis using Random Forest Classifier and Logistic Regression. The best-performing model in this test, the Logistic Regression, is then used to which features are the most important for the classification. Our research shows that the most important feature in detecting health insurance fraud is the amount of money reimbursed associated with a provider. The Logistic Regression model achieved an accuracy of 0.90, precision of 0.93, recall of 0.91, and an F1 Score of 0.90, outperforming the Random Forest model in comparative analysis.

Read the paper · More papers on PaperTik