Medical Insurance Fraud Risk Monitoring and Identification Model Based on Feature Selection and Machine Learning
Junwei Feng, Mingze Li, Yiming Feng, Yiran Xin, Yuhang Zhao, Yongxin Jiang, Yixiao Zhang, Fayuan Li · 2024
Medical insurance fraud has the characteristic of concealment and is challenging to detect promptly through manual investigation. Therefore, a dataset was collected for medical insurance monitoring and a model was designed for automatically detecting medical insurance fraud based on big data technology. This article first visualizes the dataset, removes abnormal samples, and explores the distribution and feature correlation of the dataset. Afterward, the dataset was divided into a training set and a testing set, and the training set was used to train ten machine learning algorithms. The four baseline algorithms were selected with the best performance: XGBoost, LightGBM, CatBoost, and Random Forest. A feature selection algorithm has been designed based on the concept of greed, which selects three of the most essential features from 80 features while still ensuring high accuracy. Finally, the baseline model was trained using the three selected features, and grid search was used for parameter tuning. After model fusion, an accuracy of 92.3% was achieved on the test set.