Fraud Detection in Insurance Claims Using Supervised Machine Learning Models

Ahmad Al Doulat, Oluwayemisi Elizabeth Ayo-Bali, Shehenaz Shaik · 2025

Insurance fraud poses a significant challenge to the industry, resulting in billions of dollars in financial losses annually. This study presents a comprehensive machine learning framework to detect fraudulent insurance claims, leveraging classical and advanced algorithms. The six models evaluated were Decision Trees, Random Forests, Support Vector Machines (SVM), XGBoost, k-Nearest Neighbors (k-NN), and Feed Forward Networks (FFN). The dataset underwent extensive preprocessing, including handling missing values, one-hot encoding for categorical variables, feature scaling, and class balancing using SMOTE. Models were assessed using accuracy, precision, recall, and F1-score metrics. Experimental results revealed that XGBoost and Random Forest outperformed other models, with XGBoost achieving perfect classification metrics. Decision Trees also attained perfect accuracy but showed signs of overfitting. Conversely, SVM and k-NN struggled with class imbalance, leading to lower recall rates for fraudulent claims. The FFN demonstrated promising performance with balanced metrics, highlighting the potential of deep learning in fraud detection. Feature importance analysis identified critical predictors, including income and claim history, providing actionable insights for industry practitioners. This research demonstrates the efficacy of machine learning in addressing the challenges of fraud detection, with ensemble models like XGBoost and Random Forest emerging as robust solutions. Future work will explore real-time detection systems, advanced deep learning architectures, and the integration of explainable AI to enhance model interpretability and scalability for practical applications.

Read the paper · More papers on PaperTik