Anomaly Detection in Health Insurance Cost using Ensemble Models for Claim Validation

Raihan Adam Handoyo Winarso, Yusril Falih Izzaddien, Hartawan Bahari Mulyadi, Diana Purwitasari · 2025

Health insurance claims often face challenges like inefficiencies, delays, and fraud, leading to financial losses and operational strain. Traditional detection methods struggle with the complexity of heterogeneous claims data, resulting in high false-positive rates and the need for extensive manual intervention. While machine learning (ML) techniques show promise, their application in anomaly detection for health insurance claims remains underexplored, particularly in integrating diverse models for improved robustness. This study addresses this gap by developing an ensemble framework combining Random Forest, XGBoost, and Ridge regression, optimized using SMOTE to handle data imbalance and residual error analysis for precision. Experimental results demonstrate the superior performance of the ensemble model, achieving the lowest RMSE of 0.1279 compared to individual models. The framework effectively detects anomalies within residual thresholds. For Pulmonary Disease (ICD 1), anomalies were identified with residual errors exceeding$\pm \mathbf{2. 5 2 5. 1 1 4}$. Similarly, anomalies in Heart Failure (ICD 3) were detected with errors surpassing$\pm 2.782$. 325. In contrast, no anomalies were found for Hypertension (ICD 2), which had the highest threshold range of -3.077.246 to 5.635.770. These findings validate the model's accuracy in detecting fraud and improving operational efficiency, offering a robust solution for healthcare claims management.

Read the paper · More papers on PaperTik