Health Insurance Fraud Detection Using Data Mining
Mahalakshmi Rajkumar, Salazar Marques,, Napoleon, Sarah Mai · Zenodo (CERN European Organization for Nuclear Research) · 2025
Healthcare insurance fraud remains a persistent and costly issue, contributing to billions in annual losses through fraudulent claims, waste, and abuse. This paper validates and extends the methodology proposed in "Healthcare Insurance Fraud Detection Using Data Mining" by applying its hybrid two-stage approach to a modern healthcare claims dataset (Kaggle – Enhanced Health Insurance Claims Dataset (2022–2024)). The first stage employs the Apriori association rule mining algorithm to identify frequent patterns and establish behavioral norms among patients, providers, and claims. These rules are then transformed into feature vectors and analyzed using four unsupervised anomaly detection algorithms: Isolation Forest, Cluster-Based Local Outlier Factor (CBLOF), Empirical Cumulative Distribution Outlier Detection (ECOD), and One-Class Support Vector Machine (OCSVM). Key contributions beyond the original work include: (1) comparative empirical evaluation of all four anomaly detection algorithms rather than theoretical proposal alone; (2) application to more recent data (2022–2024 vs. 2008–2010); (3) first empirical inclusion of ECOD combined with Apriori rule mining in a healthcare fraud pipeline; and (4) a multi-threshold Apriori sensitivity analysis across four configurations. Model performance is evaluated using Silhouette scores and cost-based coverage metrics. CBLOF achieved the most stable performance (Silhouette score up to 0.748), and higher-confidence lower-support rule generation consistently yielded the highest anomaly detection quality across all four algorithms.