Life Insurance Fraud Detection: A Data-Driven Approach Utilizing Ensemble Learning, CVAE, and Bi-LSTM
M. J. D. Ebinezer, Bondalapu Chaitanya Krishna · Applied Sciences · 2025
Insurance fraud detection is a significant challenge due to increasing fraudulent claims, class imbalance, and the increasing complexity of fraudulent behaviour. Traditional machine learning models often struggle to generalize effectively when applied to high-dimensional and imbalanced datasets. This study proposes a data-driven framework for intelligent fraud detection employing three distinct modelling strategies: chaotic variational autoencoders (CVAEs), idirectional long short-term memory (Bi-LSTM), and a hybrid random forest + Bi-LSTM technique. This study aims to evaluate and compare the effectiveness of generative, sequential, and ensemble-based models in identifying rare fraudulent claims within created datasets of 4000 life insurance applications containing 83 features. Following extensive preprocessing and model training, CVAEs achieved the highest accuracy (83.75%) but failed to detect many fraudulent cases due to its low recall (3.28). The Bi-LSTM model outperformed the CVAEs in recall (5.98%) and F1-score, effectively capturing temporal dependencies within the data. The hybrid RF + Bi-LSTM model matched Bi–LSTM in recall but showed more stable ROC and precision–recall curves, indicating robustness and misinterpretability. This hybrid approach balances the strengths of feature-driven and sequential modelling, making it suitable for operational deployment. While Bi–LSTM achieved the best statistical performance, the hybrid model offers enhanced reliability in threshold-sensitive fraud applications.