Unmasking Insurance Fraud: Integrating Feature Selection and Data Balancing Techniques
Jigyasha Arora, Suyash Bhardwaj · Procedia Computer Science · 2026
Information security is essential to everyone’s life since it safeguards the private information of end users. This private information is categorized into various areas, including banking, insurance, mortgage fraud, phishing, and others. Insurance claims are one of them, which are accumulating day by day, be it automobile insurance, health insurance, travel insurance, or property insurance. As the awareness about insurance is increasing, immorality is also growing. Fictitious insurance claims are those that are submitted to mislead an insurance company. The seriousness of insurance offenses varies as well, ranging from subtly exaggerating statements to deliberately creating mishaps or damage. The lives of innocent individuals are impacted by fraudulent actions in two ways: directly through deliberate or unintentional harm or damage, and indirectly through the crimes that result in increased insurance costs. Thus, insurance fraud is a serious issue, and it is indeed the need of the hour to resolve this issue utilizing machine learning techniques. We proposed a model in which feature scaling includes standardization, and one-hot encoding transforms categorical data into numerical data. Feature selection is done using the importance score from the tree-based classifier. Principal Component Analysis (PCA) is applied for feature engineering. Data visualization is done, and SMOTEENN is used for data balancing. The data is split into an 80-20% train-test ratio, and the XGBoost model is enforced, achieving a test accuracy, precision, recall, and F1-score of 100%. However, when cross-validation is applied, an accuracy of 99.6 % is attained by the model. A comparison with previously existing models was made on the same dataset to ensure the novelty of the technique used.