A Comparative Study of Model-Agnostic and Importance-Based Feature Selection Approaches

Huanjing Wang, Qianxin Liang, John Hancock, Taghi M. Khoshgoftaar · 2023

In the context of high-dimensional credit card fraud data, feature selection techniques are commonly employed by researchers and practitioners to enhance the performance of credit card fraud detection models. This study presents a comparison in model performance using the most important features selected by SHAP (SHapley Additive exPlanations) values and the model's built-in feature importance list. Both methods rank features and select the most important features for model evaluation. The performance of these feature selection techniques is assessed by building classification models using two classifiers: XGBoost and Decision Tree. The evaluation metric used to measure effectiveness is the Area under the Precision-Recall Curve (AUPRC). All experiments are conducted on the Credit Card Fraud Detection Dataset from Kaggle. The experimental results and z-tests demon-strate that, in most cases, there is no significant difference between the two feature selection methods, regardless of the feature subset size and classifier used. However, for models trained on larger datasets, it is suggested to use the model's built-in feature importance list as the preferred feature selection method over SHAP. The rationale behind this recommendation is that computing SHAP feature importance is a separate activity, and models provide feature importance as a side-effect of training, thus requiring no extra effort. Therefore, opting for the model's built-in feature importance list can offer a more efficient and manageable approach for larger datasets and more complex models.

Read the paper · More papers on PaperTik