Explainability of SMOTE Based Oversampling for Imbalanced Dataset Problems

Aum Patil, Aman Framewala, Faruk Kazi · 2020

Since the advent of Artificial Intelligence (AI), the problem of imbalanced datasets and the lack of interpretability of complex AI models has been a matter of concern for the research community. These datasets contain a very low proportion of one class (minority class) and very large proportion of another class (majority class). Even though the quantitative representation is less for minority class they have high qualitative importance as the cost associated in case of misclassification in these domains is very high. The paper presents a novel solution to deal with the issue of imbalanced dataset by using the proven method of resampling Synthetic Minority Oversampling Technique (SMOTE). Further, the interpretability of such an approach is demonstrated by some powerful eXplainable AI (XAI) techniques such as LRP, SHAP and LIME. In this paper state-of-art models like Deep Learning and Boosting classifiers were trained to classify fraud instances with high accuracy and proved to be reliable by producing explanations for their predicted instances. The results of confusion matrices and explanations showcase excellent performance and reliability of the models.

Read the paper · More papers on PaperTik