ENHANCING IMBALANCED DATA CLASSIFICATION: A CASE STUDY OF PORTUGUESE BANK MARKETING

Mahmoud Rajallah Asassfeh, Mahmoud Rajallah Asassfeh, Mohammad Rasmi, Abdullah Alqammaz, Ahmad Bany Doumi, Khaled Al-Qawasmi, Ala’a Al-Shaikh · Journal of Southwest Jiaotong University · 2023

In classification tasks, it is presumed that the number of classes of observations is balanced. While classification models usually give a heavily biased weight to the class that has higher occurrence, building an efficient classification model is likely a challenge when feeding it with imbalanced dataset observations or samples of data. This study introduces a new method for addressing imbalanced datasets in classification tasks, particularly focusing on predicting long-term deposits in banking institutions. The method involves systematic evaluation and comparison of random oversampling (ROS) and synthetic minority over-sampling technique (SMOTE) while employing meticulous feature selection to optimize classification precision. This new methodology showcased competitive performance, notably achieving an accuracy of 89.1% and a G-mean of 0.61 with SMOTE at a 500% ratio encompassing all features in Experiment 2 and an accuracy of 87.2% and a G-mean of 0.677 with ROS at a 500% ratio using the top 15 features in Experiment 3. Keywords: Synthetic Minority Oversampling Technique, Random Oversampling, Imbalanced Data, Feature Selection, Random Forest DOI: https://doi.org/10.35741/issn.0258-2724.58.6.21

Read the paper · More papers on PaperTik