Optimization Of Classification Results By Minimizing Class Imbalance On Decision Tree Algorithm

Fikri Budiman, Irwan Agus Saputro, Purwanto Purwanto, Pulung Nurtantio Andono · 2022

Imbalance in the number of training datasets for each class is a problem occurring in classification boundaries accuracy. This class imbalance causes unmaximal classification accuracy. It is necessary to increase the accuracy in class classification for minority data, so as not to cause a decrease in the accuracy of class classification with majority data. Thus an appropriate method is required to solve such minority training dataset. The method dealing with imbalance classes in this study is applied to a classification with Decision Tree algorithm. The dataset used is Breast Cancer (286 instances) with two classes, i.e. no-recurrence-events class (201 instance) and recurrence-events class (85 instances). Ensemble classifier method applied to overcome imbalance class in this study is conducted using a combination of AdaBoost, Bagging and SMOTE algorithms. The comparison of algorithms combination use tested in this study resulted in a combination of AdaBoost, Bagging and SMOTE algorithm methods which creates the highest accuracy value, i.e. 84.3%.

Read the paper · More papers on PaperTik