New approach with ensemble method to address class imbalance problem
Seyyedali Fattahi, Zalinda Othman, Zulaiha Ali Othman · 2015
An attractive research in recent years is solving class imbalance problem in imbalanced dataset. The class is imbalanced when the number of one class (majority) is more than another one (minority). The classification of this imbalanced class causes imbalanced distribution and poor predictive classification accuracy. This paper introduces a new ensemble –based method for imbalanced data set classification using Synthetic Minority Over-sampling Technique (SMOTE) and Rotation Forest algorithm to address class imbalance problem. Rotation Forest applied as ensemble classifier combines with well-known re-sampling method (SMOTE). It constructs classifiers with obtaining features by rotating subspaces of the original dataset. The advantages of Rotation Forest rather than other ensemble methods (Boosting, Bagging, Random Subspace) is that same information held as original data sets and no information lost in data sets which used to construct classifiers. Experimental results reveal the effectiveness of SMOTE and Rotation Forest performance at data level in overall accuracy, Cohen’s kappa Coefficient, False Negative rate, AUC, and RMSE compared to other related classification ensemble methods (SMOTE-Boost, SMOTE-Bagging, SMOTE-random subspace) on twenty KEEL repository imbalanced datasets (binary dataset not multi-class) which selected randomly from different ratios by implementing Java-based WEKA and STATISTICA software. SMOTE implemented for training data by values of N=100, 200, 300, and 400. Kappa-Error diagram is plotted to analysis the behavior of ensemble methods. The experimental results clarify the validness of proposed ensemble classifier.