A comparison for handling imbalanced datasets

Arif Syaripudin, Masayu Leylia Khodra · 2014

In various real case, imbalanced datasets problems are inevitable, such as in metal detecting security or diagnosis of disease. With the limitations of existing learning algorithms when faced with imbalanced datasets, the prediction error is caused by the dominance of the majority against the minority class. Various techniques have been made to address the above circumstances. This paper compares those techniques of handling imbalanced datasets with resample and ensembles. From a different standpoint, this paper examines how much influence the number of instances, number of attributes, the attributes data types, the number of the target class, and missing attribute values affect the classification results with performance analysis using f-measure. An experiment has resulted that the criteria regarding the number of attributes, attribute data types, and the number of the target class do not affect the classification results. While the missing attribute with values have an affect classification result. For better high F-measure, the experiment shows that the best performer is combination of SMOTE 5000/0 and AdaBoostMl.

Read the paper · More papers on PaperTik