New Cluster Undersampling Technique for Class Imbalance Learning

Robert Adjetey Sowah, Moses Apambila Agebure, Godfrey A. Mills, Koudjo M. Koumadi, Seth Y. Fiawoo · International Journal of Machine Learning and Computing · 2016

Adequately learning from datasets that are highly imbalance has become one of the most challenging tasks in Data Mining and Machine Learning disciplines.Most datasets from high risk application areas are often adversely affected by the class imbalance problem due to the limited occurrence of positive examples.This paper presents a new undersampling technique, called Cluster Undersampling Technique (CUST) that has the capability of further improving the performance of classification algorithms when learning from imbalance datasets.The performance of CUST is evaluated by using it to undersample 16 real world class imbalance datasets prior to building classification models using C4.5 decision tree and OneR algorithms.The performance of the models are compared to the performance of random undersampling and oversampling, synthetic minority oversampling, one-sided selection, and cluster-based undersampling.The experimental results using area under receiver-operating characteristic curve and geometric mean showed that CUST resulted in higher performance and is statistically better compared to well-known techniques.

Read the paper · More papers on PaperTik