Self–Organizing–Maps Based Undersampling for the Classification of Unbalanced Datasets
Marco Vannucci, Valentina Colla · 2018
The paper describes an approach for the preprocessing of the dataset to be used for the training of an arbitrary classifier in the context of binary classification of unbalanced datasets. The proposed method is based on the reduction of the number of frequent samples that are included in the training dataset in order to mitigate the unbalance rate and to promote the correct detection of the rare samples, which are often the most interesting ones in relation to the application the classifier is designed for. The proposed technique exploits two Self Organizing Maps that clusterize the rare and frequent samples in the original training dataset. The obtained clusterizations are exploited in order to determine the frequent samples whose removal is most convenient. The approach has been successfully tested and compared to other well known resampling techniques by exploiting different datasets coming from both the UCI repository and real industrial applications.