PROTOTYPE SELECTION IN IMBALANCED DATA FOR DISSIMILARITY REPRESENTATION - A Preliminary Study
Mónica Millán Giraldo, Vicente García, J. Salvador Sánchez · 2012
In classification problems, the dissimilarity representation has shown to be more robust than using the feature space. In order to build the dissimilarity space, a representation set of r objects is used. Several methods have been proposed for the selection of a suitable representation set that maximizes the classification performance. A recurring and crucial challenge in pattern recognition and machine learning refers to the class imbalance problem, which has been said to hinder the performance of learning algorithms. In this paper, we carry out a preliminary study that pursues to investigate the effects of several prototype selection schemes when data set are imbalanced, and also to foresee the benefits of selecting the representation set when the class imbalance is handled by resampling the data set. Statistical analysis of experimental results using Friedman test demonstrates that the application of resampling significantly improve the performance classification.