A distance-based over-sampling method for learning from imbalanced data sets

Jorge de la Calleja, Olac Fuentes · 2007

Many real-world domains present the problem of im-balanced data sets, where examples of one classes sig-nificantly outnumber examples of other classes. This makes learning difficult, as learning algorithms based on optimizing accuracy over all training examples will tend to classify all examples as belonging to the major-ity class. We introduce a method to deal with this prob-lem by means of creating a balanced data set, which allows to improve the performance of classifiers. Our method over-samples the minority class, using a ran-domized weighted distance scheme to generate syn-thetic examples in the neighborhood of each minority example.

Read the paper · More papers on PaperTik