TOP: An Efficient Two-Levels of Positive Resampling Framework for Class Imbalanced Data

Nathaniel Netirungroj, Eakasit Pacharawongsakda · 2018

In real-world applications such as fraud detection, target class values have an unequal size. This problem is called class-imbalanced data. Many strategies have been proposed to deal with this situation. Most of them focused on changing the data characteristics. For example, adjusting class distribution is one of the most popular approaches for this matter. In this work, we proposed TOP (TwO-levels of Positive resampling framework), an alternative framework to resolve such a problem. Our technique exploits DBSCAN mechanism and other resampling algorithms in order to maximize classification performance. It is able to dynamically draw two boundaries that represent similarity level between consideration positive and other instances. Many possible resampling techniques such as undersampling or over-sampling are allowed to perform inside those areas. We benchmarked TOP with three types of resampling techniques including over-sampling, down-sampling, and hybrid sampling by training eleven machine learning algorithms on fifteen datasets. As a result, our technique outperformed other techniques in several evaluation metrics.

Read the paper · More papers on PaperTik