Imbalanced Datasets: From Sampling to Classifiers
T. Ryan Hoens, Nitesh V. Chawla · 2013
Classification is one of the most fundamental tasks in the machine learning and data-mining communities. One of the most common challenges faced when trying to perform classification is the class imbalance problem. A number of sampling approaches, ranging from under-sampling to over-sampling, have been developed to solve the problem of class imbalance. This chapter provides an overview of the sampling strategies as well as classification algorithms developed for countering class imbalance. It considers the issues of correctly evaluating the performance of a classifier on imbalanced datasets and presents a discussion on various metrics. The sampling techniques discussed here include under-sampling, over-sampling, hybrid techniques and ensemble-based methods. Methods have also been developed that aim to directly combat class imbalance without the need for sampling. These methods come mainly from the cost-sensitive learning community; however, classifiers that deal with imbalance are not necessarily cost-sensitive learners. Controlled Vocabulary Terms classification; learning (artificial intelligence); sampling methods