Classification of Imbalanced Data Using Synthetic Over-Sampling Techniques

Peng Jun Huang · 2015

A dataset is considered to be imbalanced if the classication objects are notapproximately equally represented. The classication problems of imbalanceddataset have brought growing attention in the recent years. It is a relatively newchallenge in both industrial and academic elds because many machine learn-ing techniques do not have a good performance. Often the distribution of thetraining data may be dierent than that of the testing data. Typically, samplingmethods would be used in imbalanced learning applications, which modies thedistribution of the training samples by some mechanisms in order to obtain a rel-atively balanced classier. A novel synthetic sampling technique, SMOTE (Syn-thetic Minority Over-sampling Technique), has showed a great deal of success inmany applications. Soon after this powerful methods was introduced, some otherSMOTE-based sampling methods such as SMOTEboost , Border-line SMOTEand ADASYN (Adaptive Synthetic Sampling) have been developed. This pa-per reviews and compares some of these synthetic sampling methods for learningimbalanced datasets.

Read the paper · More papers on PaperTik