Improving Learner Performance with Data Sampling and Boosting

Chris Seiffert, Taghi M. Khoshgoftaar, Jason Van Hulse, Amri Napolitano · 2008

Learning from imbalanced datasets is a well known problem in the data mining community. Many techniques have been proposed to alleviate the problems associated with class imbalance, including data sampling and boosting. While data sampling has received the bulk of the attention from the research community, our results show that boosting often results in better classification performance than even the best data sampling techniques. In this work, we compare the performance of data sampling and boosting on ten datasets from various application domains using two commonly used learners. In addition, we propose the use of both data sampling and boosting in an attempt to combine the strengths of these techniques and achieve even better classification performance.

Read the paper · More papers on PaperTik