Predictive Data Mining for Highly Imbalanced Classification

Madhuri Agrawal, Gajendra Singh, Ravindra Kumar Gupta, Sri Satya · 2012

Abstract — The paper addresses some theoretical and practical aspects of data mining, focusing on predictive data mining, where two central types of prediction problems are discussed: classification and regression. Further accent is made on predictive data mining, where the time-stamped data greatly increase the dimensions and complexity of problem solving. The main goal is through processing of data (records from the past) to describe the underlying dynamics of the complex systems and predict its future. Traditional classification algorithms can be limited in their performance on highly imbalanced datasets. A popular stream of work for countering the problem of class imbalance has been application of a sundry of sampling strategies. In this work, we focus on the problem of class imbalance. We incorporate different “rebalance ” heuristics.

Read the paper · More papers on PaperTik