Multilevel Regression Models for Learning in the Presence of Rare Data

Srinath Ravindran, Dennis Bahler · 2012

Learning with imbalanced datasets has been a major topic of study for many years. In this paper, we focus on a type of imbalance called imbalance due to rare instances. Such imbalances occur in a variety of domains. Rare instances have received less focus in prediction problems and we wish to draw attention to how accuracy can be improved in the presence of rare data. We discuss an approach to regression tasks, where the training instances are first grouped by similarity and a group of heterogeneous models is applied to each of these groups. This approach enables better prediction on unseen or rare instances when compared to existing approaches. We present results that show performance across datasets from different domains. Our approach was found to provide better prediction than common approaches on rare unseen instances without affecting the overall performance on prediction tasks.

Read the paper · More papers on PaperTik