A Framework for Scalable Cost-sensitive Learning Based on Combining Probabilities and Benefits

Wei Fan, Haixun Wang, Philip S. Yu, Sal Stolfo · 2002

We present a general framework for scalable cost-sensitive learning based on ensembles of classifiers. To compute ensembles for scalable learning, base models are computed from data subsets and their predictions are combined. Previous research mainly focus on combining class labels for cost-insensitive problems. In our study, we concentrate on how to combine the probability and benefit outputs (such as charity donation amount) for scalable cost-sensitive learning. We identify a few properties of both averaging and cost-sensitive optimal decision-making. Due to these properties, averaging has great potential for higher benefit as well as its obvious advantage of scalability. Experiments have shown that the averaging methods achieve a benefit level as good as or even better than the respective single classifier trained from the same dataset as a whole. However, the averaging methods exhibit linear speedup and scaled speedup that are independent of learning algorithms. Besides the averaging methods, we have experimented with tree-structured methods using variations of meta-learning and regression. Nonetheless, these approaches do not lead to higher benefits but consume higher overhead, hence are less scalable. Compared with tree-structured methods, the averaging methods are more efficient to classify examples, and easer to implement and verify.

Read the paper · More papers on PaperTik