Efficient and Distributed Model-Based Boosting for Large Datasets
Daniel Schalk · Open access LMU (Ludwid Maxmilian's Universitat Munchen) · 2018
Component-wise boosting applies the boosting framework to statistical models, e. g., general additive models using component-wise smoothing splines.Boosting these kinds of models maintains interpretability and enables unbiased model selection in high dimensional feature spaces.A well-known implementation of this principle is the R package mboost.The R package compboost is an alternative implementation of component-wise boosting written in C++ to obtain high runtime performance and full memory control.The main idea is to provide a modular class system which can be extended without editing the source code.Therefore, it is possible to use R functions as well as C++ functions for custom base-learners, losses, logging mechanisms or stopping criteria.The main goal of this performant implementation is to enable model fitting on large datasets which can be troublesome with the mboost package.In terms of runtime, compboost is three to ten times faster then mboost and uses, depending on the baselearner, less memory.Nevertheless, compboost has much unused potential like using sparse data matrices or implementing parallel computations.These enhancements will be implemented soon.