Automatic Gradient Boosting
Coors, Stefan · Open access LMU (Ludwid Maxmilian's Universitat Munchen) · 2018
The technological progress of the last decades, especially in the computational technology, allowed the storage and fast analysis of a continuously increasing number of datasets.This development is widely omnipresent leading to a situation where the job of a data scientist is "the sexiest job of the 21st century" ( Harvard Business Review).However, well-qualified data scientist are not a dime a dozen.Instead, employees being not much familiar with data analysis are often called to do the job.Automatic machine learning can help those persons to perform predictive modeling with high performing machine learning tools without having much experience.This is achieved by making those applications parameter-free, i.e.only the data is required as input.The rest is done automatically.Projects like Auto-WEKA or auto-sklearn aim to solve the Combined Algorithm Selection and Hyperparameter optimization (CASH) problem resulting in a huge optimization space.However, for most real world applications, only few different learning algorithms are required to deliver superior performances.autoxgboost simplifies this idea one step further and the CASH problem to taking Gradient Boosting as a single learning algorithm in combination with intelligent model based hyperparameter tuning.It is based on the XGBoost R-Package and also supports categorical variables due to special inbuilt factor feature encoding.After describing the main concepts of gradient boosting and the autoxgboost package, several benchmarks are done to improve the package.This includes multiclass threshold tuning as well as the factor encoding of categorical features.Thereafter, autoxgboost is compared to Auto-WEKA and auto-sklearn in a benchmark looking on the predictive performance.We find out that even though autoxgboost only uses one learner instead of a whole library, it provides comparable or even better performances on some datasets.However, when limitations of the computational resources are not an issue, auto-sklearn still provides superior performance.