Improving Bayesian Optimization for Machine Learning using Expert Priors
Kevin Swersky · TSpace (University of Toronto) · 2017
Deep neural networks have recently become astonishingly successful at many machine learning problems such as object recognition and speech recognition, and they are now also being used in many new and creative ways. However, their performance critically relies on the proper setting of numerous hyperparameters. Manual tuning by an expert researcher has been a traditionally effective approach, however it is becoming increasingly infeasible as models become more complex and machine learning systems become further embedded within larger automated systems. Bayesian optimization has recently been proposed as a strategy for intelligently optimizing the hyperparameters of deep neural networks and other machine learning systems; it has been shown in many cases to outperform experts, and provides a promising way to reduce both the computational and human time required. Regardless, expert researchers can still be quite effective at hyperparameter tuning due to their ability to incorporate contextual knowledge and intuition into their search, while traditional Bayesian optimization treats each problem as a black box and therefore cannot take advantage of this knowledge. In this thesis, we draw inspiration from these abilities and incorporate them into the Bayesian optimization framework as additional prior information. These extensions include the ability to transfer knowledge between problems, the ability to transform the problem domain into one that is easier to optimize, and the ability to terminate experiments when they are no longer deemed to be promising, without requiring their training to converge. We demonstrate in experiments across a range of machine learning models that these extensions significantly reduce the cost and increase the robustness of Bayesian optimization for automatic hyperparameter tuning.