Improving regression estimation: Averaging methods for variance reduction with extensions to general convex measure optimization
Michael Perrone · 1993
A general theoretical framework for Monte Carlo averaging methods of improving regression estimates is presented with application to neural network classification and time series prediction. Given a population of regression estimators, it is shown how to construct a hybrid estimator which is as good as or better than, in the MSE sense, any estimator in the population. It is argued that the ensemble method presented has several properties: It efficiently uses all the regressors of a population--none need be discarded. It efficiently uses all the available data for training without over-fitting. It inherently performs regularization by smoothing in functional space which helps to avoid over-fitting. It utilizes local minima to construct improved estimates whereas other regression algorithms are hindered by local minima. It is ideally suited for parallel computation. It leads to a very useful and natural measure of the number of distinct estimators in a population. The optimal parameters of the ensemble estimator are given in closed form. It is shown that this result derives from the notion of convexity and can be applied to a wide variety of optimization algorithms including: Mean Square Error, a general class of $L\sb{p}$-norm cost functions, Maximum Likelihood Estimation, Maximum Entropy, Maximum Mutual Information, the Kullback-Leibler Information (Cross Entropy), Penalized Maximum Likelihood Estimation and Smoothing Splines. The connection to Bayesian Inference is discussed. Experimental results on the NIST OCR database, the Turk and Pentland human face database and sunspot time series prediction are presented which demonstrate that the ensemble method dramatically improves regression performance on real-world classification tasks.