Cross Validation Can Estimate How Well Prediction Variance Correlates with Error
Felipe Antonio Chegury Viana, Raphael T. Haftka · AIAA Journal · 2009
T HE use of surrogates for facilitating optimization and statistical analysis of computationally expensive simulations has become commonplace [1–4]. They offer easy-to-compute prediction and in some cases (such as kriging [5,6] and polynomial response surfaces [7,8]), they also furnish the prediction variance as a measure of uncertainty [9]. Figure 1a illustrates the concepts of prediction and prediction variance. Adaptive sampling and optimization methods use the prediction variance to select the next sampling point. For example, the Efficient Global Optimization (EGO) [10] and the Enhanced Sequential Optimization [11] algorithms use the kriging prediction variance to seek the point of maximum expected improvement as the next simulation for the optimization process. For such methods, it is important to assess the accuracy of the prediction variance; but presently, this is not available (although there is work on how to improve the uncertainty structure [12]). Cross validation is a standard tool for estimating the mean square errors (see the Appendix), thus the quality of the fit; and it can be used for selecting surrogates in a set [13–15]. Cross validation divides a set of p data points into k subsets. The surrogate isfit to all subsets except one, and the error is checked in the subset that was left out. This process is repeated for all subsets to produce a vector of cross-validation errors, eXV . Figure 1b illustrates cross validation when only one point is omitted. We propose using cross validation for estimating the correlation between the prediction variance and the errors. Specifically we propose to use the correlation between the absolute values of the cross-validation errors and the square root of the prediction variance at the points that were left out. II. Correlation Between Square Root of Prediction Variance and Absolute Errors