Cross-validation, learning set transformations, and generalization

David H. Wolpert · University of North Texas Digital Library (University of North Texas) · 1990

This paper discusses using cross-validation as an aid to generalization. It starts by showing how to use cross-validation to decide amongst a set of generalizer when the learning set consists of examples of the text-to-phoneme problem. In addition to reproducing the learning set perfectly, the generalizer so chosen by cross-validation has an error rate on the testing set (7%) close to that which NETtalk has on the learning set (5%). This paper then presents an example of using cross-validation as part of a front-end to generalizes, i.e., as part of an algorithm for pre-processing a learning set of input/output examples before trying to generalize from it. The results of thirty-six comparisons between the performance of a generalizer and the performance of the generalizer with this front-end are presented. These comparisons involve numerical, Boolean, and visual tasks. In all but one of the comparisons the front-end improves the generalization performance, often inducing perfect generalization. (The average ratio of the generalization error rate with the front-end to the generalization error rate without the front-end is 0.23, {plus minus} 0.05.) Finally, this paper ends by discussing some of the subtler mathematical issues involved in using cross-validation to help generalization. 7 refs., 3 figs.

Read the paper · More papers on PaperTik