Estimation of generalization error: random and fixed inputs
Junhui Wang, Xiaotong T. Shen · 2006
Abstract: In multicategory classication, an estimated generalization error is often used to quantify a classier's generalization ability. As a result, quality of estimation of the generalization error becomes crucial in tuning and combining classiers. This article proposes an estimation methodology for the generalization error, permitting a treatment of both xed and random inputs, which is in contrast to the conditional classi cation error commonly used in the statistics literature. In particular, we derive a novel data perturbation technique, that jointly perturbs both inputs and outputs, to estimate the generalization error. We show that the proposed technique yields optimal tuning and combination, as measured by generalization. We also demonstrate via simulation that it outperforms cross-validation for both xed and random designs, in the context of margin classication. The results support utility of the proposed methodology.