Generalization Performance of Support Vector Machines and Other Pattern Classifiers
Peter L. Bartlett, John S. Shawe-Taylor · The MIT Press eBooks · 1998
this paper has been twofold. Firstly, we have stated the known results for high confidence bounds on the generalization error of SVMs in terms of the margin and number of support vectors. Secondly, we wanted to highlight that these results can only be obtained from a data-dependent analysis relying as they do on using some measure to estimate how favourable the input distribution is in relation to the target function. This type of analysis is relatively novel [9], but we feel that its potential for motivating algorithms that are able to take advantage of collusions between distribution and target is far from being exhausted. Indeed, we believe that this is frequently an ingredient in successful learning systems which has been exploited by accident. By more careful analysis of this phenomenon it may well be possible to motivate key ingredients in the Support Vector arsenal, such as choice of kernel function, the bound used in the soft-margin approach and so on. We have also given examples to show that the style of analysis is not limited to SVMs but applies to many other learning machines including two of the most effective techniques, boosting and Bayesian methods. Acknowledgements John Shawe-Taylor was supported in part by the EPSRC research grant number GR/K70366. Peter Bartlett was supported by the Australian Research Council. Appendix: Proof of Theorem 1.6