The golden chain
Peter J. Bickel, J Rosen Milton, Paradise Lost · 2003
Jiang, Lugosi and Vayatis, and Zhang in part explicitly and in part implicitly, have done a great deal in explaining the nature of boosting from a statistical point of view. The problem all consider is that of finding classifiers that approximate the Bayes classifier using only a training sample (Xi, Yi), i = 1, . . . , n, (Xi, Yi) ∼ (X,Y ), with Y = ±1 (for simplicity). The Bayes classifier is described as sgn (Fp(X)), where Fp(X) = q ◦ log ( p[Y = 1 ∣∣ X]/P [Y = −1 ∣∣ X]), for any strictly increasing function q with q(0) = 0. The methods of approximation discussed by these and previous authors cited in their papers have the common setting that the approximating values are sgn (F (X)), where F ∈ F ≡ ∞k=1Fk, Fk = {∑k j=1 λjhj : h1, . . . , hk ∈ H, λ1, . . . , λk ∈ R } and H is a set of base classifiers, h : X → {−1, 1}. All methods are based on the following two observations: