Persistent issues in learning and estimation
George Nagy · 2002
This is a review, from an intuitive rather than a mathematical perspective, of the statistical foundations of adaptive recognition. Key considerations are priors, sample size and sampling strategy, labels, statistical dependencies, and dimensionality. The small-sample bias and variance of maximum likelihood, MAP and Bayes estimators are compared in a small concrete case. Iterative expectation maximization for estimating the sufficient statistics of mixtures is illustrated in a simple setting. It is shown that correlation among features is sometimes unjustly maligned. A counterintuitive increase in the error rate after adding a second feature is traced to the curse of dimensionality. Adaptive classification is presented in the context of both parametric and nonparametric (nearest neighbors and neural nets) estimation. Some recent theoretical results and older experimental observations on hybrid classification are summarized.