Empirical Risk Minimization is an Incomplete Inductive Principle

Tom Minka · 2000

Empirical Risk Minimization (ERM) only utilizes the loss function defined for the task and is completely agnostic about sampling distributions. Thus it only covers half of the story. Furthermore, ERM is equivalent to Bayesian decision theory with a particular choice of prior. 1 ERM is incomplete Suppose we are given a data set D = {x1,...,xN} and we want to predict future data in a way that incurs minimum expected loss. That is, we want to pick a number θ to minimize R = p(x|D)L(θ,x)dx (1) x where L is the loss in predicting θ when the real value is x and p(x|D) represents what we know about future x after having observed D. This is the decision-theoretic method of estimation: (1) find the predictive density p(x|D) and (2) choose the estimate which minimizes loss on that density. The first step employs a sampling model for the data and is irrespective of the loss. The second step employs a loss function for the task and is irrespective of the sampling model. Empirical Risk Minimization instead chooses θ to minimize the average loss we would have incurred on the data set D: ER = 1 ∑ L(θ,xi) (2) N which is seen as a Monte Carlo estimate of (1). The problem is that this method completely discards any information we may have about the sampling distribution of x. The following examples demonstrate. i 1 1.1 Gaussian data with squared loss Let the loss function be quadratic and let the data be independent Gaussian with unit variance and unknown mean: L(θ,x) = (θ−x) 2 p(x|m) ∼ N(m,1) (4) 1 = √ exp(−

Read the paper · More papers on PaperTik