Unbiased Assessment of Learning Algorithms

Tobias Scheffer, Ralf Herbrich · 1997

In order to rank the performance of machine learning algorithms, many researchers conduct experiments on benchmark data sets. Since most learning algorithms have domain-specific parameters, it is a popular custom to adapt these parameters to obtain a minimal error rate on the test set. The same rate is then used to rank the algorithm, which causes an optimistic bias. We quantify this bias, showing, in particular, that an algorithm with more parameters will probably be ranked higher than an equally good algorithm with fewer parameters. We demonstrate this result, showing the number of parameters and trials required in order to pretend to outperform C4.5 or FOIL, respectively, for various benchmark problems. We then describe out how unbiased ranking experiments should be conducted. 1 Introduction Estimating the accuracy of a classifier is a topic that has experienced much attention in the ML community. One of the main results is that N -fold cross validation provides a bias-free [ Sto74...

Read the paper · More papers on PaperTik