Overfit bounds for classification algorithms

Yoram Gat, Peter J. Bickel · 2000

A major issue in machine learning is managing the overfit of a learning algorithm. The overfit of an algorithm is the degree to which the concept learned is representative of the data available at the time the learning takes place, but not of the mechanism which generated the data. In the context of classification, the overfit is expressed as the difference between the degree of success of classification of the training set and that of the classification of a test set. This dissertation deals with the analysis of the overfit behavior of several classes of classification algorithms. The analysis provides insight into the the sources of overfit and yields bounds which can be used to control the overfit of some well-known algorithms, such as classification trees, the perceptron and edited nearest neighbors. The structure of the dissertation is as ...

Read the paper · More papers on PaperTik