Online Learning from Finite Training Sets: An Analytical Case Study

Peter Sollich, David Barber · 1996

We analyse online learning from finite training sets at noninfinitesimal learning rates j. By an extension of statistical mechanics methods, we obtain exact results for the time-dependent generalization error of a linear network with a large number of weights N . We find, for example, that for small training sets of size p ß N , larger learning rates can be used without compromising asymptotic generalization performance or convergence speed. Encouragingly, for optimal settings of j (and, less importantly, weight decay ) at given final learning time, the generalization performance of online learning is essentially as good as that of offline learning. 1 INTRODUCTION The analysis of online (gradient descent) learning, which is one of the most common approaches to supervised learning found in the neural networks community, has recently been the focus of much attention [1]. The characteristic feature of online learning is that the weights of a network (`student') are updated each time a n...

Read the paper · More papers on PaperTik