Dynamics of Training

Siegfried Bös, Manfred Opper · 2007

A new method to calculate the full training process of a neural network is introduced. No sophisticated methods like the replica trick are used. The results are directly related to the actual number of training steps. Some results are presented here, like the maximal learning rate, an exact description of early stopping, and the necessary number of training steps. Further problems can be addressed with this approach. 1 INTRODUCTION Training guided by empirical risk minimization does not always minimize the expected risk. This phenomenon is called overfitting and is one of the major problems in neural network learning. In a previous work [B¨os 1995] we developed an approximate description of the training process using statistical mechanics. To solve this problem exactly, we introduce a new description which is directly dependent on the actual training steps. As a first result we get analytical curves for empirical risk and expected risk as functions of the training time, like the ones...

Read the paper · More papers on PaperTik