Stochastic dynamics of learning with momentum in neural networks

Wim Wiegerinck, Andrzej Komoda, Tom Heskes · Journal of Physics A Mathematical and General · 1994

We study on-line learning with a momentum term for nonlinear learning rules. Through introduction of auxiliary variables, we show that the learning process can be described by a Markov process. For small learning parameters eta and momentum parameters alpha close to 1, such that gamma = eta /(1- alpha ) 2 is finite, the time-scales for the evolution of the weights and the auxiliary variables are the same. In this case Van Kampen's expansion can be applied in a straightforward manner. We obtain evolution equations for the average network state and the fluctuations around this average. These evolution equations depend (after rescaling a of the time and fluctuations) only on gamma : all combinations ( eta , alpha ) with the same value of gamma give rise to similar behaviour. The case with alpha constant and eta small requires a completely different analysis. There are two different time-scales: a fast time-scale on which the auxiliary variables equilibrate and a slow time-scale for the change of the weights. By projection on the space of slow variables the fast variables can be eliminated. We find that, for small learning parameters eta and finite momentum parameters alpha , learning with momentum is equivalent to learning without a momentum term with a rescaled learning parameter eta = eta /(1- alpha ). Simulations with the nonlinear Oja learning rule confirm the theoretical results.

Read the paper · More papers on PaperTik