Scaling properties of on-line learning with momentum

Tom Heskes, Wim Wiegerinck, Andrzej Komoda · 1994

We study online learning with momentum term for nonlinear learning rules. Through introduction of auxiliary variables, we show that the learning process can still be described by a first-order Markov process. For small learning parameters /spl eta/ and momentum parameters /spl alpha/ close to 1 (we consider the case /spl alpha/=1-/spl radic/(/spl eta///spl lambda/) for small /spl eta/), Van Kampen's expansion can be applied in a straightforward manner. We obtain evolution equations for the average network state and the fluctuations around this average. These evolution equations depend (after rescaling of time and fluctuations) only on /spl lambda/=/spl eta//(1-/spl alpha/)/sup 2/: all combinations (/spl eta/,/spl alpha/) with the same value of /spl lambda/ give rise to similar graphs. For small /spl lambda/, i.e., /spl eta//spl Lt/(1-/spl alpha/)/sup 2/, learning with momentum term is equivalent to learning without momentum term with rescaled learning parameter /spl eta//spl tilde/=/spl eta//(1-/spl alpha/). Simulations with the nonlinear Oja learning rule confirm our theoretical results.>

Read the paper · More papers on PaperTik