Stability analysis of gradient-based training algorithms of discrete-time recurrent neural network

Yilei Wu · 2008

vii 4.2 Closed loop dynamics of hidden layer training of RNN . . . . . . . .4.3 Illustration of a local attractor basin of the RNN against a scalar estimated weight Ŵ (k) . . . . . . . . . . . . . . . . . . . . . . . . . .4.4 Flow chart of the RAGD algorithm for MISO RNN . . . . . . . . . .5.1 Time sequences of signal and noise and respective power spectrums .5.2 Squared training errors of the first 100 steps with the same set of random initializations for different algorithms . . . . . . . . . . . . .5.3 Squared training errors of full 3000 steps for different algorithms . . .5.4 Traces of normalization factors ρ v (k) and ρ w (k) . . . . . . . . . . . .5.5 Traces of learning rate α v (k) and α w (k) . . . . . . . . . . . . . . . . .5.6 Traces of hybrid adaptive learning rate β v (k) and β w (k) . . . . . . . .5.7 Traces of the Frobenius norms of RNN weights with the RAGD training 5.8 Sequences of the time series for training and evaluation . . . . . . . .5.9 Squared training errors of the first 100 steps with the same set of random initializations for different algorithms . . . . . . . . . . . . .5.10 Squared training errors of full 3000 steps for different algorithms . . .5.11 Traces of normalization factors ρ v (k) and ρ w (k) . . . . . . . . . . . .5.12 Traces of learning rate α v (k) and α w (k) . . . . . . . . . . . . . . . . .5.13 Traces of hybrid learning rate β v (k) and β w (k) . . . . . . . . . . . . .5.14 Traces of the Frobenius norms of RNN weights with the RAGD training101 5.15 Trace of model input: AMPRP sequence . . . . . . . . . . . . . . . .

Read the paper · More papers on PaperTik