Non-Random Weight Initialisation in Deep Learning Networks for Repeatable Determinism

Richard N M Rudd-Orthner, Lyudmila S. Mihaylova · 2019

This research is examining the change in weight values of deep learning networks after learning. These research experiments require to make measurements and comparisons from a stable set of known weights and biases before and after learning is conducted, such that comparisons after learning are repeatable and the experiment is controlled. As such the current accepted schemes of random number initialisations of the weight values may need to be deterministic rather than stochastic to have little run to run varying effects, so that the weight value initialisations are not a varying contributor. This paper looks at the viability of non-random weight initialisation schemes, to be used in place of the random number weight initialisations of an established well understood test case. The viability of non-random weight initialisation schemes in neural networks may make a network more deterministic in learning sessions which is a desirable property in mission and safety critical systems. The paper will use a variety of schemes over number ranges and gradients and will achieve a 97.97% accuracy figure just 0.18% less than the original random number scheme at 98.05%. The paper may highlight that in this case it may be the number range and not the gradient that is effecting the achieved accuracy most dominantly, although there may be a coupling of number range with activation functions used. Unexpectedly in this paper, an effect of numerical instability will be discovered from run to run when run on a multi-core CPU. The paper will also show the enforcement of consistent deterministic results on an multi-core CPU by defining atomic critical code regions aiding repeatable Information Assurance (IA) in model fitting (or learning sessions).

Read the paper · More papers on PaperTik