Information Theoretical Analysis of Artificial Neural Network training dynamics
Pieter Bouwman · Utrecht University Repository (Utrecht University) · 2020
In Shwartz-Ziv & Tishby (2017), during training of a feed forward neural network (FFNN), the mutual information between input layer X and output of the ith hidden layer Ti (an internal representation), I(X;Ti), was empirically observed to increase at the beginning of training and to decrease towards the end of training. Goldfeld et al. (2019), however, showed that in case of continuous modelling of Ti, mutual information, I(X;Ti), can only be a constant. Controversy ensued, which this thesis attempts to help solving, thereby meaning to contribute to the ongoing research into the training dynamics of artificial neural networks from an information theoretical perspective. This thesis attempts to get a theoretical hold on what was empirically observed in Shwartz-Ziv & Tishby (2017) about the behaviour of I(X;Ti), by modelling Ti as a discrete random variable. Modelling Ti in this way makes it possible to look into the kind of I(X;Ti) behavior that was empirically observed in Shwartz-Ziv & Tishby (2017). This modelling set-up is then used to investigate what factors may have an influence on the behavior of I(X;Ti) during FFNN training, in order to get a more precise idea of what gave rise to the pattern in the behavior of I(X;Ti) that was empirically observed in Shwartz-Ziv & Tishby (2017). In this modelling setup, I(X;Ti) will be calculated precisely. In all previous work I(X;Ti) was estimated. An explanatory variable is introduced, the amount of non-injectiveness in X → Ti, which can be seen to play a central role in the story by bridging between the mutual information I(X;Ti) and a number of factors that may be influencing I(X;Ti). Three factors, namely the representational precision of Ti, FFNN architectural choices and a clustering phenomenon, are identified that could possibly have influenced estimates of I(X;Ti) in previous work, thereby explaining the observations in Shwartz-Ziv & Tishby (2017) and Goldfeld et al. (2019). It is made plausible that these three factors influence I(X;Ti) by theoretically considering the effect the factors have on the aforementioned explanatory variable, the amount of non-injectiveness in X → Ti. Two of these factors are experimentally shown to influence I(X;Ti). The third factor, namely clustering in the internal representations of the hidden layers Ti of an FFNN, was found to influence I(X;Ti) in a systematic way in Goldfeld et al. (2019). Our experimental results show that there is a relation between this clustering factor and I(X;Ti), but this relation seems to be more complex than was found in Goldfeld et al. (2019).