Learnable weight initialization in neural networks
Aninda Bhattacharya · Research Repository (Delft University of Technology) · 2020
A new method of initializing the weights in deep neural networks is proposed. The method follows two steps. First, consider each layer as a model and perform a linear regression to keep the mean of the layer output to zero and variance after the data is passed through the activation function to one. Once each layer converges to the target mean and variance, initialize the weights of the original model with the learned weights. Performance is evaluated on LeNet and ResNet18 architectures on FashionMNIST and Imagenette datasets. The activation functions used to analyze the performance are sigmoid, tanh and ReLU. Findings show that the learned weights can perform similarly, and for certain scenarios, better than the different types of weight initializers used frequently in the field of deep learning. It is important to mention that this method requires the weights to be learned independently of the training of the model. Thus, there is a small time overhead. Moreover, it is required to adjust the hyperparameters(learning rate, epochs etc) to find the optimal weights. The findings from this thesis can be used in the future to better understand how the gradient flow could be controlled through the network and finding a more generic approach towards the vanishing and exploding gradient problem. The method requires to learn the weights followed by training the network. Learning the weights involves tweaking the hyperparameters(learning rate, number of epochs, etc). For future work, these aspects could be automated for the optimal performance of the network.