An Application of Normalizer Free Neural Networks on the SVHN Dataset
P. S. S. Madhulika, Nalini Sampath · 2022 International Conference on Applied Artificial Intelligence and Computing (ICAAIC) · 2022
Batch normalization is a process in which it applies the normal distribution on the previous layer by subtracting its mean over that particular batch and then divides it by the standard deviation. During the training of the input data, the values of the weights of the input layers and the outputs from the activation function are constantly changing. The estimate changes from layers to layers. It has to be updated in order to reach the final stage and obtain results. This process is totally time consuming and also a lot of computations are needed to be done to achieve the output. When training the dataset in smaller batches it tends to ignore the importance of some parameters resulting in the underflow or overflow of weights and the model developed will not be stable and whatever the dataset applied it will either overfit or underfit the data resulting into a futile one. Hence adaptive gradient clipping technique is used along with the Normalizer free networks to prevent the problem of the Batch Normalization. The proposed approach uses normalizer free residual neural networks which are 8.7x times faster than the regular networks. Normalizer free networks are used to overcome the problems of batch normalization and the gradient clipping is used to prevent the exploding of the gradients. The Nfnets are sensitive to alpha value hence introduced a modified form called as adaptive gradient technique which clips the value at every step.