Training Deep Neural Networks
Mohamed Abdel‐Basset, Nour Moustafa, Hossam Hawash · 2022
This chapter describes the backpropagation algorithm of training neural networks and its implementation. It explains how deep neural networks are trained. First, gradient descent is revisited by exploring different variants including the difference between them, and their advantages and disadvantages. The chapter looks at the gradient-vanishing and the gradient-exploding problems, which are the major challenges facing the successful training of any deep networks. It investigates the main methods (such as gradient clipping, nonlinear activations, etc.) to be adopted in order to avoid the problems during the training. The chapter provides a deep dive into the state-of-the-art parameter initialization methods. It also investigates the state-of-the-art optimization algorithms for deep networks along with the distinct characteristics of each of them. The chapter provides a detailed exploration of the state-of-the-art optimizers including momentum optimizers, RMSProp optimizers, AdaGrad optimizers, and Adam optimizers.