Delay Compensated Asynchronous Adam Algorithm for Deep Neural Networks
Naiyang Guan, Lei Shan, Canqun Yang, Weixia Xu, Minxuan Zhang · 2017
Deep neural network (DNN) learns hierarchical representations from big data in a multilayer network structure and has achieved great successes in many fields such as computer vision and speech analysis. Since DNN usually contains several billions of parameters, the asynchronous stochastic gradient descent (ASGD) algorithm is often used to train an effective DNN model on a computer cluster. However, as the increase of computing nodes and data size, ASGD suffers from serious slow convergence deficiency because the parameters might be wrongly updated by long-term delayed gradients. In this paper, we propose a delay compensated asynchronous Adam (DC-Adam) algorithm to train DNN. In particular, DC-Adam updates the parameters with the moment increment which is the division of the first and the second moments to take the advantage of the original Adam algorithm, and compensates the gradient with the first-order component in its Taylor expansion. Since the delay compensation technique reduces the error of delayed gradients, and the moment increment further counteracts the influence of approximated compensation, DC-Adam converges much more rapidly than ASGD on a computer cluster with moderate computing nodes. We theoretically analyze the Ergodic convergence rate of DC-Adam and compare with DC-ASGD. We implement our DC-Adam algorithm on 61 computing nodes in a computer cluster, and conduct image classification by using LeNet and ResNet, respectively, on the MNIST and CIFAR-10 datasets. The experimental results demonstrate that DC-Adam greatly accelerates the training progress and achieves almost linear speedup rate as increasing the computing nodes.