Dual-Way Gradient Sparsification for Asynchronous Distributed Deep Learning
Zijie Yan, Danyang Xiao, Mengqiang Chen, Jieying Zhou, Weigang Wu · 2020
Distributed parallel training using computing clusters is desirable for large scale deep neural networks. One of the key challenges in distributed training is the communication cost for exchanging information, such as stochastic gradients, among training nodes. Recently, gradient sparsification techniques have been proposed to reduce the amount of data exchanged and thus alleviate the network overhead. However, most existing gradient sparsification approaches consider only synchronous parallelism and cannot be applied in asynchronous distributed training.