GT-SGD: A Novel Gradient Synchronization Algorithm in Training Distributed Recurrent Neural Network Language Models

Xiaoci Zhang, Naijie Gu, Robail Yasrab, Hong Ye · 2017

The emergence of Recurrent Neural Networks (RNNs) has resulted in the state-of-the-art efficiency and performance in many fields including language modeling, machine translation and speech recognition. RNNs are difficult to apply distributed training because of the communication bottleneck problem, which is caused by the large number of trainable weights. In this paper, a Gradient Truncated Stochastic Gradient Descent (GT-SGD) algorithm is proposed for distributed RNN training. This algorithm can aggressively reduce the communication workload across compute nodes, by truncating gradients that are below a certain threshold. We implement the GT-SGD algorithm on a cluster with multiple Graphic Processing Units (GPUs) for training a large-scale language modeling task. Experimental results demonstrate that the GT-SGD algorithm significantly improves the training efficiency with no performance loss in the obtained RNN model.

Read the paper · More papers on PaperTik