Training distributed deep recurrent neural networks with mixed precision on GPU clusters

Alexey Svyatkovskiy, Julian Kates‐Harbeck, William Tang · 2017

In this paper, we evaluate training of deep recurrent neural networks with half-precision floats. We implement a distributed, data-parallel, synchronous training algorithm by integrating TensorFlow and CUDA-aware MPI to enable execution across multiple GPU nodes and making use of high-speed interconnects. We introduce a learning rate schedule facilitating neural network convergence at up to O(100) workers.

Read the paper · More papers on PaperTik