Keynote Lecture : Gradient compression for efficient distributed deep learning

Nikos Deligiannis · 2021

Summary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. Recent successful results in the field of artificial intelligence and machine learning are achieved with deep learning models that contain a large number of parameters and are trained using a massive amount of data. Training such deep networks in a single machine (given a fixed set of hyperparameters) can take weeks. An answer to this problem is data-parallel distributed training, where a deep model is replicated into several computational nodes that have access to different chunks of the data. This approach, however, entails high communication rates and latency because of the computed gradients that need to be shared among nodes at every iteration. We will elaborate on various gradient compression strategies proposed to address this bottleneck within distributed training, including gradient sparsification, quantization, and entropy encoding. We will also discuss error correction techniques that compensate for the errors introduced by gradient compression. Furthermore, we will present new communication strategies that explore the correlation of gradients across distributed nodes to achieve further improvements in reducing the communication rate and latency.

Read the paper · More papers on PaperTik