Compressing Key-Value Pairs for Efficient Distributed SGD Algotithm

Yazhou Ruan, Xuexin Zheng, Weiwei Zhao, Changhui Hu · 2023

When performing SGD in a distributed environment, a large number of local gradients need to be exchanged through the network, so the communication cost becomes a bottleneck of distributed machine learning, and compressing the transmitted gradients can effectively reduce the communication cost. In this paper, we propose a gradient compressing algotithm for distributed machine learning. In the method, gradients are stored as key-value pairs. For gradient values, it first classifies raw gradient values into buckets, converts them to bucket index, and then compresses the bucket index using the MinMaxSketch algorithm. Moreover, we introduce momentum and error accumulation to realize convergence. For gradient keys, it uses an incremental binary encoding method to reduce the storage cost. Finally, we prove the convergence of the compressing method. We experimentally demonstrate that our compression method has the best performance compared to existing methods.

Read the paper · More papers on PaperTik