DHC: Distributed Homomorphic Compression for Gradient Aggregation in Allreduce

Lida Liao, Zhengli Lin, Haodong Chen, Longlong Zhu, Hongyan Liu, Jiashuo Yu, Dong Zhang, Chunming Wu · 2025

Distributed training is critical for efficiently developing deep neural networks (DNNs) on tasks like image classification and natural language processing. However, as model and dataset sizes continue to grow, high communication overhead during gradient exchanges has become a major bottleneck in distributed training. Although existing homomorphic compression frameworks effectively reduce communication overhead, their reliance on centralized architectures makes them unsuitable for the mainstream decentralized AllReduce architecture. To address this, we propose DHC, a framework for homomorphic gradient compression in AllReduce architectures. Its key idea is HG-Sketch, which leverages multi-level index tables for direct in-network aggregation of compressed gradients, thereby eliminating additional computational overhead. Additionally, DHC introduces an index-sharing method to optimize memory usage on programmable switches. Furthermore, we establish an Integer Linear Programming (ILP) model to optimize the deployment strategy of programmable switches, further enhancing in-network aggregation capabilities. Experimental results demonstrate that DHC achieves a$3.8 \times$increase in aggregation speed and a$4.2 \times$improvement in aggregation throughput.

Read the paper · More papers on PaperTik