Zebra: Accelerating Distributed Sparse Deep Training With in-Network Gradient Aggregation for Hot Parameters

Heng Pan, Penglai Cui, Zhenyu Li, Ru Jia, Penghao Zhang, Leilei Zhang, Ye Yang, Jiahao Wu, Mathy Lauren, Gaogang Xie · 2024

Distributed sparse deep learning has been widely used in many Internet-scale applications. Network communication is one of the major hurdles for training performance. In-network gradient aggregation on programmable switches is a promising solution for speeding up the performance. Nevertheless, existing in-network aggregation solutions are designed for the dense deep training, and fall short when used for the sparse training. To address this gap, we present Zebra based on our key observation on the extremely biased update frequency of parameters in distributed sparse deep training. Specifically, Zebra offloads only the aggregation for “hot” parameters that are updated frequently onto programmable switches. To enable this offloading and achieve high aggregation throughput, we propose solutions to address the challenges related to hot parameter identification, parameter orchestration and gradient aggregation as well as system reliability. We implemented Zebra on Intel Tofino switches and integrated it with PS-lite. Finally, we evaluate Zebra's performance through extensive experiments and show that it can speed up the gradient aggregation by$1.5 \sim 4 \times$and the end-to-end performance by$1.4 \sim 2.6 \times$.

Read the paper · More papers on PaperTik