Enhancing In-network Aggregation with Adaptive Gradient Quantization for Multi-tenant Learning
Huifeng Xing, Yinfan Hu, Hao Wang, Yang Chen, Zixuan Chen, Sen Liu, Yang Xu · 2025
With the increasing popularity of distributed training applications, the growth in network traffic has become an impediment to the communication among worker nodes in the system. In-network aggregation (INA) has emerged as a solution to improve communication efficiency by offloading gradient aggregation to switches. However, in multi-tenant scenarios, INA switch memory capacity has been identified as a main bottleneck, leading to reduced network throughput and slower training processes. To address this, we propose Adaptive Gradient Quantization (AGQ) on the switch. AGQ reduces the quantization bit-width of gradients, allowing for storage of more gradients within the limited switch memory while maintaining training accuracy. Compared to quantization on hosts, AGQ can swiftly adapt to the available memory on switches and offers an improved balance between minimizing precision loss and enhancing training throughput. We implement AGQ on a P4 switch testbed, and experimental results demonstrate that enabling AGQ can achieve an up to 100% increase in training throughput without explicit drop of training accuracy compared with existing INA solutions like ATP and host-based quantization methods like THC.