DirectReduce: A Scalable Ring AllReduce Offloading Architecture for Torus Topologies
Lihuan Hui, Wang Yang, Fan Wu, Yanbo Wang, Feng Lyu, Yaoxue Zhang · IEEE Internet of Things Journal · 2025
The all-reduce operation is critically important for communication-intensive workloads emerging at the convergence of High-Performance Computing (HPC) and Internet of Things (IoT) applications. However, existing optimization efforts primarily concentrate on offloading the all-reduce onto network switches, known as In-Network Aggregation, which are incompatible with switchless torus topologies. Driven by our systematic analysis, we identified two key factors that impact the performance of the standard ring all-reduce operation: i. The all-reduce computation process frequently interrupts the GPU/CPU’s computation tasks; ii. The GPU/CPU is, in fact, indifferent to intermediate computational results. Based on this insight, we propose DirectReduce, a fully offloading ring all-reduce architecture that is comprised of three components: (i) the GateKeeper module, responsible for evaluating outgoing data to decide its progression-either directing it to the Protocol Engine for packetization or intercepting it for reduction (e.g., sum, maximum); (ii) the DataDirector module, which classifies the incoming data either is for intermediate result reduction or final result storage; and (iii) the ComputeEnhancer module, designed to execute reduction operations directly on the SmartNIC. Extensive simulation results show that DirectReduce can reduce the ring all-reduce latency by up to 1.98X in a ring (1D-torus) topology, 1.97X in a 3D-torus topology, and 1.75X in a 6D-torus topology compared to the standard ring all-reduce.