Revisiting the In-Network Aggregation in Distributed Machine Learning

Haowen Zhu, Zehua Guo · IEEE Transactions on Networking · 2025

Distributed Machine Learning (DML) is proposed to accelerate machine learning model training by utilizing multiple training nodes to train models in parallel. Recent studies apply emerging In-Network Aggregation (INA) to further improve training efficiency by offloading the gradient aggregation process from hosts to programmable switches. However, existing INA solutions neither provide high training performance due to inefficient gradient aggregation with a single switch nor are easily deployed because of modifying the protocol stack in hosts. In this paper, we propose an easily deployable INA-based solution called Hierarchical In-Network Aggregation (HINA) to accelerate DML training process by hierarchically performing multiple aggregations in the Data Center Network (DCN). We formulate the gradient aggregation problem as the Joint Gradient Routing and Sending Rate (JGRSR) problem, which is a Mixed Integer Linear Programming (MILP) problem with high computation complexity. In addition, we propose HINA using progressive rounding and randomized rounding to determine the paths of gradient flows and the sending rates of training nodes to simplify and solve the JGRSR problem. Simulation results show that HINA reduces communication time by 61%-92% and decreases network load by 39%-66%, compared with state-of-the-art solutions.

Read the paper · More papers on PaperTik