ABS: Adaptive Bounded Staleness Converges Faster and Communicates Less

Qiao Tan, Feng Zhu, Jingjing Zhang, Shengyun Liu · Research Square · 2023

Abstract The convergence time and communication rounds are critical performance metrics in distributed learning with a parameter-server (PS) setting. Local stochastic gradient descent (SGD) is a widely used method for improving communication efficiency in distributed learning. However, synchronous local SGD can experience substantial slowdowns, forcing stragglers to perform multiple local updates and the other nodes to keep waiting. Asynchronous methods are immune to these slowdowns but can suffer from gradient staleness, leading to suboptimal results or even divergence. To address these challenges, we propose a novel asynchronous strategy named adaptive bounded staleness (ABS). ABS leverages two key enablers. Firstly, the number of workers that the PS waits for per round for gradient aggregation is adaptively selected to strike a balance between straggling and staleness. Secondly, workers with relatively high staleness are prompted to initiate a new round of computation, alleviating the negative effects of staleness. Simulation results demonstrate that ABS outperforms state-of-the-art schemes in terms of wall-clock time and communication rounds.

Read the paper · More papers on PaperTik