Serene: Handling the Effects of Stragglers in In-Network Machine Learning Aggregation
Diego Cardoso Nunes, Bruno Loureiro Coelho, Ricardo Parizotto, Alberto Egon Schaeffer-Filho · 2023
Achieving high-performance aggregation is essential to scale data-parallel distributed machine learning (ML) training. Recent research efforts in the area of in-network computing have shown that offloading the aggregation to the network data plane can accelerate the aggregation process compared to traditional server-only approaches, reducing the propagation delay and consequently speeding up distributed training. However, the existing literature on in-network aggregation does not provide ways to deal with slower workers (called stragglers). The presence of stragglers can negatively impact distributed training, increasing the time it takes to complete. In this paper, we present Serene, an in-network aggregation system capable of circumventing the effects of stragglers. Serene coordinates the ML workers to cooperate with a programmable switch according to a hybrid synchronization approach. We also employ an efficient data structure for managing synchronization. We implemented and evaluated a prototype using BMv2 and realistic ML workloads, including a neural network trained for image classification. Our preliminary results show that Serene can speed up training by up to 40% in emulation scenarios.