Blade: Pushing the Performance Envelope of Asynchronous Federated Learning

Chen Ying, Baochun Li, Bo Li · 2024

Asynchronous federated learning (FL) has been proposed to decrease the training time in conventional FL where the communication paradigm is synchronous. Instead of aggregating after receiving updates from all the selected clients, an asynchronous FL server conducts aggregation without waiting for slow clients. Though superior to synchronous FL, the performance of existing works in asynchronous FL — measured by the wall-clock time of global training — leaves much to be desired, as the staleness of client updates may degrade the performance substantially. In this paper, we propose Blade, a new stalenessaware framework that seeks to push the performance envelope of asynchronous FL by designing new mechanisms in all important design aspects of FL training, including client selection, adaptive pruning, quantization, and update aggregation. Blade selects clients based on their staleness and the quality of their previous updates. Before reporting to the server, every client prunes its update with a pruning amount related to its staleness and quantizes the pruned update. When aggregating updates, Blade tunes the aggregation weight of each update according to its staleness and divergence from the previous global model. In an extensive array of performance evaluations with six benchmark datasets, Blade consistently showed its substantial performance superiority over its state-of-the-art competitors. It decreased the wall-clock training time by up to 64.6%.

Read the paper · More papers on PaperTik