T-Flash: Topology-flexible latency-aware scheduling for hierarchical decentralized federated learning

Sudad H. Abed, Nasser R. Sabar, Abdun Naser Mahmood · Information Fusion · 2025

• T-FLASH reduces the convergence time by dynamically constructing a latency-aware hierarchical communication schedule that optimizes both computation and communication durations. • T-FLASH accelerates model convergence by leveraging an optimized communication schedule combined with weighted FedAvg aggregation, ensuring that contributions from all clients are incorporated in each round. • T-FLASH reduces the communication overhead on the client compared to fully P2P methods in different network topologies. Decentralized Federated Learning (DFL) is an emerging and increasingly popular approach for obtaining a generalizable AI model by simultaneously training several copies of the model across multiple distributed data. It improves privacy, scalability, and robustness by facilitating direct peer-to-peer (P2P) model training without a central server where clients holding the data communicate repeatedly to build an effective model despite the variations in the distributed data. However, in real-world scenarios such as healthcare systems, the heterogeneity of network topologies and client resources pose significant challenges in terms of computation and communication efficiency, leading to high latency and slow model convergence. This paper addresses these challenges by proposing T-FLASH, a new framework that optimizes DFL in a heterogeneous network topology. T-FLASH dynamically schedules clients across the network to construct a latency-aware hierarchical communication plan. It utilizes a shortest path finding algorithm for efficient client scheduling and weighted federated averaging for model aggregation. It accounts for arbitrary connectivity patterns, communication, and computation latency among clients, such as those observed in the healthcare sector, minimizing the time required for the global model to achieve convergence. The framework was validated using three benchmark datasets of varying sizes and features, demonstrating its effectiveness in improving model convergence time and reducing communication overhead on clients. The results highlight substantial efficiency improvements in model convergence, achieving average speedups of 2.7x over Gossip based communication and 1.7x over fully P2P in randomly connected networks, and 2.56x over Gossip based communication and 1.36x over fully P2P in fully connected networks, while also achieving higher accuracy. These findings underscore the potential of our approach, particularly in sensitive domains like healthcare, where data privacy is critical.

Read the paper · More papers on PaperTik