Optimizing Fault Tolerance and Latency of Federated Learning Using Edge Servers and Pre-Trained Model
Rahul Haripriya, Amit Kumar, Manish Kumar Pandey, Nilay Khare, Jaytrilok Choudhary, Dhirendra Pratap Singh, Surendra Solanki, Akash Haripriya · IEEE Access · 2025
Federated learning (FL) presents a promising paradigm for decentralized machine learning, particularly well-suited for data-sensitive cyber-physical systems (CPS) where privacy preservation and low-latency inference are paramount. However, FL deployments at the edge face acute challenges in fault tolerance and communication efficiency due to the inherent unpredictability and heterogeneity of edge environments. This research addresses these issues through the development of a hierarchical federated learning framework that strategically combines clustered edge servers, the transfer-learning capabilities of pre-trained EfficientNet-B0 and TabNet models, and an enhanced aggregation approach based on FedProx to deliver superior robustness and efficiency. The proposed architecture introduces a multi-cluster aggregation mechanism, where edge servers within clusters coordinate local model updates before participating in higher-level aggregation, reducing inter-device communication overhead and overall latency. A distinctive contribution of this work is the explicit modeling and evaluation of system reliability by defining a “paralysis ratio” to quantify failed or unresponsive edge servers, and conducting real-time simulation experiments with incremental data delivery. Through this, the framework’s resilience is demonstrated, maintaining model accuracy above 93% even under up to 40% server paralysis. In comparative benchmarks against the established Multi-Cluster Hierarchical FL (MCHFL) baseline, the framework achieves a 4–6% increase in classification accuracy and a 25% reduction in average communication latency. Beyond quantitative improvements, the study analyzes hyperparameter trade-offs, notably the impact of the FedProx constraint ($\mu $), and highlights the practical value of real-time transfer learning EfficientNet-B0 for images and TabNet for complex time-series accelerating convergence under non-IID conditions. Collectively, these innovations advance FL by delivering a fault-tolerant, latency-aware, and transferable aggregation strategy validated through real-time simulation, significantly broadening the applicability of FL in complex, real-world edge and CPS deployments.