Improving Federated Transfer Learning for Different Levels of Non-IID Data with Multi-Round Transfer and Forgetting Mitigation
Boyuan Zhang, Mohammad Shikh‐Bahaei · 2025
Federated Learning (FL) enhances data privacy by allowing clients to share only model parameters. However, it suffers from slow convergence speed and low accuracy under non-Independent and Identically Distributed (non-IID) data conditions due to the statistical heterogeneity across clients. To address this issue, Federated Transfer Learning (FTL) was introduced, which leverages pre-trained feature extraction layers from a subset of clients and transfers them to others, thereby improving convergence and reducing communication costs in lowly and moderately non-IID settings. Nevertheless, FTL remains ineffective under highly non-IID conditions, as the transferred representations struggle to generalize across diverse data patterns. In this paper, we systematically evaluate the performance of FL and FTL under varying non-IID levels and identify their limitations in extreme scenarios. We then propose two novel approaches: Multiple-round Federated Transfer Learning (MFTL) and Multiple-round Federated Transfer Continual Learning (MFTCL). MFTL employs iterative rounds of Transfer Learning (TL) to enhance robustness, while MFTCL integrates Continual Learning (CL) strategies by a replaying mechanism to mitigate catastrophic forgetting. Experimental results on a radar sensing image dataset, where varying numbers of clients are used during the pre-training phase to control feature learning under different levels of non-IID conditions, demonstrate that our methods outperform traditional FL and FTL in terms of convergence speed, accuracy and communication burden.