Federated Learning With Selective Knowledge Distillation Over Bandwidth-constrained Wireless Networks

Gad Gad, Zubair Md. Fadlullah, Mostafa M. Fouda, Mohamed I. Ibrahem, Nei Kato · 2024

Artificial Intelligence (AI) applications on Internet of Things (IoT) networks often involve relaying generated data to a server for deep learning training, which poses security risks to users' data. Federated Learning (FL) offers a distributed model training paradigm in which local data are kept at the edge and locally trained models are exchanged and aggregated by a server over several rounds to produce a global model. While successful, standard FL algorithms do not support heterogeneous local model design, an essential requirement, especially for resource-limited edge devices. Recently, Knowledge Distillation-based FL algorithms have provided model-agnostic FL to enable clients to independently design their local model and share soft labels instead of model parameters. KD-based FL algorithms are computationally expensive due to additional distillation training. We propose Federated Learning with Selective Knowledge Distillation (FedSKD) to address the limitations of system heterogeneity; and computation and communication demands. We evaluate different aspects of the proposed algorithm relative to baseline FL algorithms. Results show that FedSKD incurs significantly less per-round computation time and communication overhead relative to the considered model-based and KD-based FL algorithms.

Read the paper · More papers on PaperTik