Adaptive Knowledge Transfer for Federated Learning With Large Models on Edge IoT Devices

Benteng Zhang, Yingchi Mao, Haotian Zheng, Xiang Li, Jianxin Huang, Xiaoming He, Jie Wu · IEEE Internet of Things Journal · 2025

Cloud-centric deployment of large models brings privacy concerns and communication burdens to IoT devices. Federated Learning (FL) can help numerous IoT devices collaboratively train a large model in a distributed manner to provide intelligent services and applications (AIoT). However, due to the limited storage capacity of IoT devices, fresh data collected by IoT devices often overwrites outdated data and establishes Non-IID data distributions. Moreover, state-of-the-art studies have indicated that FL tends to use fresh data for model training. This causes the global model to forget outdated data’s characteristics (i.e., catastrophic forgetting). Some methods utilize Knowledge Distillation (KD) to extract and integrate characteristics from both fresh data and outdated data. However, existing KD-based methods use fixed distillation temperatures for different IoT devices, which overlooks that fixed distillation temperatures cannot match the data distributions on different IoT devices and may degrade global model accuracy. To this end, we propose a Federated Dynamic Decoupled Distillation method based on Logits distribution (Fed3DL). Specifically, Fed3DL utilizes decoupled distillation to mitigate catastrophic forgetting. To dynamically adjust the distillation temperature of each IoT device, Fed3DL novelly builds an adaptive temperature-aware mechanism based on the Logits distribution. Furthermore, Fed3DL introduces a regularization term into the local distillation loss to improve global model accuracy. Experiments on four datasets show that Fed3DL can improve the global model accuracy by an average of 3.19%, reduce the forgetting rate by an average of 4.46%, and achieve the lowest inter-class accuracy disparity.

Read the paper · More papers on PaperTik