FedKDShap: Enhancing Federated Learning via Shapley Values Driven Knowledge Distillation on Non-IID Data

Nazmus Shakib Shadin, Xinyue Zhang · 2025

Federated Learning (FL) has achieved significant popularity in privacy-preserving distributed learning, wherein data remains on edge devices, ensuring data security and user privacy. Regardless of its advantages, FL has considerable challenges, including the non-independent and identically distributed (non-IID) nature of data across clients, degradation in performance compared to centralized learning methods, and communication efficiency issues due to frequent exchanges of large model updates between clients and the server. To overcome these challenges, FL combined with knowledge distillation (KD), which is a decentralized learning methodology that facilitates collaborative model training across several devices or clients with data privacy. Traditional KD generally focuses on the transfer of knowledge through logits; however, this methodology omits the importance of intermediate feature representations within the model. To address this limitation, we propose incorporating Shapley Additive Explanations (SHAP) or Shapley values into knowledge distillation (KD) methods. Shapley values quantify feature importance, enabling the transfer of critical feature contributions, and thereby enhancing the effectiveness of KD. In this work, we propose a novel decentralized machine learning approach, named FedKDShap, refined federated learning with Shapley values-informed KD which prioritizes performance, interpretability, resource efficiency, and feature importance within the distillation process to optimize knowledge transfer from a high-capacity teacher model to a lightweight student model. This integration not only reduces communication demands on resource-constrained devices but also enhances model convergence in non-IID data settings by embedding Shapley values into the KD loss function. Our experiment leverages benchmark datasets to simulate real-world non-IID data distribution which also demonstrates that the FedKDShap method enhances model accuracy and outperforms state-of-the-art architectures. Our code repository is available at: https://github.com/shadhin39/FedKDShap

Read the paper · More papers on PaperTik