DCFL: Non-IID Awareness Dataset Condensation Aided Federated Learning
Xingwang Wang, Shaohan Sha, Yafeng Sun · 2024
Federated learning (FL) is a decentralized learning paradigm wherein a central server iteratively trains a global model by utilizing clients who possess a certain amount of private datasets. The main challenge of FL lies in the fact that the client-side private data may not be identically and independently distributed (Non-IID), significantly impacting the accuracy of the global model. Existing methods tend to overlook analysis and utilize the characteristics of data complementary among clients due to privacy constraints. Intuitively, utilizing statistical distinctions among private data on the client side can help mitigate the Non-IID degree. Besides, the recent advancements in dataset condensation technology have inspired us to investigate its potential applicability in addressing Non-IID issues while maintaining privacy. Motivated by this, we propose DCFL which divides clients into groups by using weight similarity measurement method, like Centered Kernel Alignment (CKA), to approximate and represent data similarity. The private data from clients within the same group be complementary and then can use dataset condensation methods with Non-IID awareness to complement clients. Additionally, filtering mechanism, data enhancement techniques are incorporated to efficiently utilize condensed data, enhance model performance, and minimize communication time. Experimental results demonstrate that DCFL achieves competitive performance on popular federated learning benchmarks including MNIST, Fashion-MNIST, SVHN, and CIFAR-10 with existing FL algorithms.