A Comparative Study of Clustering Distance Metrics for Personalized Federated Learning in Human Activity Recognition

Xiaoxu Wen, Chenggang Lu, Ruixiang Hu, Yingrui Geng, Yan Wang, Hongnian Yu, Qiangsong Zhao · 2024

In Federated Learning (FL), Non-Independent and Identically Distributed (Non-IID) data across different clients presents a significant challenge for a global model to effectively generalize to all clients, thus affecting fairness across the system. To address this, the Clustered Personalized Federated Learning (CPFL) framework groups clients with similar data distributions and trains models that are tailored to the characteristics within each cluster. In this study, we experimentally evaluate the performance of four distance metrics - Cosine distance, Euclidean distance, Manhattan distance, and Mahalanobis distance - across varying numbers numbers of clusters within the CPFL framework, focusing on their impact on model accuracy and fairness across client models. Our results demonstrate that Euclidean distance not only improves model accuracy but also significantly reduces performance disparities between clients, exhibiting the lowest variance and enhancing fairness. In contrast, the other distance metrics introduce larger variance in certain scenarios, leading to greater differences in client model performance. This study provides empirical support for the use of distance metrics in CPFL, showing that selecting appropriate clustering methods can enhance both overall performance and fairness across client model.

Read the paper · More papers on PaperTik