Generalization vs Personalization: A Trade-off for better Data Heterogeneity impact Mitigation in FL

Sinda Besrour, Gael S. Mubibya, Chayma Ben Abdeljelil, Jalal Almhana · 2024

Federated learning (FL) was introduced recently as a new machine learning (ML) paradigm. It is a distributed network of client nodes that train ML and deep learning (DL) models on their local data without sharing them to preserve data privacy (DP). However, these data are heterogeneous by nature as they are collected in different contexts using various sources such as IoT devices. Consequently, data heterogeneity (DH) in FL has brought new performance-related challenges. Few of these challenges have been addressed in the literature; moreover, context heterogeneity and balance rate were not explored at all. In this paper, we introduce an FL approach in which a trade-off between personalization and generalization is achieved to mitigate the impact of DH and obtain better performance. We focus on three DH challenges: context, non-independent and identically distributed (non-IID) data, and balance rate. For the implementation, fall detection (FD) data is used to demonstrate the potential of our approach in improving the FL system’s performance. FD is an important subject and is particularly prevalent for the safety of elderly people. Hence, we collected fall data from two sensors: accelerometer (ACC) and heart rate (HR), then, we used two ML models to evaluate our approach. We utilized XGBoost (XGB) for balanced and unbalanced clients and One-Class Support Vector Machine (OC-SVM) for one-label clients. Our approach achieved an average F1-score of 88%. A comparative study was also conducted with previous works on FD. Our results showed a performance improvement which exceeded 94.30% on average.

Read the paper · More papers on PaperTik