Quality-Aware Node Selection for Efficient Federated Learning Based on a Global Perspective
Nawraz Saeed Mohamed, Mohamed Ashour, Maggie Ezzat Mashaly · 2024
The quality of training data is a critical determinant of the performance, reliability, and predictive accuracy of any AI model. Federated learning provides a promising framework for developing global AI models in distributed computing environments, where individual data nodes can maintain data privacy while contributing to a shared model. Although each node can locally enhance the quality of its data, the distributed nature of the system does not guarantee global data quality. One of the main challenges arises from heterogeneity introduced by federated learning is data quality, which often result in global data degradation. Despite the increasing focus on the challenge of Non Independently and Identically Distributed (NIID) datasets, data quality is often regarded as a subset of this problem and not fully explored in its own right. In this paper, we introduce a comprehensive node scoring approach designed to enhance quality awareness without relaying on a shared reference dataset. The scoring approach focuses on assessing global data quality with respect to variations in data volume, label imbalance, global duplicates, and label preference skew. The proposed Quality scoring is utilized to either determine nodes respective weights in Fedq aggregation method. It is used also in sorting and sampling nodes for further selections. To evaluate the proposed node scoring method, the paper introduces a simulation model based on the MNIST dataset. The introduced model simulates extreme low level of data quality nature of the federated data and uses this data to evaluate the proposed scoring approach versus the baseline FedAvg and Irrelevance score in terms of global model accuracy, model stability and model convergence with particular attention to the content diversity score, which quantifies global duplicity.