Assessing the Implications of Data Heterogeneity on Privacy-Enhanced Federated Learning: A Comprehensive Examination Using CIFAR-10
Phoebe Joanne Go, Victor Calinao, Melchizedek Alipio · 2023
In the context of an increasingly digital society with pressing privacy concerns, our research investigates the effectiveness of privacy-preserving artificial intelligence solutions like Federated Learning. This study focuses on three main areas: Federated and Centralized Learning applications, the influence of data heterogeneity on client data accuracy, and the evaluation of contemporary federated algorithms in scenarios of extreme heterogeneity. Federated Learning, using the FedAvg algorithm, demonstrated superior testing accuracy (88.54%) over Centralized Learning (87.98%) on the non-heterogeneous CIFAR-l0 dataset, indicating its potential as an efficient, privacy-preserving solution for various machine learning applications. Additionally, our findings highlight an inverse relationship between data heterogeneity and Federated Learning model accuracy, underscoring the need for strategies to mitigate this challenge and boost model performance. Upon evaluating several federated learning algorithms under high data heterogeneity (alpha=1.0), SCAFFOLD and FedOpt outperformed FedAvg and FedProx, demonstrating the significance of algorithm design in addressing data heterogeneity. SCAFFOLD and FedOpt showcased greater communication efficiency, attributed to their faster convergence and fewer required communication rounds. This study offers invaluable insights into addressing data heterogeneity, improving communication efficiency, and enhancing federated learning's performance and applicability in real-world scenarios, thereby furthering privacy-preserving artificial intelligence research.