Data Distribution-Aware Model Aggregation for non-IID Data in a Federated Learning Framework

Deepali Kushwaha, Ananya Mehrotra, Rajesh Mahanand Hegde · 2024

An increase in dependence on data-driven technologies has raised user privacy concerns. Utilizing federated learning allows for the parallel processing of extensive amounts of data on edge devices, resulting in minimized latency and data privacy preservation. However, the issue arises when multiple devices with diverse datasets are involved, and conventional aggregation methods prove ineffective in achieving optimal global performance. Thus, an aggregation method is required to account for both heterogeneity in datasets and diverse data distributions to improve global model performance. This work presents a novel local model aggregation method for federated learning that recognizes the deviation between local and global data distribution to define local model aggregation weight. The Kullback-Leibler (KL) divergence measures the deviation in data distributions across classes. The proposed data distribution-aware aggregation (DDAA) method is evaluated on the Google Speech Commands (GKWS) dataset with three types of data distribution among devices: Independent and Identical Distribution (IID), semi-non-IID, and non-IID. Compared to the conventional Federated Average (FedAvg) aggregation method, the experimental results indicate reasonable improvements in classification accuracy for all three distributions. The proposed work is compared with several aggregation methods present in the literature, and the results show an improvement in F1 scores ranging from 0.02 to 0.13. Apart from performance improvements, the significance of the DDAA method is also elucidated by providing insights into the computational complexity compared to existing methods.

Read the paper · More papers on PaperTik