Large Scale Delocalized Federated Learning Over a Huge Diversity of Devices in Emerging Next-Generation Edge Intelligence Environments

Mahdi Morafah, Hojin Chang, Bill Lin · 2024

Prior research in Federated Learning (FL) has primarily focused on a setting in which all client models are identical (model-homogeneous). However, in practice, client devices may be diverse and heterogeneous, ranging possibly from small devices (e.g. IoTs) to large devices (e.g. supercomputers), each device class possessing distinct computing power, memory capacities, and configurations (device-heterogeneity). In this paper, we take a substantial step forward by defining a new problem wherein clusters of heterogeneous devices, each possessing distinctive models, memory capacities, computational resources and FL configurations, collaborate to enhance each other's global model within the federated learning paradigm. In particular, we propose FedHD (Federated Learning for Heterogeneous Devices), a knowledge distillation (KD)-based approach to address this new problem setting. Our approach leverages heterogeneous ensembles from all device clusters, to transfer knowledge between these clusters using a publicly available unlabeled dataset hosted at the server. To maximize knowledge exchange, we introduce an adaptive weighting strategy that assigns weights to each cluster's ensembles (as teachers) based on their model sizes relative to others. The core insight is that larger ensembles are more effective as teachers when transferring knowledge to smaller models, while smaller models are less effective in instructing larger models. Our experimental results demonstrate that FedHD achieves state-of-the-art performance compared to several baselines across various datasets and model architectures.

Read the paper · More papers on PaperTik