Resilient, Federated Large Language Models over Wireless Networks: Why the PHY Matters
Vlad C. Andrei, Aladin Djuhera, Xinyang Li, Ullrich J. Mönich, Walid Saad, Holger Boche · 2024
In this paper, the problem of training large language models (LLMs) in split federated learning over real-world wireless networks is investigated. In the considered system, the embedding layers of an LLM are first computed at a client and then trans-mitted over a wireless MIMO-OFDM link to a server instance for further processing, continuing the forward- and initiating the backpropagation of the training to the originating client. Due to channel impairments and adversarial attacks, the server needs to compute the model losses and gradients using corrupted parameters such as embeddings in LLMs. The computation of the corresponding model losses is rigorously characterized using such perturbed embeddings and a direct connection to the communication mean-squared error (MSE) for models beyond simple neural networks is established. Subsequently, the communication errors are modeled as part of the training process, and a method to design beamforming, scheduling and power allocation is proposed, ensuring high task performance and model convergence even in the case of worst-case jamming. Results on two natural language processing tasks using different LLM architectures confirm the validity of the theoretical analysis and prove the effectiveness of the proposed wireless system design in terms of accuracy and F1 score.