Deploying Large AI Models on Resource-Limited Devices With Split Federated Learning

Xianke Qiang, Hongda Liu, Xinran Zhang, Zheng Chang, Ying‐Chang Liang · IEEE Transactions on Mobile Computing · 2026

Large Artificial Intelligence Models (LAMs) have delivered impressive capabilities, but deploying and fine-tuning them on resource-limited mobile edge devices remains challenging due to data privacy concerns, limited computation and memory resources, and prohibitive communication overhead. This paper proposes a novel framework, named Quantized Split Federated Fine-Tuning Large AI Model (SFLAM). By splitting the model across edge devices and an edge server, SFLAM assigns only lightweight computation to devices and offloads the remaining training to the server, substantially reducing the device-side memory footprint while keeping raw data local. However, high-dimensional intermediate activations impose a heavy uplink burden. SFLAM addresses this through activation quantization and derives a convergence bound showing that the upper bound decreases with increasing quantization bit-width. We further formulate an accuracy–energy efficiency objective jointly optimize transmit power, bandwidth allocation, and quantization bit-width to improve training efficiency over wireless links. Simulations under heterogeneous data and wireless conditions show that SFLAM improves training efficiency and scalability over baselines.

Read the paper · More papers on PaperTik