Menos: Split Fine-Tuning Large Language Models with Efficient GPU Memory Sharing

Chenghao Hu, Baochun Li · 2024

Fine-tuning of pre-trained large language models has become increasingly popular, yet existing fine-tuning methods are typically centralized, requiring users to send local data to centralized servers, or model owners to open-source their models. However, data and models are valuable assets that few enterprises and users wish to share. In this paper, we deviate from conventional wisdom and advocate the use of split learning for fine-tuning models with private data, local to each of the clients. The most formidable challenge to split fine-tuning is the size of large language models: when multiple clients start their fine-tuning tasks, their use of GPU memory will overwhelm a GPU-equipped server, especially as the number of clients scales up. To address this challenge, we present Menos, the first memory-efficient split fine-tuning framework designed to optimize the server GPU footprint through spatial and temporal sharing. Specifically, Menos utilizes the adapter-based nature of modern fine-tuning techniques, and proposes to spatially share the base model parameters among multiple clients. It also schedules memory-intensive operations during the communication gaps of split learning, thereby temporally sharing limited GPU memory at runtime. Comprehensive real-world evaluations using state-of-the-art large language models demonstrate the effectiveness of Menos, reducing GPU memory consumption by up to 72%, yet incurring negligible overhead.

Read the paper · More papers on PaperTik