Is Split Learning Privacy-Preserving for Fine-Tuning Large Language Models?

Dixi Yao, Baochun Li · IEEE Transactions on Big Data · 2024

With the success of pre-trained large language models in various tasks, users, individuals and enterprises alike, may need to fine-tune these models with their own datasets. Split learning was proposed to divide the model and place a portion on each user's own device, and intermediate results in each iteration of training will be sent to the server to complete the forward pass. There were concerns in the literature about whether private data can be leaked by sending such intermediate results from the training process. In this paper, we conduct empirical studies on typical large language models, such as GPT-2, OPT, Llama, and Qwen, to show that in most situations, an honest-but-curious server is not able to reconstruct private data using such intermediate results. To find out the reason why large language models preserve data privacy better in these situations, we present our theoretical analyses on these empirical observations. In one special case, where a state-of-the-art existing attack can reconstruct data in the first iteration, we show that it can be easily defended with a simple but effective solution leveraging publicly accessible data.

Read the paper · More papers on PaperTik