FLUTE: FSS-Based Secure Two-Party LLM Inference Using Partial Transformer Encryption
Yujie Xue, Lin Liu, Yuchuan Luo, Bing Sun, Shaojing Fu · IEEE Transactions on Dependable and Secure Computing · 2025
Recently, transformer-based large language models (LLMs) have become mainstream, particularly when used as agents. However, users with exceedingly sensitive data cannot benefit due to a lack of high-performance devices to train LLMs or limited access to large companies' LLM APIs. Secure two-party computing (2PC), especially the recently popular function secret sharing (FSS), enables secure inference of LLMs by protecting both users' inputs and LLM parameters from leakage. In this paper, we propose${\sf FLUTE}$, the first FSS-based 2PC inference framework for LLMs with partial transformer encryption. We first identify a subset of transformer core blocks by simulating an adversary$\mathcal {A}$attempting to recover model parameters layer by layer, and then design GPU-friendly FSS protocols for each module within the core blocks, optimizing communication using the matrix multiplication protocol of${\sf ABY2.0}$. Security analysis shows that our partial encryption scheme provides security comparable to encrypting the entire LLM, thereby enhancing the performance-security trade-off of the entire end-to-end secure inference. The experimental results in the latest LLM, Llama 3.1-8B, show that${\sf FLUTE}$outperforms the SOTA (SIGMA) by$3\times$in both latency and communication, and it achieves even greater advantages in larger LLMs, such as Llama 3.1-70B.