PrivTF: Efficient and Privacy-Preserving Transformer Inference With Optimized Protocols for Secure NLP
Ziwei Peng, Kaiping Xue, Bin Zhu, Yaxuan Huang, Jian Li, Meng Hao, Tianwei Zhang · IEEE Transactions on Network Science and Engineering · 2025
As many natural language processing services employ language models like Transformer in the cloud, privacy concerns arise for both users and model owners. Secure inference is proposed in the literature to perform the computing service securely. While prior schemes have studied convolutional neural networks and analogous models, they do not easily apply to Transformer because it requires more efficient protocols of different layers with large-sized weight matrices and complex non-linear functions on high-dimensional vectors. Therefore, we propose a privacy-preserving scheme PrivTF for secure transformer inference. PrivTF consists of protocols for Transformer-unique layers (encoder/decoder layer and their attention sub-layer) with our specialized design. Specifically, we present protocols of base sub-layers to address the heavy performance overhead: a new protocol of embedding layer, protocols of matrix multiplication and softmax that make use of mixed bitwidths, and protocols of GeLU functions and normalization. Analysis and experiments demonstrate that our protocols outperform previous secure inference works and maintain accuracy on practical inference tasks.