PPQformer: Privacy-Preserving Quantized Transformer for Efficient and Secure Inference.
Benchang Dong, Zhili Chen · 2025
In the context of the popularity of Transformer model services, the issue of privacy protection has gradually gained attention. However, the primary issues with existing private inference solutions lie in memory constraints and suboptimal computational performance. In this work, we introduce PPQformer, a framework that leverages Replicated Secret Sharing (RSS) and quantization techniques to enable efficient and secure inference of Transformer models in a Multi-Party Computation (MPC) setting. By integrating quantization into the MPC protocols, PPQformer significantly reduces memory footprint and enhances computational performance while preserving data and model privacy. Experimental results demonstrate the effectiveness of our approach, achieving competitive accuracies with reduced communication cost and inference time compared to existing methods.