BMP-SD: Marrying Binary and Mixed-Precision Quantization for Efficient Stable Diffusion Inference

Cheng Gu, Gang Li, Xiaolong Lin, Jiayao Ling, Jian Cheng, Xiaoyao Liang · 2025

Stable Diffusion (SD) is an emerging deep neural network (DNN) model that has demonstrated impressive capabilities in generative tasks such as text-to-image generation. However, the iterative denoising stage of the SD model is extremely expensive in both computations and memory accesses, making it challenging for fast and energy-efficient edge deployment. To alleviate the overhead of denoising, in this paper we propose BMP-SD, a post-training quantization framework for hardware-efficient SD inference. BMP-SD employs binary weight quantization to significantly reduce the computational complexity and memory footprint of iterative denoising, along with dynamic step-aware mixed-precision activation quantization, based on the observation that not all denoising steps are equally important for a specific input prompt. Experiments on the text-to-image generation task show that BMP-SD achieves mixed-precision (W1.73A4.87) with minimal accuracy loss on MS-COCO 2014 dataset. We also evaluate the BMP-SD quantized model on three state-of-the-art bit-flexible DNN accelerators, results reveal that our method can deliver up to 5.14× performance and 3.85×energy efficiency improvements compared to W8A8 quantization.

Read the paper · More papers on PaperTik