Radical quantitative acceleration and block-dimension quantization precision analysis for stable video diffusion

Zelin Zhang, Liqi Jiang, Junmin Wu · 2025

Represented by the Stable Diffusion model family, in recent years, diffusion models (DMs), have become the mainstream choice for image generation-related tasks, including but not limited to text-to-image and image-to-video. Since video is one dimension higher than images, image-to-video tasks are the most time-consuming. Meanwhile, in today's diffusion model pipelines, almost all the time consumption comes from the U-Net. Therefore, to accelerate the inference speed of image-to-video, it is necessary to reduce the time consumption of the U-net. From the era of DNNs to the era of LLMs, quantization has been proven to be an effective acceleration method. This study reveals that the U-Net quantization schemes suitable for text-to-image or image-to-image are not suitable for the image-to-video U-Net family. We proposed an acceleration method that leverages the bandwidth advantages post-quantization combined with matrix block multiplication, and further achieves additional acceleration through decoupling, achieving a 10%-20% acceleration on different machines. Furthermore, a scheme is proposed to measure the effect and performance of quantization by blocks. The quantization method used is int8 static PTQ. SVD is selected as the research subject of this paper, using TensorRT 10.3.0 as the inference engine, with experiments conducted on Nvidia L40 and Nvidia A100.

Read the paper · More papers on PaperTik