Quantizing Diffusion Models for Scalable and Efficient Generative Inference Across Diverse Hardware Platforms

Markus Feldner, Jiawei Zhou, Leila Foroughi, Samuel Nartey, Erika Nishimura, Tomás Herrera, Nadine Meier, Chand Aline · 2025

Diffusion models have recently emerged as a dominant framework in generative modeling, achieving unprecedented performance in high-fidelity image synthesis, video generation, and multimodal tasks. Despite their success, these models remain computationally intensive and memory-heavy, which significantly hinders their deployment in real-world scenarios, particularly on edge devices and in latency-sensitive applications. Quantization—the process of reducing the numerical precision of model weights, activations, or gradients—offers a promising avenue to mitigate these limitations by enabling more efficient inference with reduced resource consumption. However, quantizing diffusion models poses unique challenges that differ markedly from those encountered in traditional classification or language models. These challenges arise from the multi-step nature of the generative process, the sensitivity of score-based sampling to numerical approximation, and the architectural complexity of components such as timestep-conditioned U-Nets and attention mechanisms. This review provides a comprehensive and detailed exploration of the current landscape of diffusion model quantization. We systematically examine the theoretical underpinnings of diffusion processes and how they interact with various quantization schemes, including post-training quantization, quantization-aware training, mixed-precision quantization, and adaptive bitwidth techniques. We analyze the trade-offs between model accuracy, perceptual fidelity, computational efficiency, and hard-ware compatibility, drawing on extensive empirical evidence and recent benchmarks. Moreover, we highlight the interplay between quantization and emerging areas such as architectural co-design, distillation, calibration techniques, and compiler toolchain development. Special emphasis is placed on evaluating quantization methods in light of practical deployment constraints, robustness to distributional shifts, and ethical considerations such as fairness and accessibility. We also articulate a set of future research directions that include quantization-aware generative training objectives, learning bitwidth allocation strategies, and building scalable, energy-efficient inference pipelines for diffusion models. Finally, we explore the societal and environmental implications of quantization, arguing that it is not merely an engineering optimization but a key enabler of responsible, inclusive, and sustainable generative AI. Through this review, we aim to provide both a foundational understanding and a forward-looking perspective on quantizing diffusion models, guiding researchers and practitioners toward more efficient and ethical deployment of generative technologies.

Read the paper · More papers on PaperTik