Quantization as a Foundation for Deployable High Performance Diffusion Models within the Landscape of Large Scale Generative AI
Mikkel Sørensen, Freja Holm, Anders S. Kristensen, Sofie Jensen, Lars Peter Madsen, Emilie Rasmussen, Chand Aline · 2025
Diffusion models have emerged as a cornerstone of large-scale generative artificial intelligence, underpinning some of the most advanced text-to-image, text-to-video, and multimodal generation systems. Their remarkable capabilities, however, come at the cost of immense computational and memory demands, rendering them difficult to deploy outside of specialized high-performance computing infrastructures. Quantization, the process of reducing numerical precision of model parameters, activations, and intermediate computations, has consequently attracted growing interest as a strategy for making diffusion models more efficient without prohibitive loss in generative quality. Yet, unlike in discriminative or autoregressive generative models, quantization in diffusion-based architectures presents unique challenges due to the iterative, stochastic nature of the sampling process and the sensitivity of score estimation to numerical perturbations. As such, the study of quantization in diffusion models is not a straightforward extension of existing compression paradigms but rather a new frontier that requires theoretical, algorithmic, and systems-level innovation. This review synthesizes the state of the art in diffusion model quantization, tracing developments across algorithmic techniques, empirical benchmarks, hardware considerations, and emerging applications. We begin by situating quantization within the broader history of efficient deep learning, highlighting how the particular characteristics of denoising diffusion probabilistic models introduce new forms of sensitivity to quantization noise. We then provide a detailed examination of quantizationaware training, post-training quantization, mixed-precision methods, adaptive bit-width allocation, and hybrid compression strategies, emphasizing their respective advantages, trade-offs, and limitations. Special attention is given to the mathematical underpinnings of error accumulation in iterative stochastic processes, the role of scale-normalization mechanisms in mitigating instability, and the interplay between quantization granularity and multimodal alignment fidelity. Beyond algorithmic advances, we survey hardware and systems-level aspects, underscoring how deployment efficiency depends critically on accelerator architectures, memory bandwidth, and kernel optimization. We further identify key open challenges, including the absence of standardized evaluation protocols, the difficulty of theoretically modeling quantization-induced error propagation, the limited generalization of methods to modalities beyond images, and the need for hardware–algorithm co-design. Moreover, we discuss broader societal and ethical implications, considering how efficient deployment may democratize access to generative AI while simultaneously raising risks of misuse and amplifying hidden biases. By synthesizing these dimensions, this review positions diffusion model quantization not as a peripheral technical concern but as a central question at the intersection of algorithm design, hardware optimization, and responsible AI deployment. Ultimately, we argue that the future of quantized diffusion models lies in moving beyond heuristic-driven approaches toward principled, automated, and hardware-aware quantization pipelines that integrate ethical safeguards. The convergence of quantization with pruning, distillation, and architectural reparameterization—together with new forms of accelerator design—offers a path toward generative AI systems that are not only powerful but also efficient, accessible, and sustainable. As such, diffusion model quantization represents both a practical necessity for real-world deployment and a conceptual catalyst for rethinking how generative AI can scale responsibly in the years to come.