Aligning Text-to-Image Diffusion Models With Noise-Conditioned Perception
Alexander Gambashidze, Anton Kulikov, Yuriy Sosnin, Ilya Makarov · IEEE Access · 2025
Human preference optimization, originally developed for Language Models, has shown promise in improving text-to-image Diffusion Models by enhancing prompt alignment, visual appeal, and user preference. However, Diffusion Models are typically optimized in pixel or VAE space, which often misaligns with human perception, resulting in slower and less efficient training during fine-tuning and preference optimization. In this work, we demonstrate that using a perceptual objective significantly enhances both training speed and overall model quality. We fine-tune Stable Diffusion 1.5 and XL using Direct Preference Optimization (DPO), Contrastive Preference Optimization (CPO), and supervised fine-tuning (SFT) within this perceptual embedding space. Our approach significantly outperforms standard latent-space implementations across various metrics, including quality and computational cost, when training on the widely-used Pick-a-Pic dataset. For SDXL, our method achieves a 64.6% general preference over the baseline DPO on the PartiPrompts dataset while significantly reducing compute to reach comparable DPO performance. Additionally, we enhance the Pick-a-Pic dataset, making it approximately 10 times smaller while training models that surpass the original published versions in just 12 and 3.5 GPU hours for SDXL and SD1.5, respectively. This paper demonstrates that the overall quality of Diffusion models after fine-tuning can be significantly improved while also being more efficient and requiring far less data.