Enhancing Denoising Models Performance Through Timestep-Aware Predictor for Latent Diffusion-Based Image Generation
Jong-U Kim, Dong‐Joong Kang · 2023
In various deep learning domains, transformer-based models have shown high performance, particularly in the field of generative modeling, where stable diffusion models have become prominent. In diffusion models, the process begins by applying noise to the original images using a scheduler. The denoising process is then trained to generate clean images from the noisy inputs. And the noise application step is not learned, focusing solely on training the denoising process to improve image generation. One notable denoising model is the time-embedded U-Net, which incorporates time embedding into its architecture. In this study, we propose enhancing the performance of the time-embedded denoising U-Net by introducing a predictor specifically designed for timesteps. The input images are compared based on the timestep error, while the output images are compared with a fixed timestep of 0, analyzing the timestep features of the images. By incorporating the predictor, we observed that even for unlearned timesteps, the similarity with the learned timesteps remained consistent. This modification enables the denoising U-Net to better seize the sequential flow of timesteps.