LCGD: Enhancing Text-to-Video Generation via Contextual LLM Guidance and U-Net Denoising

Muhammad Waseem, Muhammad Usman Ghani Khan, Syed Khaldoon Khurshid · IEEE Access · 2025

Diffusion models have emerged as a leading solution in computer vision and they excel at audio, image, and video generation by utilizing the Markov chain to map complex latent spaces. These models outperform other generative models such as GANs and VAEs, with their noising and denoising processes modeled after U-Net architecture, enabling high-quality text-to-image and text-to-video synthesis. However, existing research has largely focused on application rather than improving underlying architectures, leading to limitations, such as oversmoothing in approaches like FreeU. To address these gaps, we introduce LLM Contextual Guided Diffusion (LCGD), which integrates large language models (LLMs) into the noising and denoising phases to enhance semantic understanding, noise modulation, and feature selection. This approach improves the output realism and coherence, as demonstrated by our results, where SD+LCGD achieved 89.91% compared to 85.88% for SD+FreeU.

Read the paper · More papers on PaperTik