AI Driven LLM integrated diffusion models for Image Enhancement- A Survey
Md Sadiq Z Pattankudi, Abdul Rafay Attar, Kashish Jewargi, Samarth Uppin, Ashlesha Khanapure, Uday Kulkarni · 2024
Text-to-image generation has emerged as a prominent research problem in computer vision and natural language processing, resulting from recent progress in generative models and LLMs. This paper reviews the latest research on Text-to-image generation models integrated with Large Language Models (LLMs). The paper focuses on the analysis of six LLM integrated models namely DiffusionGPT, LLM Grounded Diffusion Model, IN-STRUCTCV, ECLIPSE, Self-Correcting LLM-Controlled Diffusion Models and SUR adapter. Each model is evaluated based on its architecture, methodology and results. Our analysis identifies the advantages and disadvantages of different models and provides recommendations for further research directions. At the end of this paper is a summary of the current advancements in the field and suggestions for future study possibilities.