Text To Image Generation By Using Stable Diffusion Model With Variational Autoencoder Decoder
Usharani Budige, Srikar Goud Konda · International Journal for Research in Applied Science and Engineering Technology · 2023
Abstract: Imagen is a text-to-image diffusion model with a profound comprehension of language and an unmatched level of photorealism. Imagen relies on the potency of diffusion models for creating high-fidelity images and draws on the strength of massive transformer language models for comprehending text. Our most important finding is that general large language models, like T5, pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: expanding the language model in Imagen improves sample fidelity and image to text alignment much more than expanding the image diffusion model.