Attn-VAE-GAN:Text-Driven High-Fidelity Image Generation Model with Deep Fusion of Self-Attention and Variational Autoencoder

Chen Yang · 2024

In the current realm of research on text-to-image (T2I) transformation, the daunting task of translating natural language descriptions into visually realistic images is apparent. This intricate undertaking requires models to deeply comprehend cross-modal information and achieve precise semantic mapping. Despite the commendable progress made by GANs in image synthesis—crafting high-resolution, lifelike visual effects from random noise—persistent challenges like training instability, mode collapse, and fidelity issues persist, prompting the need for innovative solutions.This study introduces an avant-garde text-driven high-fidelity image generation model, labeled as Attn-VAE-GAN. The initial integration of a generator module featuring a fused Self-Attention mechanism empowers the model to thoroughly grasp and accurately capture intricate semantic information from the input text. Moreover, the assimilation of a Variational Autoencoder (VAE) module taps into its latent representation learning prowess to optimize both image quality and diversity. To amplify the realism and detail expression of the generated images, a meticulously designed comprehensive loss function combines intrinsic VAE loss and hinge loss.Empirical results underscore the substantial enhancement achieved by the Attn-VAE model in both the quality and diversity of the generated images across diverse publicly available benchmark datasets.

Read the paper · More papers on PaperTik