Text-to-Image Synthesis with Spatial-Semantic Modulation GAN
Junpeng Liu, Qingfeng Wu · 2024
Generating realistic images that are semantically consistent with text descriptions is a challenging cross-modal task. Existing text-to-image generation methods usually perform text-conditioned affine transformations equally within the image region. There is a limitation to this: while it is possible to generate images related to the description of the sentence as a whole, local areas usually do not correspond to the words in the sentence. To address this, we propose a novel Spatial-Semantic Modulation GAN for synthesizing images from input text. Specifically, we propose an efficient Spatial-Semantic Modulation Block, which consists of two main parts: (1) the Spatial Soft Mask Predictor, which enhances feature processing and uses important features of images for modeling; and (2) the Semantic Adaptive Perceptual module, which aligns the semantic information of the text with the reconstructed image to establish a close semantic correlation. Additionally, we complement these with the Channel Attention, which adaptively learns the importance of each channel. Extensive experimental results and ablation studies demonstrate that our proposed SSM-GAN can effectively improve the quality of the generated images.