RaT2IGen: Relation-aware Text-to-image Generation via Learnable Prompt
Zhaoda Ye, Xiangteng He, Yuxin Peng · ACM Transactions on Multimedia Computing Communications and Applications · 2025
Text-to-image generation is to generate photo-realistic images according to the given text descriptions by users. Current methods have achieved promising performance. However, these methods still fail to generate the correct relation of the objects in the text descriptions, which cannot correctly reflect the users’ intention and hinders the application of text-to-image generation. This is due to two aspects: Firstly, they focus on the attributes and concepts of the object and have difficulty in transferring/mapping object relation knowledge into tasks. Secondly, they lack an effective mechanism to apply the semantic information from the text to guide spatial layout of the object in the generation stage. Thus, this article proposes RaT2IGen (Relation-aware Text-to-image Generation), which aims to improve the semantic consistency of the generation model of object relation with the proper spatial layout. The main contribution can be summarized as follows: (1) The learnable relation prompt is proposed to capture the semantic information of the relation of the objects, which aims to strengthen the generation model to understand the object relations. (2) A bidirectional condition generation mechanism is proposed to generate the condition vectors of the image patches in both horizontal and vertical flattening, which helps the model effectively control the spatial layout of the objects according to the text description. (3) This article proposes a new metric to evaluate the consistency of the spatial layout of the generated images with the real images. The experiments conducted on MS-COCO and LN-COCO datasets demonstrate the effectiveness of our approach in achieving the best performance.