Text-Guided Attribute Enhancement Framework for Composed Image Retrieval

Zhi Ma, Yizi Huang, Di Wang, Bo Wan, Lin Zhao, Quan Wang · 2025

Composed image retrieval is a challenging multimodal task that refers to the process of retrieving target image by taking advantage of both complementary and synergistic image and text input. Existing efforts often focus on designing interaction models to fuse global query image and text features. However, these approaches struggle to capture fine-grained semantic association information between query image and text, especially when it comes to identifying specific objects or attributes in the query image that need to be modified in the text. In addition, these methods fail to adequately model cross-modal attention when dealing with composed query and target image, resulting in the model's inability to accurately map the semantic information from the composed query to the corresponding regions in the target image. To address these challenges, we propose a Text-Guided Attribute Enhancement Framework for Composed Image Retrieval (TAE-CIR). Our approach consists of three key modules: (a) Multi-granularity vision aggregation module, which extracts multi-granularity visual features and captures fine-grained object-level features related to the query text, refining object and attribute representations for more precise retrieval; (b) Multi-level fusion interaction module, which facilitates deep cross-modal interactions between the composed query and target image features, effectively capturing complex semantic relationships from the composed query to target image; (c) Composed feature alignment, which fuses multi-granularity visual features with the text using a text-guided Q-Former and contrastive learning to ensure accurate alignment between the composed query and the target image. Our extensive experiments on benchmark datasets FashionIQ and CIRR demonstrate the superiority of our proposed method.

Read the paper · More papers on PaperTik