Adaptive Visual Semantic Embedding with Adaptive Pooling and Instance-Level Interaction

Jiawen Chen, Yonghua Zhu, Hong Yao, Wenjun Zhang · 2023

To enhance the effectiveness of image-text matching, this study introduces a novel approach known as Adaptive Visual Semantic Embedding with Adaptive Pooling and Instance-Level Interaction (AdIVSE). AdIVSE employs an adaptive pooling technique to autonomously determine the most suitable pooling strategy for aggregating visual and textual embeddings. Additionally, it incorporates an instance-level interaction module and a dynamic hard-negative loss function to facilitate the discrimination of subtle distinctions within similar image-text pairs, thereby improving fine-grained image-text matching. Qualitative and ablation experiments conducted on the Flickr30k dataset demonstrate that the proposed network exhibits the capability to discern similar instances. This capability translates into superior performance in fine-grained matching scenarios.

Read the paper · More papers on PaperTik