Multi-Attention Cascade Model Based on Multi-Head Structure for Image-Text Retrieval

Haotian Zhang, Wei Biao Wu, Meng Zhang · 2022 International Joint Conference on Neural Networks (IJCNN) · 2022

Image-text retrieval aims to find semantically relevant results from another modality when given an image or a sentence as a query. It remains challenging because the non-semantic information in the current image representation affects the alignment of the two-modal feature. In addition, there is a lack of topic information combined with textual features in the text representation. To solve these problems, we propose a novel Multi-Attention Cascade Model (MACM) for cross-modal image-text retrieval. Specifically, we first design a Multi-Attention Cascade Transformer Encoder (MACTE) to infer less semantically relevant information in images, which cascades the self-attention mechanism, the channel attention mechanism and the spatial attention mechanism based on the multi-head structure. Therefore, it can filter unimportant regions in the image in different dimensions and highlight important regions. Then, to solve the problem of losing topic information due to too long sentences, we design a Central Word Enhancer (CWE), which uses Gate Recurrent Unit and Transformer Encoder to extract multi-granularity text features jointly capture long-distance dependencies, thereby effectively preventing the deviation of finding central words which means topic information. Finally, we introduce a graph structure as an inter-modality interaction module to represent the features of a single modality respectively. Experiments on Flickr30k and MSCOCO datasets demonstrate the effectiveness of our method, which achieves the state-of-the-art result for text-image retrieval. The Recall@1 on Flickr30k by our method improves image-to-text retrieval by 5% and text to image retrieval by 5.1%. The Recall@1 on MSCOCO improves image-to-text retrieval by 2.5% and text to image retrieval by 2.8%.

Read the paper · More papers on PaperTik