Multimodal Graph Learning for Cross-Modal Retrieval
Jingyou Xie, Zishuo Zhao, Zhenzhou Lin, Ying Shen · Society for Industrial and Applied Mathematics eBooks · 2023
Cross-modal retrieval has attracted much attention lately for its various applications in Internet data mining. Existing approaches mainly adopt the projection function learning paradigm to construct dual-stream models, which suffer from two limitations: 1) They only utilize the correlations provided by cross-modal data pairs but the multiple correlations among data are unexplored. 2) They typically face the challenge of abstractness of semantics, which means that an instance may have distinct semantic information in different scenarios. In this paper, we propose a novel graph learning based framework termed Multimodal Graph Learning for cross-modal retrieval (MGL), which aims to fully exploit multiple correlations embedded in multimodal data and leverage a graph neural network to capture complementary information to alleviate the information sparsity and abstractness of semantics. First, we propose a graph construction algorithm to explore diverse multimedia information. Second, a modal feature projector is designed to learn modality-shared information, and a co-attention mechanism module is proposed to capture complementary information and perform dynamic feature integration. Third, a fusion and gate module is proposed to fully aggregate captured information and perform denoising. Furthermore, we employ a graph sampling algorithm to make our approach flexible to large-scale scenarios. Experimental results on three benchmark datasets prove the effectiveness of MGL.