Multimodal Information Integration and Retrieval Framework Based on Graph Neural Networks
Yuping Yuan, Haozhong Xue · 2025
With the rapid development and prosperity of multimodal data, ranging from text, image, audio, and more, how to better integrate and retrieve information across different modalities becomes an extremely hot topic. In this paper, we propose a novel Graph Neural Network -based Multimodal Information Integration and Retrieval framework. The objective is to achieve a better fusion effect and cross-modal retrieval performance of heterogeneous data. Our model creatively uses a graph structure to complete the description of multimodal relationship, which is an innovation on the basis of the multimodal fusion method. More specifically, we introduce a hierarchically structured graph while taking node as modality and edge as one of relation within/dependent on modality. When processed by a Graph Convolutional Network, the Graph Convolutional Network can aggregate features from neighboring nodes for multiple levels to optimize the multimodal joint representation. A cross-modal attention mechanism is combined additionally, dynamically learning the importance of different modality modalities under one specific query to further optimize retrieval accuracy. Our framework can be trained end to end, which can efficiently train multimodal representation learning and increase the generalization ability of the model to retrieve it. Our experimental results show that our model significantly improves retrieval accuracy and recall on the benchmark dataset compared with the existing multimodal retrieval models.