Multimodal Knowledge Graph Inference Method Based on Cross-Attention Mechanism
Zijian Han, Changbo Hou · 2025
Existing multimodal knowledge graph reasoning methods often rely on simple weighted averaging or attention mechanisms for multimodal feature fusion, which frequently results in the loss of modality-specific information and fails to fully exploit the complementary relationships between modalities. To address these issues, this paper proposes a multimodal knowledge graph reasoning method based on a cross-attention mechanism. The model employs the BERT model to encode textual information, introduces a VGG-16 model with a multi-scale attention module to encode visual information, and utilizes a graph attention mechanism to extract the structural information of the knowledge graph. Additionally, a contrastive learning framework is introduced to reduce semantic inconsistencies between modalities, and a cross-attention mechanism is used to achieve dynamic feature fusion across modalities. To validate the effectiveness of the proposed model, comparative experiments were conducted on the FB15k-237 and DB15K datasets. The results show that the proposed model significantly outperforms existing benchmark models on key evaluation metrics such as MRR, Hit@1, Hit@3, and Hit@10, demonstrating its effectiveness and robustness in multimodal knowledge graph reasoning.