DLGP: grounded multimodal named entity recognition based on dynamic learnable graph prompt

Shiyuan Liu, Zhaogong Zhang, Haiyang Jiang, Xin Guan · 2025

Grounded Multimodal Named Entity Recognition (GMNER) is essential for extracting comprehensive entity information from both text and images, with significant applications in information retrieval, question answering, and knowledge graph construction. However, current methods face challenges in managing the complex relationships within cross-modal data. To address this, we propose an innovative approach that combines a pre-trained BART model with Graph Convolutional Networks (GCNs) and prompt learning to better capture structural information in multimodal data, thereby enhancing GMNER task performance. Specifically, our method dynamically learns relational graphs between entities and objects across text and image modalities, constructs multi-scale prompts using multi-level GCNs, and introduces these prompts via dynamic gating to improve multimodal information fusion. Experimental results on benchmark datasets demonstrate the effectiveness and superiority of our approach.

Read the paper · More papers on PaperTik