GMNER-LF: Generative Multi-modal Named Entity Recognition Based on LLM with Information Fusion

Hui-Yun Hu, Junda Kong, Fei Wang, Hongzhi Sun, Yang Ge, Bo Xiao · 2024

Multi-modal Named Entity Recognition (MNER) leverages visual information to enhance the effectiveness of text-only Named Entity Recognition (NER). Currently, many methods are based on sequence label, but the model architecture is relatively complex. Large Language Model (LLM) has recently demonstrated powerful generative and comprehension abilities, so we propose GMNER-LF, a new paradigm for MNER with generative method. Firstly, we retrieve relevant image-text pairs to provide prior knowledge for the recognition. Secondly, we construct our task into a MRC task so that the LLM can better understand the problem. In addition, we design a multi-modal fusion module and add a gating mechanism to help filter noise information in the image to obtain high-quality fusion representations. The multi-modal fusion module is injected into the LLM block to achieve deep fusion of LLM and multi-modal representations and fully explore the internal knowledge of LLM. The proposed method not only fully explores the internal knowledge of LLM, but also filters important modal information through the gating mechanism. Experimental results show that compared with other generative methods, the proposed method has improved performance on both Twitter-2015 and Twitter-2017 datasets.

Read the paper · More papers on PaperTik