A Multimodal Named Entity Recognition Approach Based on Multi-Perspective Contrastive Learning

Huafu Liu, Yongli Wang, Dongmei Liu · 2024

Multimodal Named Entity Recognition (MNER) is a vision-language task aimed at detecting the scope of entities and determining their types in given text-image pairs. Although existing MNER models have made some progress, they often encounter issues like mismatched feature dimensions and inconsistent feature representations when interacting with text sequences. This makes it difficult to bridge the semantic gap between images and text. Therefore, this paper proposes a multimodal named entity recognition method based on multiperspective contrastive learning (MPCL-NER). Initially, the ALBERT pretrained language model and the ViT visual model are used to extract features from text and images, respectively. Based on this, image feature representations are enhanced through multi-perspective contrastive learning, which also promotes the integration of image and text information. Finally, the two types of features are merged through a multi-head adaptive attention mechanism and a gated fusion mechanism. This is followed by entity tagging with a CRF (Conditional Random Field) decoder. Experimental results show that this method outperforms other baseline methods on the public dataset Twitter-2017, achieving an F1 score of 86.11%, demonstrating the effectiveness of this approach.

Read the paper · More papers on PaperTik