Chinese Multimodal Named Entity Recognition in Conversational Scenarios

Bolin Chen, Rui Liu, Ding Cao, Guiquan Liu · 2024

In this paper, we primarily investigate Multimodal Named Entity Recognition (MNER) within conversational scenarios. Existing MNER methods predominantly focus on integrating textual and visual modalities; however, visual information is often scarce in typical conversational scenarios such as telephone customer service, meeting conversations, and business discussions. In contrast, acoustic information can provide a wealth of insights beyond the textual content in conversational dialogues. In this work, we specifically focus on Chinese multimodal NER using text and speech modalities. To address the issue of insufficient emphasis on acoustic information by existing MNER methods, we propose a novel MNER model, termed the Bidirectional Perception Multimodal Interaction Model (BiPMM). Specifically, to achieve the alignment of speech and text at token level, we designed a multimodal information alignment module. Moreover, to foster bidirectional interaction between textual and acoustic modalities, we introduced a bi-directional perception fusion module for multimodal information interaction. Finally, we constructed a conversational scenario-based multimodal NER dataset and conducted experiments on it. The results demonstrate that our model achieved optimal performance, thereby proving its effectiveness.

Read the paper · More papers on PaperTik