SEPA: An Semantic Projection Alignment Framework for Multimodal Named Entity Recognition

Guohui Ding, Ying Kong, Xinlei Li · 2025

Multimodal Named Entity Recognition (MNER) leverages visual information to assist in the localization and classification of named entities within textual data. The essence of this process resides in cross-modal alignment, with the current mainstream research generally focusing on correlating entities within textual content to their corresponding visual regions in images, thereby achieving object-level word-region alignment. However, a significant drawback of this alignment approach is its neglect of semantic relationships and contextual connections between entities. In response to the aforementioned issues, we propose an MNER framework for semantic-level alignment, termed Semantic Projection Alignment Framework (SEPA).Specifically, we have devised a semantic affinity mapping mechanism based on cross-modal attention calibration, aimed at effectively capturing the semantic correspondence between text and images. This mechanism establishes semantic connections both within the same modality and across different modalities, and it evaluates the semantic distance between the two modalities by calculating the contextual affinity score, thereby assessing the degree of their alignment. Furthermore, we employ the Fourier transform to map text and image features into the frequency domain within the context of the MNER task. This approach offers the advantage of capturing global structural information such as color distribution and shape, effectively addressing the issue of global information loss caused by feature heterogeneity during the multimodal feature fusion process. Simultaneously, the characteristic of the Fourier transform that focuses on low-frequency components effectively mitigates unnecessary informational interference, thereby addressing the issue of noise. We conducted extensive experiments on two of the most popular datasets in the field, and the results indicate that our model consistently surpasses the baseline, achieving state-of-the-art performance.

Read the paper · More papers on PaperTik