Semantic Web-Enabled Multimodal Entity Recognition Using Collaborative Cross-Attention Mechanism
Dongxiu Wang, Yulan Wen, Zhengxiang Qiu · International Journal on Semantic Web and Information Systems · 2025
Named entity recognition (NER) is crucial in tasks such as information extraction, question and answer systems, and opinion analysis, but the existing methods still have deficiencies in cross-modal feature alignment and fusion, and it is difficult to make full use of visual information and structured knowledge to improve the recognition accuracy. To this end, this paper proposes CoAtt-NER, a multimodal NER model supporting semantic web, which combines textual representation generated by ALBERT with knowledge graph embedding to enrich the semantic information of entities; and adopts CLIP-ViT for better visual feature extraction. In addition, Co-Attention is proposed to establish two-way interaction between text and visual modalities to achieve dynamic modelling and deep fusion of information. Experiments on Twitter-2015 and Twitter-2017 datasets show that the F1 scores of CoAtt-NER reach 76.25% and 87.31%, respectively, which achieve significant improvement compared with existing methods, verifying the effectiveness of this study in multimodal entity recognition tasks.