Multimodal Entity Recognition and Relation Extraction via Dynamic Visual-Textual Enhanced Fusion for Social Media Applications in Consumer Electronics
Qingchuan Zhang, Zihan Li, Jianlei Kong, Min Zuo · IEEE Transactions on Consumer Electronics · 2024
Accurate entity and relation analysis are crucial for enhancing the user experience of consumer electronics products with integrated social media capabilities. An open gap in this research field is that existing approaches present lower predictive performance due to information conflicts and mismatch among multimodal data. Conversely, this work presents a novel neural network architecture for multimodal entity recognition and relation extraction, which combines the power of data discrimination and dynamic gating fusion with deep spatiotemporal learning. The idea is to infer inherent relationships between co-occurring entities by overcoming the limitations of multimodal mismatch and parallel task. Extensive experiments demonstrate that our method outperforms advanced approaches on three benchmark datasets, signifying a major performance enhancement due to the effective integration of visual and textual information. Ablation experiments and case studies further demonstrate the potential of the architecture to improve the intelligent social media capabilities of consumer electronics. The detailed source codes are available athttps://github.com/katouMegumiH/CIModel.