Enhancing Multimodal Object-Entity Relation Extraction via Multi-Aspect Contrastive Learning in Large Multimodal Models

Yaming Zhang, Jianfei Yu, Wenya Wang, Li Yang, Yang Jia, Rui Xia · IEEE Transactions on Audio Speech and Language Processing · 2025

Multimodal Object-Entity Relation Extraction (MORE) is an emerging task in information extraction, which aims to extract object-entity relational facts from text and image data. Despite obtaining promising results, previous works are based on small language models or small vision-language models, and the potential of using the power of generative large multimodal models (LMMs) for the MORE task still remains unclear. Moreover, most existing studies focus on learning the multimodal representation merely based on the supervision signals of each sample, failing to consider the semantic similarities and differences between samples. To tackle these two issues, in this paper, we propose a multi-AsPect cOntrastive Learning-enhanced Large multimOdal model named APOLLO, which formulates MORE as a generation problem and employs a parameter-efficient fune-tuning method LoRA to inject task-specific knowledge into a representative LMM named InstructBLIP, followed by enhancing the multimodal representation to capture the inter-sample relationship with a multi-aspect contrastive learning algorithm. Experimental results on a benchmark dataset demonstrate the superiority of APOLLO over existing multimodal approaches.

Read the paper · More papers on PaperTik