A unified prompt-based framework for few-shot multimodal language analysis

Xiaohan Zhang, Runmin Cao, Yifan Wang, Songze Li, Hua Xu, Kai Gao, Lunsong Huang · Intelligent Systems with Applications · 2025

Multimodal language analysis is a trending topic in NLP. It relies on large-scale annotated data, which is scarce due to its time-consuming and labor-intensive nature. Multimodal prompt learning has shown promise in low-resource scenarios. However, previous works either cannot handle semantically complex tasks, or involve too few modalities. In addition, most of them only focus on prompting language modality, disregarding the untapped potential of other modalities. We propose a unified prompt-based framework for few-shot multimodal language analysis . Specifically, based on pretrained language model, our model can handle semantically complex tasks involving text, audio and video modalities. To enable more effective utilization of video and audio modalities by the language model, we introduce semantic alignment pre-training to bridge the semantic gap between them and the language model, alongside effective fusion method for video and audio modalities. Additionally, we introduce a novel effective prompt method—Multimodal Prompt Encoder—to prompt the entirety of multimodal information. Extensive experiments conducted on six datasets under four multimodal language subtasks demonstrate the effectiveness of our approach. • We first propose a unified prompt-based framework that involves text, video, and audio modalities, which can be applied to various subtasks of multimodal language analysis in few-shot setting. • We introduce Semantic Alignment Pretraining for video and audio encoders, aiming to narrow the semantic gaps between them and PLM. • Based on the idea of MAG, we propose an effective fusion method for video and audio modality. • We propose an effective prompt method for multimodal information - Multimodal Prompt Encoder. • We conduct experiments on six datasets across four different multimodal language subtasks to verify the effectiveness of our model.

Read the paper · More papers on PaperTik