Fine-Tuning CLIP With Dynamic Prompt Tuning and Cross-Modal Contrastive Alignment for Multimodal Sentiment Analysis
Ju Qin, Yuntao Sun · IEEE Access · 2026
Adapting CLIP to multimodal sentiment analysis has largely relied on optimizing textual prompts, yet such single-branch tuning fails to account for the evolving interplay between visual and textual representations, leading to suboptimal cross-modal alignment. To overcome this limitation, we propose a dual-branch dynamic prompt learning framework that jointly optimizes vision–language features through a shared pool of trainable prompts. The core novelty lies in a cross-modal synchronized prompting mechanism that explicitly couples the adaptation of CLIP’s language and vision branches via a learnable prompt-to-prompt mapping, preventing asymmetric representation drift during parameter-efficient fine-tuning. A similarity-driven gating mechanism is introduced to dynamically select contextually appropriate prompts for each input, enabling more adaptive and fine-grained representation learning than static prompt sets. To further strengthen cross-modal consistency, we incorporate supervised contrastive learning that constructs positive pairs from image–text samples sharing the same sentiment label and negative pairs from different classes. This contrastive objective effectively pulls semantically aligned multimodal features closer while pushing apart sentiment-inconsistent representations, thereby enhancing discriminative alignment across modalities. We evaluate the proposed framework on three benchmark datasets—MVSA-Single, MVSA-Multiple, and HFM. Experimental results show consistent and substantial improvements over state-of-the-art baselines in both accuracy and F1 score. Comprehensive ablation studies confirm the complementary contributions of dynamic prompt selection and joint vision–language optimization, highlighting their synergistic effect in stabilizing training, strengthening cross-modal alignment, and improving overall sentiment classification performance.