Dimension-Wise Gated Cross-Attention for Multimodal Sentiment Analysis

Mohammed Jahangir Hossain, Md. Mithun Hossain, Sudipto Chaki, M. F. Mridha, Md. Saifur Rahman, Mohammad Ali Moni · 2025

Multimodal sentiment analysis necessitates the seamless integration of textual and visual signals for the precise interpretation of user-generated material. In this paper, we introduce Dimension-Wise Gated Cross-Attention (DGCA). This new fusion mechanism fine-tunes the interaction between language and images more precisely than prior methods. Our method uses a bidirectional cross-attention module to iteratively enhance text and image features. We use a dimension-wise gating technique in which each latent dimension independently learns to weigh contributions from text or image signals using softmax-normalized modality gates. The approach uses selective per-dimension fusion to highlight important cues from one modality while minimizing less useful characteristics from another. On the SemEval-2020 Memotion dataset, DGCA outperformed the state-of-the-art (SOTA) baselines by 2.27%, highlighting its ability to detect subtle affective cues. In summary, DGCA improves performance and interpretability, enabling fine-grained and context-aware multimodal sentiment analysis.

Read the paper · More papers on PaperTik