Dynamic Weighted Gating for Enhanced Cross-Modal Interaction in Multimodal Sentiment Analysis

Nan Wang, Qi Wang · ACM Transactions on Multimedia Computing Communications and Applications · 2024

Advancements in Multimodal Sentiment Analysis (MSA) have predominantly focused on leveraging the interdependence of text, acoustic, and visual modalities to enhance sentiment prediction. However, efficiently integrating these modalities remains a challenge. In response, we investigate the design of a practical approach to control the information flow more accurately and propose DWGCI. This multimodal model employs systematic cross-modal interactions and gating mechanisms to analyze sentiment. It ensures comprehensive feature extraction across modalities by utilizing BERT for text features, COVAREP for acoustic features, and CNNs for visual features. To improve the common feature expression of each modality and increase the dynamic interaction between modalities, we introduced a Text-Guided Multimodal Fusion Module (TGMFM). Another integral part of our model is a unique Dynamic Gating Module (DGM) strategically positioned to follow cross-modal interactions. This mechanism represents a significant innovation that can adapt to modality-specific sentiment cues with unprecedented precision dynamically. The model can make nuanced distinctions between modalities, which significantly enhances the accuracy of sentiment prediction. Our DWGCI has been subjected to rigorous testing on two widely used datasets. It demonstrated its superior performance, established a benchmark in MSA for gate fusion innovation, and highlighted its transformative potential.

Read the paper · More papers on PaperTik