DualStreamAttentionNet: A Multimodal Approach for Audio and Text sentiment analysis with Self and Cross Modal Attention Mechanisms

Rui Jiang, Zhan Wen, Hantao Liu, Chang Liu · 2023

Sentiment analysis is a crucial task in natural language processing. While unimodal sentiment analysis relies on one modality, multimodal sentiment analysis combines features from various modalities to improve accuracy. However, many current models emphasize complex fusion techniques and neglect the importance of enhancing modality feature extraction. They also often overlook the interaction between modalities. In response, we introduce the DualStreamAttentionNet(DSAN), an innovative model tailored for joint text and audio sentiment analysis. Our architecture employs a Transformer encoder for robust textual feature extraction and processes audio using 1D convolutions followed by a bidirectional LSTM for enhanced audio feature representation. Crucially, by leveraging self and cross-attention mechanisms, our model not only amplifies the intrinsic features of each modality but also effectively captures the interrelated dynamics between text and audio. This dual-focus approach leads to a synergistic fusion of features, resulting in heightened emotion classification accuracy.

Read the paper · More papers on PaperTik