Multimodal Sarcasm Recognition by Fusing Textual, Visual and Acoustic content via Multi-Headed Attention for Video Dataset

Sajal Aggarwal, Ananya Pandey, Dinesh Kumar Vishwakarma · 2023

Multimodal sarcasm recognition uses a combination of acoustic, video, and text-based cues to detect sarcasm. Though multimodal approaches have outperformed unimodal detection by providing a more comprehensive description of the speaker’s sentiment, they are particularly challenging as cues may not be consistent across the multiple modalities. In our research study, we propose a system that first extracts multimodal features from the input provided and then applies a bimodal multi-head attention mechanism to them. Subsequently, the features are concatenated and passed through a softmax layer for detection. The proposed model is evaluated on the MUStARD dataset for multimodal sarcasm recognition. For the speaker-dependent configuration, the proposed model beats cutting-edge methods in terms of accuracy, precision, recall, and F1-score by 0.75%, 2.9, 2.82, and 3.1, respectively.

Read the paper · More papers on PaperTik