A Deep Learning Based Contrastive Attention Model For Multi-Modal Sarcasm Detection

Vidyullatha Sukhavasi, Venkatesulu Dondeti · 2025

Due to its reliance on contextual and non-verbal information and heavy focus on textual scenarios, sarcasm detection has been a challenging topic to solve in the last decade. The majority of previous research has been on one of two areas: audio feature detection or text feature detection for sarcasm in audio data. Recent years have seen a surge in interest in analysis in multi-modal scenarios, thanks to the proliferation of video communication. As a result, there is a lot of excitement in the subject of natural-language-processing(NLP) and multi-modal analysis surrounding multi-modal sarcasm recognition, which seeks to classify sarcasm in video discussions. In this work, we build a design called Contrastive-Attention-based Sarcasm Detection (ConAttSD). It utilizes an inter-modality contrastive attention procedure to mine multiple contrastive features for an utterance, taking into account that sarcasm is regularly transported over incongruity amid modalities. For example, script could be uttering a accolade while auditory tone could indicate a complaint. When data from two different modalities are inconsistent with one another, this is called a contrastive characteristic. Our results show that the ConAttSD model is effective on the benchmark multi-modal sarcasm dataset MUStARD..

Read the paper · More papers on PaperTik