Multimodal Sarcasm Detection Method Using RNN and CNN

Bouchra Azahouani, Hanane Elfaik, El Habib Nfaoui, Said El Garouani · 2024

In the current digital environment, sarcasm is preva-lent on social media platforms, characterized by a combination of verbal and non-verbal cues, such as prosodic variations, phonetic inflections, and textual markers including lexical selection, irony, and exaggeration. Previous research has primarily focused on sarcasm detection in either audio or text data independently. This paper presents a novel deep learning approach to detect sarcasm in conversational data by integrating both textual and auditory elements. Our method employs a bidirectional Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) network for text processing, capturing sequential dependencies and contextual information. For the audio component, we use a Convolutional Neural Network (CNN) to extract key features from speech. The fusion of these modalities is achieved by combining the extracted features into a composite vector, which improves the detection of sarcasm. Experimental evaluations on the MUStARD Extended dataset show that our hybrid model significantly outperforms unimodal models, achieving an F1-score of 74.67%.

Read the paper · More papers on PaperTik