A Deep Learning based approach for MultiModal Sarcasm Detection
Aruna B. Bhat, Aditya Chauhan · 2022
Sarcasm detection is used to single out natural language statements where intended meaning differs from what the surface meaning implies. A number of tasks in natural language processing areas like sentiment analysis and also mining of opinions use sarcasm detection underneath. Many of the key research in the area of sarcasm detection primarily focus on only text-based input. In the present day scenario , there has been a sudden explosion in the amount of multimodal data mainly due to social media. As a result of that, users these days are not just limited to text while expressing themselves , but also make heavy use of visuals like in images and videos. The objective of our paper is to incorporate multimodal data so as to enhance the performance of present sarcasm detection algorithms. Multiple methods that leverage data in the form of image and text, and also as a combination of both, have been presented thus far.We present a unique architecture which works on the Robustly Optimized BERT pre-training approach(RoBERTa) which is a facebook modified version of well known model BERT with a co¬attention layer on the top for including the context incongruency between input text and attributes of the image. Using Gated Recurrent Unit(GRU), the text-based features are extracted and the features of the image are retrieved by conditioning the image through Feature-wise Linearly Modulated ResNet Blocks. These combined with CLS token from the model RoBERTa are used to obtain the final predictions. In the end, we show performance enhancements from our model when it is stacked against the other latest models with competitive results which make use of the same publicly available Twitter dataset.