Multi-Modal Sarcasm Detection Based on Cross-Modal Composition of Inscribed Entity Relations

Lingshan Li, Di Jin, Xiaobao Wang, Fengyu Guo, Longbiao Wang, Jianwu Dang · 2023

Sarcasm, a linguistic technique employed to express emotions opposite to their literal meaning, has garnered significant attention from researchers due to the rise of social media. Detecting sarcasm in a multi-modal context has become a focal point in recent studies. However, existing research primarily relies on identifying inconsistencies between text semantics and image semantics, often lacking a deep understanding of images. Consequently, capturing inconsistencies between images and texts poses a challenge in many cases. In this paper, we propose the Entity-Relational Graph Convolutional Network (ERGCN) as a solution to detect sarcasm by examining the relationship between entities within images. Our approach involves extracting entities and text descriptions from each image, which provides valuable entity information. Subsequently, we employ external knowledge to construct a cross-modal graph for each text and image pair, emphasizing the presence of internal contradictory information. Finally, we utilize the graph convolutional network to identify inconsistent information across modalities and successfully detect sarcasm. Experimental results demonstrate that our model achieves state-of-the-art performance on a widely used multimodal Twitter dataset.

Read the paper · More papers on PaperTik