Optimizing Multimodal Sarcasm Detection with a Dual Perceiving Network and Adaptive Fusion

Yi Zhai, Jin Liu, Yuxin Liu, Xiantao Jiang · 2024

Nowadays, posting sarcastic text or visual content on platforms like WhatsApp, Twitter, and Facebook has become a popular style, allowing individuals to avoid directly expressing pessimism, thereby indirectly conveying their thoughts or intentions. As more people express their opinions online, the need to develop effective models for accurate multimodal sarcasm detection is growing. However, most existing methods have deficiencies in extracting effective features from non-text modalities, which can result in the neglect of some sarcasm-related contextual information. Additionally, during the fusion phase of different modalities, simply using straightforward feature concatenation may introduce redundant information and noise. To address these challenges, a new model called Dual-Perceiving Network (DPN) was developed. It includes a dual-view perception module, which enables the text modality to fully perceive both the local and global features of the image modality, thereby obtaining context and background information conducive to constructing inconsistencies. Additionally, it employs an adaptive fusion method to effectively reduce the impact of redundant information. Extensive experiments comparing the proposed DPN with other multimodal sarcasm detection methods demonstrate its superior performance on public datasets.

Read the paper · More papers on PaperTik