Adaptive Fusion of Single-stream and Dual-stream Models based on Cross-modal Association Strength for Harmful Meme Detection
Lin Meng, Qinghao Huang, Xingyi Yu, Xianjing Guo, Tao Guo · 2023
With the rapid evolution of social media, internet memes have become a prominent means for harmful speech dissemination, prompting the need for multi-modal hate speech detection. Recent predominant studies focus on designing elaborate architectures of pre-trained models for multi-modal fusion, notably the single-stream model ViLT and the dual-stream model ALBEF. The text and image of a meme may be related explicitly or implicitly, showing diverse cross-modal association strength. Through a pilot experiment we observe that single-stream models are good at detecting explicitly related memes, while dual-stream models perform better at implicitly related ones. However, the problem of fusing single-stream and dual-stream models remains underexplored so far. In this paper we present an novel approach called Adaptive Fusion of Single-stream and Dual-stream models (AFSD). Our model flexibly fuses single-stream and dual-stream models based on cross-modal association strength in memes, which is defined as the semantic similarity between the generated image caption and original text in a meme. We conduct comprehensive experiments and experimental results show that our model achieves new state-of-the-art results on two widely-used datasets.