Toxic Memes Recognition Through with Multimodal Bidirectional Cross-Attention

Alex Hashim, Natalie Coleman, Jannat Roy · Preprints.org · 2024

Despite the considerable advancements achieved with machine learning techniques in identifying hate speech, numerous technical obstacles persist that hinder these models from reaching human-level accuracy. Challenges such as understanding nuanced context, detecting sarcasm, and effectively interpreting the interplay between visual and textual elements in memes complicate the detection process. This study delves into the comprehensive evaluation of several cutting-edge visual-linguistic Transformer architectures, including VL-BERT, VLP, UNITER, and LXMERT, to assess their capabilities and limitations in handling the multifaceted nature of hateful content within memes. Building upon these evaluations, we introduce significant enhancements aimed at boosting their efficacy in this domain by developing a novel bidirectional cross-attention mechanism. This mechanism facilitates a more seamless integration of visual and textual information, enabling the model to better capture the subtle cues that distinguish hateful memes from benign ones. In addition to the architectural improvements, we leverage deep ensemble strategies to aggregate predictions from multiple model instances, thereby enhancing the robustness and reliability of the detection system. By combining the strengths of diverse models, the ensemble approach mitigates individual weaknesses and reduces the likelihood of false positives and negatives. Our proposed framework not only addresses the existing shortcomings of single-model approaches but also markedly surpasses existing baseline performances by a large margin, achieving higher AUROC and accuracy scores. The refined model demonstrates superior capability in discerning hateful content within multimodal memes, offering a more robust and reliable tool for mitigating the proliferation of harmful online material. Furthermore, the scalability of our approach ensures its applicability to evolving online threats, providing a sustainable solution for automated hate speech detection. These advancements signify a meaningful step towards enhancing the effectiveness of machine learning models in creating a safer and more respectful online environment.

Read the paper · More papers on PaperTik