Cross-modal consistency types in multimodal social data
Lorenzo Vaiani, Luca Cagliero, Paolo Garza, Jason Ravagli · Knowledge-Based Systems · 2025
Social media content, such as internet memes or tweets, are nowadays largely or mainly multimodal. Machine learning models often need to jointly process images and text to solve complex tasks such as hate speech detection or sentiment analysis. For example, the misogyny of a meme cannot be accurately predicted while considering the visual and textual modalities separately. Similarly, sentiment annotations for tweets’ images and text can be discordant. Detecting the samples with inconsistent modality contributions is particularly relevant to analyze machine learning model performance and explain classification errors. In this paper, we formalize the types of cross-modal consistency by differentiating between consistent cases and not. Cross-modal consistency denotes whether all modalities agree on the label (i.e., full consistency) or not (i.e., inconsistency). When the visual and textual modalities are discordant, we distinguish the cases in which a joint analysis of multimodal features is sufficient to solve the issue from those requiring a human agreement (i.e., NOR consistency). We also propose a CLIP-based architecture to predict the cross-modal consistency types and identify the modalities causing the inconsistency. The results achieved on benchmark datasets show that cross-modal consistency annotation is cost-effective, i.e., it provides relevant insights into model predictions while requiring a limited extra human effort.