On Multimodal Semantic Consistency Detection of News Articles with Image Caption Pairs

Yuwei Chen, Ming‐Ching Chang · 2022 IEEE International Conference on Consumer Electronics - Taiwan · 2022

Recently multi-modal consistency detection has been proposed as method to combat disinformation. However, state-of-the-art methods lack the fine grained localized evidence needed for consistency detection. Current methods also struggle on longer texts that are present in news articles. We propose an ensemble method that combines a series of visual detectors and BERT based NLP models to compute consistency between modalities. The proposed method is effective in both detecting the standard image-caption pairs and news articles containing multiple paragraphs. Our method can localize and provide fined grained evidence towards its given responses. We evaluate our method on a MSCOCO image-caption subset and image & text inconsistency evaluation of news articles from the U.S. DARPA SemaFor program. We achieved an 83% AUC on the gathered MSCOCO dataset, which shows the effectiveness of our method.

Read the paper · More papers on PaperTik