Towards Multimodal Semantic Consistency Analysis of Long Form Articles
Yuwei Chen, Ming‐Ching Chang · 2022
With the rise of misinformation in news sites and social media, multi-modal machine learning methods that can identify fake news by analyzing inconsistencies within the articles have become increasingly important. Current state-of-the-art methods based on traditional image-caption models can only process captions within 1 to 2 sentences. Existing models struggle on analyzing long articles as they were not trained for such purposes. The main limitation is the lack of fine-grained localized evidence needed for consistency detection. We propose an ensemble method combining a bank of visual detectors and BERT-based NLP models that can effectively compute the consistency among the image(s) and paragraph(s) of texts. Our method is effective in both detecting the standard image-caption pairs and longer form news articles. Our method is able to process longer form of multi-modal media via the localization of fine-grained evidence with modularity and explainability. Evaluation is performed on a MS COCO data subset and a news article benchmark of the DARPA SemaFor program. We achieved 83% AUC on the COCO subset as well as a competitive result within the SemaFor evaluation.