Fine-grained Structural Hallucination Detection for Unified Visual Comprehension and Generation in Multimodal LLM
Hao Fei, Meng Luo, Jundong Xu, Shengqiong Wu, Wei Ji, Mong Li Lee, Wynne Hsu · 2024
Multimodal large language models (MLLMs) are evolving rapidly but suffer from significant challenges, such as hallucinations, which compromise the reliability and utility of MLLMs and are primarily caused by inadequate fine-grained visual representations. While some existing studies on multimodal hallucination detection have demonstrated effectiveness, they still fall short in two critical areas: the absence of fine-grained semantic representations, and the failure to address hallucinations in both visual comprehension and generation processes. To combat this, this paper introduces a novel framework for detecting and mitigating these multimodal hallucinations by employing structured visual scene graph representations. We propose a dual-process approach utilizing a cross-graph siamese network to semantically differentiate the fine-grained discrepancies between the visual scene graph (VSG) and the textual scene graph (TSG) representations in both image-to-text (I2T) and text-to-image (T2I) processes. We also introduce a semantic-centric heterogeneous graph contrastive learning strategy, enhancing the detection capabilities of our model. Extensive evaluation on the MHaluBench data demonstrates a significant improvement of our system over existing baselines by clear margins. This work is expected to pioneer fine-grained structured multimodal hallucination detection, presenting an advancement in the field of unified understanding and generation of MLLMs.