Mitigating Hallucinations in Large Vision-Language Models via Reasoning Uncertainty-Guided Refinement
Shenshen Li, Xing Xu, Wenxin Meng, Jingkuan Song, Chong Saw Peng, Heng Tao Shen · IEEE Transactions on Multimedia · 2025
Despite demonstrating impressive capabilities in comprehending multi-modal contexts, large vision-language models (LVLMs) are invariably prone to generate unreliable answers, i.e., hallucinations. Existing methods mainly mitigate this hallucination by introducing specific designed datasets or employing contrastive decoding techniques. However, these methods heavily rely on the quality of constructed datasets and negative samples, overlooking the inherent ambiguity in reasoning caused by over-reliance on linguistic priors and data complexity, termed reasoning uncertainty. This oversight hinders the models from effectively identifying the causal relationships behind each token, increasing their susceptibility to hallucinations. To address this issue, we propose a novel framework namedReasoningUncertainty-guidedRefinement (RUR)for mitigating hallucinations in LVLMs from an uncertainty perspective. Specifically, unlike conventional uncertainty quantification methods, we first extract the causal reasoning relationships between tokens by exploiting the link between structural causal models and the Transformer architecture. Based on this relationship, we then employ the Subjective Logic principle to model the reasoning uncertainty at both token and sentence levels, which reflects the unreliability degree of generated tokens and sentences. Finally, guided by reasoning uncertainty, we develop multi-level uncertainty-based adjustment to eliminate deceptive tokens exhibiting severe uncertainty and mitigate potential hallucinations in sentences. Extensive experiments demonstrate that our RUR method consistently achieves state-of-the-art performance on five benchmarks.